Commit Graph
46 Commits
Author SHA1 Message Date
wassnameandClaudypoo 339c67fefe scratch: readout UAT/calibration scripts
calib_pos_C (fluent +C knee), smoke_sweep (coherence_sweep sampling gives
ans_std>0), proto_lens_slice (reference compute_slice readout prototype:
auto-tracked tokens incl 巴黎, Paris rank->0 at the model row).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-11 11:03:20 +08:00
wassnameandClaudypoo 9149e8c1d1 demo: quantitative readouts -- rubric coherence-sweep + reference lens-rank
Two readouts added to the steering demo, both fresh-eyes signed off:

- rubric_score + coherence_sweep + plot_sweep: the model thinks then answers a
  forced {"ans": N} slot; we read the logprob-weighted expected digit and its
  pmass coherence. coherence_sweep walks C outward from 0 both ways, stopping a
  side when the answer slot goes incoherent (pmass<floor), averaging n_samples
  seeded traces (BMA) so the dose-response isn't single-sample noise. plot_sweep
  colours points by pmass with the ramp anchored to [floor-0.15,1] (0-1 washed
  every point one colour) and a red cutoff line. show_steer gains a `rubric` arg.

- lens_slice_ranks + plot_lens_slice: render jlens's own compute_slice output
  (the reference's auto token selection over the full layer grid + full-vocab
  ranks + J=I model row) as a table + rank-vs-depth plot, rather than reimplement
  it. Rank, not raw lens-logit, is comparable across layers. CJK-first font so
  multilingual tokens (e.g. the auto-surfaced 巴黎) render in the legend.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-11 11:03:13 +08:00
wassnameandClaudypoo 2062bcbb1f persona_steering_v2 notebook: soft add+clamp, masked topk, pinv, mean_diff, rubric readout
Generated by scripts/scratch/build_persona_v2.py; queued for headless
execution (outputs committed after the run).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-11 08:05:09 +08:00
wassnameandClaudypoo 038a651197 persona refinements: soft-cotangent vector, pinv tangent transport, topk word-like mask
- persona_soft_vector: w = W_U^T (softmax(u_pos/T) - softmax(u_neg/T)) over
  word-like tokens; gradient of the expected-logprob contrast (genuine
  cotangent), full-vocab replacement for hard top-k; logs TV distance.
- persona_pinv_vector: h_diff is a tangent, so solve J delta = h_diff
  (ridge lstsq) instead of the J^T type error; logs per-layer residual.
- persona_topk_vector: mask emoji/special tokens out of the contrast
  (they were the degenerate emit-targets behind the C=1.5 emoji spam).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-11 08:05:01 +08:00
wassnameandClaudypoo 143e9add80 demo: rubric readout -- think then {"ans":N}, logprob-weighted expected digit per C
show_steer gains a rubric= param and rubric_score(): the model rates a 0-9 axis,
we force the {"ans": slot and read the logprob-weighted expected digit. guided.py's
mechanism reduced to one scalar for the demo (rigorous K-way debiased version stays
in moral-maps). UAT (scripts/scratch/uat_rubric.py) on happy/joy: in the coherent
window ans rises 3.52->4.99->8.06 across C=-0.5,0,+0.5 (pmass=1.00); at the
degeneration extremes (C=+-1.5) pmass collapses to ~0 and the number is correctly
flagged meaningless.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-11 07:40:38 +08:00
wassnameandClaudypoo fa302a3ac6 word_steering: fix clamp fluency overclaim (fresh-eyes catch)
Claimed 'C~6 stays fluent, C>=8 spams', but the committed C=6 output degenerates
into 'Ihopeyouarehappy!' repetition after a coherent opening (my calib probe only
read the first 180 chars and missed the tail). Corrected: C~3 is fluent, C~6 reads
happy then collapses, clamp's clean window is narrow (<=~4); C=6 shown as the
degeneration edge (like add's C=1.5). Outputs unchanged (comment-only edit).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 21:08:44 +08:00
wassnameandClaudypoo 92e05478ce persona_steering: executed headless with j-thoughts labels + cowsay + delivery fix
UAT: persona_topk now logs 'j-thoughts (content of mental workspace)' with
contrastive positive [❤ 😊 happy ...] vs negative [Worse 绝望 Panic ...] tokens
(the contrast-before-topk fix), all three methods run, mean_diff baseline no
longer crashes on the delivery tag.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 21:05:12 +08:00
wassnameandClaudypoo f65e39ed66 demo: omit delivery tag when cfg has no apply_mode (fixes MeanDiffC baseline crash)
show_steer's header assumed vec.cfg.apply_mode, but steering-lite's own configs
(MeanDiffC persona baseline) don't have it -- only jsteer's configs do. getattr
-> None -> no tag, so the mean_diff cell in persona_steering runs again.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 21:00:41 +08:00
wassnameandClaudypoo 554af789e1 word_steering: calibrate clamp Cs=(0,3,6), drop degenerate replace_last demo
Ran scripts/scratch/calib_delivery.py: clamp C~6 reads clearly happy while fluent
(C>=8 spams); replace_last is gibberish at every C (0.05..0.25) because with span=1
it overwrites every generated token's residual across the band, so it can't build
coherent text. Dropped its demo cell, documented why in the markdown (it's a
fixed-prompt-span injection tool, not a generation-steering one). Executed headless:
cowsay readout + raw special-token output + clamp/add_last coherent.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:55:38 +08:00
wassnameandClaudypoo a091cc7527 persona_topk: demote per-layer |J^T w| trace to debug; add short README demo
- the pre-norm per-layer norm line is a fit-health trace, not demo output -> logger.debug
- README: short 'Persona j-thoughts' section showing the contrast-first extraction
  (clean positive/negative tokens), kept in the experimental/untested-specificity frame

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:37:05 +08:00
wassnameandClaudypoo d52b7ece7d persona_topk: log j-thoughts as positive:/negative: (was pos>neg/neg>pos)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:35:29 +08:00
wassnameandClaudypoo 07759f73ce demo: apply_mode kwarg for delivery-mode demos + cthulhu-mini cowsay readout
- show_steer(apply_mode=, apply_span=) swaps delivery (add|clamp|add_last|
  replace_last) by rebuilding the cfg, no re-extraction -- delivery is decoupled
  from extraction (applies.py), so the demo layer is where you pick the mode
- j-space readout now speaks from a mini cowsay bubble (^(;,;)^)
- word_steering.ipynb: new 'Delivery modes' section, one demo per mode, each
  with its C=0 semantics called out (clamp C=0=ablation, replace_last C=0=zero)

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:34:01 +08:00
wassnameandClaudypoo 00395ce398 persona_topk: contrast logits BEFORE top-k, not top-k of each then contrast
User: 'we need to contrast then take the top k, otherwise we just get the'. Both
persona means unembed to the same generic high-freq tokens (\n, ' I', ' The'),
so topk(pos) ~= topk(neg) and the contrast collapses to null. Take the top-k of
(logits_pos - logits_neg) instead: the tokens each persona evokes MORE than the
other, where the actual persona signal lives.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:19:28 +08:00
wassnameandClaudypoo bd84c77ceb demo: print raw generation (skip_special_tokens=False), drop split_think
User: dont fabricate think tokens; show raw so it's debuggable. The chat
template already emits real <think>/</think> -- parsing them out and re-wrapping
with my own tags hid the raw text and invented tokens. Now decode with special
tokens on and print the model's output verbatim (real <think>, <|im_end|>).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:17:32 +08:00
wassnameandClaudypoo 45b8aa66aa demo: clear C-block separators, honest unclosed-<think> handling, 512 tok default
- header carries name/method/prompt once; per-C blocks only vary in C (was repeated every block)
- split_think returns `closed`; unclosed reasoning is labelled UNCLOSED instead of masquerading as the answer
- max_new_tokens 256->512 so Qwen3 <think> actually closes and an answer appears

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 20:16:02 +08:00
wassnameandClaudypoo 99b50695bd config: docstrings say default is pre-fitted RAW lens, chat_corpus is fallback-only
Fresh-eyes caught stale framing: the module/chat_corpus docstrings still sold chat-
templated fitting as the corpus convention. The default demo path now loads the
authors' pre-fitted raw-wikitext lens; chat_corpus only feeds the local-fit fallback,
and chat-vs-raw was never compared head-to-head (run-524 used chat), so it's flagged
unresolved rather than claimed better.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 19:24:51 +08:00
wassnameandClaudypoo 220f5018ac scratch: pre-fitted UAT + C-calibration probes + run_nb tool; drop dead 9B/guard scripts
Adds this session's evidence: uat_prefitted_4b.py (proves the n1000 lens steers,
word>random), calib_c_prefitted.py + calib_persona.py (how the demo Cs were chosen:
word knee ~0.5, mean_diff ~1), run_nb.py (nbclient notebook executor, bypasses the
broken global nbconvert config). Removes u4_step3_guard.sh / retry.sh; fit.py gains
--out for scratch fits to non-canonical paths.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 19:22:42 +08:00
wassnameandClaudypoo 6b28b72651 persona_steering: load pre-fitted lens, all 3 variants, calibrated C
Loads the same n1000 Hub lens, steers on steer_band. Shows all variants at honest
calibrated coefficients: persona_vector (C~0.5, drifts to CJK/off-topic by 1.5),
persona_topk (pos/neg top-k collapse to generic starters -> ~null contrast),
mean_diff baseline (needs C~1, cleanest steered+fluent of the three). Markdown SHOULDs
updated to match. UAT: nbclient executes end-to-end, all three variants render.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 19:22:13 +08:00
wassnameandClaudypoo d4b74d329d word_steering + README: load pre-fitted lens, recalibrate C to the steep knee
Notebook and hello-world now load the n1000 Hub lens (no local fit) and steer on
jac.steer_band(model). The clean pre-fitted direction has a much steeper coherence
knee than the old coarse fit: C~0.5 shifts tone while staying fluent, C~1 degenerates.
UAT: nbconvert/nbclient executes end-to-end; C=0.5 lens shows Happy climbing, C=1.5
over-drives to joyjoy, masked lens resolves Eiffel Tower city->Paris.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 19:20:07 +08:00
wassnameandClaudypoo a43630cec2 feat: load authors' pre-fitted n1000 Hub lenses; steer_band + masked lens_topk
config.HUB_LENS_FILE maps HF model -> the authors' pre-fitted Jacobian lens on the
Hub (neuronpedia/jacobian-lens, raw Salesforce-wikitext, n=1000). Loading one beats
fitting locally: same estimator, 1000 prompts, zero compute. Our Jacobian already
wraps jlens.JacobianLens, so their .pt loads through Jacobian.from_pretrained with no
format change (verified: n1000 4B loads, d_model=2560, layers [0..30]).

jacobian.py:
- steer_band(model, lo=0.3, hi=0.9): pre-fitted lenses span every layer; steering all
  of them over-drives the residual, so restrict to the mid-depth band run-524 used.
- lens_topk reuses jlens.vis._meaningful_token_mask so j-space readouts hide
  punctuation/single-char/special tokens (per the walkthrough these trail the
  interesting word tokens on Qwen). Verified: Eiffel Tower resolves city->Paris clean.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 19:20:06 +08:00
wassnameandClaudypoo 69bd650dc5 scratch: move u4_step3 loop-close scripts (fit4b/guard/retry) out of scripts/ top
The U4 loop-close is a separate finished-enough goal from the demo; guard killed.
scripts/ top level is now just fit.py + smoke.py.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 16:31:47 +08:00
wassnameandClaudypoo 8fd82bf731 fit: tqdm progress bar + first-prompt trace + summary; config configures loguru on import
- Jacobian.fit wraps prompts in tqdm (jlens has no bar; safe since fit only
  enumerate/len's them), logs the full first prompt (special tokens on, SHOULD
  line) and a done-summary -- token-efficient-logging style, both tqdm intervals set
- config.py sets up loguru on import (compact single-char icons, routed through
  tqdm.write so bars survive), so every script/notebook importing config gets it
- notebooks drop their manual logger setup and import config in cell 1

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 16:31:25 +08:00
wassnameandClaudypoo 69d7d2f1d9 demo: apply gpt-5.5 review -- float-C fix, trim comments, soften overclaims
External review (docs/reviews/code_demo.md) triaged scout-mindset:
- FIX float-C crash: C={C:+g} not {:+d} (steering coeffs are floats)
- ACCEPT trim: shorter demo.py/config/fit docstrings (user also flagged verbosity)
- ACCEPT soften "fit J where we steer" -> "closer to the chat distribution" (most
  fitted positions are user/doc tokens, not assistant <think>; run-524 went further)
- ACCEPT soften "what the model thinks" -> "lens readout (linear approx)"
- ADD seed to show_steer so per-C blocks are comparable under sampling
- REJECT "</think> stripped by skip_special_tokens" -- verified false: decode keeps
  think tags (they're added, not registered-special tokens), split_think works

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 15:02:42 +08:00
wassnameandClaudypoo 57a0276ae6 U4 step3 guard: back off when GPU/RAM in use, yield to the live user
Discovered the user is actively developing jsteer (commit 80353ed 90s ago,
live 16GB VRAM kernel from their demo work) -- they are NOT afk. The blind
requeue guard would compete with their interactive work and risk oomd-killing
THEIR process. Now the guard only launches the fit when GPU free >=13GB and
host avail >=20GB, so it fills genuinely-idle windows (overnight) and never
fights the human for their own machine. Killed the competing fit (555) to
yield the GPU now; checkpoint preserved at n_done=69.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:53:35 +08:00
wassnameandClaudypoo 17f5310660 demo: resumable notebook fits (checkpoint_path) + README to chat-template show_steer
- both notebook fit cells pass checkpoint_path so a multi-hour 4B fit survives an OOM
- README hello-world uses show_steer + qwen3.5-4b.jac (was raw greedy + 0.6b cache),
  coefficient note made model-generic

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:52:51 +08:00
wassnameandClaudypoo bba61807b2 demo: persona_steering.ipynb to 3.5-4B + chat template + show_steer
Same rewire as word_steering: chat_corpus fit, show_steer j-space/<think>/answer
display, dim_batch=4. Markdown 0.6B-specific null-result claims softened to
expectations (results re-run on 4B); the specificity-control finding kept.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:50:32 +08:00
wassnameandClaudypoo 80353ed62c demo: chat-template extract+generate, j-space+<think> display (jsteer.demo.show_steer)
The verified run-524 vectors were fit on chat-templated prompts (u4_prompts.json),
so fitting raw WikiText diverged from what worked. Now:
- config.chat_corpus wraps jlens WikiText in the chat template (fit J where we steer)
- jsteer.demo.show_steer generates through apply_chat_template(enable_thinking) with
  the model's own generation_config sampling, splits </think>, shows lens_topk j-space
  readout + reasoning + answer as Tufte small-multiples per C
- word_steering.ipynb rewired to Qwen3.5-4B, dim_batch=4 (3090-safe 4B), show_steer
- fit.py defaults to Qwen3.5-4B + chat_corpus

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:48:46 +08:00
wassnameandClaudypoo defbe0c483 U4 step3: external oomd-resilient requeue guard
systemd-oomd kills the fit's whole pueue task-scope (task 554 died with bash
+ python together, no retry line), defeating the in-cgroup retry wrapper. This
guard runs detached in its own session/cgroup (~0 RAM, so oomd ignores it) and
keeps the resumable fit queued until u4_loopclose.txt appears. Grinds through
the user's active-hours GPU bursts, finishes clean overnight. Layered with the
retry wrapper (in-cgroup CUDA-OOM retry) for both kill modes.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:45:18 +08:00
wassnameandClaudypoo 57c8d4b166 add clamp apply mode: pin v-component to C instead of accumulating
clamp: y += (C - <y,v_hat>)v_hat at all positions -- bounded perturbation
regardless of generation length, vs add's per-step accumulation via KV cache.
C=0 is directional ablation. Smoke (Qwen3-0.6B, happy/joy): clamp C=+20 stays
coherent and on-concept (drifts to 'happiness and joy of my childhood', in
Chinese) while add C=+8 already degenerates to 'joyjoyjoy...'.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:42:22 +08:00
wassname 474f74ac33 wip 2026-07-10 14:36:55 +08:00
wassnameandClaudypoo 265a0789c9 U4 step3: auto-resume wrapper to self-heal through user's GPU bursts
Three OOM kills (551/552/553) from the user's bursty live VS Code kernel, but
jlens resumes from checkpoint each time (n_done 36->45->64, monotonic). Rather
than predict the bursts, relaunch until exit 0. MAX_RETRIES=40 caps a genuine
no-progress bug; n_done logged per retry to distinguish OOM from a real crash.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:31:20 +08:00
wassnameandClaudypoo 4c93aaddb5 U4 step3: dim_batch 8->4 + expandable_segments after 2nd OOM
552 CUDA-OOM'd at n_done=45: the user's VS Code GPU kernel grew to 8.18GB
while this fit's 13.23GB hit the 23.5GB ceiling (44MB free, fragmentation).
dim_batch=4 shrinks the fit to ~10.5GB (polite co-tenant, leaves user ~13GB)
and expandable_segments:True defragments (the OOM's own suggestion). Still
only changes the backward schedule, not the Jacobian. Resumes from n_done=45.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 14:10:39 +08:00
wassnameandClaudypoo ebb912452d U4 step3: dim_batch 16->8 to survive OOM contention with user's VS Code kernel
Run 551 was OOM-killed at n_done=36: the user's VS Code Jupyter kernel
(jsteer venv, PID 3214401) co-loaded ~1.5GB VRAM + 1.9GB RAM while the fit
sat at the 22.4/24.6GB ceiling. Clean SIGKILL with no CUDA traceback = host
OOM killer, not a CUDA OOM. dim_batch=8 halves the fit's peak footprint;
it changes only the backward schedule, not the accumulated Jacobian, so U4
exactness is preserved. Resumes from checkpoint (n_done=36), lossless.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:56:17 +08:00
wassnameandClaudypoo beec9189da U4 step 2 PASS: jsteer VJP == regenerated run-524 vector, cos +1.000000 all 21 layers
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:20:53 +08:00
wassnameandClaudypoo 261a79ebdd fix: _to_vector forces CPU fp32 -- vjp path returned cuda Vectors, crashing U4 step-2 cosine vs cpu reference
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:19:54 +08:00
wassnameandClaudypoo 6f59986139 apply fresh-eyes UAT: jlens availability stated in README install, persona -C degeneration narrated honestly
U2 verdict: hello-world PASS verbatim zero-edit; fixes are the PARTIAL FAIL
(jlens admission was buried in a pyproject comment) and the notebook SHOULD
promising '-C more negative' when persona_vector at -2 actually degenerates.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:19:03 +08:00
wassnameandClaudypoo 2a05dbc799 U4 loop-close scripts: regenerate ref-524 vector, port check, 4B fit (pueue 549-551)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:13:50 +08:00
wassnameandClaudypoo 0e7126ba6e README: pitch, install, verbatim-tested hello-world, API table, evidence
Hello-world block extracted and run verbatim from repo root (passes).
Evidence section: 3/5 moral foundations vs norm-matched random on ONE
model/eval; persona variants marked experimental (failed specificity).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:10:15 +08:00
wassnameandClaudypoo dbd526bed0 notebook: persona steering (EXPERIMENTAL framing, executed)
persona_vector and mean_diff baseline both move tone at C=2;
persona_topk is an honest null on this setup (both personas evoke the
same generic sentence starters, contrast=0) and the notebook says so.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:10:15 +08:00
wassnameandClaudypoo a2bf5c31a0 notebook: word steering hello-world (executed on real 0.6B cache)
C sweep shows coherence/strength tradeoff (C=1 fluent+happy, C=8 spam),
greedy+sampled at C=1, negative steering, lens_topk layer progression
(city slot -> candidates -> Paris), Vector save/load. Adds nbconvert to
notebooks extra; ignores notebook-produced safetensors.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:10:15 +08:00
wassnameandClaudypoo 4fb1cfd63e apply external review (deepseek): reject float bands post-fit, assert->ValueError, 2 clarifying comments
Review verdict APPROVE; rejected findings (empty-prompts guard = preemptive
defensive check, from_hf dedup = ms-scale, lm.forward swap = loses attention
mask on padded batches) documented in docs/reviews/code.md triage.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:05:27 +08:00
wassnameandClaudypoo e2a28a9c5d fit script: 64-prompt reproducible Qwen3-0.6B Jacobian cache
Hardcoded diverse web-text prompts (min 21 tokens), layers 0.3-0.9,
resumable checkpoint, saves artifacts/qwen3-0.6b.jac.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:03:18 +08:00
wassnameandClaudypoo f22bbeb0d2 parity(U1): cached-J pullback vs direct VJP, all layers cos>0.999
min cos 0.999801 (layer 8), rising to 0.999996; gap is fp16 cache
storage as expected. Gate PASS.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 12:52:18 +08:00
wassnameandClaudypoo 6bf51a75cc smoke: fit+steer Qwen3-0.6B on happy/joy end-to-end
+C collapses to "joy", -C to negative tone: sign correct. Coherence
breaks at |C|=8 uncalibrated, as expected.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 12:51:08 +08:00
wassnameandClaudypoo aa44834cb3 uv: switch to local editable path deps; add tabulate
jlens git source is not fetchable; both deps are locally co-developed.
tabulate is imported by steering_lite.calibrate at module load but only
declared in its test extras.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 12:48:47 +08:00
wassnameandClaudypoo 1f458fd53e core algo: Jacobian fit/cache/pullback (jlens) + word/persona vectors (steering-lite runtime)
Ported from the verified j-steer-dev experiment (word pullback beat random
on 3/5 moral foundations, Qwen3-4B n=3). jacobian.py wraps jlens fit/save/
load and derives steering vectors by CPU matvec against the cached J;
vjp.py is the one-backward parity path; applies.py registers the methods
and delivery modes into steering-lite.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 12:08:07 +08:00