Files
wassnameandClaudypoo a428af4413 docs: oracle (deepseek-v4-pro) review + triage on why steering fails the verdict
Oracle's sharp reads adopted: ans_mass collapse is the primary signal not a nuisance;
||J^T w|| pre-norm magnitude is the missing diagnostic (already logged, we discard it
by unit-normalizing); no-think zero-shot sweep is the cheapest test of CoT-buffering vs
off-target-direction; meandiff tie => Jacobian adds nothing for persona-contrast on a
verdict. AGENTS records the result + next steps.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-12 16:29:46 +08:00

111 lines
7.1 KiB
Markdown

# jsteer -- agent notes
Plan of record: /home/wassname/.claude/plans/review-specs-00-minimal-experiment-md-an-peppy-sky.md
Evidence base: ../j-steer-dev/docs/RESEARCH_JOURNAL.md (verified 3/5 word-steering result).
## What this is
repeng-style UX for Jacobian pullback steering. Core algo files are WRITTEN
(by the main agent, ported from verified j-steer-dev code -- do not rewrite
the math, it is parity-gated against the verified experiment):
- `jsteer/jacobian.py` -- Jacobian.fit/save/load (wraps jlens, the
researchers' verified primary code) + word/persona/persona_topk/random
vectors -> steering_lite.Vector.
- `jsteer/applies.py` -- steering-lite method registration + delivery modes.
- `jsteer/vjp.py` -- direct VJP path; the parity reference for the cache.
Runtime is steering-lite: `with v(model, C=8): model.generate(...)`.
## Status (shipped + verified)
- Core API built: `Jacobian.fit/save/load/from_pretrained` + word / persona /
persona_topk / random vectors, one shared pullback path (`jacobian.py`);
delivery modes (add / add_last / replace_last) in `applies.py`.
- Smoke (`scripts/smoke.py`) and any-model fit (`scripts/fit.py --model ...`,
prompts from jlens's WikiText corpus) green; `config.py` holds slug/paths.
- U1 parity gate PASS -- cache pullback == direct VJP, cos > 0.999:
`docs/evidence/parity_u1.txt`.
- U4 port check PASS -- jsteer VJP == run-524 reference vector, cos +1.0:
`docs/evidence/u4_step2_vjp_parity.txt`.
- Notebooks: `word_steering` (verified), `persona_steering` (experimental --
failed specificity controls in j-steer-dev, framing kept honest).
- README at classic-repeng length with an honest evidence section.
## Open
- U4 loop-close (`scripts/u4_step3_fit4b.py`): full 4B fit -> cached word
vector must match the VJP and run-524 vectors (cos > 0.999). Resumable from
`artifacts/qwen3-4b-authority.ckpt`; writes `artifacts/u4_loopclose.txt`.
- One-off validation scripts live in `scripts/scratch/` (u4_step1/2, parity_u1).
- TODO eval notebook: steer -authority, tinymfv fast (N=16, tokens=16, mfq-2)
vs unsteered baseline.
- OPEN: steering-demo calibration is not method-comparable yet. To compare
methods you want the same OFF-TARGET budget, then read the on-target effect.
The current search finds each method's own "max coherent" edge, which does NOT
equalize off-target -- but the CAUSE (checked in artifacts/steering_demo_results
.json) is NOT that `rep` is insensitive. It is that the dual gate
`rep<0.35 AND ans_mass>0.5` let `ans_mass` pre-empt: 11 of 14 edges stopped on
`ans_mass<0.5` with rep still 0.00-0.02, nowhere near its gate. rep never got
to fire, so we can't conclude it fails as a calibration axis.
`ans_mass` is answer-commitment (confidence, same family as the rejected
`pmass`), a READOUT-VALIDITY concern, not off-target coherence; folding it into
the search is what broke comparability. (Interesting: ans_mass drops before rep
rises -> the steer makes the model hedge/refuse the answer BEFORE its reasoning
goes incoherent. Real effect, keep it as a per-row flag.)
FIX (simpler than a graded measure): calibrate on `rep` ALONE (iso-rep budget,
each method to rep~=0.35), demote `ans_mass` to a per-row "is P(YES) valid"
flag out of the search, then compare on-target at iso-rep. If the readout is
invalid at the rep-budget for a method, that is itself a finding.
0.35 is anchored to the empirical coherent/degenerate gap (rep<~0.3 vs >~0.6,
rep_metric_check.py), not to a base degradation; base C=0 rep=0.00, gap is wide
so anything ~[0.35,0.55] gives the same edge.
Caveat: the "rep non-monotone in C" anomaly (word -0.35 rep=1.0 vs -0.70
rep=0.34) is UNCHECKED -- read the traces qualitatively; likely a short-trace/
seed artifact, not real.
- OPEN (scale, blocks cross-method comparison): the coefficient C is NOT on a
comparable scale across methods -- each vector v has its own norm, so C=0.5 for
`word` != C=0.5 for `persona_pinv`. Report the scale-invariant perturbation
instead: rho = ||C*v|| / ||h|| (fraction of the residual-stream norm at the
steered layers). Until then the C*+/C*- columns are per-method, not comparable.
Also: `max_C` should never bind (raised to 1e5, safety only); the real search
limiter is `budget` (~6 evals -> Illinois edge is +-~20% of the rep budget, and
robust methods cap out via too-few step-outs, flagged at_budget=False). rep is
single-seed noisy too. So the current swing/score numbers are directionally
useful but NOT yet a meaningful comparable scale -- fix rho + raise budget +
multi-seed before trusting cross-method ranks. (min-C floor: also consider,
raised by wassname, TBD.)
- RESULT (task 42, rep-only calibration): it OVER-STEERS. At the rep=0.35 edge
ans_mass has collapsed to 0.00-0.40 (base 0.56) for ~every method, so P(YES)
there is read off dead answers -- the big swings (0.65-0.97) are ARTIFACTS.
score (validity-weighted) correctly nukes them to ~1e-5..1e-19. Finding: for a
YES/NO VERDICT readout, answer-commitment (ans_mass) dies BEFORE repetition
(rep) breaks, so the binding off-target budget for a readable verdict is
ans_mass, not rep. => Partly REVERSES "rep alone": the dual-gate STRUCTURE
min(rep, ans_mass) was right, the error was calling ans_mass "coherence" (it is
READOUT-VALIDITY). FIX (needs wassname nod, reverses his rep-only call): edge =
min(rep-margin, ans_mass_valid-margin) with ans_mass base-anchored at 0.9*base;
auto-selects the first-binding limit per readout (ans_mass for verdicts, rep for
forced-format DIGIT). Evidence: artifacts/steering_demo_results.json (task 42).
- RESULT (task 43, dual-gate regenerated): the gate fix works structurally (word's fake
swing now nulled: score -0.06, readout_ok=False) but the substantive finding is NEGATIVE:
NONE of the 7 methods flip the deliberated YES/NO while keeping the answer alive. Inside
the valid budget every method's swing is noise-level (<=0.06, vs random 0.03; baseline
P(YES)=0.107) and the promoted tokens are function-word/cross-lingual JUNK, not lie/honest
-- the concept direction isn't reaching the output at a safe C. Only word moves P(YES)
(0.97) and only by derailing (ans_mass 0.14, rambles off-task). readout_ok=False for 6/7,
but for the personas that is a single-seed boundary miss (persona_soft 0.895 vs 0.90 floor)
NOT over-steer -- the instrument is single-seed greedy and noisy near the edge. meandiff
ties the Jacobian variants => the Jacobian adds nothing for persona-contrast on a verdict.
Oracle (deepseek-v4-pro) review + triage: docs/reviews/oracle_steering.md. NEXT (priority):
(1) surface per-method ||J^T w|| pre-norm magnitude (already logged at DEBUG, jacobian.py
:240; predicted word>>personas~0 -- unit-norm amplifies dead persona pullbacks to noise);
(2) no-think zero-shot P(YES) sweep to separate "CoT buffers the offset" from "direction is
off-target"; (3) multi-seed the edge (readout_ok flickers on single-seed noise). Deliverable:
nbs/steering_demo.py (dual-gate table + Claude qualitative read + per-method generations).
## Style
Fail fast, no defensive programming, loguru, no LLM-tell prose in README.
Comments marked as Claude-authored where opinionated.