Files
wassnameandClaudypoo a428af4413 docs: oracle (deepseek-v4-pro) review + triage on why steering fails the verdict
Oracle's sharp reads adopted: ans_mass collapse is the primary signal not a nuisance;
||J^T w|| pre-norm magnitude is the missing diagnostic (already logged, we discard it
by unit-normalizing); no-think zero-shot sweep is the cheapest test of CoT-buffering vs
off-target-direction; meandiff tie => Jacobian adds nothing for persona-contrast on a
verdict. AGENTS records the result + next steps.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-12 16:29:46 +08:00

7.1 KiB

jsteer -- agent notes

Plan of record: /home/wassname/.claude/plans/review-specs-00-minimal-experiment-md-an-peppy-sky.md Evidence base: ../j-steer-dev/docs/RESEARCH_JOURNAL.md (verified 3/5 word-steering result).

What this is

repeng-style UX for Jacobian pullback steering. Core algo files are WRITTEN (by the main agent, ported from verified j-steer-dev code -- do not rewrite the math, it is parity-gated against the verified experiment):

  • jsteer/jacobian.py -- Jacobian.fit/save/load (wraps jlens, the researchers' verified primary code) + word/persona/persona_topk/random vectors -> steering_lite.Vector.
  • jsteer/applies.py -- steering-lite method registration + delivery modes.
  • jsteer/vjp.py -- direct VJP path; the parity reference for the cache.

Runtime is steering-lite: with v(model, C=8): model.generate(...).

Status (shipped + verified)

  • Core API built: Jacobian.fit/save/load/from_pretrained + word / persona / persona_topk / random vectors, one shared pullback path (jacobian.py); delivery modes (add / add_last / replace_last) in applies.py.
  • Smoke (scripts/smoke.py) and any-model fit (scripts/fit.py --model ..., prompts from jlens's WikiText corpus) green; config.py holds slug/paths.
  • U1 parity gate PASS -- cache pullback == direct VJP, cos > 0.999: docs/evidence/parity_u1.txt.
  • U4 port check PASS -- jsteer VJP == run-524 reference vector, cos +1.0: docs/evidence/u4_step2_vjp_parity.txt.
  • Notebooks: word_steering (verified), persona_steering (experimental -- failed specificity controls in j-steer-dev, framing kept honest).
  • README at classic-repeng length with an honest evidence section.

Open

  • U4 loop-close (scripts/u4_step3_fit4b.py): full 4B fit -> cached word vector must match the VJP and run-524 vectors (cos > 0.999). Resumable from artifacts/qwen3-4b-authority.ckpt; writes artifacts/u4_loopclose.txt.
  • One-off validation scripts live in scripts/scratch/ (u4_step1/2, parity_u1).
  • TODO eval notebook: steer -authority, tinymfv fast (N=16, tokens=16, mfq-2) vs unsteered baseline.
  • OPEN: steering-demo calibration is not method-comparable yet. To compare methods you want the same OFF-TARGET budget, then read the on-target effect. The current search finds each method's own "max coherent" edge, which does NOT equalize off-target -- but the CAUSE (checked in artifacts/steering_demo_results .json) is NOT that rep is insensitive. It is that the dual gate rep<0.35 AND ans_mass>0.5 let ans_mass pre-empt: 11 of 14 edges stopped on ans_mass<0.5 with rep still 0.00-0.02, nowhere near its gate. rep never got to fire, so we can't conclude it fails as a calibration axis. ans_mass is answer-commitment (confidence, same family as the rejected pmass), a READOUT-VALIDITY concern, not off-target coherence; folding it into the search is what broke comparability. (Interesting: ans_mass drops before rep rises -> the steer makes the model hedge/refuse the answer BEFORE its reasoning goes incoherent. Real effect, keep it as a per-row flag.) FIX (simpler than a graded measure): calibrate on rep ALONE (iso-rep budget, each method to rep~=0.35), demote ans_mass to a per-row "is P(YES) valid" flag out of the search, then compare on-target at iso-rep. If the readout is invalid at the rep-budget for a method, that is itself a finding. 0.35 is anchored to the empirical coherent/degenerate gap (rep<~0.3 vs >~0.6, rep_metric_check.py), not to a base degradation; base C=0 rep=0.00, gap is wide so anything ~[0.35,0.55] gives the same edge. Caveat: the "rep non-monotone in C" anomaly (word -0.35 rep=1.0 vs -0.70 rep=0.34) is UNCHECKED -- read the traces qualitatively; likely a short-trace/ seed artifact, not real.
  • OPEN (scale, blocks cross-method comparison): the coefficient C is NOT on a comparable scale across methods -- each vector v has its own norm, so C=0.5 for word != C=0.5 for persona_pinv. Report the scale-invariant perturbation instead: rho = ||Cv|| / ||h|| (fraction of the residual-stream norm at the steered layers). Until then the C+/C*- columns are per-method, not comparable. Also: max_C should never bind (raised to 1e5, safety only); the real search limiter is budget (~6 evals -> Illinois edge is +-~20% of the rep budget, and robust methods cap out via too-few step-outs, flagged at_budget=False). rep is single-seed noisy too. So the current swing/score numbers are directionally useful but NOT yet a meaningful comparable scale -- fix rho + raise budget + multi-seed before trusting cross-method ranks. (min-C floor: also consider, raised by wassname, TBD.)
  • RESULT (task 42, rep-only calibration): it OVER-STEERS. At the rep=0.35 edge ans_mass has collapsed to 0.00-0.40 (base 0.56) for ~every method, so P(YES) there is read off dead answers -- the big swings (0.65-0.97) are ARTIFACTS. score (validity-weighted) correctly nukes them to ~1e-5..1e-19. Finding: for a YES/NO VERDICT readout, answer-commitment (ans_mass) dies BEFORE repetition (rep) breaks, so the binding off-target budget for a readable verdict is ans_mass, not rep. => Partly REVERSES "rep alone": the dual-gate STRUCTURE min(rep, ans_mass) was right, the error was calling ans_mass "coherence" (it is READOUT-VALIDITY). FIX (needs wassname nod, reverses his rep-only call): edge = min(rep-margin, ans_mass_valid-margin) with ans_mass base-anchored at 0.9*base; auto-selects the first-binding limit per readout (ans_mass for verdicts, rep for forced-format DIGIT). Evidence: artifacts/steering_demo_results.json (task 42).
  • RESULT (task 43, dual-gate regenerated): the gate fix works structurally (word's fake swing now nulled: score -0.06, readout_ok=False) but the substantive finding is NEGATIVE: NONE of the 7 methods flip the deliberated YES/NO while keeping the answer alive. Inside the valid budget every method's swing is noise-level (<=0.06, vs random 0.03; baseline P(YES)=0.107) and the promoted tokens are function-word/cross-lingual JUNK, not lie/honest -- the concept direction isn't reaching the output at a safe C. Only word moves P(YES) (0.97) and only by derailing (ans_mass 0.14, rambles off-task). readout_ok=False for 6/7, but for the personas that is a single-seed boundary miss (persona_soft 0.895 vs 0.90 floor) NOT over-steer -- the instrument is single-seed greedy and noisy near the edge. meandiff ties the Jacobian variants => the Jacobian adds nothing for persona-contrast on a verdict. Oracle (deepseek-v4-pro) review + triage: docs/reviews/oracle_steering.md. NEXT (priority): (1) surface per-method ||J^T w|| pre-norm magnitude (already logged at DEBUG, jacobian.py :240; predicted word>>personas~0 -- unit-norm amplifies dead persona pullbacks to noise); (2) no-think zero-shot P(YES) sweep to separate "CoT buffers the offset" from "direction is off-target"; (3) multi-seed the edge (readout_ok flickers on single-seed noise). Deliverable: nbs/steering_demo.py (dual-gate table + Claude qualitative read + per-method generations).

Style

Fail fast, no defensive programming, loguru, no LLM-tell prose in README. Comments marked as Claude-authored where opinionated.