mirror of
https://github.com/wassname/jsteer.git
synced 2026-09-09 11:25:03 +08:00
Oracle's sharp reads adopted: ans_mass collapse is the primary signal not a nuisance; ||J^T w|| pre-norm magnitude is the missing diagnostic (already logged, we discard it by unit-normalizing); no-think zero-shot sweep is the cheapest test of CoT-buffering vs off-target-direction; meandiff tie => Jacobian adds nothing for persona-contrast on a verdict. AGENTS records the result + next steps. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
7.1 KiB
7.1 KiB
jsteer -- agent notes
Plan of record: /home/wassname/.claude/plans/review-specs-00-minimal-experiment-md-an-peppy-sky.md Evidence base: ../j-steer-dev/docs/RESEARCH_JOURNAL.md (verified 3/5 word-steering result).
What this is
repeng-style UX for Jacobian pullback steering. Core algo files are WRITTEN (by the main agent, ported from verified j-steer-dev code -- do not rewrite the math, it is parity-gated against the verified experiment):
jsteer/jacobian.py-- Jacobian.fit/save/load (wraps jlens, the researchers' verified primary code) + word/persona/persona_topk/random vectors -> steering_lite.Vector.jsteer/applies.py-- steering-lite method registration + delivery modes.jsteer/vjp.py-- direct VJP path; the parity reference for the cache.
Runtime is steering-lite: with v(model, C=8): model.generate(...).
Status (shipped + verified)
- Core API built:
Jacobian.fit/save/load/from_pretrained+ word / persona / persona_topk / random vectors, one shared pullback path (jacobian.py); delivery modes (add / add_last / replace_last) inapplies.py. - Smoke (
scripts/smoke.py) and any-model fit (scripts/fit.py --model ..., prompts from jlens's WikiText corpus) green;config.pyholds slug/paths. - U1 parity gate PASS -- cache pullback == direct VJP, cos > 0.999:
docs/evidence/parity_u1.txt. - U4 port check PASS -- jsteer VJP == run-524 reference vector, cos +1.0:
docs/evidence/u4_step2_vjp_parity.txt. - Notebooks:
word_steering(verified),persona_steering(experimental -- failed specificity controls in j-steer-dev, framing kept honest). - README at classic-repeng length with an honest evidence section.
Open
- U4 loop-close (
scripts/u4_step3_fit4b.py): full 4B fit -> cached word vector must match the VJP and run-524 vectors (cos > 0.999). Resumable fromartifacts/qwen3-4b-authority.ckpt; writesartifacts/u4_loopclose.txt. - One-off validation scripts live in
scripts/scratch/(u4_step1/2, parity_u1). - TODO eval notebook: steer -authority, tinymfv fast (N=16, tokens=16, mfq-2) vs unsteered baseline.
- OPEN: steering-demo calibration is not method-comparable yet. To compare
methods you want the same OFF-TARGET budget, then read the on-target effect.
The current search finds each method's own "max coherent" edge, which does NOT
equalize off-target -- but the CAUSE (checked in artifacts/steering_demo_results
.json) is NOT that
repis insensitive. It is that the dual gaterep<0.35 AND ans_mass>0.5letans_masspre-empt: 11 of 14 edges stopped onans_mass<0.5with rep still 0.00-0.02, nowhere near its gate. rep never got to fire, so we can't conclude it fails as a calibration axis.ans_massis answer-commitment (confidence, same family as the rejectedpmass), a READOUT-VALIDITY concern, not off-target coherence; folding it into the search is what broke comparability. (Interesting: ans_mass drops before rep rises -> the steer makes the model hedge/refuse the answer BEFORE its reasoning goes incoherent. Real effect, keep it as a per-row flag.) FIX (simpler than a graded measure): calibrate onrepALONE (iso-rep budget, each method to rep~=0.35), demoteans_massto a per-row "is P(YES) valid" flag out of the search, then compare on-target at iso-rep. If the readout is invalid at the rep-budget for a method, that is itself a finding. 0.35 is anchored to the empirical coherent/degenerate gap (rep<~0.3 vs >~0.6, rep_metric_check.py), not to a base degradation; base C=0 rep=0.00, gap is wide so anything ~[0.35,0.55] gives the same edge. Caveat: the "rep non-monotone in C" anomaly (word -0.35 rep=1.0 vs -0.70 rep=0.34) is UNCHECKED -- read the traces qualitatively; likely a short-trace/ seed artifact, not real. - OPEN (scale, blocks cross-method comparison): the coefficient C is NOT on a
comparable scale across methods -- each vector v has its own norm, so C=0.5 for
word!= C=0.5 forpersona_pinv. Report the scale-invariant perturbation instead: rho = ||Cv|| / ||h|| (fraction of the residual-stream norm at the steered layers). Until then the C+/C*- columns are per-method, not comparable. Also:max_Cshould never bind (raised to 1e5, safety only); the real search limiter isbudget(~6 evals -> Illinois edge is +-~20% of the rep budget, and robust methods cap out via too-few step-outs, flagged at_budget=False). rep is single-seed noisy too. So the current swing/score numbers are directionally useful but NOT yet a meaningful comparable scale -- fix rho + raise budget + multi-seed before trusting cross-method ranks. (min-C floor: also consider, raised by wassname, TBD.) - RESULT (task 42, rep-only calibration): it OVER-STEERS. At the rep=0.35 edge ans_mass has collapsed to 0.00-0.40 (base 0.56) for ~every method, so P(YES) there is read off dead answers -- the big swings (0.65-0.97) are ARTIFACTS. score (validity-weighted) correctly nukes them to ~1e-5..1e-19. Finding: for a YES/NO VERDICT readout, answer-commitment (ans_mass) dies BEFORE repetition (rep) breaks, so the binding off-target budget for a readable verdict is ans_mass, not rep. => Partly REVERSES "rep alone": the dual-gate STRUCTURE min(rep, ans_mass) was right, the error was calling ans_mass "coherence" (it is READOUT-VALIDITY). FIX (needs wassname nod, reverses his rep-only call): edge = min(rep-margin, ans_mass_valid-margin) with ans_mass base-anchored at 0.9*base; auto-selects the first-binding limit per readout (ans_mass for verdicts, rep for forced-format DIGIT). Evidence: artifacts/steering_demo_results.json (task 42).
- RESULT (task 43, dual-gate regenerated): the gate fix works structurally (word's fake swing now nulled: score -0.06, readout_ok=False) but the substantive finding is NEGATIVE: NONE of the 7 methods flip the deliberated YES/NO while keeping the answer alive. Inside the valid budget every method's swing is noise-level (<=0.06, vs random 0.03; baseline P(YES)=0.107) and the promoted tokens are function-word/cross-lingual JUNK, not lie/honest -- the concept direction isn't reaching the output at a safe C. Only word moves P(YES) (0.97) and only by derailing (ans_mass 0.14, rambles off-task). readout_ok=False for 6/7, but for the personas that is a single-seed boundary miss (persona_soft 0.895 vs 0.90 floor) NOT over-steer -- the instrument is single-seed greedy and noisy near the edge. meandiff ties the Jacobian variants => the Jacobian adds nothing for persona-contrast on a verdict. Oracle (deepseek-v4-pro) review + triage: docs/reviews/oracle_steering.md. NEXT (priority): (1) surface per-method ||J^T w|| pre-norm magnitude (already logged at DEBUG, jacobian.py :240; predicted word>>personas~0 -- unit-norm amplifies dead persona pullbacks to noise); (2) no-think zero-shot P(YES) sweep to separate "CoT buffers the offset" from "direction is off-target"; (3) multi-seed the edge (readout_ok flickers on single-seed noise). Deliverable: nbs/steering_demo.py (dual-gate table + Claude qualitative read + per-method generations).
Style
Fail fast, no defensive programming, loguru, no LLM-tell prose in README. Comments marked as Claude-authored where opinionated.