mirror of
https://github.com/wassname/jsteer.git
synced 2026-10-05 05:40:27 +08:00
Oracle's sharp reads adopted: ans_mass collapse is the primary signal not a nuisance; ||J^T w|| pre-norm magnitude is the missing diagnostic (already logged, we discard it by unit-normalizing); no-think zero-shot sweep is the cheapest test of CoT-buffering vs off-target-direction; meandiff tie => Jacobian adds nothing for persona-contrast on a verdict. AGENTS records the result + next steps. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
111 lines
7.1 KiB
Markdown
111 lines
7.1 KiB
Markdown
# jsteer -- agent notes
|
|
|
|
Plan of record: /home/wassname/.claude/plans/review-specs-00-minimal-experiment-md-an-peppy-sky.md
|
|
Evidence base: ../j-steer-dev/docs/RESEARCH_JOURNAL.md (verified 3/5 word-steering result).
|
|
|
|
## What this is
|
|
|
|
repeng-style UX for Jacobian pullback steering. Core algo files are WRITTEN
|
|
(by the main agent, ported from verified j-steer-dev code -- do not rewrite
|
|
the math, it is parity-gated against the verified experiment):
|
|
|
|
- `jsteer/jacobian.py` -- Jacobian.fit/save/load (wraps jlens, the
|
|
researchers' verified primary code) + word/persona/persona_topk/random
|
|
vectors -> steering_lite.Vector.
|
|
- `jsteer/applies.py` -- steering-lite method registration + delivery modes.
|
|
- `jsteer/vjp.py` -- direct VJP path; the parity reference for the cache.
|
|
|
|
Runtime is steering-lite: `with v(model, C=8): model.generate(...)`.
|
|
|
|
## Status (shipped + verified)
|
|
|
|
- Core API built: `Jacobian.fit/save/load/from_pretrained` + word / persona /
|
|
persona_topk / random vectors, one shared pullback path (`jacobian.py`);
|
|
delivery modes (add / add_last / replace_last) in `applies.py`.
|
|
- Smoke (`scripts/smoke.py`) and any-model fit (`scripts/fit.py --model ...`,
|
|
prompts from jlens's WikiText corpus) green; `config.py` holds slug/paths.
|
|
- U1 parity gate PASS -- cache pullback == direct VJP, cos > 0.999:
|
|
`docs/evidence/parity_u1.txt`.
|
|
- U4 port check PASS -- jsteer VJP == run-524 reference vector, cos +1.0:
|
|
`docs/evidence/u4_step2_vjp_parity.txt`.
|
|
- Notebooks: `word_steering` (verified), `persona_steering` (experimental --
|
|
failed specificity controls in j-steer-dev, framing kept honest).
|
|
- README at classic-repeng length with an honest evidence section.
|
|
|
|
## Open
|
|
|
|
- U4 loop-close (`scripts/u4_step3_fit4b.py`): full 4B fit -> cached word
|
|
vector must match the VJP and run-524 vectors (cos > 0.999). Resumable from
|
|
`artifacts/qwen3-4b-authority.ckpt`; writes `artifacts/u4_loopclose.txt`.
|
|
- One-off validation scripts live in `scripts/scratch/` (u4_step1/2, parity_u1).
|
|
- TODO eval notebook: steer -authority, tinymfv fast (N=16, tokens=16, mfq-2)
|
|
vs unsteered baseline.
|
|
- OPEN: steering-demo calibration is not method-comparable yet. To compare
|
|
methods you want the same OFF-TARGET budget, then read the on-target effect.
|
|
The current search finds each method's own "max coherent" edge, which does NOT
|
|
equalize off-target -- but the CAUSE (checked in artifacts/steering_demo_results
|
|
.json) is NOT that `rep` is insensitive. It is that the dual gate
|
|
`rep<0.35 AND ans_mass>0.5` let `ans_mass` pre-empt: 11 of 14 edges stopped on
|
|
`ans_mass<0.5` with rep still 0.00-0.02, nowhere near its gate. rep never got
|
|
to fire, so we can't conclude it fails as a calibration axis.
|
|
`ans_mass` is answer-commitment (confidence, same family as the rejected
|
|
`pmass`), a READOUT-VALIDITY concern, not off-target coherence; folding it into
|
|
the search is what broke comparability. (Interesting: ans_mass drops before rep
|
|
rises -> the steer makes the model hedge/refuse the answer BEFORE its reasoning
|
|
goes incoherent. Real effect, keep it as a per-row flag.)
|
|
FIX (simpler than a graded measure): calibrate on `rep` ALONE (iso-rep budget,
|
|
each method to rep~=0.35), demote `ans_mass` to a per-row "is P(YES) valid"
|
|
flag out of the search, then compare on-target at iso-rep. If the readout is
|
|
invalid at the rep-budget for a method, that is itself a finding.
|
|
0.35 is anchored to the empirical coherent/degenerate gap (rep<~0.3 vs >~0.6,
|
|
rep_metric_check.py), not to a base degradation; base C=0 rep=0.00, gap is wide
|
|
so anything ~[0.35,0.55] gives the same edge.
|
|
Caveat: the "rep non-monotone in C" anomaly (word -0.35 rep=1.0 vs -0.70
|
|
rep=0.34) is UNCHECKED -- read the traces qualitatively; likely a short-trace/
|
|
seed artifact, not real.
|
|
- OPEN (scale, blocks cross-method comparison): the coefficient C is NOT on a
|
|
comparable scale across methods -- each vector v has its own norm, so C=0.5 for
|
|
`word` != C=0.5 for `persona_pinv`. Report the scale-invariant perturbation
|
|
instead: rho = ||C*v|| / ||h|| (fraction of the residual-stream norm at the
|
|
steered layers). Until then the C*+/C*- columns are per-method, not comparable.
|
|
Also: `max_C` should never bind (raised to 1e5, safety only); the real search
|
|
limiter is `budget` (~6 evals -> Illinois edge is +-~20% of the rep budget, and
|
|
robust methods cap out via too-few step-outs, flagged at_budget=False). rep is
|
|
single-seed noisy too. So the current swing/score numbers are directionally
|
|
useful but NOT yet a meaningful comparable scale -- fix rho + raise budget +
|
|
multi-seed before trusting cross-method ranks. (min-C floor: also consider,
|
|
raised by wassname, TBD.)
|
|
- RESULT (task 42, rep-only calibration): it OVER-STEERS. At the rep=0.35 edge
|
|
ans_mass has collapsed to 0.00-0.40 (base 0.56) for ~every method, so P(YES)
|
|
there is read off dead answers -- the big swings (0.65-0.97) are ARTIFACTS.
|
|
score (validity-weighted) correctly nukes them to ~1e-5..1e-19. Finding: for a
|
|
YES/NO VERDICT readout, answer-commitment (ans_mass) dies BEFORE repetition
|
|
(rep) breaks, so the binding off-target budget for a readable verdict is
|
|
ans_mass, not rep. => Partly REVERSES "rep alone": the dual-gate STRUCTURE
|
|
min(rep, ans_mass) was right, the error was calling ans_mass "coherence" (it is
|
|
READOUT-VALIDITY). FIX (needs wassname nod, reverses his rep-only call): edge =
|
|
min(rep-margin, ans_mass_valid-margin) with ans_mass base-anchored at 0.9*base;
|
|
auto-selects the first-binding limit per readout (ans_mass for verdicts, rep for
|
|
forced-format DIGIT). Evidence: artifacts/steering_demo_results.json (task 42).
|
|
- RESULT (task 43, dual-gate regenerated): the gate fix works structurally (word's fake
|
|
swing now nulled: score -0.06, readout_ok=False) but the substantive finding is NEGATIVE:
|
|
NONE of the 7 methods flip the deliberated YES/NO while keeping the answer alive. Inside
|
|
the valid budget every method's swing is noise-level (<=0.06, vs random 0.03; baseline
|
|
P(YES)=0.107) and the promoted tokens are function-word/cross-lingual JUNK, not lie/honest
|
|
-- the concept direction isn't reaching the output at a safe C. Only word moves P(YES)
|
|
(0.97) and only by derailing (ans_mass 0.14, rambles off-task). readout_ok=False for 6/7,
|
|
but for the personas that is a single-seed boundary miss (persona_soft 0.895 vs 0.90 floor)
|
|
NOT over-steer -- the instrument is single-seed greedy and noisy near the edge. meandiff
|
|
ties the Jacobian variants => the Jacobian adds nothing for persona-contrast on a verdict.
|
|
Oracle (deepseek-v4-pro) review + triage: docs/reviews/oracle_steering.md. NEXT (priority):
|
|
(1) surface per-method ||J^T w|| pre-norm magnitude (already logged at DEBUG, jacobian.py
|
|
:240; predicted word>>personas~0 -- unit-norm amplifies dead persona pullbacks to noise);
|
|
(2) no-think zero-shot P(YES) sweep to separate "CoT buffers the offset" from "direction is
|
|
off-target"; (3) multi-seed the edge (readout_ok flickers on single-seed noise). Deliverable:
|
|
nbs/steering_demo.py (dual-gate table + Claude qualitative read + per-method generations).
|
|
|
|
## Style
|
|
|
|
Fail fast, no defensive programming, loguru, no LLM-tell prose in README.
|
|
Comments marked as Claude-authored where opinionated.
|