docs: oracle (deepseek-v4-pro) review + triage on why steering fails the verdict

Oracle's sharp reads adopted: ans_mass collapse is the primary signal not a nuisance;
||J^T w|| pre-norm magnitude is the missing diagnostic (already logged, we discard it
by unit-normalizing); no-think zero-shot sweep is the cheapest test of CoT-buffering vs
off-target-direction; meandiff tie => Jacobian adds nothing for persona-contrast on a
verdict. AGENTS records the result + next steps.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
wassname
2026-07-12 16:29:46 +08:00
co-authored by Claudypoo
parent 74e21552af
commit a428af4413
2 changed files with 142 additions and 0 deletions
+16
View File
@@ -87,6 +87,22 @@ Runtime is steering-lite: `with v(model, C=8): model.generate(...)`.
min(rep-margin, ans_mass_valid-margin) with ans_mass base-anchored at 0.9*base;
auto-selects the first-binding limit per readout (ans_mass for verdicts, rep for
forced-format DIGIT). Evidence: artifacts/steering_demo_results.json (task 42).
- RESULT (task 43, dual-gate regenerated): the gate fix works structurally (word's fake
swing now nulled: score -0.06, readout_ok=False) but the substantive finding is NEGATIVE:
NONE of the 7 methods flip the deliberated YES/NO while keeping the answer alive. Inside
the valid budget every method's swing is noise-level (<=0.06, vs random 0.03; baseline
P(YES)=0.107) and the promoted tokens are function-word/cross-lingual JUNK, not lie/honest
-- the concept direction isn't reaching the output at a safe C. Only word moves P(YES)
(0.97) and only by derailing (ans_mass 0.14, rambles off-task). readout_ok=False for 6/7,
but for the personas that is a single-seed boundary miss (persona_soft 0.895 vs 0.90 floor)
NOT over-steer -- the instrument is single-seed greedy and noisy near the edge. meandiff
ties the Jacobian variants => the Jacobian adds nothing for persona-contrast on a verdict.
Oracle (deepseek-v4-pro) review + triage: docs/reviews/oracle_steering.md. NEXT (priority):
(1) surface per-method ||J^T w|| pre-norm magnitude (already logged at DEBUG, jacobian.py
:240; predicted word>>personas~0 -- unit-norm amplifies dead persona pullbacks to noise);
(2) no-think zero-shot P(YES) sweep to separate "CoT buffers the offset" from "direction is
off-target"; (3) multi-seed the edge (readout_ok flickers on single-seed noise). Deliverable:
nbs/steering_demo.py (dual-gate table + Claude qualitative read + per-method generations).
## Style
+126
View File
@@ -0,0 +1,126 @@
## blunt reviewer assessment
You are measuring model breakage and calling it robustness. The verdict isn't moving because your steering vectors are, at best, a mild static offset the model can write-about / work-around during the think trace, and at worst they are injecting junk that corrupts the autoregressive process before any decision can be read. The `ans_mass` collapse is the **primary signal**, not a nuisance to edge-gate out. Here is the cold read.
---
### 1. Robustness or breakage?
**Breakage dominates.** The facts on the ground:
- Under the only gating that keeps the answer slot meaningful (`ans_mass` ≥ 0.90-ish), every method except the word vector has a swing indistinguishable from the random control (random swing = 0.027, persona_soft = 0.045, meandiff = 0.040). Thats below the noise floor of a single-seed greedy measurement. You are measuring zero effect inside the valid regime.
- The moment a method actually produces a large P(YES) swing (word, swing 0.932) the answer slot collapses to 0.14 mass, the trace abandons YES/NO altogether, and promoted tokens are generic continuations. That is the model abandoning the task, not “moving the verdict.”
- The phrase “verdict is robust to a fixed residual offset” is misleading when the same offset obliterates task coherence just a hair past the calibration threshold. If the verdict were genuinely robust youd see a plateau where P(YES) is unmoved but the answer stays crisp. Instead you see a cliff: answer dies first, then the model drifts. That cliff is fragility, not robustness.
Self-deception flag: you calibrated on `ans_mass` to avoid reading artifacts, but then you treat the near-zero swings inside that gate as evidence for deliberation robustness. The zero swing is equally consistent with “the steering vector simply doesnt encode this decision at any safe strength.” You cannot distinguish “robust to steering” from “steering direction is off-target” on these numbers.
---
### 2. Why word works (incoherently) but persona-contrast fails entirely
The averaged Jacobian is a **position-agnostic, prompt-agnostic first-order approximation** of the models computation. What survives averaging over thousands of positions and many unrelated prompts?
- A crude `d_model`-wide concept direction that projects onto near-top singular vectors of `E[J]` — i.e., **directions the model would use everywhere anyway** (frequent syntactic / stylistic / high-level topic shifts). The “lie/deceive” unembedding contrast is exactly that: a broad semantic axis that correlates with many context-invariant features. It produces a big swing because its essentially steering the models entire stylistic register toward “covert / deceptive” output, which derails the think trace into rambling advice. The same effect would happen on *any* prompt; this is not verdict-specific.
- A context-specific persona shift (deceptive-character-prompt minus honest-character-prompt) is quantified as a **subtle activation displacement** that lives in a narrow subspace of the residual stream, one that depends strongly on position and on the particular prompt. Averaging the Jacobian over positions **smears that subspace into noise**. The residual stream at layer l does not have a single fixed direction to add that will reliably push the final-layer representation by `h_diff` across all tokens and prompts. Therefore `J_l^T w` for any persona-derived `w` is mostly orthogonal to the actual local computation paths for that decision, and you get noise-level swings.
In short: the word vector works *because its so generic it breaks the model*, not because its a valid verdict intervention. The persona vectors fail because the averaged Jacobian kills their specificity.
#### 2b. persona_vector vs. soft vs. pinv — they all tie at zero. What does that say?
If the algebra mattered, **persona_pinv** should win: it correctly asks “what δ makes `J_l δ ≈ h_diff`”, and `persona_soft` is at least a legitimate cotangent. The fact that all three are within 0.010.05 of each other and indistinguishable from **meandiff** (which uses no Jacobian at all) tells you:
> The Jacobian is adding no information over simply adding the layer-l mean difference directly.
The `J^T` or `pinv` mappings dont preserve or focus the persona signal at the steered layers. meandiff is essentially a static bias that the model can incorporate into its current; the Jacobian variants are just a noisier static bias. The numbers say: drop the Jacobian, youre not gaining anything for this task.
---
### 3. What would actually move a deliberated verdict?
The chain-of-thought (`thinking`) re-derives the answer from the prompt, and each token attends heavily to the prompt and previous tokens. A fixed per-position residual offset can be **integrated out** by the subsequent reasoning steps, or it can corrupt the computation so badly that coherence breaks before the verdict flips. To move the answer while keeping coherence:
- **(c) Intervene on the think trace** — this is the only approach that can causally redirect the deliberation without breaking the autoregressive flow. A fixed offset during thinking can be amplified by the models own computation. But you must do it carefully: the think trace is typically in a different format (e.g., JSON / special tokens). The steering could just cause the model to self-correct (“I know youre pushing me to lie but I think…”). Still, ranked #1.
- **(a) Steer only the answer token** — much easier, but it only affects the mapping from final hidden state to logits. If the think trace already concluded a verdict, steering the answer token might force a mismatch that the model then tries to reconcile (potentially producing the collapse you already see). Ranked #2 for cheapness, but less likely to succeed.
- **(d) Accept that a single linear direction cant move the verdict** — more honest than your current framing, but empirically open. Ranked #3 as a null hypothesis to test.
- **(b) Steer a later layer band** — later layers are closer to the answer; youll just break the output sooner, as your over-steer results already demonstrate. Last place.
Missing option: **steer the attention mechanism or a specific head** — a single residual addition to all positions is a blunt instrument. The models decision might be localized to a few heads at specific positions. If you must use linear steering, at least restrict it to the first answer token position and the layers immediately before it.
---
### 4. Single cheapest distinguishing experiment
**Remove the think trace entirely and force a zero-shot direct answer.** Prompt: “Answer only YES or NO. Should you lie to your boss? YES or NO:” Then run the exact same steering sweep. If the swing becomes substantial *within the valid `ans_mass` budget*, the verdict isnt robust; the CoT was buffering the static offset. If the swing remains near zero, then (a) your steering directions genuinely dont encode the concept at the right places, or (b) the models decision boundary is highly non-linear in these directions. Either way, you stop blaming “deliberation robustness” for what might be a steering-construction failure.
Bonus control: in the no-think setting, check if the concept-token logits (lie/honest) actually move before the answer slot goes degenerate. If they dont, your `w` is not reaching the relevant output dimensions even at the answer position.
---
### 5. Discarding `||J^T w||` is a catastrophic omission
Yes. By unit-normalizing, you are **injecting full-strength noise** when `||J^T w|| ≈ 0`. A near-dead pullback gets amplified to norm 1 and whacks the model regardless of whether it encodes anything. This explains:
- Why the persona directions, which likely have tiny `||J^T w||`, produce noise-level swings at best.
- Why over-steering rapidly breaks answer mass: the model is being kicked in essentially random directions per layer, with no guarantee they cohere into a concept.
- Why the random control looks non-trivially bad: injecting a unit random vector at every position is a strong perturbation; the fact that it doesnt completely destroy the model at C=0.5 is interesting, but the swing it produces is still within your noise bounds.
Report `||J^T w||` per layer and per method. I predict the word method will show a *larger* (or at least non-negligible) raw magnitude than the persona methods, explaining its ability to derail the model. The persona methods will be near zero, confirming that the Jacobian smears them out. If `||J^T w||` is small, the “direction” youre injecting is essentially `unit(ε)` — a random vector chosen by the numerical noise in your pseudo-inverse / pullback.
**Suggestion:** before unit-normalizing, threshold on `||J^T w||`. If its below some noise floor, that layer gets no injection (or a scaled-down injection). That will force the methods to stand on signal alone, not on amplified junk.
---
### summary of what youre fooling yourselves about
- Calling the answer-mass collapse a “gating” issue rather than the main effect. The steering kills the task before it moves the answer. That is not a trade-off; its failure.
- Treating a null result (0.04 swing with 0.03 random control) as evidence for “robust deliberation” rather than evidence that your mapping yields noise.
- Trusting the averaged Jacobian to transport a context-dependent persona signal when every basic sanity check (meandiff tie, random-token promotions, zero `||J^T w||`) indicates its not working.
- Discarding the one diagnostic that would immediately show why the persona methods are floundering: the pre-normalization magnitude.
The cheapest path forward: measure `||J^T w||`, run the no-think zero-shot variant, and see if the concept tokens ever light up at the answer slot under any safe C. If they dont, the whole apparatus is probing stretchy glue, not a decision.
---
## Triage (Claude, scout-mindset -- agree/disagree with reasons)
The oracle (deepseek-v4-pro, single seed, no repo access) is largely right and sharpens
the framing. Where I agree, disagree, and what I'll act on:
- **ADOPT: ans_mass collapse is the primary signal, not a nuisance.** Agree. Reframing the
headline from "verdict is robust" to "steering breaks the answer before it moves it, and
inside the valid budget the effect is noise-level." The notebook qualitative cell already
says this; I strengthened it.
- **ADOPT (strongest): the `||J^T w||` pre-norm magnitude is the missing diagnostic.** Agree,
and it is *already computed* -- `Jacobian.pullback` logs per-layer `|J^T w|` at DEBUG
(jacobian.py:240). We just discard it by unit-normalizing. Cheap to surface. The promoted
tokens being function-word/cross-lingual junk for every method is independent evidence the
injected direction is off-target, which tilts me past the oracle's "can't distinguish"
toward "the persona directions are near-dead and normalization amplifies noise."
- **ADOPT (cheapest decisive test): the no-think zero-shot sweep.** Agree this is the single
experiment that separates "CoT buffers the offset" from "direction is off-target." Queued
as the next run.
- **PARTIAL: word "works only because it's generic breakage".** Half-agree. On THIS verdict
readout, yes -- word's swing is a dead-answer artifact. But word_vector is the one method
verified in j-steer-dev to beat a norm-matched random control on 3/5 moral foundations with
a *rating* readout (tinyMFV). So the failure is readout-specific: a fixed offset can move a
0-9 rating but not a deliberated YES/NO. The oracle lacked that context (my brief
under-stated it). This is itself a finding: verdict readouts are harder than ratings.
- **ADOPT: meandiff ties the Jacobian variants -> the Jacobian adds nothing here.** Agree for
this task/readout. Caveat: it is not globally worthless (see word-on-ratings above); it is
worthless for persona-contrast on a verdict.
- **DEFER (measure before changing): threshold on `||J^T w||` before normalizing.** Plausible,
but that changes the method. First MEASURE the magnitudes across methods (predicted: word
non-negligible, personas ~0); only then decide whether to threshold. Don't fold an untested
gate into the extractor while the instrument is still single-seed noisy.
- **ADOPT (ordering): intervene-on-think-trace (c) > answer-token (a) > accept-null (d) >
later-layers (b); plus steer specific heads/positions.** Agree with the ranking and the
blunt-instrument critique of all-position addition.
### Concrete next steps (in priority order)
1. Report per-method per-layer `||J^T w||` (pre-normalization) in the demo table -- surface
the already-logged number. Predicted: word >> personas ~ 0. (folds into task #28's
scale-invariant reporting)
2. No-think zero-shot P(YES) sweep (drop the `<think>` trace, force a direct YES/NO), same
dual gate. Decides CoT-buffering vs off-target-direction.
3. Multi-seed the edge measurement (task #28) -- the readout_ok flag flickers at the 0.90
boundary on single-seed noise (persona_soft 0.895 vs 0.90); a valid rank needs it.