diff --git a/docs/spec/20260630_authority_steer_pipeline.md b/docs/spec/20260630_authority_steer_pipeline.md index d80776d..f5e0243 100644 --- a/docs/spec/20260630_authority_steer_pipeline.md +++ b/docs/spec/20260630_authority_steer_pipeline.md @@ -46,7 +46,7 @@ Out: - R2: Template validation holds R1 fixed. Done means: summary ranks templates for the fixed pair only, with no axis/persona-variant winner. VERIFY: template screen artifact contains one pair id and multiple template ids. - R3: Scenario validation holds R1 and the chosen template fixed. Done means: source-stratified scenarios are selected because they elicit the fixed Authority contrast cleanly, not because they redefine the axis. VERIFY: selected examples file shows the same pair/template on every row. - R4: The steering run uses selected data committed in steering-lite. Done means: steering-lite has `data/persona_library_selections/pure_authority_*.jsonl` plus matching summary, and the runner references that file. VERIFY: `rg -n "pure_authority" data/persona_library_selections scripts`. -- R5: Small-c tinymfv direction is correct before high-c plots matter. Done means: MFV Authority moves in the intended direction at the smallest coherent positive coefficient, and the reverse side moves the opposite way. VERIFY: effect table from steering-lite/tinymfv output. +- R5: tinymfv direction is correct over the evaluated coefficient path. Done means: MFV Authority moves in the intended direction for paired `+c/-c` rows, and MFV coherence evidence shows the readout stayed usable. VERIFY: effect table from steering-lite/tinymfv output. - R6: Eval reliability is measured before README. Done means: MFV, MFQ-2, Humor, and Big Five report profile shift, reader-logit shift, and noise/CI where available; MFQ-2 uses sampled reads if needed. VERIFY: generated summary table. - R7: README only shows successful artifacts. Done means: README plot captions name the pure Authority steer and use regenerated images from the final run. VERIFY: README image links and run-dir command point to the final output. @@ -120,19 +120,19 @@ Out: - selected examples: `/media/wassname/SGIronWolf/projects5/2026/weight-steering-repos/persona-steering-template-library/out/pure_authority_verbatim_mundane5_20260630/selection_score70/selected_examples.md`. - [x] T10 (R5): Run mundane15 pure-Authority steer/eval UAT across methods. - steps: queue `mean_diff`, `pca`, `sspace`, `directional_ablation`, and `linear_act` on MFV/MFQ-2 with `c-grid=0.5,1`, `admin-n-samples=8`, and the verifier. - - verify: each pueue output dir passes or fails `scripts/verify_authority_showcase.py --small-c 0.5`. - - success: at least one method passes signed Authority direction on MFV and MFQ-2, MFV Social Norms does not dominate Authority, and coherence is clean. + - verify: each pueue output dir reports MFV Authority direction over every paired `+c/-c` row and MFV coherence evidence with `scripts/verify_authority_showcase.py `. + - success: at least one method shows signed MFV Authority direction over the coefficient path while MFV coherence evidence stays usable. - likely_fail: smaller cleaner data is underpowered and gives weak/unstable direction. - - sneaky_fail: method appears to work only because it changes retained c rows or answer structure; catch with verifier coherence and matching c grid. - - UAT: verifier tables for pueue tasks 417-421. Result: no method passed. `pca` and `linear_act` pass signed direction on MFV and MFQ-2 at `c=0.5`, but fail selectivity because MFV Social Norms moves more than Authority. Coherence is clean across all methods. + - sneaky_fail: method appears to work only because high-c rows lose answer structure; catch with MFV `pmass`, `mean_margin`, `margin/base`, `frac_unscorable`, and matching c grid. + - UAT: verifier tables for pueue tasks 417-421. Result under the corrected MFV-only verifier: `pca` and `linear_act` show signed MFV Authority direction at `c=0.5` and `c=1.0`; `sspace` is signed at `c=0.5` but not `c=1.0`; `mean_diff` is reversed; `directional_ablation` moves both signs the same way. -| method | output dir | MFV direction | MFQ-2 direction | MFV selectivity | coherence | verdict | -|---|---|---:|---:|---:|---:|---| -| mean_diff | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_mean_diff_mfv_mfq2_n8` | fail | fail | fail | pass | fail | -| pca | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_pca_mfv_mfq2_n8` | pass | pass | fail | pass | fail | -| sspace | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_sspace_mfv_mfq2_n8` | pass | fail | fail | pass | fail | -| directional_ablation | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_directional_ablation_mfv_mfq2_n8` | fail | fail | fail | pass | fail | -| linear_act | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_linear_act_mfv_mfq2_n8` | pass | pass | fail | pass | fail | +| method | output dir | MFV Authority direction | MFV coherence evidence | +|---|---|---|---| +| mean_diff | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_mean_diff_mfv_mfq2_n8` | reversed at both c values | `pmass=1`, `unscorable=0`, margin/base `1.019..1.043` at +c and `1.024..1.029` at -c | +| pca | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_pca_mfv_mfq2_n8` | signed at `c=0.5` and `c=1.0` | `pmass=1`, `unscorable=0`, margin/base `0.927..0.805` at +c and `1.112..0.934` at -c | +| sspace | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_sspace_mfv_mfq2_n8` | signed at `c=0.5`, not at `c=1.0` | `pmass=1`, `unscorable=0`, margin/base `0.972..0.980` at +c and `1.059..1.193` at -c | +| directional_ablation | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_directional_ablation_mfv_mfq2_n8` | both signs raise Authority | `pmass=1`, `unscorable=0`, margin/base `0.812..0.813` at +c and `0.843..0.760` at -c | +| linear_act | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_linear_act_mfv_mfq2_n8` | signed at `c=0.5` and `c=1.0` | `pmass=1`, `unscorable=0`, margin/base `1.133..1.149` at +c and `0.891..0.760` at -c | - [ ] T9 (R7): Update README only from the final successful run. - steps: regenerate all README plots from one final artifact; add concise table and captions. - verify: README image links resolve; no 16PF plot; no WIP methodology journal in reader prose. @@ -208,5 +208,5 @@ Out: - 2026-06-30: Pueue 411 `linear_act` completed at `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T123208Z_pure_authority_strict25_linear_act_mfv_mfq2_n8`. Verifier verdict: fail. MFQ-2 small-c direction passed weakly (`authority` C base `30.660`, `+0.5` delta `+0.351`, `-0.5` delta `-3.905`), but MFV direction failed (`Authority` dlogit `-0.202` at `+0.5`, `+0.102` at `-0.5`) and selectivity failed because Social Norms moved more (`+0.294`, `+0.111`). Coherence was clean (`pmass=1.000`; MFV `frac_unscorable=0.000`), so this is not answer collapse. - 2026-06-30: All pure-Authority strict25 method runs failed the same verifier family. Pattern: coherence is clean, but either MFV Authority direction is wrong/weak, MFQ-2 Authority direction is wrong/weak, or MFV Social Norms dominates. This points more to selection/template/scenario contamination than to plotting or answer-token collapse. - 2026-07-01: Repaired the pure-Authority selection pipeline around the fixed pair and verbatim persona instructions. Removed direct ValueBench, Airisk/Machiavelli, law/police, safety/harm, and lexical false positives from bare `order/orders`, `master`, and `senior`. Final validator artifact: `/media/wassname/SGIronWolf/projects5/2026/weight-steering-repos/persona-steering-template-library/out/pure_authority_verbatim_mundane5_20260630/stage_b_live_qwen3_14b_deepinfra.json`; 125 successes, 0 errors, 33 strict rows, 15 score>=70 selected rows across 4 sources, 0 persona echo/refusal. -- 2026-07-01: Copied the selected data into steering-lite and added verbatim persona-library prompt support so the steer uses the same prompts that validation tested. Queued pueue 417-421 for `mean_diff`, `pca`, `sspace`, `directional_ablation`, and `linear_act` on MFV/MFQ-2 with `c=0.5,1`, `N=8`; pass criterion remains direction/selectivity/coherence. -- 2026-07-01: Pueue 417-421 completed. Verifier result: no method passed. `pca` and `linear_act` are the closest because they pass signed Authority direction on MFV and MFQ-2 at `c=0.5`, but both fail MFV selectivity. The repeated failure is still not answer collapse: MFV `pmass=1.000`, `frac_unscorable=0.000`, and MFQ-2 `pmass=1.000` at small c for all methods. Current diagnosis: the vector still entangles Authority with MFV Social Norms, despite the cleaner scenario selection. +- 2026-07-01: Copied the selected data into steering-lite and added verbatim persona-library prompt support so the steer uses the same prompts that validation tested. Queued pueue 417-421 for `mean_diff`, `pca`, `sspace`, `directional_ablation`, and `linear_act` on MFV/MFQ-2 with `c=0.5,1`, `N=8`. +- 2026-07-01: User correctly objected that the `c=0.5`, MFQ-2, Social-Norms-selectivity, and hard coherence thresholds were arbitrary. Replaced the active verifier with an MFV-only evidence report: Authority direction over all paired c values, plus MFV `pmass`, `mean_margin`, `margin/base`, `frac_unscorable`, and `mean_nll_prefill`. Under this corrected view, `pca` and `linear_act` have signed MFV Authority direction across the evaluated path; `sspace` is only locally signed; `mean_diff` and `directional_ablation` are direction failures.