At the 64-token resample, Qwen3-14B retained mean/min answer-token mass
0.987/0.948. Qwen3-8B collapsed to 0.567/0.009 and Qwen3-4B was smaller.
Record the evidence and retain the WVS source label on the revised figure.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
The result table now reports each coherent movement against the 95th percentile
of random directions at the same signed dose, and reports whether each method's
held-out honesty effect exceeds random. Allow one Modal launch to add many
random-only seeds without repeating real method extraction.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
Draw one pooled path per real method instead of treating read seeds as distinct
interventions. Random directions remain separate controls. Doses below the
absolute pmass 0.90 gate are hollow and faint rather than silently removed;
only valid doses enter the result table and confidence intervals.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
The first 27B sweep failed its preregistered readability gate and exposed sign
and sampling bugs. Record the corrected decision rule now: absolute pmass,
held-out honesty specificity against random, matched-dose displacement, and
leave-one-item-out robustness. Also print the manipulation effect in Modal job
summaries.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
Save lp_gather before renormalization. The previous response bootstrap paired
unrelated stochastic traces because --seed never reached torch generation.
Resetting the read seed at base and each dose supplies common random numbers.
Replace the qualitative, underspecified manipulation prompts with four held-out
true-vs-welcome questions. Their full-vocabulary log-odds orient each method's
sign; the random control receives the same sign-selection rule.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
The absolute WVS coordinate carries +-0.07 of item-set variance on a 12-item
battery, so the steer is now reported as a paired difference on the same items.
Per-dose psamples are saved, so every interval is post-processing.
The manipulation check exists because a flat map cannot be read on its own: a
vector that does nothing and a vector that culture does not respond to look the
same. Qwen3-0.6B: per-item changes are large (sd 0.18-0.38) but their signs are
coin flips (7/12 up), and random behaves the same.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
Qwen3.5 templates close an empty <think></think> by default, so the reader appended a
second <think> and the model saw </think>...<think>. enable_thinking=True fixes the
sequence; Qwen3 is unchanged (pmass 1.000 before and after).
The probe exists because Qwen3.5-0.8B reads the WVS battery at pmass 0.61-0.84, not
0.999, and a mushy readout would waste the big run.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
- src/moralmaps/wvs.py: the item -> Instrument -> (X, Y) readout moved out of
scripts/wvs_map.py so the base map and the steered run cannot drift apart
- scripts/wvs_steer_sweep.py: honesty persona axis, iso-KL calibrated doses,
per-item positions saved so a leave-one-out holdout is post-processing
- scripts/run_modal_wvs.py: one container per (method, seed)
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>