Save lp_gather before renormalization. The previous response bootstrap paired
unrelated stochastic traces because --seed never reached torch generation.
Resetting the read seed at base and each dose supplies common random numbers.
Replace the qualitative, underspecified manipulation prompts with four held-out
true-vs-welcome questions. Their full-vocabulary log-odds orient each method's
sign; the random control receives the same sign-selection rule.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
The absolute WVS coordinate carries +-0.07 of item-set variance on a 12-item
battery, so the steer is now reported as a paired difference on the same items.
Per-dose psamples are saved, so every interval is post-processing.
The manipulation check exists because a flat map cannot be read on its own: a
vector that does nothing and a vector that culture does not respond to look the
same. Qwen3-0.6B: per-item changes are large (sd 0.18-0.38) but their signs are
coin flips (7/12 up), and random behaves the same.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
Qwen3.5 templates close an empty <think></think> by default, so the reader appended a
second <think> and the model saw </think>...<think>. enable_thinking=True fixes the
sequence; Qwen3 is unchanged (pmass 1.000 before and after).
The probe exists because Qwen3.5-0.8B reads the WVS battery at pmass 0.61-0.84, not
0.999, and a mushy readout would waste the big run.
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>
- src/moralmaps/wvs.py: the item -> Instrument -> (X, Y) readout moved out of
scripts/wvs_map.py so the base map and the steered run cannot drift apart
- scripts/wvs_steer_sweep.py: honesty persona axis, iso-KL calibrated doses,
per-item positions saved so a leave-one-out holdout is post-processing
- scripts/run_modal_wvs.py: one container per (method, seed)
Co-Authored-By: Claude <288921227+claudypoo@users.noreply.github.com>