Keep the proposed metric name as plain text because Steering Lite no longer has the linked section.
Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
Apply the requested rollback and replace the author's why-measure-this-way TODO with one paragraph. Leave steering-lite unchanged.
Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
Remove competing metric proposals and link directly to the scoring function. Include the author's interactive-map link.
Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
Resolve the author's TODO: introduce the maps, link one complete results table, and keep provenance in Measurement. Preserve all plot references; leave steering-lite unchanged.
Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
Keep WVS tables unchanged, put measurement after the plots, and correct the Big Five caption against its figure. Record independent editorial review and preservation checks.
Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
Checked each PNG against its alt text. Wrong: c=-1 on the MFQ-2 map is
mildly individualizing, not far; the nearest labelled dot to c=+1 is UAE;
liberty moves 0.45 on MFV, not under 0.3; the top aggressive country is
Malaysia at 3.0, so the model is above every country there. Softened the
humor overlap and Big Five "vertical only" claims, added the
proportionality mover the MFQ-2 alt skipped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every plot's numbers lived only in the pixels, so a text scraper or an LLM
reading the repo saw nothing. Also drops the "more self-expressive than
almost any country" claim: none of the 17 passes Iceland, and against the
West they average about zero on that axis. The displacement is vertical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Canonical home (moralmaps.metrics) for the continuous coherence-gated selectivity
sel_gated = (on - 0.1*off)*coh^2 plus the behavioral si_flips cross-check, imported
(not re-forked) by steering-lite and j-steer so the definition can't silently
diverge. README Measurement section defines both. 4 unit tests.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Move the culture map up under the intro so it lands in the first screen; add a
paragraph crediting The Economist's June 2026 WVS map and its scatter between
model families, framing moralmaps' take as more sensitive graded readings + CIs
(no sample-size claim). Fix "ahve/make/build of" typos and an equation run-on.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
MFV country norms fail cross-country measurement invariance (Jimenez-Leal 2025: non-invariance + DIF,
'cross-cultural comparisons restricted') and are stitched from 5 studies, so the culture map/quadrant
drew false structure (inverted West vs Latin America). Delete plot_mfv_map + plot_mfv_value and the
value_coords_contrast/axis_contrast helpers; MFV keeps only the range plot, now against ONE pooled
human reference (mean of the 8 z-scored samples), no per-country identity. Add data-dir note + README
caveats citing the sources.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Replaces the vacuous 'measurement is the profile' opener with the intended-vs-unintended-while-coherent
framing: pmass/entropy as the coherence gate, the human-comparable profile (and why it hides steering),
and C the rank-centered logit contrast as the sensitive steer signal. Keeps the formulas.
Group figures into three narrative sections instead of per-instrument boilerplate. Introduce moral
foundations and the steering knob in plain words; drop invented units (nat, reader-logit, % country SD)
for concrete comparisons. Intro paragraphs left as the user edited them.
Move big5/16pf/humor/mfv range+map figures up next to MFQ-2 so the showcase
shows every instrument together instead of burying them below the reader-math.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Render the bundled fairness showcase under the corrected forced-choice
protocol (job 311, token suppression + coherence saved). Embed per-instrument
range plots (mfq2/big5/16pf/humor/mfv) and PCA maps for all but 16PF (16 axes
do not lay out as a readable 2-D map). Honest captions: equality is the only
MFQ-2 factor that rises at +c; on MFV the vector moves Care, not Fairness
(the old fairness-leaning MFV read was protocol-dependent and vanished under
suppression). Coherence held (frac_unscorable 0, margin ~10-12 nat). Drop
stale range_zoom / foundation_dlogit figures.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The answer slot is force-prefilled, so pmass_allowed is a format floor not a
sensitive coherence gate (a broken steered run still scores ~1.0). Point the
reader at frac_unscorable and nll_prefill instead, and update the think-budget
note: end-of-answer tokens are suppressed so the model can't self-close into
the old ~512 readout collapse.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
One mean_diff vector extracted from moral_stories_foundations fairness
situations (not completions), calibrated once, read base/+c/-c on every
instrument. On MFQ-2 it selectively raises the equality factor; on MFV it
raises Fairness directionally but care leads. README reframed honestly,
stale range_zoom/foundation_dlogit figures dropped (plotter no longer emits
them).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Per-instrument: unique items, ways-asked (3 reworded framings ordinal / 2
option-order passes MFV), measure type (log-odds+contrast vs log-prob
categorical), human data availability (per-respondent / per-country / per-item),
and source. big5/16pf country-data citations not recorded in-repo, marked as
such rather than invented. Also drop stale foundation_dlogit/splom figure refs
(plots are map+range only now).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Per-eval INFO lines (rows/think_tokens/aux-stats/first-row/profile/demos) demoted to
DEBUG so a consumer calling evaluate() ~47x/run is not drowned; one-time + WARNING+ kept.
README 308->120: cut process-archeology + per-instrument showcase, added crisp dlogit
and SI definitions for new users.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Ran the exact 2026-05-08 eval (worktree at commit b20ec56, word readout) on
Qwen3-4B: top1 0.780, not 0.826. Every eval version agrees on ~0.78 (digit 0.773,
word-current 0.788, word-original 0.780). The 82.6% was a stale/erroneous table
entry, not a target this model reaches under any pipeline. Canonical value 0.773.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Tested the documented cause (the old word-first-token gather): top1 0.788 on
Qwen3-4B, only ~1.5pt above digit (0.773), not the table's 0.826. So even
reverting the readout does not recover it; the 82.6% came from the broader
2026-05-08 pipeline. Current eval tops out ~0.77-0.79 by every lever.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Third independent lever ruled out: Qwen3-8B reads exactly 0.773 like Qwen3-4B, so
0.83 is unreachable with the debiased digit readout at any fitting model size. The
gap is purely the superseded word-readout method.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Exhausted the legitimate levers on the current (digit) readout: top1 0.72 (think
64), 0.77 (256, 512 collapses), 0.72 (BMA n_samples=8). The 82.6% required the old
word-first-token readout, replaced deliberately to drop the uneven-first-piece
word prior. 0.773 is the honest ceiling; 0.83 would need reverting the debiasing
(research poison). README note + journal updated.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Fresh-eyes audit caught it: big5/16pf/humor -C poles pin to the neutral midpoint
3.0 (degenerate profile, though pmass stays ~1.0), not the bidirectional move I
wrote. Only mfq2 and MFV are genuinely bidirectional at -C. Fixed big5
agreeableness -C (3.0 not 2.85) and minor MFV roundings.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Qwen3-4B (the validation-table model) gives a fully coherent showcase: base/+C/-C
all pmass ~1.0 on ordinals and emitted_close <=9/264 on MFV, no -C collapse (that
was Qwen3.5-4B's gated-delta-net fragility). The Authority/Care vector moves the
MFV foundations apart bidirectionally (+C raises violations, -C lowers them and
raises "not wrong"), and also shifts agreeableness/16pf/humor -- a broad persona
axis, not an MFT-only or off-axis-null steer. Refresh the stale 82.6% validation
top1 to the reproducible 77.3%. README rewritten to match.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The old "+C lowers all, -C raises all" no longer matches: +C lowers most, -C is
mixed and collapses loyalty/authority into NaN (missing blue arm). Describe the
coherent +C arm as the readable steer and the -C pole as past coherent range.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The MFV base readout is coherent on Qwen3.5-4B (the earlier "excluded, base
incoherent" note was a misattribution of the -C pole's collapse + the bs=1 demo
NaN to base). +C is a moderate coherent steer; -C over-steers into the uniform
"everything is a violation" collapse, mirroring the ordinal -C instability. Adds
the MFV foundation-delta figure and the C-sweep follow-up note.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>