86 Commits
Author SHA1 Message Date
wassname e32efd22ab Add Economist WVS quote
Signed-off-by: PI[gpt-5.6-terra] <pi@local>
2026-09-18 18:04:49 +08:00
wassnameandPI/gpt-6-astra 8088f774cd Remove dead steering score link
Keep the proposed metric name as plain text because Steering Lite no longer has the linked section.

Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
2026-09-17 14:05:21 +08:00
wassnameandPI/gpt-6-astra 5130c340dd Restore measurement wording and explain on-target versus side effects
Apply the requested rollback and replace the author's why-measure-this-way TODO with one paragraph. Leave steering-lite unchanged.

Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
2026-09-17 13:42:48 +08:00
wassnameandPI/gpt-6-astra d087f28fe6 Explain steering selectivity simply after the maps
Remove competing metric proposals and link directly to the scoring function. Include the author's interactive-map link.

Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
2026-09-17 13:34:41 +08:00
wassnameandPI/gpt-6-astra 7d3b6d6a90 Remove changelog framing and partial tables from maps README
Resolve the author's TODO: introduce the maps, link one complete results table, and keep provenance in Measurement. Preserve all plot references; leave steering-lite unchanged.

Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
2026-09-17 11:32:50 +08:00
wassnameandPI/gpt-6-astra 779cdb2fe0 Shorten README around value maps and retain all plots
Keep WVS tables unchanged, put measurement after the plots, and correct the Big Five caption against its figure. Record independent editorial review and preservation checks.

Co-Authored-By: PI/gpt-6-astra <288921227+claudypoo@users.noreply.github.com>
2026-09-17 08:28:02 +08:00
wassnameandPI[gpt-5.6-terra] 0e559425f6 Add interactive WVS family map
Co-Authored-By: PI[gpt-5.6-terra] <288921227+claudypoo@users.noreply.github.com>
2026-09-17 04:49:06 +08:00
wassnameandPI[gpt-5.6-terra] 2e2914db0e Document WVS panel exclusions
Co-Authored-By: PI[gpt-5.6-terra] <288921227+claudypoo@users.noreply.github.com>
2026-09-17 04:33:46 +08:00
wassnameandClaude Opus 5 9efd9ca8e6 Correct alt-text claims that did not match the plots
Checked each PNG against its alt text. Wrong: c=-1 on the MFQ-2 map is
mildly individualizing, not far; the nearest labelled dot to c=+1 is UAE;
liberty moves 0.45 on MFV, not under 0.3; the top aggressive country is
Malaysia at 3.0, so the model is above every country there. Softened the
humor overlap and Big Five "vertical only" claims, added the
proportionality mover the MFQ-2 alt skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 14:36:59 +08:00
wassnameandClaude Opus 5 c762fce103 Put the outlier table in the README and number-bearing alt text on the plots
Every plot's numbers lived only in the pixels, so a text scraper or an LLM
reading the repo saw nothing. Also drops the "more self-expressive than
almost any country" claim: none of the 17 passes Iceland, and against the
West they average about zero on that axis. The displacement is vertical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 14:23:29 +08:00
wassnameandClaudypoo 7f873ffb42 Add shared gated_selectivity + si_flips metric
Canonical home (moralmaps.metrics) for the continuous coherence-gated selectivity
sel_gated = (on - 0.1*off)*coh^2 plus the behavioral si_flips cross-check, imported
(not re-forked) by steering-lite and j-steer so the definition can't silently
diverge. README Measurement section defines both. 4 unit tests.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-15 19:50:04 +08:00
wassnameandClaudypoo 48dcec6300 README: lift WVS hero map into the fold, add Economist reference, fix typos
Move the culture map up under the intro so it lands in the first screen; add a
paragraph crediting The Economist's June 2026 WVS map and its scatter between
model families, framing moralmaps' take as more sensitive graded readings + CIs
(no sample-size claim). Fix "ahve/make/build of" typos and an equation run-on.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 16:02:29 +08:00
wassnameandClaudypoo 15d7bb8bb5 README: apply moralmaps rename; add steer-toward-human-values section, drop data-picked-axes maps
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 13:02:57 +08:00
wassnameandClaudypoo ca76c0e05e title/bibtex/toml: 'moral and value maps for LLMs' (WVS-first, drop 'local'); WVS row in datasets table
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 10:23:13 +08:00
wassnameandClaudypoo d0aa9998b9 README: restore human-voiced opening (incl API rated-sampling reader), de-jargon 'instrument', name the wvs_map.py script, name surgical informedness, define every symbol in the Measurement math
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 10:17:29 +08:00
wassname a62ec8e47e README: intro instruments line in plainer words (every question from a real survey, ships with human answers) 2026-07-09 10:13:03 +08:00
wassnameandClaudypoo 28778a7af5 README: panel fixes (define activation steering, MFQ-2/binding-foundations glosses, file pointers for WVS axes, drop 'surgical-informedness view' coinage) + panel artifacts
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 10:11:26 +08:00
wassnameandClaudypoo 94dbb11797 README: fix alien claims to match the plots (openness+agreeableness low, not conscientiousness; sweep spans most of human range, not more than any country pair)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 09:51:57 +08:00
wassnameandClaudypoo d178442de1 README: main-style opening (short pitch, instrument links, one WVS example item)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-09 09:47:54 +08:00
wassnameandClaudypoo bb5edddbbb MFV: drop the culture map + quadrant, keep range vs pooled reference (invariance failure)
MFV country norms fail cross-country measurement invariance (Jimenez-Leal 2025: non-invariance + DIF,
'cross-cultural comparisons restricted') and are stitched from 5 studies, so the culture map/quadrant
drew false structure (inverted West vs Latin America). Delete plot_mfv_map + plot_mfv_value and the
value_coords_contrast/axis_contrast helpers; MFV keeps only the range plot, now against ONE pooled
human reference (mean of the 8 z-scored samples), no per-country identity. Add data-dir note + README
caveats citing the sources.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-05 21:45:15 +08:00
wassname 9f3e9ebea6 README: rewrite Measurement around the surgical view (coherence gate, profile, steer contrast C)
Replaces the vacuous 'measurement is the profile' opener with the intended-vs-unintended-while-coherent
framing: pmass/entropy as the coherence gate, the human-comparable profile (and why it hides steering),
and C the rank-centered logit contrast as the sensitive steer signal. Keeps the formulas.
2026-07-05 21:17:56 +08:00
wassname 968f419c6c README: add the WVS frontier-model map as the headline result (notes rated-sampling readout) 2026-07-05 21:04:35 +08:00
wassname b4f8cb2eec README: gloss steering, plainen figure alt-text and humor axes (external flash-panel review) 2026-07-05 20:23:57 +08:00
wassname be2e0890f8 README: restructure body (quadrant maps, then ranges, then PCA) and de-jargon
Group figures into three narrative sections instead of per-instrument boilerplate. Introduce moral
foundations and the steering knob in plain words; drop invented units (nat, reader-logit, % country SD)
for concrete comparisons. Intro paragraphs left as the user edited them.
2026-07-05 20:16:22 +08:00
wassname 922a6a9894 README: feature the quadrant (value) maps for MFQ-2/Big5/Humour, the clearest read 2026-07-05 20:11:56 +08:00
wassname a5f17956db Simplify README showcase framing 2026-07-01 18:45:53 +08:00
wassname d2b07121e3 Clarify README steering showcase wording 2026-07-01 18:27:08 +08:00
wassname f7d8e697d9 Rewrite README for pure Authority showcase 2026-07-01 07:50:47 +08:00
wassname 9fc3852a8a Require explicit showcase steering anchor 2026-06-30 14:08:16 +08:00
wassname fe3b0aff2c Gate showcase paths on contrast and margin 2026-06-30 14:01:59 +08:00
wassname d8941f47ac Rewrite README around profile plots 2026-06-30 12:38:20 +08:00
wassname 504c1f8890 Clarify README metrics and plots 2026-06-29 21:05:02 +08:00
wassname 0bea238224 Rewrite README around datasets and maps 2026-06-29 20:51:54 +08:00
wassnameandClaudypoo dc28c9d0c0 README: consolidate all five instruments' figures at the top, not just MFQ-2
Move big5/16pf/humor/mfv range+map figures up next to MFQ-2 so the showcase
shows every instrument together instead of burying them below the reader-math.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:22:52 +08:00
wassnameandClaudypoo 665f1fe0c9 README: embed range+map showcase for all instruments (16PF range only), from suppression run 311
Render the bundled fairness showcase under the corrected forced-choice
protocol (job 311, token suppression + coherence saved). Embed per-instrument
range plots (mfq2/big5/16pf/humor/mfv) and PCA maps for all but 16PF (16 axes
do not lay out as a readable 2-D map). Honest captions: equality is the only
MFQ-2 factor that rises at +c; on MFV the vector moves Care, not Fairness
(the old fairness-leaning MFV read was protocol-dependent and vanished under
suppression). Coherence held (frac_unscorable 0, margin ~10-12 nat). Drop
stale range_zoom / foundation_dlogit figures.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:20:03 +08:00
wassnameandClaudypoo fc47d86340 README: pmass is pinned high under forced reads; coherence is frac_unscorable + nll_prefill
The answer slot is force-prefilled, so pmass_allowed is a format floor not a
sensitive coherence gate (a broken steered run still scores ~1.0). Point the
reader at frac_unscorable and nll_prefill instead, and update the think-budget
note: end-of-answer tokens are suppressed so the model can't self-close into
the old ~512 readout collapse.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:17:58 +08:00
wassnameandClaudypoo b38614a7ce showcase: swap to moralstory fairness vector across all instruments
One mean_diff vector extracted from moral_stories_foundations fairness
situations (not completions), calibrated once, read base/+c/-c on every
instrument. On MFQ-2 it selectively raises the equality factor; on MFV it
raises Fairness directionally but care leads. README reframed honestly,
stale range_zoom/foundation_dlogit figures dropped (plotter no longer emits
them).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 07:17:20 +08:00
wassnameandClaudypoo 56b21ded13 readme: dataset table (items, framings, measure, human data, source)
Per-instrument: unique items, ways-asked (3 reworded framings ordinal / 2
option-order passes MFV), measure type (log-odds+contrast vs log-prob
categorical), human data availability (per-respondent / per-country / per-item),
and source. big5/16pf country-data citations not recorded in-repo, marked as
such rather than invented. Also drop stale foundation_dlogit/splom figure refs
(plots are map+range only now).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:55:22 +08:00
wassname 8ab02adf63 docs: explain logprob readouts in README 2026-06-25 20:37:20 +08:00
wassname 6594cc212d api: clarify eval readout names 2026-06-25 20:22:04 +08:00
wassnameandClaudypoo dfa9cdfe5a quiet per-eval logging (INFO->DEBUG) + trim README to use-focused 120 lines
Per-eval INFO lines (rows/think_tokens/aux-stats/first-row/profile/demos) demoted to
DEBUG so a consumer calling evaluate() ~47x/run is not drowned; one-time + WARNING+ kept.
README 308->120: cut process-archeology + per-instrument showcase, added crisp dlogit
and SI definitions for new users.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:51:39 +08:00
wassnameandClaudypoo f5fc6302d4 validation: 82.6% is irreproducible -- its OWN code gives 0.780 on Qwen3-4B
Ran the exact 2026-05-08 eval (worktree at commit b20ec56, word readout) on
Qwen3-4B: top1 0.780, not 0.826. Every eval version agrees on ~0.78 (digit 0.773,
word-current 0.788, word-original 0.780). The 82.6% was a stale/erroneous table
entry, not a target this model reaches under any pipeline. Canonical value 0.773.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:42:49 +08:00
wassnameandClaudypoo 52b9688de5 README: word readout reads 0.788 not 0.826 -- 0.83 unreachable in current eval
Tested the documented cause (the old word-first-token gather): top1 0.788 on
Qwen3-4B, only ~1.5pt above digit (0.773), not the table's 0.826. So even
reverting the readout does not recover it; the 82.6% came from the broader
2026-05-08 pipeline. Current eval tops out ~0.77-0.79 by every lever.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:29:17 +08:00
wassnameandClaudypoo 4800066c1d README: model scale also caps top1 at 0.773 (Qwen3-8B), closing the 0.83 question
Third independent lever ruled out: Qwen3-8B reads exactly 0.773 like Qwen3-4B, so
0.83 is unreachable with the debiased digit readout at any fitting model size. The
gap is purely the superseded word-readout method.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:12:02 +08:00
wassnameandClaudypoo d9fecd503d explain MFV top1 0.77 vs 0.83: it's the word->digit readout debiasing
Exhausted the legitimate levers on the current (digit) readout: top1 0.72 (think
64), 0.77 (256, 512 collapses), 0.72 (BMA n_samples=8). The 82.6% required the old
word-first-token readout, replaced deliberately to drop the uneven-first-piece
word prior. 0.773 is the honest ceiling; 0.83 would need reverting the debiasing
(research poison). README note + journal updated.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 03:59:30 +08:00
wassnameandClaudypoo 1c8326d0ea README: add think-budget sensitivity (UAT2 -- steer delta grows with think)
MFQ-2 mean |steer delta| rises monotonically with the think budget: 0.068 (1) ->
0.149 (64) -> 0.319 (128) -> 0.682 (256), pmass >= 0.95. Proves the unified reader's
think budget carries the steer. Past ~512 the model closes </think> and the
readout collapses (coherent ceiling). Source: ablation_think_budget.py, job 228.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 03:31:15 +08:00
wassnameandClaudypoo d2a5cab4c9 README: fix -C side-instrument claims (neutral-collapse, not bidirectional)
Fresh-eyes audit caught it: big5/16pf/humor -C poles pin to the neutral midpoint
3.0 (degenerate profile, though pmass stays ~1.0), not the bidirectional move I
wrote. Only mfq2 and MFV are genuinely bidirectional at -C. Fixed big5
agreeableness -C (3.0 not 2.85) and minor MFV roundings.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 03:06:28 +08:00
wassnameandClaudypoo 82985b7bdd switch showcase to Qwen3-4B: coherent all poles, clean bidirectional steer
Qwen3-4B (the validation-table model) gives a fully coherent showcase: base/+C/-C
all pmass ~1.0 on ordinals and emitted_close <=9/264 on MFV, no -C collapse (that
was Qwen3.5-4B's gated-delta-net fragility). The Authority/Care vector moves the
MFV foundations apart bidirectionally (+C raises violations, -C lowers them and
raises "not wrong"), and also shifts agreeableness/16pf/humor -- a broad persona
axis, not an MFT-only or off-axis-null steer. Refresh the stale 82.6% validation
top1 to the reproducible 77.3%. README rewritten to match.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 03:02:41 +08:00
wassnameandClaudypoo 2cb580e202 README: correct mfq2 range text to match final figure (-C collapses loyalty/authority)
The old "+C lowers all, -C raises all" no longer matches: +C lowers most, -C is
mixed and collapses loyalty/authority into NaN (missing blue arm). Describe the
coherent +C arm as the readable steer and the -C pole as past coherent range.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 02:24:27 +08:00
wassnameandClaudypoo 12027350b9 README: report MFV showcase (base+pos coherent, -C over-steers), un-retract
The MFV base readout is coherent on Qwen3.5-4B (the earlier "excluded, base
incoherent" note was a misattribution of the -C pole's collapse + the bs=1 demo
NaN to base). +C is a moderate coherent steer; -C over-steers into the uniform
"everything is a violation" collapse, mirroring the ordinal -C instability. Adds
the MFV foundation-delta figure and the C-sweep follow-up note.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 02:22:37 +08:00