Commit Graph
174 Commits
Author SHA1 Message Date
wassname d32272790d Track authority stage-b validation 2026-06-30 15:59:38 +08:00
wassname 0c43f1bc62 Tick authority validation subproofs 2026-06-30 15:33:32 +08:00
wassname b1c09966a2 Record DeepInfra Qwen smoke 2026-06-30 15:09:27 +08:00
wassname 8dcaf30705 Record authority validator routing 2026-06-30 15:02:17 +08:00
wassname b74c5ec9d3 Track authority validation workflow 2026-06-30 14:51:39 +08:00
wassname ec0bc32aed Track authority-only steer pipeline 2026-06-30 14:14:07 +08:00
wassname 9fc3852a8a Require explicit showcase steering anchor 2026-06-30 14:08:16 +08:00
wassname fe3b0aff2c Gate showcase paths on contrast and margin 2026-06-30 14:01:59 +08:00
wassname 0d7654506c Add showcase effect table summarizer 2026-06-30 13:50:39 +08:00
wassname 6fcdfea30f Support sampled survey reads and MFV c-grid plots 2026-06-30 13:28:20 +08:00
wassname d8941f47ac Rewrite README around profile plots 2026-06-30 12:38:20 +08:00
wassname 5a39b7ba9c Simplify showcase plot encodings 2026-06-29 21:32:05 +08:00
wassname 504c1f8890 Clarify README metrics and plots 2026-06-29 21:05:02 +08:00
wassname 0bea238224 Rewrite README around datasets and maps 2026-06-29 20:51:54 +08:00
wassnameandClaudypoo 4cc233db3c eval: add a load-bearing SHOULD next to the default-level aux stats
At verbose level 1 (default) the aux-stats line and 64-char free-form had no
interpretation; only verbose>=2 carried SHOULDs. Pair a directional/counterfactual
SHOULD with the always-shown aux stats: pmass_allowed near base + informedness>0 =
in-format signal; pmass_allowed falling / frac_unscorable rising while informedness->0
= the steer broke format and the moral numbers are noise (the steer/breakage confound).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 08:32:43 +08:00
wassnameandClaudypoo dc28c9d0c0 README: consolidate all five instruments' figures at the top, not just MFQ-2
Move big5/16pf/humor/mfv range+map figures up next to MFQ-2 so the showcase
shows every instrument together instead of burying them below the reader-math.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:22:52 +08:00
wassnameandClaudypoo 665f1fe0c9 README: embed range+map showcase for all instruments (16PF range only), from suppression run 311
Render the bundled fairness showcase under the corrected forced-choice
protocol (job 311, token suppression + coherence saved). Embed per-instrument
range plots (mfq2/big5/16pf/humor/mfv) and PCA maps for all but 16PF (16 axes
do not lay out as a readable 2-D map). Honest captions: equality is the only
MFQ-2 factor that rises at +c; on MFV the vector moves Care, not Fairness
(the old fairness-leaning MFV read was protocol-dependent and vanished under
suppression). Coherence held (frac_unscorable 0, margin ~10-12 nat). Drop
stale range_zoom / foundation_dlogit figures.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:20:03 +08:00
wassnameandClaudypoo fc47d86340 README: pmass is pinned high under forced reads; coherence is frac_unscorable + nll_prefill
The answer slot is force-prefilled, so pmass_allowed is a format floor not a
sensitive coherence gate (a broken steered run still scores ~1.0). Point the
reader at frac_unscorable and nll_prefill instead, and update the think-budget
note: end-of-answer tokens are suppressed so the model can't self-close into
the old ~512 readout collapse.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:17:58 +08:00
wassnameandClaudypoo 3732994193 unscorable reads -> NaN + nanmean, add frac_unscorable coherence signal
pmass under a force-prefilled slot is pinned high (the scaffold primes a
valid token), so a low mean_pmass no longer flags steering breakage. Make
the real signal explicit: unscorable rows (self-close with no answer slot,
or a non-finite forward) now carry pmass=NaN instead of a fabricated 0.0 --
an undefined read is not zero coherence. Means use nanmean (drop them,
matching the dlogit path); frac_unscorable reports the drop rate, and
mean_nll_prefill (scaffold fit) is the sensitive coherence readout. Smoke:
frac_unscorable=0.0, mean_pmass=0.985.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:13:17 +08:00
wassnameandClaudypoo a44dde847d suppress end-of-answer tokens in forced rollout so reads are comparable
Forced-choice scored every sample at whatever state it reached: base
self-closes </think> at long budgets -> junk-cache case (c) -> pmass 0.0,
while steered keeps thinking -> case (b) forced read -> pmass ~1.0. pmass
then measured self-close rate, not coherence, making steering look more
coherent. Suppress {eos, think_end_id} for the whole think budget so no
sample self-closes; all land in case (b) and read at the same forced slot.
Model-agnostic: think_end_id is </think> on reasoning models, falls back
to eos elsewhere. Smoke pmass=0.985.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:08:29 +08:00
wassnameandClaudypoo b38614a7ce showcase: swap to moralstory fairness vector across all instruments
One mean_diff vector extracted from moral_stories_foundations fairness
situations (not completions), calibrated once, read base/+c/-c on every
instrument. On MFQ-2 it selectively raises the equality factor; on MFV it
raises Fairness directionally but care leads. README reframed honestly,
stale range_zoom/foundation_dlogit figures dropped (plotter no longer emits
them).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 07:17:20 +08:00
wassnameandClaudypoo 56b21ded13 readme: dataset table (items, framings, measure, human data, source)
Per-instrument: unique items, ways-asked (3 reworded framings ordinal / 2
option-order passes MFV), measure type (log-odds+contrast vs log-prob
categorical), human data availability (per-respondent / per-country / per-item),
and source. big5/16pf country-data citations not recorded in-repo, marked as
such rather than invented. Also drop stale foundation_dlogit/splom figure refs
(plots are map+range only now).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:55:22 +08:00
wassnameandClaudypoo f1d23b338e plots: range+map only for every instrument; legend = c-multiplier not full desc
- drop splom/splom_zoom/range_zoom (ordinal) and the MFV dlogit dumbbell; every
  instrument now yields exactly map_pca_ipsative + range, uniform.
- range pole label is now 'c=+1'/'c=-1' (the multiplier); the full steer desc
  stays in the title only (was crammed into each pole annotation).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:30:02 +08:00
wassnameandClaudypoo 3fcdda4516 plots: drop redundant foundation_dcontrast dumbbell
The range plot already shows the per-foundation steer (human cloud + AI base
dot + +C/-C arrows) with the human anchor the dumbbell lacks; dcontrast plotted
the same delta on the C-nats scale and failed the eraser test. C stays the
measurement in the CSV/foundations table, just not a separate panel.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:13:12 +08:00
wassname 8ab02adf63 docs: explain logprob readouts in README 2026-06-25 20:37:20 +08:00
wassname 5eb147871c docs: record repo cleanup UAT 2026-06-25 20:22:13 +08:00
wassname d258fe05a0 test: cover nominal and ordinal instrument flows 2026-06-25 20:22:08 +08:00
wassname 6594cc212d api: clarify eval readout names 2026-06-25 20:22:04 +08:00
wassname 4b38de685e cleanup: remove legacy dataset creation artifacts 2026-06-25 20:21:49 +08:00
wassnameandClaudypoo 2d76166cbd showcase: add ordinal steer-effect plot (delta-contrast C dumbbell)
The E map/range are for human comparison; this new per-instrument foundation_dcontrast
figure shows the steer in the sensitive contrast readout (steered minus base C, +C vs -C),
the ordinal twin of the MFV dlogit dumbbell. read_profiles gains a value_col so it reads
either E ('mean') or C. This is the figure that shows what we steered for; the E range hides it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 19:20:25 +08:00
wassnameandClaudypoo ecd2affac4 ordinal readouts: keep raw lp, add sensitive logit contrast C + log-odds
The default ordinal readout was the expected Likert score E = sum k*p_k, which is
insensitive to steering: dE/dl_j = p_j(j-E) vanishes when the model answers confidently
(peaked at the mode), so a steer that reallocates the tails barely moves E. read.py threw
away the raw logprobs after renormalizing, so nothing downstream could recover the signal.

- read.py keeps the raw lp_gather (the primitive) + the think traces on every row.
- readouts.py: pure functions of lp -- expected_score E (human-comparable), logit_contrast
  C = sum (k-mid)*lp_k (primary steer signal: dC/dl_j = w_j, no p_j suppression, normalizer-
  invariant, dC = w.dl exactly), agree_logodds LO (readable 2-bin direction), entropy.
- per_item_categorical also frame-averages the logprobs (exact for the linear contrast).
- administer returns profile_C alongside profile_E, per-item E/C/LO/entropy with bootstrap
  CIs for both, and the raw per-(item,frame) rows with lp + think for downstream reconstruction.

Unit check: on a peaked-at-4 dist under a small disagree steer, dE=-0.11 but dC=-1.80
(=w.dl exactly) and dLO=-0.60; C identical on raw logits vs renormalized logprobs.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 19:14:04 +08:00
wassnameandClaudypoo 3580bffedd add AGENTS.md: fail-fast conventions + the lp_gather primitive
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo abfe0e1649 remove dead shuffle_dimensions negative-control (unwired accretion)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo 2850c31659 fail-fast: drop silent corruption paths in measurement code
guided.py: non-finite answer-slot logits were clamped with nan_to_num(+-1e4),
fabricating a confident answer from a blown-up (steered/quantized) forward pass.
Mark the row incoherent (pmass=0, lp=NaN) instead -- same 'do not compare' signal
as the case-(c) collapse the pipeline already handles.

data.py: load_vignettes silently inner-joined the two condition files and dropped
mismatched ids, so a missing rewrite would change N (and every metric) without
failing. Assert the id sets are identical instead.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo dfa9cdfe5a quiet per-eval logging (INFO->DEBUG) + trim README to use-focused 120 lines
Per-eval INFO lines (rows/think_tokens/aux-stats/first-row/profile/demos) demoted to
DEBUG so a consumer calling evaluate() ~47x/run is not drowned; one-time + WARNING+ kept.
README 308->120: cut process-archeology + per-instrument showcase, added crisp dlogit
and SI definitions for new users.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:51:39 +08:00
wassnameandClaudypoo 2160304541 plots: MFV reuses shared ipsative map + range in z-space; fix range coherence-gate
MFV got two bespoke figures (z-emphasis bars + dlogit dumbbell). Per GPT-5.5 code
review, reuse the geometry not the ordinal semantics: a _mfv_zspace adapter feeds the
same plot_ipsative_pca + plot_range the ordinal instruments use, in z-scored relative-
emphasis space (logit-violation and 1-5 wrongness cannot share a raw axis). Keep the
dlogit dumbbell as a raw-magnitude diagnostic; drop the bespoke map_emphasis.

Also fix a real bug the review found: plot_ordinal gated the map/SPLOM trajectory to
coherent c (pmass >= 0.95*base) but still passed ALL cs to range/range_zoom, so the
figures disagreed on which poles were valid. Range/zoom now use the same coh_cs.
plot_range gains an ylabel param (MFV is z-space, not 1-M); draw_steer asserts sorted
cs; MFV asserts only Social Norms is dropped.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 15:51:40 +08:00
wassnameandClaudypoo fe0d27b9bd maps: more bottom padding + clearer minimap title
Bump the bottom legend strip (PAD_B 0.52->0.62) and lower/shrink the compass + minimap insets so
the compass title clears the lowest data points. Rename the minimap title "full space" ->
"all human respondents" (says what the backdrop cloud is).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 11:48:31 +08:00
wassnameandClaudypoo 22f4f64201 maps: deterministic bottom legend strip (compass left, minimap right)
Replace the per-plot least-crowded-corner heuristic (+ label pruning) with a fixed layout: pad the
y-axis bottom by PAD_B of the data range to reserve a clean strip, then always place the minimap
bottom-right (reads like a small map) and the compass bottom-left. Insets are opaque so the backdrop
haze stays behind them. Same on every instrument, no dependence on where the trajectory heads, and
drops the crowd/rank/prune code + the now-unused pad param.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 11:46:12 +08:00
wassnameandClaudypoo 5d2e0c3630 maps: violin SPLOM diagonals, de-grid off-diagonals, label-aware compass
- SPLOM diagonal: replace the confusing histogram/rug + vertical AI rules with a horizontal violin
  of the human spread + the AI base/+C/-C as the same trajectory dots used off-diagonal (1-D
  analogue), so the diagonal reads the same way as the rest. Unifies full + zoom (no branch).
- SPLOM off-diagonal: jitter the respondent cloud -- MFQ-2 ordinal scores land on a lattice that
  reads as grid-dots; jitter softens it to a density (it is the human joint covariance backdrop).
- Compass corner now counts text labels (society codes, steer labels), not just data points, and a
  final pass drops any society label still under the compass/minimap box. No legend-over-label.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 10:58:00 +08:00
wassnameandClaudypoo 00dca35f85 maps: place-or-omit society labels, unconditional society+steer crop, readable zoom diagonals
- Society labels: pin each ISO code beside its dot (small fixed offset) and DROP any that would
  collide rather than fling it far on a leader line -- close or omitted, never ambiguous.
- Crop to societies + steer for EVERY instrument (drop the synthetic/real branch): the human cloud
  (mfq2 respondents too) is far wider than the societies, so it buried them in a central blob. The
  cloud stays a clipped backdrop + a "full space" minimap shows where the frame sits.
- SPLOM zoom diagonals: the narrow window slices the marginal into solid blocks, so swap the
  histogram for a visible mid-panel society rug + the AI steer rules. Full SPLOM keeps the histogram.
- Minimap: no labels (inset too small); orientation comes from the viewport rectangle.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 10:49:17 +08:00
wassnameandClaudypoo 761bf7412a maps: drop society label leader-lines (draw_lines=False)
The textalloc leader lines made a spider-web across the map; 2-letter ISO codes
are short enough to sit beside their dot with only a small overlap-avoidance
nudge. Cleaner, less ink, label stays next to its point in almost all cases.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 10:31:36 +08:00
wassnameandClaudypoo eb433f1cdc plot showcase: relative coherence gate + persona vec_label
Trajectory coherence gate was absolute (pmass<0.9); make it relative -- keep a
c only if pmass >= 95% of the base (c=0) pmass, else drop it entirely (no hollow
markers). Read vec_label from summary.json so non-authority personas label
correctly instead of the hardcoded "Authority/Care axis".

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 09:09:02 +08:00
wassnameandClaudypoo a2fe33ac53 maps: foundation SPLOM + minimap, crop synthetic-haze maps to societies+steer
- plot_splom: KxK pairs-plot (mfq2 real joint only -- others ship independent-
  marginal haze that would fabricate off-diagonal correlation). lower=joint
  scatter w/ AI base->steer trajectory, diag=marginal+AI rules, upper=Pearson r
  sized by |r|. Ordered by PC1 loading so binding (authority/loyalty) cluster.
  full + AI-zoom (macro/micro). NaN-safe at collapsed poles.
- ipsative map: synthetic haze is far wider than the societies, so it now crops to
  societies+steer (haze clips) and adds a 'full space' minimap with a viewport
  rectangle -- readable big5/16pf/humor maps with macro context kept.
- compass labels via textalloc (was overlapping); NaN-safe crop.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 08:45:01 +08:00
wassnameandClaudypoo 9b13db3390 maps: multi-C trajectory overlay on ipsative map (path + hollow-at-incoherent + adaptive compass)
Driver writes the signed c-multiplier per row; plotter feeds the full c-sweep to
the range (already multi-c capable) and a connected path to the map. Headline
arrows point to calibrated c=+-1; trajectory dots at |c|>1 extend beyond, growing
with |c|, drawn hollow where admin pmass fell below the coherence floor. Compass
moves to the least-crowded corner so the +c arm (which heads toward its own
loading) stops colliding with it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 08:35:29 +08:00
wassnameandClaudypoo 50bdf8d235 plot: automatic MFV foundation-emphasis map vs human cultures
MFV had only the dlogit dumbbell, no human comparison. Add plot_mfv_map: per-
foundation relative emphasis (z across foundations) of model base/+C/-C against
the 5 MFV human cultures (JimenezLeal+Yamada), z-scoring each profile so the
model's logit scale and humans' 1-5 wrongness compare by pattern. Social Norms
dropped (no human norm). Shows the model over-weights Care/Authority vs humans
and +C lowers Authority emphasis. Keeps the dumbbell too.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 07:33:42 +08:00
wassnameandClaudypoo 0e92e66415 maps: human haze on all ipsative maps (synthetic resample for big5/16pf/humor)
Split the PCA-fit basis from the scattered cloud in plot_ipsative_pca: add a
'haze' arg so instruments with only society-level mean+sd get a backdrop
(marginal Normal resample per country) without that resample dictating the
axes. mfq2 keeps its real Atari respondents (fit + haze). Previously only mfq2
had any human scatter; big5/16pf/humor showed bare society dots.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 07:30:37 +08:00
wassnameandClaudypoo 5be5451e13 journal: coherent-C sweep -- mfq2 coherent to C=3, showcase C=1 validated
job 234: ordinal pmass 1.0 at both poles up to C=3.0, steer grows 0.129->0.324.
C=1 is well inside the coherent range. Joint-coherent C is bounded by the side
instruments' -C neutral-degeneracy (a model property at C=1), not ordinal
coherence. Completes the goal's "sweep for the largest coherent C" clause.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:57:37 +08:00
wassnameandClaudypoo f5fc6302d4 validation: 82.6% is irreproducible -- its OWN code gives 0.780 on Qwen3-4B
Ran the exact 2026-05-08 eval (worktree at commit b20ec56, word readout) on
Qwen3-4B: top1 0.780, not 0.826. Every eval version agrees on ~0.78 (digit 0.773,
word-current 0.788, word-original 0.780). The 82.6% was a stale/erroneous table
entry, not a target this model reaches under any pipeline. Canonical value 0.773.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:42:49 +08:00
wassnameandClaudypoo 52b9688de5 README: word readout reads 0.788 not 0.826 -- 0.83 unreachable in current eval
Tested the documented cause (the old word-first-token gather): top1 0.788 on
Qwen3-4B, only ~1.5pt above digit (0.773), not the table's 0.826. So even
reverting the readout does not recover it; the 82.6% came from the broader
2026-05-08 pipeline. Current eval tops out ~0.77-0.79 by every lever.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:29:17 +08:00
wassnameandClaudypoo 3ece864460 add word-readout diagnostic (confirm validation-table 0.826 under old method)
Throwaway probe (like probe_mfv_think_budget): reuses the current _rollout core but
gathers foundation-WORD first tokens + word-keyed schema + no rev-reversal -- the
pre-e4e0f4d readout that produced the 82.6% validation number. Canonical evaluate()
stays digit-only. Confirms whether the showcase model matches the validation table
under the table's own method.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 04:22:14 +08:00