Commit Graph
188 Commits
Author SHA1 Message Date
wassname a8933a98b4 Record directional ablation authority verifier failure 2026-06-30 23:10:53 +08:00
wassname 9eb32d0569 Record sspace authority verifier failure 2026-06-30 22:38:26 +08:00
wassname 9dd7a9a24a Record pca authority verifier failure 2026-06-30 21:35:22 +08:00
wassname 23829d21bb Record mean_diff authority verifier failure 2026-06-30 21:03:58 +08:00
wassname df320d670b Track pure authority method queue 2026-06-30 20:32:46 +08:00
wassname 2d8a736935 Record pure authority UAT gates 2026-06-30 19:52:03 +08:00
wassname 8241551ec1 Replan pure authority steering workflow 2026-06-30 19:02:34 +08:00
wassname e9572ae00a Fix authority pipeline to hold persona pair fixed 2026-06-30 18:59:39 +08:00
wassname 713b3003fd Track authority method comparison queue 2026-06-30 18:43:50 +08:00
wassname 754da269ba Record authority steer verification verdict 2026-06-30 18:15:48 +08:00
wassname bbe59be7e3 Record MFQ2 path failure 2026-06-30 17:36:06 +08:00
wassname 81fda62ecf Track unbuffered authority UAT run 2026-06-30 17:22:24 +08:00
wassname 6788e84738 Track authority fixed-C UAT run 2026-06-30 17:11:58 +08:00
wassname 4862e6f24f Record authority score60 export 2026-06-30 16:29:34 +08:00
wassname d32272790d Track authority stage-b validation 2026-06-30 15:59:38 +08:00
wassname 0c43f1bc62 Tick authority validation subproofs 2026-06-30 15:33:32 +08:00
wassname b1c09966a2 Record DeepInfra Qwen smoke 2026-06-30 15:09:27 +08:00
wassname 8dcaf30705 Record authority validator routing 2026-06-30 15:02:17 +08:00
wassname b74c5ec9d3 Track authority validation workflow 2026-06-30 14:51:39 +08:00
wassname ec0bc32aed Track authority-only steer pipeline 2026-06-30 14:14:07 +08:00
wassname 9fc3852a8a Require explicit showcase steering anchor 2026-06-30 14:08:16 +08:00
wassname fe3b0aff2c Gate showcase paths on contrast and margin 2026-06-30 14:01:59 +08:00
wassname 0d7654506c Add showcase effect table summarizer 2026-06-30 13:50:39 +08:00
wassname 6fcdfea30f Support sampled survey reads and MFV c-grid plots 2026-06-30 13:28:20 +08:00
wassname d8941f47ac Rewrite README around profile plots 2026-06-30 12:38:20 +08:00
wassname 5a39b7ba9c Simplify showcase plot encodings 2026-06-29 21:32:05 +08:00
wassname 504c1f8890 Clarify README metrics and plots 2026-06-29 21:05:02 +08:00
wassname 0bea238224 Rewrite README around datasets and maps 2026-06-29 20:51:54 +08:00
wassnameandClaudypoo 4cc233db3c eval: add a load-bearing SHOULD next to the default-level aux stats
At verbose level 1 (default) the aux-stats line and 64-char free-form had no
interpretation; only verbose>=2 carried SHOULDs. Pair a directional/counterfactual
SHOULD with the always-shown aux stats: pmass_allowed near base + informedness>0 =
in-format signal; pmass_allowed falling / frac_unscorable rising while informedness->0
= the steer broke format and the moral numbers are noise (the steer/breakage confound).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 08:32:43 +08:00
wassnameandClaudypoo dc28c9d0c0 README: consolidate all five instruments' figures at the top, not just MFQ-2
Move big5/16pf/humor/mfv range+map figures up next to MFQ-2 so the showcase
shows every instrument together instead of burying them below the reader-math.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:22:52 +08:00
wassnameandClaudypoo 665f1fe0c9 README: embed range+map showcase for all instruments (16PF range only), from suppression run 311
Render the bundled fairness showcase under the corrected forced-choice
protocol (job 311, token suppression + coherence saved). Embed per-instrument
range plots (mfq2/big5/16pf/humor/mfv) and PCA maps for all but 16PF (16 axes
do not lay out as a readable 2-D map). Honest captions: equality is the only
MFQ-2 factor that rises at +c; on MFV the vector moves Care, not Fairness
(the old fairness-leaning MFV read was protocol-dependent and vanished under
suppression). Coherence held (frac_unscorable 0, margin ~10-12 nat). Drop
stale range_zoom / foundation_dlogit figures.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-27 06:20:03 +08:00
wassnameandClaudypoo fc47d86340 README: pmass is pinned high under forced reads; coherence is frac_unscorable + nll_prefill
The answer slot is force-prefilled, so pmass_allowed is a format floor not a
sensitive coherence gate (a broken steered run still scores ~1.0). Point the
reader at frac_unscorable and nll_prefill instead, and update the think-budget
note: end-of-answer tokens are suppressed so the model can't self-close into
the old ~512 readout collapse.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:17:58 +08:00
wassnameandClaudypoo 3732994193 unscorable reads -> NaN + nanmean, add frac_unscorable coherence signal
pmass under a force-prefilled slot is pinned high (the scaffold primes a
valid token), so a low mean_pmass no longer flags steering breakage. Make
the real signal explicit: unscorable rows (self-close with no answer slot,
or a non-finite forward) now carry pmass=NaN instead of a fabricated 0.0 --
an undefined read is not zero coherence. Means use nanmean (drop them,
matching the dlogit path); frac_unscorable reports the drop rate, and
mean_nll_prefill (scaffold fit) is the sensitive coherence readout. Smoke:
frac_unscorable=0.0, mean_pmass=0.985.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:13:17 +08:00
wassnameandClaudypoo a44dde847d suppress end-of-answer tokens in forced rollout so reads are comparable
Forced-choice scored every sample at whatever state it reached: base
self-closes </think> at long budgets -> junk-cache case (c) -> pmass 0.0,
while steered keeps thinking -> case (b) forced read -> pmass ~1.0. pmass
then measured self-close rate, not coherence, making steering look more
coherent. Suppress {eos, think_end_id} for the whole think budget so no
sample self-closes; all land in case (b) and read at the same forced slot.
Model-agnostic: think_end_id is </think> on reasoning models, falls back
to eos elsewhere. Smoke pmass=0.985.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 21:08:29 +08:00
wassnameandClaudypoo b38614a7ce showcase: swap to moralstory fairness vector across all instruments
One mean_diff vector extracted from moral_stories_foundations fairness
situations (not completions), calibrated once, read base/+c/-c on every
instrument. On MFQ-2 it selectively raises the equality factor; on MFV it
raises Fairness directionally but care leads. README reframed honestly,
stale range_zoom/foundation_dlogit figures dropped (plotter no longer emits
them).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 07:17:20 +08:00
wassnameandClaudypoo 56b21ded13 readme: dataset table (items, framings, measure, human data, source)
Per-instrument: unique items, ways-asked (3 reworded framings ordinal / 2
option-order passes MFV), measure type (log-odds+contrast vs log-prob
categorical), human data availability (per-respondent / per-country / per-item),
and source. big5/16pf country-data citations not recorded in-repo, marked as
such rather than invented. Also drop stale foundation_dlogit/splom figure refs
(plots are map+range only now).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:55:22 +08:00
wassnameandClaudypoo f1d23b338e plots: range+map only for every instrument; legend = c-multiplier not full desc
- drop splom/splom_zoom/range_zoom (ordinal) and the MFV dlogit dumbbell; every
  instrument now yields exactly map_pca_ipsative + range, uniform.
- range pole label is now 'c=+1'/'c=-1' (the multiplier); the full steer desc
  stays in the title only (was crammed into each pole annotation).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:30:02 +08:00
wassnameandClaudypoo 3fcdda4516 plots: drop redundant foundation_dcontrast dumbbell
The range plot already shows the per-foundation steer (human cloud + AI base
dot + +C/-C arrows) with the human anchor the dumbbell lacks; dcontrast plotted
the same delta on the C-nats scale and failed the eraser test. C stays the
measurement in the CSV/foundations table, just not a separate panel.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-26 04:13:12 +08:00
wassname 8ab02adf63 docs: explain logprob readouts in README 2026-06-25 20:37:20 +08:00
wassname 5eb147871c docs: record repo cleanup UAT 2026-06-25 20:22:13 +08:00
wassname d258fe05a0 test: cover nominal and ordinal instrument flows 2026-06-25 20:22:08 +08:00
wassname 6594cc212d api: clarify eval readout names 2026-06-25 20:22:04 +08:00
wassname 4b38de685e cleanup: remove legacy dataset creation artifacts 2026-06-25 20:21:49 +08:00
wassnameandClaudypoo 2d76166cbd showcase: add ordinal steer-effect plot (delta-contrast C dumbbell)
The E map/range are for human comparison; this new per-instrument foundation_dcontrast
figure shows the steer in the sensitive contrast readout (steered minus base C, +C vs -C),
the ordinal twin of the MFV dlogit dumbbell. read_profiles gains a value_col so it reads
either E ('mean') or C. This is the figure that shows what we steered for; the E range hides it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 19:20:25 +08:00
wassnameandClaudypoo ecd2affac4 ordinal readouts: keep raw lp, add sensitive logit contrast C + log-odds
The default ordinal readout was the expected Likert score E = sum k*p_k, which is
insensitive to steering: dE/dl_j = p_j(j-E) vanishes when the model answers confidently
(peaked at the mode), so a steer that reallocates the tails barely moves E. read.py threw
away the raw logprobs after renormalizing, so nothing downstream could recover the signal.

- read.py keeps the raw lp_gather (the primitive) + the think traces on every row.
- readouts.py: pure functions of lp -- expected_score E (human-comparable), logit_contrast
  C = sum (k-mid)*lp_k (primary steer signal: dC/dl_j = w_j, no p_j suppression, normalizer-
  invariant, dC = w.dl exactly), agree_logodds LO (readable 2-bin direction), entropy.
- per_item_categorical also frame-averages the logprobs (exact for the linear contrast).
- administer returns profile_C alongside profile_E, per-item E/C/LO/entropy with bootstrap
  CIs for both, and the raw per-(item,frame) rows with lp + think for downstream reconstruction.

Unit check: on a peaked-at-4 dist under a small disagree steer, dE=-0.11 but dC=-1.80
(=w.dl exactly) and dLO=-0.60; C identical on raw logits vs renormalized logprobs.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 19:14:04 +08:00
wassnameandClaudypoo 3580bffedd add AGENTS.md: fail-fast conventions + the lp_gather primitive
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo abfe0e1649 remove dead shuffle_dimensions negative-control (unwired accretion)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo 2850c31659 fail-fast: drop silent corruption paths in measurement code
guided.py: non-finite answer-slot logits were clamped with nan_to_num(+-1e4),
fabricating a confident answer from a blown-up (steered/quantized) forward pass.
Mark the row incoherent (pmass=0, lp=NaN) instead -- same 'do not compare' signal
as the case-(c) collapse the pipeline already handles.

data.py: load_vignettes silently inner-joined the two condition files and dropped
mismatched ids, so a missing rewrite would change N (and every metric) without
failing. Assert the id sets are identical instead.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:53:11 +08:00
wassnameandClaudypoo dfa9cdfe5a quiet per-eval logging (INFO->DEBUG) + trim README to use-focused 120 lines
Per-eval INFO lines (rows/think_tokens/aux-stats/first-row/profile/demos) demoted to
DEBUG so a consumer calling evaluate() ~47x/run is not drowned; one-time + WARNING+ kept.
README 308->120: cut process-archeology + per-instrument showcase, added crisp dlogit
and SI definitions for new users.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 18:51:39 +08:00
wassnameandClaudypoo 2160304541 plots: MFV reuses shared ipsative map + range in z-space; fix range coherence-gate
MFV got two bespoke figures (z-emphasis bars + dlogit dumbbell). Per GPT-5.5 code
review, reuse the geometry not the ordinal semantics: a _mfv_zspace adapter feeds the
same plot_ipsative_pca + plot_range the ordinal instruments use, in z-scored relative-
emphasis space (logit-violation and 1-5 wrongness cannot share a raw axis). Keep the
dlogit dumbbell as a raw-magnitude diagnostic; drop the bespoke map_emphasis.

Also fix a real bug the review found: plot_ordinal gated the map/SPLOM trajectory to
coherent c (pmass >= 0.95*base) but still passed ALL cs to range/range_zoom, so the
figures disagreed on which poles were valid. Range/zoom now use the same coh_cs.
plot_range gains an ylabel param (MFV is z-space, not 1-M); draw_steer asserts sorted
cs; MFV asserts only Social Norms is dropped.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 15:51:40 +08:00