At verbose level 1 (default) the aux-stats line and 64-char free-form had no
interpretation; only verbose>=2 carried SHOULDs. Pair a directional/counterfactual
SHOULD with the always-shown aux stats: pmass_allowed near base + informedness>0 =
in-format signal; pmass_allowed falling / frac_unscorable rising while informedness->0
= the steer broke format and the moral numbers are noise (the steer/breakage confound).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Move big5/16pf/humor/mfv range+map figures up next to MFQ-2 so the showcase
shows every instrument together instead of burying them below the reader-math.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Render the bundled fairness showcase under the corrected forced-choice
protocol (job 311, token suppression + coherence saved). Embed per-instrument
range plots (mfq2/big5/16pf/humor/mfv) and PCA maps for all but 16PF (16 axes
do not lay out as a readable 2-D map). Honest captions: equality is the only
MFQ-2 factor that rises at +c; on MFV the vector moves Care, not Fairness
(the old fairness-leaning MFV read was protocol-dependent and vanished under
suppression). Coherence held (frac_unscorable 0, margin ~10-12 nat). Drop
stale range_zoom / foundation_dlogit figures.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The answer slot is force-prefilled, so pmass_allowed is a format floor not a
sensitive coherence gate (a broken steered run still scores ~1.0). Point the
reader at frac_unscorable and nll_prefill instead, and update the think-budget
note: end-of-answer tokens are suppressed so the model can't self-close into
the old ~512 readout collapse.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
pmass under a force-prefilled slot is pinned high (the scaffold primes a
valid token), so a low mean_pmass no longer flags steering breakage. Make
the real signal explicit: unscorable rows (self-close with no answer slot,
or a non-finite forward) now carry pmass=NaN instead of a fabricated 0.0 --
an undefined read is not zero coherence. Means use nanmean (drop them,
matching the dlogit path); frac_unscorable reports the drop rate, and
mean_nll_prefill (scaffold fit) is the sensitive coherence readout. Smoke:
frac_unscorable=0.0, mean_pmass=0.985.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Forced-choice scored every sample at whatever state it reached: base
self-closes </think> at long budgets -> junk-cache case (c) -> pmass 0.0,
while steered keeps thinking -> case (b) forced read -> pmass ~1.0. pmass
then measured self-close rate, not coherence, making steering look more
coherent. Suppress {eos, think_end_id} for the whole think budget so no
sample self-closes; all land in case (b) and read at the same forced slot.
Model-agnostic: think_end_id is </think> on reasoning models, falls back
to eos elsewhere. Smoke pmass=0.985.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
One mean_diff vector extracted from moral_stories_foundations fairness
situations (not completions), calibrated once, read base/+c/-c on every
instrument. On MFQ-2 it selectively raises the equality factor; on MFV it
raises Fairness directionally but care leads. README reframed honestly,
stale range_zoom/foundation_dlogit figures dropped (plotter no longer emits
them).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Per-instrument: unique items, ways-asked (3 reworded framings ordinal / 2
option-order passes MFV), measure type (log-odds+contrast vs log-prob
categorical), human data availability (per-respondent / per-country / per-item),
and source. big5/16pf country-data citations not recorded in-repo, marked as
such rather than invented. Also drop stale foundation_dlogit/splom figure refs
(plots are map+range only now).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
- drop splom/splom_zoom/range_zoom (ordinal) and the MFV dlogit dumbbell; every
instrument now yields exactly map_pca_ipsative + range, uniform.
- range pole label is now 'c=+1'/'c=-1' (the multiplier); the full steer desc
stays in the title only (was crammed into each pole annotation).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The range plot already shows the per-foundation steer (human cloud + AI base
dot + +C/-C arrows) with the human anchor the dumbbell lacks; dcontrast plotted
the same delta on the C-nats scale and failed the eraser test. C stays the
measurement in the CSV/foundations table, just not a separate panel.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The E map/range are for human comparison; this new per-instrument foundation_dcontrast
figure shows the steer in the sensitive contrast readout (steered minus base C, +C vs -C),
the ordinal twin of the MFV dlogit dumbbell. read_profiles gains a value_col so it reads
either E ('mean') or C. This is the figure that shows what we steered for; the E range hides it.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The default ordinal readout was the expected Likert score E = sum k*p_k, which is
insensitive to steering: dE/dl_j = p_j(j-E) vanishes when the model answers confidently
(peaked at the mode), so a steer that reallocates the tails barely moves E. read.py threw
away the raw logprobs after renormalizing, so nothing downstream could recover the signal.
- read.py keeps the raw lp_gather (the primitive) + the think traces on every row.
- readouts.py: pure functions of lp -- expected_score E (human-comparable), logit_contrast
C = sum (k-mid)*lp_k (primary steer signal: dC/dl_j = w_j, no p_j suppression, normalizer-
invariant, dC = w.dl exactly), agree_logodds LO (readable 2-bin direction), entropy.
- per_item_categorical also frame-averages the logprobs (exact for the linear contrast).
- administer returns profile_C alongside profile_E, per-item E/C/LO/entropy with bootstrap
CIs for both, and the raw per-(item,frame) rows with lp + think for downstream reconstruction.
Unit check: on a peaked-at-4 dist under a small disagree steer, dE=-0.11 but dC=-1.80
(=w.dl exactly) and dLO=-0.60; C identical on raw logits vs renormalized logprobs.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
guided.py: non-finite answer-slot logits were clamped with nan_to_num(+-1e4),
fabricating a confident answer from a blown-up (steered/quantized) forward pass.
Mark the row incoherent (pmass=0, lp=NaN) instead -- same 'do not compare' signal
as the case-(c) collapse the pipeline already handles.
data.py: load_vignettes silently inner-joined the two condition files and dropped
mismatched ids, so a missing rewrite would change N (and every metric) without
failing. Assert the id sets are identical instead.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Per-eval INFO lines (rows/think_tokens/aux-stats/first-row/profile/demos) demoted to
DEBUG so a consumer calling evaluate() ~47x/run is not drowned; one-time + WARNING+ kept.
README 308->120: cut process-archeology + per-instrument showcase, added crisp dlogit
and SI definitions for new users.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
MFV got two bespoke figures (z-emphasis bars + dlogit dumbbell). Per GPT-5.5 code
review, reuse the geometry not the ordinal semantics: a _mfv_zspace adapter feeds the
same plot_ipsative_pca + plot_range the ordinal instruments use, in z-scored relative-
emphasis space (logit-violation and 1-5 wrongness cannot share a raw axis). Keep the
dlogit dumbbell as a raw-magnitude diagnostic; drop the bespoke map_emphasis.
Also fix a real bug the review found: plot_ordinal gated the map/SPLOM trajectory to
coherent c (pmass >= 0.95*base) but still passed ALL cs to range/range_zoom, so the
figures disagreed on which poles were valid. Range/zoom now use the same coh_cs.
plot_range gains an ylabel param (MFV is z-space, not 1-M); draw_steer asserts sorted
cs; MFV asserts only Social Norms is dropped.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>