Reframes from MFV-only to the multi-instrument eval: leads with the LLM-vs-human-cultures
map and quickstart; documents the design choices (logprobs for sensitivity, sliding think
budget, fwd/rev position debias w/ arXiv:2308.11483, SI answer-flip metric, coherence
canary); dev (N=1 x 2 orderings, 64 think, greedy) vs full (N=4 x 2 orderings, high think,
+SI +sampling variance) modes; instrument zoo (MFV forced-choice working, Likert landed but
wiring in progress); used-in + moral_stories_foundations training labels. Preserves the
mechanism/labels/validation/citation content. Lint clean (humanizer).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Unifies forced-choice (nominal) and Likert (ordinal, expectation over the integer
distribution) on tinymfv's answer-token reader. Per scientist panel (docs/reviews/
sci_ma_*.md): canonicalize every frame's distribution to one forward orientation before
metrics+reducer (fixes the asymmetric nominal-reader/ordinal-reducer reflection), renormalize
p with pmass kept as canary, add ordinal |E-error| metric, keying applied only in the profile
reducer (proven orthogonal to framing, not a double-flip), cross-scale guard, negative-control
shuffle. 8 pure-function unit tests pass.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Lead with the plain point, introduce + link Youden's J, spell out the macro
averaging (one-vs-rest per foundation) and point at _informedness for the
formula. Fix stale "two scalars" -> "three". Drop the "flip-informedness"
coinage and "the headline" tell. External-panel comprehension pass: ready
(4.1/5), accuracy and caveats 4-5 across panelists.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
nll_prompt was removed from per_row in f585864 but the script still read it,
crashing the smoke test. Remove the dead refs and surface the new
informedness scalar alongside top1_acc.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Chance-corrected, argmax-only companion to mean_nll: moves when the answer
flips, not when confidence shifts. 0 = base-rate guessing, so it exposes
majority-class models that top1_acc flatters. Same flip-informedness family
as steering-lite's surgical informedness, anchored on the human argmax here.
README also points at the paired training set moral_stories_foundations.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
When the model emits </think> in natural generation but the answer-slot
window detection fails, that's coherence collapse — the model "finished
thinking" without producing JSON. pmass=0.0 is the honest measurement
(no probability mass on allowed tokens at a non-existent slot) and lets
the coherence canary see the failure as a real signal rather than
propagating NaN through np.mean to crash c_scan. nll_json stays NaN
since no JSON was emitted to score.
Triggered by qwen3.6-27b nf4 + LoRA at c=1.0: 1/4 samples hit case (c)
and the NaN aborted c_scan instead of letting it walk down further.
Qwen3.6-27B nf4 + adapter at c=1.0 produced a non-finite raw logit at a
single generated step in 1/4 samples (others used forced-prefill path);
the natural-path F.log_softmax propagated NaN into mean_pmass_allowed,
crashing c_scan. Bound with nan_to_num(±1e4) — leaves argmax-finite rows
unchanged.
Drop legacy-cache-bug rationale from module + function docstrings; the design
stands on its own. Rename to match what the function does (no forking).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Phase 1: batched generate with min_new_tokens=max_new_tokens so cache is uniform
length across the batch (no early stop at </think>). Phase 2: single batched
forced-suffix forward over that cache. Per-sample classification picks
gen.scores at the natural answer position (case a), forced logits (case b
interrupted), or NaN (case c emitted </think> but no answer).
Drops _slice_pkv_one + per-sample fork. The slice helper used layer.keys /
layer.values which crashes on Qwen3.5/3.6 LinearAttentionLayer (gated-delta-net
recurrent state has no .keys/.values). Uniform-length batched cache sidesteps
the cache surface entirely.
Bumps transformers>=5.7 for the Qwen3.5/3.6 gated-delta-net cached-forward
bugfix (resolves to 5.9.0).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Two paired changes the previous commit should have included.
skip_special_tokens kwarg on guided_rollout_forced_choice and evaluate()
threads into tok.decode for gen_text / gen_text_rev. Default False (return
the raw stream with </think>, chat markers, etc.) matches the "return all
the free things" principle. Callers who want stripped output strip
themselves.
emitted_close now uses a token-id match on gen_ids (`(gen_ids ==
think_end_id).any()`) instead of substring on the decoded text. On models
that mark </think> as a special token, the old substring check would
silently always return False when skip_special_tokens=True stripped it.
Qwen3 currently does NOT mark </think> as special so the bug is latent
there, but the fix is strictly more robust and decouples the detection
from the decode flag.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Lets callers ask for N sampled think rollouts per direction instead of one
greedy trace. Per direction we Bayesian-model-average the answer logprobs
across the N samples (logsumexp_n lp - log N) before the fwd/rev average.
Raw per-sample [N, K] logprob matrices stay on the result as
lp_fwd_samples / lp_rev_samples so callers can re-aggregate (log-pooling,
majority vote, etc.).
gen_text and gen_text_rev are now always list[str] of length N (even at
N=1). think_tokens, think_tokens_rev, emitted_close, emitted_close_rev are
length-N lists. At N=1 the BMA is the identity and headline numbers match
the prior greedy path bit-for-bit.
Default max_think_tokens lowered 256 -> 64 for faster default eval (was
expensive overhead on small models that rarely emit </think> anyway).
README updated to match.
Phase 1.5 / Phase 2 already operated per-row, so they extend to B*N
expanded rows without change. Added an explicit assert that the HF
num_return_sequences expansion matches len(user_prompts) * n_samples.
Smoke-tested on Qwen3-0.6B: greedy N=1 matches BMA identity; N=4
temperature=0.7 returns [4, 7] sample matrices and finite pmass; guard
raises if n_samples>1 with temperature=0. evaluate() throughput log
extended to sum fwd+rev think tokens over all samples.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The "negation" mention in pyproject + __init__ docstring was stale —
the actual second pass is internal fwd+rev enum-order debias inside
guided_rollout (position-bias cancellation), not a negation framing.
self_violate is not in Clifford 2015 classic (other-violation only).
Default `evaluate(..., conditions=...)` to ("other_violate",); callers
who want both can opt in explicitly. Halves walltime per eval.
CONDITIONS in data.py still lists both (other_violate, self_violate)
as available — the change is only the evaluate() default.
Old API returned both `think_text` (stripped at </think>) and
`gen_text_full` (everything) — confusing dual field where one was a
strict subset of the other. Library should never silently drop info;
callers can split on `_CLOSE_MARKER` themselves (one line) if they
want the pre-close subset.
Rename:
think_text -> gen_text (forward-frame full decoded gen)
think_text_rev -> gen_text_rev (reverse-frame full decoded gen)
gen_text_full -> dropped (redundant with new gen_text)
Internal `_rollout_kv_fork` now returns 3-tuples
(gen_text, n_think, emitted_close) instead of 4-tuples; suf_ids_for
closure updated. per_row dict in eval.py exposes gen_text + gen_text_rev.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Drop force_min_new_tokens — banning EOS to force a 2048-token think
generates ~250 tokens of real reasoning + </think>, then ~1800 tokens of
post-EOS sycophancy spew. Measuring pmass at the forced-answer slot with
that spew in the KV cache corrupted the coherence signal.
Replace with per-sample Phase 1.5: find each sample's first </think> in
phase1_ids, slice the batched DynamicCache (B, n_heads_kv, seq, d_head)
down to one sample × end_pos seq via _slice_pkv_one. The Phase 2 suffix
forward then runs per-sample over the rewound cache so the answer slot
sees only the coherent thinking trace.
GQA-safe (slices batch + seq, not heads). Phase 1 stays batched, Phase
1.5/2 loop adds ~5-10% wall-clock for the bs=1 forward.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Refactor _rollout_kv_fork from 3 phases (gen → prefix-forward →
prefix+suffix-forward) to 2 phases (gen → suffix-only-forward with
cached pkv). The function name finally matches what it does again.
Phase 1: generate(..., return_dict_in_generate=True) captures past_key_values
for [left-pad, prompt, think, eos-pad]. Same as before; generation already
used cache internally.
Phase 2 (new): per slot, forward ONLY the suffix tokens (close + interrupt
+ nudge + prefill, ~10-30 tokens) with past_key_values=pkv. Logits come
out at suffix positions only; pick the last real one. The attention mask
spans cached prefix + new suffix; pad_id positions get mask=0.
Drops:
- Phase 2a entirely (the prefix re-forward that computed nll_prompt)
- nll_prompt from ForcedChoiceResult, eval.py per_row, eval output dict
- All the sp_per_row / sp_ids_per_row retokenisation gymnastics + boundary-
merge edge cases (lines 107-129 in the old code) — no more text round-trip
- ~115 lines net
Per-row prompt-NLL was a free diagnostic from the prefix forward; with the
forward gone it would cost a dedicated extra forward. pmass_format is the
stronger coherence canary anyway (per AGENTS.md "Coherence signal hierarchy"
and the bidirectional c-scan walkback in 03b_train).
Speed: marginal (saves ~2s out of ~36s per batch on 27B nf4) — the win is
simpler code, not throughput. Module docstring updated to reflect 2-phase
reality.
Smoke (downstream weight-steering-lite repo, on tiny-random) PASS.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
use_cache=False on Phase 2a/2b single forwards saved zero compute — the
flag was leftover from the multi-slot KV-fork era (commit ada854c)
that was refactored away (d34dbfa). One forward per call, no cache to
reuse. Cleaner without it; behaviour identical.
eval.py: add `think_tokens` + `emitted_close` to per_row dict (data was
already in ForcedChoiceResult, just not captured). Log distribution
after eval: median/p75/p90/p99/max + emitted_close count. Lets us see
the actual think budget used vs max_think_tokens cap, to decide if the
512 bump (from 128) is paying for itself or can revert.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
ForcedChoiceResult now carries pmass_format (sum prob mass on the K
foundation answer tokens at the JSON answer slot, averaged across fwd
and rev framings). eval.py aggregates it as mean_pmass_format in both
the headline return dict and the info subdict, and propagates per-row
for sweep/audit consumers.
Direct coherence canary: drops when steering pushes the model toward
non-foundation tokens (gibberish, refusal, format collapse). Independent
of which foundation is picked — complementary to top1_acc (label-
agreement; intentional target shift) and mean_nll_prompt (teacher-forced
prompt nll; falls under steering even when generations break).
Surfacing this lets downstream callers (weight-steering-lite walkback,
report dashboards) gate on actual coherence rather than misusing top1
as a budget.
The 1-row demo block (prompt + think + nudge + prefill + scored token +
64-token free continuation) used logger.info, which meant any caller
wrapping the function in an INFO-level sink (e.g. an agent harness)
got the full trace in their stdout. Downgrade to logger.debug so it
still lands in the user's verbose log but doesn't leak to agents.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Replace 3 parallel scoring paths (guided_rollout / _batch / _multibool)
with a single internal `_rollout_kv_fork` core: phase-1 batched think,
one cached prefix forward, N forked suffix forwards (one per scoring
slot). Binary case is just N_slots=1.
Drops from GuidedResult (no callers): answer_text, raw_full_text,
rep_ratio_think, prompt_nll. eval.py updated accordingly. _ngram_rep_ratio
and _scoring_text helpers removed -- their logic folded into the core.
Verbose=True now logs the full conversation (prefix + suffix the model
sees) plus a 64-token free-form generate continuation, so format issues
are obvious from one slot's log.
File shrinks from ~600 to ~340 lines. smoke_batch_parity passes (bf16
max Δp_true=0.098 within 0.20 tol; pre-existing batched-greedy drift).
Replace mid-turn splice (`I should answer now.</think>{prefill}`) with a
clean turn close + user nudge + fresh assistant prefill. Mirrors what a
chat UI emits when a human interrupts a partial assistant turn, which is
on-policy in chat-tuned training data. Empirically: pmass_format ~0.987
on smoke set vs the OOD splice path.
Close marker is probed from the tokenizer's chat template (sentinel
diff), so it works on Qwen/ChatML, Llama3 (`<|eot_id|>`), etc -- no
hardcoded `<|im_end|>`.
Drops:
- emitted_prefill field (no callers)
- try/except TypeError around apply_chat_template (defensive)
- enable_thinking=False kwarg (some templates reject it; complete
assistant messages auto-strip the think block anyway)
- `\nI should answer now.` fallback in multibool
Adds:
- verbose=True flag on guided_rollout to log scoring_text for debugging
- _assistant_close(tok) sentinel-probe helper
Note: prompt_nll magnitudes shift since scoring_text now includes the
user-nudge tokens. Not comparable to pre-refactor saved results.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
pmass=1.000, 0/132 low-pmass, inter-foundation |r|=0.154 (down from 0.51).
4B fully resolves the foundation-conflation problem seen at 0.6B.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- suf_ids_for() now uses interrupt-msg format: close assistant turn, inject
per-foundation user question + {"Answer": prefill -- fixes authority pmass
(0.344->0.898) by avoiding JSON string-priming from key names
- Add _FOUNDATION_DESCS and _DEFAULT_MULTIBOOL_HINT rubric for discrimination
- Low-pmass diagnostic: first occurrence now runs .generate(max_new_tokens=32)
to show what model actually produces; subsequent cases log top-5 tokens
- suf_ids_per stored so diagnostic generate can reconstruct full input
- Journal: add inter-foundation correlation results (mean |r|=0.51); note
Spearman vs human raters is not a valid metric (exclusive vs independent labels)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>