Commit Graph
27 Commits
Author SHA1 Message Date
wassname d34dbfa9f8 Refactor guided rollout scoring to use flat prefix+suffix approach and remove full attention assertion 2026-05-14 11:35:03 +00:00
wassname 9abddaeac5 return pmass 2026-05-08 16:36:36 +08:00
wassname 0442935279 clean 2026-05-08 15:33:29 +08:00
wassname 8dfaf299ca rename 2026-05-08 15:30:06 +08:00
wassname b12770cb78 fixes, naming 2026-05-08 15:23:54 +08:00
wassname c96d02a675 refactor 2026-05-08 15:15:14 +08:00
wassname d796df85c8 improved to have better airisk, better eval that distinguished factors 2026-05-08 14:04:56 +08:00
wassname 0a769d71a8 Merge feat/batched-guided-rollout: KV-fork unified guided rollout
- guided.py turn-boundary close+nudge scoring (chat-template probed)
- unified guided_rollout / _batch / _multibool onto _rollout_kv_fork core
- dropped prompt_nll, answer_text, raw_full_text, rep_ratio_think (no callers)
- verbose=True now logs full convo + 64-tok generate continuation
- file shrinks ~600->~340 lines

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

# Conflicts:
#	src/tinymfv/core.py
2026-05-06 13:26:09 +08:00
wassname 053009b22e core: drop prompt_nll plumbing in analyse()
Follow-up to e8b51b2 which removed prompt_nll from GuidedResult. Nothing
in eval.py or analyse callers reads it anymore.
2026-05-06 13:17:46 +08:00
wassname e8b51b2b48 guided: unify binary path onto multibool KV-fork core
Replace 3 parallel scoring paths (guided_rollout / _batch / _multibool)
with a single internal `_rollout_kv_fork` core: phase-1 batched think,
one cached prefix forward, N forked suffix forwards (one per scoring
slot). Binary case is just N_slots=1.

Drops from GuidedResult (no callers): answer_text, raw_full_text,
rep_ratio_think, prompt_nll. eval.py updated accordingly. _ngram_rep_ratio
and _scoring_text helpers removed -- their logic folded into the core.

Verbose=True now logs the full conversation (prefix + suffix the model
sees) plus a 64-token free-form generate continuation, so format issues
are obvious from one slot's log.

File shrinks from ~600 to ~340 lines. smoke_batch_parity passes (bf16
max Δp_true=0.098 within 0.20 tol; pre-existing batched-greedy drift).
2026-05-06 13:16:02 +08:00
wassnameandClaude Opus 4.7 a59003d2d0 guided: turn-boundary scoring text via chat-template probe
Replace mid-turn splice (`I should answer now.</think>{prefill}`) with a
clean turn close + user nudge + fresh assistant prefill. Mirrors what a
chat UI emits when a human interrupts a partial assistant turn, which is
on-policy in chat-tuned training data. Empirically: pmass_format ~0.987
on smoke set vs the OOD splice path.

Close marker is probed from the tokenizer's chat template (sentinel
diff), so it works on Qwen/ChatML, Llama3 (`<|eot_id|>`), etc -- no
hardcoded `<|im_end|>`.

Drops:
- emitted_prefill field (no callers)
- try/except TypeError around apply_chat_template (defensive)
- enable_thinking=False kwarg (some templates reject it; complete
  assistant messages auto-strip the think block anyway)
- `\nI should answer now.` fallback in multibool

Adds:
- verbose=True flag on guided_rollout to log scoring_text for debugging
- _assistant_close(tok) sentinel-probe helper

Note: prompt_nll magnitudes shift since scoring_text now includes the
user-nudge tokens. Not comparable to pre-refactor saved results.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-06 12:39:10 +08:00
wassnameandClaude Sonnet 4.6 2ca6aa9c6c multibool: interrupt-msg fork, generate trace on low-pmass, journal update
- suf_ids_for() now uses interrupt-msg format: close assistant turn, inject
  per-foundation user question + {"Answer": prefill -- fixes authority pmass
  (0.344->0.898) by avoiding JSON string-priming from key names
- Add _FOUNDATION_DESCS and _DEFAULT_MULTIBOOL_HINT rubric for discrimination
- Low-pmass diagnostic: first occurrence now runs .generate(max_new_tokens=32)
  to show what model actually produces; subsequent cases log top-5 tokens
- suf_ids_per stored so diagnostic generate can reconstruct full input
- Journal: add inter-foundation correlation results (mean |r|=0.51); note
  Spearman vs human raters is not a valid metric (exclusive vs independent labels)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-06 11:29:33 +08:00
wassnameandClaude Opus 4.7 ada854c499 guided_rollout_multibool: switch to 12 single-slot KV-forks + assert full-attention
The chained-fill design (one suffix with all foundations, two passes for true/false)
hit a non-recoverable conv-state issue on hybrid linear-attention layers (Qwen3.5):
splitting prefix and suffix forwards via past_key_values silently produced
wrong logits (pmass dropping to 0.04, top token leaking to ' "' = 0.72).

Switched to 12 independent single-slot completions per prompt:
  for (frame, foundation) in {is_violation, is_ok} × foundations:
    cache scoring_prefix once, fork suffix `\n{"<frame>": {"<f>":`,
    read logits at the last token (predicting `true|false`).
  final[f] = 0.5 * (lr_violation[f] - lr_ok[f])

Framing flip cancels per-key prior bias the same way true/false fill did,
without the chained-slot causality that interacts badly with split forwards.

Added _assert_full_attention(): checks model.config.layer_types and fails
loudly on hybrid models. Verified parity vs flat forward on Qwen3-0.6B
(Δ ≤ 0.13 nats; signal of interest is ≫1 nat) and assert fires on Qwen3.5-0.8B.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-05 22:19:11 +08:00
wassname 1650aa9191 add prompt_nll (free coherence proxy) + eval summary line
Per-token NLL over the scoring text is free since we already compute
full-sequence logits in guided_rollout / guided_rollout_batch (just
gather instead of slicing [:, -1]). Higher NLL = model less coherent
on this prompt under whatever steering is attached.

eval.py logs pmass/ppl/nll aggregate at end of guided eval so the
next run shows degradation at a glance instead of buried tqdm noise.
analyse() now exposes raw_nll and info.prompt_nll_mean.
2026-05-05 06:34:24 +08:00
wassname 8f39fb1462 immediately 2026-05-04 06:09:56 +08:00
wassnameandClaude Sonnet 4.6 50efa6063c fix: use valid JSON Schema in eval frame prompts
`boolean` is not valid JSON; use `{"type": "boolean"}` so models
produce true/false rather than the integer shorthand 1.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-03 14:16:42 +08:00
wassnameandClaude Sonnet 4.6 b879596c2b fix: use valid JSON Schema in eval frame prompts
`boolean` is not valid JSON; use `{"type": "boolean"}` so models
produce true/false rather than the integer shorthand 1.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-03 14:16:21 +08:00
wassname 076b859b9a fix: remove HuggingFace fallback from load_vignettes for strict fail-first behavior 2026-05-03 13:00:11 +08:00
wassname bcbdb9cc6f feat: multi-label moral foundation ratings with z-scored frame averaging and human calibration
- Add scripts/07_multilabel.py: LLM judge rates all 7 foundations per vignette
  using violation (forward) and acceptability (reverse) frames
- Foundation definitions drawn from Clifford et al. (2015) survey rubric
- Z-score each frame per foundation before averaging to cancel range bias
- Calibrate LLM Likert → human % via per-foundation OLS (classic set only)
- Add scripts/07a_merge_labels.py: merges llm_* and calibrated_* into vignette files
- Update README and HF dataset card with methodology and calibration quality table
- Classic set: 80.3% dominant-foundation accuracy, Pearson r 0.69-0.89 per foundation
2026-05-03 12:48:14 +08:00
wassname cee50a6f9d pmass 2026-05-03 11:49:58 +08:00
wassname 881ac16c24 API improvements: rename clifford->classic, default load_vignettes to all, add dual-axis docs, and update HF upload script 2026-05-03 07:01:28 +08:00
wassname b7f92dccb1 lower pmass-warn threshold 0.9 -> 0.5 in batched (match sequential) 2026-05-03 06:51:21 +08:00
wassname addf47c5a0 quiet pmass-low warning: one summary per batch
Was emitting `logger.warning("pmass=0.XX<0.9 — top-5: ...")` per-row, which
spammed the log heavily during heavy-steering eval (many rows go OOD at once).
Now collects all low-pmass rows in the batch and emits one summary line with
the worst-case top-5, e.g.:

    pmass<0.9 on 7/16 rows in this batch; worst=0.412 top-5: '1'=0.40, ...

Same diagnostic signal, ~16× fewer log lines per batch.
2026-05-03 06:50:19 +08:00
wassname e996d57051 batch the eval: guided_rollout_batch + 4× speedup
Sequential eval was the bottleneck (~12s/vignette × 131 = 27 min/pass; with
bidirectional ±C × 14 methods, that projected to ~17 h). Three model calls
per row (phase1 generate, scoring forward, cosmetic continuation) became one
phase1 + one scoring per *batch*; continuation generate dropped (callers
only use p_true + pmass_format).

Parity smoke (scripts/smoke_batch_parity.py): float32 is bit-exact (max
Δp_true=0.0000 over 16 prompts, 4.3× speedup at limit=4). bf16 drifts on
individual rows — greedy argmax flips at near-tie tokens then phase1
diverges — but pmass agrees within 0.03 (scoring forward correct) and
aggregates over 131 vignettes will average out the per-row noise.

Other changes shipped in this commit:
- core.py: analyse() now returns raw_pmass dict alongside raw p_true (callers
  needed per-(vid,cond,frame) pmass for diagnostic warnings).
- guided.py guided_rollout: warn + log top-5 when pmass<0.9 (catches OOD
  steering / format-broken vignettes without a separate audit pass).
- eval.py: pre-tokenize a sample to log expected prompt+cache budget so OOM
  is predictable from the SHOULD line; group items by frame so each batch
  shares schema_hint + prefill.
2026-05-03 06:42:17 +08:00
wassname 0f8048d5d9 Implement N-token evaluation with guided rollouts
- Refactored evaluation logic in `src/tinymfv/eval.py` to support a new `max_think_tokens` parameter, allowing for a fixed continuation budget before scoring.
- Introduced `guided_rollout` function in `src/tinymfv/guided.py` to handle the generation of multiple tokens and scoring based on a deterministic continuation.
- Updated the CLI in `scripts/03_eval.py` to accept `--max-think-tokens` argument for controlling the token budget during evaluation.
- Created a new specification document `docs/spec/20260501_n_token_eval.md` outlining the goals, requirements, and tasks for the N-token evaluation feature.
- Simplified the record creation in `scripts/02_rewrite.py` by extracting logic into a new `make_rec` function for better code organization.
2026-05-01 21:44:14 +08:00
wassname 252e62abb7 decent 2026-04-30 21:22:07 +08:00
wassname a155f5594b valdiation 2026-04-30 20:08:12 +08:00