wassnameandClaudypoo 2d76166cbd showcase: add ordinal steer-effect plot (delta-contrast C dumbbell)
The E map/range are for human comparison; this new per-instrument foundation_dcontrast
figure shows the steer in the sensitive contrast readout (steered minus base C, +C vs -C),
the ordinal twin of the MFV dlogit dumbbell. read_profiles gains a value_col so it reads
either E ('mean') or C. This is the figure that shows what we steered for; the E range hides it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 19:20:25 +08:00
2026-05-08 16:14:07 +08:00
2026-05-08 16:06:41 +08:00

tinymfv (tiny moral/value eval for local LLMs)

A fast, sensitive eval that measures a local model's moral profile and whether an intervention (weight steering, a prompt, a fine-tune) moves it. Instead of sampling and parsing an answer, it prefills the answer slot and reads the next-token logprobs over the seven moral foundations, so a small steering vector shows up as a shift in nats long before it would flip a sampled argmax.

The default instrument is the Clifford moral-foundation vignettes (MFV): 132 scenarios, each with a human distribution over the foundations care / fairness / loyalty / authority / sanctity / liberty / social (Clifford et al. 2015). Three configs (classic, scifi, ai-actor), two framings each (other_violate, self_violate). [HF dataset]

LLM vs 19 human societies on the moral-foundations map

Install

uv pip install git+https://github.com/wassname/tinymfv

Core API

from transformers import AutoModelForCausalLM, AutoTokenizer
from tinymfv import load_vignettes, evaluate

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda()

vignettes = load_vignettes("classic")              # list[dict], one per scenario
report = evaluate(model, tok, vignettes=vignettes) # dev mode: N=1, 64 think tokens
print(report["top1_acc"], report["mean_pmass_allowed"])
print(report["profile"])                           # mean p[foundation] vs the human profile

load_vignettes(name) returns the scenarios (each a dict with the prompt per framing and the human label distribution). evaluate(model, tok, ...) runs a forced-choice probe per (vignette, condition) and returns a dict with:

  • profile: mean p[foundation] across vignettes, on the same 7-way simplex as the human profile.
  • top1_acc, mean_js, mean_nll_T: agreement vs the human label (None if the config is unlabeled).
  • mean_pmass_allowed: the coherence canary -- mean probability mass on valid answer tokens at the answer slot. It drops when the model refuses, rambles, or format-collapses, so a degenerate intervention is visible independent of which answer it picks.
  • per_row (with return_per_row=True): the per-row 7-vec p, raw score (nats), pmass_allowed, top1, margin. This is what the steering metrics below consume.

To measure a steering intervention, run evaluate twice (base vs steered, same vignettes) and diff the reports. The steering-lite package wraps this as evaluate_with_vector(model, tok, vector=v), which returns raw_logratios[vid|cond][foundation] = logit(p[foundation]) for the two metrics below.

The two steering metrics

Both compare a steered report against a base report, per foundation.

dlogit / dlogprob (dlogit_per_foundation). The paired delta in the logit of the foundation probability:

\Delta(\text{vid},\text{cond},f) = \mathrm{logit}\,p_{\text{steer}}[f] - \mathrm{logit}\,p_{\text{base}}[f]

averaged over all (vignette, condition) pairs. Units are nats. Positive means the steer made the model more likely to call that foundation the violation; negative, less likely. It is paired (same vignette base vs steer) and calibration-free, so it does not saturate the way a probability delta would near 0 or 1. This is the continuous effect-size signal.

SI -- Surgical Informedness (si_per_foundation in steering-lite, the canonical implementation). A bidirectional, reference-anchored score that asks "did the steer move the model toward the intended direction at +C and away from it at -C, without breaking the rows that were already right?" Per foundation:

\mathrm{SI} = \mathrm{mean}(\mathrm{SI_{fwd}}, \mathrm{SI_{rev}}) \times \text{pmass\_scale}

where SI_fwd = fix_rate - k * broke_rate (fixes are rows the steer flipped toward intent; broke are rows it flipped away, penalized kx), and SI_rev is the same on the -C pole. The reference is the base model's per-row decision at threshold logit=0 (p=0.5), so SI > 0 always means "moved toward intent at +C and away at -C". pmass_scale = tanh(min margin)^2 softly drops methods whose K-way decision has collapsed. SI moves only when an answer flips, so it is less sensitive than dlogit but more robust: use dlogit for effect size, SI for "did the steer do the intended surgical thing".

Design notes

  • Logprobs, not sampled answers. Prefill the answer slot, read the next-token distribution; small interventions register in nats before changing an argmax.
  • Position-bias control. Each row is scored twice (options forward and reversed) and the logprob vectors averaged, cancelling option-order effects (Pezeshkpour & Hruschka 2023).
  • A sliding think budget. max_think_tokens (0 / 64 dev default / 4096 / unbounded) is a knob you sweep: steering accrues over the think trace, so the same vector moves the profile more with more think, up to the point (~512) where the model closes </think> on its own and the readout collapses.
  • Two modes. dev (N=1, greedy, 64 think) is fast and granular, the default. full (N=4 sampled traces per ordering, Bayesian-model-averaged, high think) adds sampling-variance uncertainty and SI.

Instruments

The reader is answer-space-agnostic: it gathers logprobs over a set of answer tokens at a prefilled slot (src/tinymfv/instrument.py). Forced-choice (nominal, the MFV default) reads a foundation choice; Likert (ordinal) reads a 1..M scale point for MFQ-2 / Big-Five / 16PF / humor-styles (spec and reducers landed; wiring through evaluate() is in progress).

Scope

A fast sensitive eval for small steering interventions on local models, not a full moral-reasoning evaluation. For behaviour-heavy evals see machiavelli, AIRiskDilemmas, ethics_expression_preferences.

Used in steering-lite, lora-lite, w2schar-mini.

Citation

@misc{clark2026tinymfv,
  title = {tinymfv: tiny moral/value eval for local LLMs},
  author = {Michael Clark},
  year = {2026},
  url = {https://github.com/wassname/tinymfv/}
}
S
Description
tiny moral foundations vignettes. logprob eval for steering
Readme
42 MiB
Languages
Python 99.9%
Just 0.1%