tinymfv (tiny moral/value eval for local LLMs)

tinymfv reads answer-token logprobs from local chat LMs and turns them into moral/value profiles. It is meant for fast steering experiments: prefill the answer slot, read the next-token distribution, and compare base vs steered runs in nats instead of waiting for sampled answers to flip.

There are two instrument kinds:

  • Nominal MFV vignettes: the answer is a foundation category. The profile is the mean probability of care / fairness / loyalty / authority / sanctity / liberty / social.
  • Ordinal questionnaires: the answer is a scale point 1..M. MFQ-2, Big Five, 16PF, and Humor Styles all use the same reader and reduce to per-factor Likert profiles.

The default nominal instrument is the Clifford moral-foundation vignettes (MFV): 132 scenarios, each with a human distribution over foundations (Clifford et al. 2015). Three configs (classic, scifi, ai-actor), two framings each (other_violate, self_violate). [HF dataset]

MFQ-2 culture map: LLM base and steered points against human societies

MFQ-2 range plot: human society ranges beside the model steer path

Install

uv pip install git+https://github.com/wassname/tinymfv

For maps:

uv pip install "tiny-mfv[maps] @ git+https://github.com/wassname/tinymfv"

For repo development:

git clone https://github.com/wassname/tinymfv
cd tinymfv
uv sync --extra maps --dev
just smoke

Use it

Nominal MFV vignettes use evaluate:

from transformers import AutoModelForCausalLM, AutoTokenizer
from tinymfv import load_vignettes, evaluate

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda()

vignettes = load_vignettes("classic")              # list[dict], one per scenario
report = evaluate(model, tok, vignettes=vignettes, return_per_row=True)
print(report["mean_pmass_allowed"])
print(report["per_row"][0]["score"])               # logprob score per foundation, nats
print(report["profile"])                           # mean p[foundation], easier to read

Ordinal questionnaires use administer:

from transformers import AutoModelForCausalLM, AutoTokenizer
from tinymfv import administer, get_instrument

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda()

instr = get_instrument("mfq2")       # "mfq2", "big5", "16pf", or "humor_styles"
report = administer(model, tok, instr)
print(report["dimensions"])
print(report["per_item_frame"][0]["lp"])           # raw logprobs for the answer tokens
print(report["profile_E"])           # human-comparable factor means
print(report["profile_C"])           # steer-sensitive logit contrast per factor
print(report["mean_pmass_allowed"])

From a checkout, the small functional run is:

just smoke
just eval Qwen/Qwen3-0.6B classic

The plotting helpers live in tinymfv.maps; scripts/plot_steer_showcase.py shows the current map/range pipeline for steering-lite outputs.

Reader math

For an answer set A = \{a_1,\ldots,a_K\}, tinymfv gathers the full-vocab logprobs at the answer tokens:

\ell_k = \log P(a_k \mid \text{prompt}, \text{think trace}, \text{assistant prefill})

The raw gathered vector \ell is the primitive. Everything else is a pure readout:

\mathrm{pmass\_allowed} = \sum_{k=1}^{K} \exp(\ell_k) p_k = \frac{\exp(\ell_k)}{\mathrm{pmass\_allowed}}

pmass_allowed is answer-format coherence: high means the model put probability mass on valid answer tokens at the answer slot. Low means it wanted prose, punctuation, a refusal, or another out-of-space token. It does not say which valid answer is right.

nll_prefill measures whether the forced assistant prefill itself fits the model:

\mathrm{nll\_prefill} = -\frac{1}{J}\sum_{j=1}^{J} \log P(u_j \mid \text{context}, u_{<j})

where u_1,\ldots,u_J are the prefill tokens before the answer token. This catches scaffold friction before the answer slot; pmass_allowed catches the slot itself.

What to use

Most research code should use the logprob-level outputs. They are sensitive and easy to interpret in relative terms: positive means the steer moved the answer up on that axis, negative means down, and the units are nats. For MFV, use per-row score or the paired dlogit_per_foundation. For ordinal questionnaires, use raw per-frame lp when you want the primitive, or profile_C when you want a factor-level steer readout.

The value summaries are less sensitive but easier to interpret. MFV profile is mean foundation probability. Ordinal profile_E is the expected Likert score, in the human scale direction:

E = \sum_{k=1}^{M} k p_k

For reverse-keyed items, keyed_E = M + 1 - E. E is bounded, so it is good for comparing to human survey means but can hide small steering effects near confident answers. That is why the ordinal steering readout is the rank-centered logit contrast:

C = \sum_{k=1}^{M} \left(k - \frac{M + 1}{2}\right)\ell_k

The weights sum to zero, so any constant shift in the logprobs cancels. C is invariant to renormalizing over the allowed answer tokens and linear in logprob changes:

\Delta C = \sum_{k=1}^{M} \left(k - \frac{M + 1}{2}\right)(\ell^{\mathrm{steer}}_k - \ell^{\mathrm{base}}_k)

For MFV, the equivalent sensitive readout is the paired change in foundation logit:

\Delta_f(i,c) = \mathrm{logit}\,p^{\mathrm{steer}}_{i,c,f} - \mathrm{logit}\,p^{\mathrm{base}}_{i,c,f}

averaged over vignette i and condition c. Positive means the steer made the model more likely to call foundation f the violation; negative means less likely.

logodds_agree is the easier ordinal direction summary:

$$\mathrm{logodds_agree} = \log\sum_{k \in \mathrm{agree}} \exp(\ell_k)

\log\sum_{k \in \mathrm{disagree}} \exp(\ell_k)$$

It drops the neutral middle option on odd scales. It is more readable than C, but it throws away rank information.

Full return schemas live in the TypedDicts in src/tinymfv/eval.py and src/tinymfv/administer.py.

Steering plots

MFV steer effect: per-foundation delta logit for positive and negative poles

si_per_foundation lives in steering-lite. It is a stricter, thresholded flip metric: did the steer move rows toward the intended foundation at +C and away at -C, while penalizing rows that were already right and got broken?

\mathrm{SI} = \mathrm{mean}(\mathrm{SI_{fwd}}, \mathrm{SI_{rev}}) \times \text{pmass\_scale}

Use delta logit or delta C for effect size. Use SI only when you care about thresholded decisions.

Design notes

  • Logprobs, not sampled answers. Prefill the answer slot, read the next-token distribution; small interventions register in nats before changing an argmax.
  • Position-bias control. Each row is scored twice (options forward and reversed) and the logprob vectors averaged, cancelling option-order effects (Pezeshkpour & Hruschka 2023).
  • A sliding think budget. max_think_tokens (0 / 64 dev default / 4096 / unbounded) is a setting you sweep: steering accrues over the think trace, so the same vector moves the profile more with more think, up to the point (~512) where the model closes </think> on its own and the readout collapses.
  • Two modes. dev (N=1, greedy, 64 think) is fast and granular, the default. full (N=4 sampled traces per ordering, Bayesian-model-averaged, high think) adds sampling-variance uncertainty and SI.

Instruments

The reader is answer-space-agnostic: it gathers logprobs over answer tokens at a prefilled slot (src/tinymfv/instrument.py).

  • Nominal instruments, the MFV vignettes, read a foundation category and reduce to mean category probability.
  • Ordinal instruments, MFQ-2 / Big-Five / 16PF / humor-styles, read a 1..M scale point and reduce to keyed expected score E, logit contrast C, logodds_agree, entropy, and pmass_allowed.

The bundled public map references are:

  • docs/img/showcase/mfq2/map_pca_ipsative.png: culture map, model base and steer poles against human societies.
  • docs/img/showcase/mfq2/range.png: per-factor human ranges beside the model steer path.
  • docs/img/showcase/mfv/foundation_dlogit.png: MFV per-foundation steer effect in nats.

Scope

A fast sensitive eval for small steering interventions on local models, not a full moral-reasoning evaluation. For behaviour-heavy evals see machiavelli, AIRiskDilemmas, ethics_expression_preferences.

Used in steering-lite, lora-lite, w2schar-mini.

Citation

@misc{clark2026tinymfv,
  title = {tinymfv: tiny moral/value eval for local LLMs},
  author = {Michael Clark},
  year = {2026},
  url = {https://github.com/wassname/tinymfv/}
}
S
Description
tiny moral foundations vignettes. logprob eval for steering
Readme
42 MiB
Languages
Python 99.9%
Just 0.1%