diff --git a/README.md b/README.md index b9e549b..53b1ede 100644 --- a/README.md +++ b/README.md @@ -1,16 +1,25 @@ # tinymfv (tiny moral/value eval for local LLMs) -A fast, sensitive eval that measures a local model's moral profile and whether an intervention -(weight steering, a prompt, a fine-tune) moves it. Instead of sampling and parsing an answer, it -prefills the answer slot and reads the next-token logprobs over the seven moral foundations, so a -small steering vector shows up as a shift in nats long before it would flip a sampled argmax. +tinymfv reads answer-token logprobs from local chat LMs and turns them into moral/value profiles. +It is meant for fast steering experiments: prefill the answer slot, read the next-token +distribution, and compare base vs steered runs in nats instead of waiting for sampled answers to +flip. -The default instrument is the Clifford moral-foundation vignettes (MFV): 132 scenarios, each with a -human distribution over the foundations care / fairness / loyalty / authority / sanctity / liberty / -social (Clifford et al. 2015). Three configs (`classic`, `scifi`, `ai-actor`), two framings each -(`other_violate`, `self_violate`). [[HF dataset](https://huggingface.co/datasets/wassname/tiny-mfv)] +There are two instrument kinds: -![LLM vs 19 human societies on the moral-foundations map](docs/img/showcase/mfq2/map_pca_ipsative.png) +- Nominal MFV vignettes: the answer is a foundation category. The profile is the mean probability + of care / fairness / loyalty / authority / sanctity / liberty / social. +- Ordinal questionnaires: the answer is a scale point 1..M. MFQ-2, Big Five, 16PF, and Humor Styles + all use the same reader and reduce to per-factor Likert profiles. + +The default nominal instrument is the Clifford moral-foundation vignettes (MFV): 132 scenarios, each +with a human distribution over foundations (Clifford et al. 2015). Three configs (`classic`, +`scifi`, `ai-actor`), two framings each (`other_violate`, `self_violate`). +[[HF dataset](https://huggingface.co/datasets/wassname/tiny-mfv)] + +![MFQ-2 culture map: LLM base and steered points against human societies](docs/img/showcase/mfq2/map_pca_ipsative.png) + +![MFQ-2 range plot: human society ranges beside the model steer path](docs/img/showcase/mfq2/range.png) ## Install @@ -18,7 +27,24 @@ social (Clifford et al. 2015). Three configs (`classic`, `scifi`, `ai-actor`), t uv pip install git+https://github.com/wassname/tinymfv ``` -## Core API +For maps: + +```bash +uv pip install "tiny-mfv[maps] @ git+https://github.com/wassname/tinymfv" +``` + +For repo development: + +```bash +git clone https://github.com/wassname/tinymfv +cd tinymfv +uv sync --extra maps --dev +just smoke +``` + +## Use it + +Nominal MFV vignettes use `evaluate`: ```python from transformers import AutoModelForCausalLM, AutoTokenizer @@ -28,55 +54,120 @@ tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda() vignettes = load_vignettes("classic") # list[dict], one per scenario -report = evaluate(model, tok, vignettes=vignettes) # dev mode: N=1, 64 think tokens -print(report["top1_acc"], report["mean_pmass_allowed"]) -print(report["profile"]) # mean p[foundation] vs the human profile +report = evaluate(model, tok, vignettes=vignettes, return_per_row=True) +print(report["mean_pmass_allowed"]) +print(report["per_row"][0]["score"]) # logprob score per foundation, nats +print(report["profile"]) # mean p[foundation], easier to read ``` -`load_vignettes(name)` returns the scenarios (each a dict with the prompt per framing and the human -label distribution). `evaluate(model, tok, ...)` runs a forced-choice probe per (vignette, -condition) and returns a dict with: +Ordinal questionnaires use `administer`: -- `profile`: mean `p[foundation]` across vignettes, on the same 7-way simplex as the human profile. -- `top1_acc`, `informedness`, `mean_nll_T`: agreement vs the human label (None if the config is unlabeled). -- `mean_pmass_allowed`: mean full-vocab probability mass on the allowed answer tokens at the answer - slot. This is answer-slot format coherence: it drops when the model wants to emit prose, refusal, - punctuation, or another out-of-space token, independent of which valid answer is top. -- `mean_nll_prefill`: mean NLL/token of the forced assistant prefill that leads into the answer slot. -- `per_row` (with `return_per_row=True`): the per-row 7-vec `p`, raw `score` (nats), `pmass_allowed`, - `nll_prefill`, `top1`, `margin`. This is what the steering metrics below consume. +```python +from transformers import AutoModelForCausalLM, AutoTokenizer +from tinymfv import administer, get_instrument -To measure a steering intervention, run `evaluate` twice (base vs steered, same vignettes) and diff -the reports. The steering-lite package wraps this as `evaluate_with_vector(model, tok, vector=v)`, -which returns `raw_logratios[vid|cond][foundation] = logit(p[foundation])` for the two metrics below. +tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B") +model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda() -## The two steering metrics +instr = get_instrument("mfq2") # "mfq2", "big5", "16pf", or "humor_styles" +report = administer(model, tok, instr) +print(report["dimensions"]) +print(report["per_item_frame"][0]["lp"]) # raw logprobs for the answer tokens +print(report["profile_E"]) # human-comparable factor means +print(report["profile_C"]) # steer-sensitive logit contrast per factor +print(report["mean_pmass_allowed"]) +``` -Both compare a steered report against a base report, per foundation. +From a checkout, the small functional run is: -**dlogit / dlogprob** (`dlogit_per_foundation`). The paired delta in the logit of the foundation -probability: +```bash +just smoke +just eval Qwen/Qwen3-0.6B classic +``` -$$\Delta(\text{vid},\text{cond},f) = \mathrm{logit}\,p_{\text{steer}}[f] - \mathrm{logit}\,p_{\text{base}}[f]$$ +The plotting helpers live in `tinymfv.maps`; `scripts/plot_steer_showcase.py` shows the current +map/range pipeline for steering-lite outputs. -averaged over all (vignette, condition) pairs. Units are nats. Positive means the steer made the -model more likely to call that foundation the violation; negative, less likely. It is paired -(same vignette base vs steer) and calibration-free, so it does not saturate the way a probability -delta would near 0 or 1. This is the continuous effect-size signal. +## Reader math -**SI -- Surgical Informedness** (`si_per_foundation` in steering-lite, the canonical implementation). -A bidirectional, reference-anchored score that asks "did the steer move the model toward the intended -direction at `+C` and away from it at `-C`, without breaking the rows that were already right?" Per -foundation: +For an answer set $A = \{a_1,\ldots,a_K\}$, tinymfv gathers the full-vocab logprobs at the answer +tokens: + +$$\ell_k = \log P(a_k \mid \text{prompt}, \text{think trace}, \text{assistant prefill})$$ + +The raw gathered vector $\ell$ is the primitive. Everything else is a pure readout: + +$$\mathrm{pmass\_allowed} = \sum_{k=1}^{K} \exp(\ell_k)$$ + +$$p_k = \frac{\exp(\ell_k)}{\mathrm{pmass\_allowed}}$$ + +`pmass_allowed` is answer-format coherence: high means the model put probability mass on valid +answer tokens at the answer slot. Low means it wanted prose, punctuation, a refusal, or another +out-of-space token. It does not say which valid answer is right. + +`nll_prefill` measures whether the forced assistant prefill itself fits the model: + +$$\mathrm{nll\_prefill} = -\frac{1}{J}\sum_{j=1}^{J} \log P(u_j \mid \text{context}, u_{ 0 always means "moved toward -intent at `+C` and away at `-C`". `pmass_scale = tanh(min margin)^2` softly drops methods whose K-way -decision has collapsed. SI moves only when an answer flips, so it is less sensitive than dlogit but -more robust: use dlogit for effect size, SI for "did the steer do the intended surgical thing". +Use delta logit or delta $C$ for effect size. Use SI only when you care about thresholded decisions. ## Design notes @@ -100,6 +191,13 @@ The reader is answer-space-agnostic: it gathers logprobs over answer tokens at a - Ordinal instruments, MFQ-2 / Big-Five / 16PF / humor-styles, read a 1..M scale point and reduce to keyed expected score `E`, logit contrast `C`, `logodds_agree`, entropy, and `pmass_allowed`. +The bundled public map references are: + +- `docs/img/showcase/mfq2/map_pca_ipsative.png`: culture map, model base and steer poles against + human societies. +- `docs/img/showcase/mfq2/range.png`: per-factor human ranges beside the model steer path. +- `docs/img/showcase/mfv/foundation_dlogit.png`: MFV per-foundation steer effect in nats. + ## Scope A fast sensitive eval for small steering interventions on local models, not a full moral-reasoning diff --git a/docs/spec/20260625_repo_cleanup.md b/docs/spec/20260625_repo_cleanup.md index 57a6a72..ca9ef39 100644 --- a/docs/spec/20260625_repo_cleanup.md +++ b/docs/spec/20260625_repo_cleanup.md @@ -15,6 +15,7 @@ Out: rewriting the model rollout core, changing steering-lite metrics. - R5: Python files compile. Done means: `python -m py_compile src/tinymfv/*.py scripts/*.py`. UAT: `/tmp/tinymfv_pycompile_cleanup.log`. - R6: The main functional path still runs. Done means: `just smoke` loads a real tiny HF model, evaluates 5 vignettes, and writes a result JSONL. UAT: `/tmp/tinymfv_smoke_cleanup.log` and `data/results/forced_choice_classic.jsonl`. - R7: Public prose avoids obvious AI-generated phrasing. Done means: humanizer lint has no OVER checks on README/spec. UAT: `/tmp/tinymfv_humanizer_lint_final.log`. +- R8: README explains what tinymfv measures, how to run it, why logprob readouts are the main research handle, and how the map/range figures relate to the easier summaries. Done means: README contains `evaluate`, `administer`, `pmass_allowed`, `nll_prefill`, `score`, `lp`, `profile_E`, `profile_C`, `dlogit_per_foundation`, `si_per_foundation`, and tracked image links. UAT: `/tmp/tinymfv_readme_docs_check.log`. ## Tasks - [x] T1 (R1): Remove the dead `shuffle_dimensions` test and add nominal categorical reducer coverage. @@ -24,11 +25,13 @@ Out: rewriting the model rollout core, changing steering-lite metrics. - [x] T5 (R4): Update public docs and comments away from legacy JS headline and stale "wiring in progress" text. - [x] T6 (R5/R6): Run UAT commands and record paths. - [x] T7 (R7): Run humanizer on README/spec prose and fix mechanical tells. +- [x] T8 (R8): Rewrite README around the measured quantities: raw logprobs first, easier value summaries second, then showcase figures. ## Log - 2026-06-25: `shuffle_dimensions` was removed in commit `abfe0e1`, but tests still import it. - 2026-06-25: Root and packaged vignette JSONL currently hash-identical, so deleting root duplicates should not change eval data. - 2026-06-25: Branch `creating-tinymfv` exists before deleting the OpenRouter creation scripts from main. +- 2026-06-25: README had the MFQ-2 culture-map image and MFV steering metrics, but it did not show the range plot or the ordinal `administer` path. ## UAT - Human-readable evidence file: `docs/spec/20260625_repo_cleanup_uat.md`. @@ -37,6 +40,8 @@ Out: rewriting the model rollout core, changing steering-lite metrics. - `just smoke`: PASS, 5 Qwen/Qwen3-0.6B vignette rows, `mean_pmass_allowed=0.985396534204483`, result at `data/results/forced_choice_classic.jsonl`, see `/tmp/tinymfv_smoke_cleanup.log`. - `python - <<'PY' ... data/results/forced_choice_classic.jsonl`: PASS, first result row includes `p`, `score`, `label`, `pmass_allowed`, `nll_prefill`, `top1`, and `margin`. - `python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md`: PASS, no OVER checks, see `/tmp/tinymfv_humanizer_lint_final.log`. +- `python - <<'PY' ... README.md`: PASS, required terms and all README image links present, see `/tmp/tinymfv_readme_docs_check.log`. +- `python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md docs/spec/20260625_repo_cleanup_uat.md`: PASS, no OVER checks, see `/tmp/tinymfv_humanizer_lint_readme_followup.log`. - `rg -n 'mean_js|median_js|max_js|nll_json|mean_nll_json|agree_logodds|logodds\b|keyed_logodds\b|openrouter|OpenRouter|OPENROUTER|openrouter_request|shuffle_dimensions|\bJS\b|Jensen|jensen' src scripts tests README.md pyproject.toml justfile`: PASS, output empty. - `rg -n 'ROOT / "data" / f"vignettes|ROOT / "data" / "vignettes' scripts src`: PASS, output empty. - `git branch --list creating-tinymfv`: PASS, branch exists. diff --git a/docs/spec/20260625_repo_cleanup_uat.md b/docs/spec/20260625_repo_cleanup_uat.md index 04d99d9..cb868b5 100644 --- a/docs/spec/20260625_repo_cleanup_uat.md +++ b/docs/spec/20260625_repo_cleanup_uat.md @@ -82,6 +82,58 @@ docs/spec/20260625_repo_cleanup_uat.md: 19 prose lines check count budget status ``` +## README follow-up + +Command: + +```sh +python - <<'PY' 2>&1 | tee /tmp/tinymfv_readme_docs_check.log +... +PY +``` + +Result excerpt: + +```text +required_terms_missing: [] +image_links: ['docs/img/showcase/mfq2/map_pca_ipsative.png', 'docs/img/showcase/mfq2/range.png', 'docs/img/showcase/mfv/foundation_dlogit.png'] +missing_image_links: [] +tracked_image_links: ['docs/img/showcase/mfq2/map_pca_ipsative.png', 'docs/img/showcase/mfq2/range.png', 'docs/img/showcase/mfv/foundation_dlogit.png'] +README docs check: PASS +``` + +Humanizer follow-up command: + +```sh +python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md docs/spec/20260625_repo_cleanup_uat.md +``` + +Result excerpt from `/tmp/tinymfv_humanizer_lint_readme_followup.log`: + +```text +README.md: 108 prose lines +check count budget status +tricolon 1 - info + +docs/spec/20260625_repo_cleanup.md: 41 prose lines +check count budget status + +docs/spec/20260625_repo_cleanup_uat.md: 25 prose lines +check count budget status +``` + +Fresh-eyes review: + +```text +no blockers. +README diff answers the latest ask well: raw answer-token logprobs are the primitive; +pmass_allowed/nll_prefill are defined; sensitive relative readouts are score, lp, +dlogit_per_foundation, and delta C; easier value summaries are profile, profile_E, +and logodds_agree. Artifact sweeps found no tracked/staged __pycache__, .pyc, +data/results, logs, root duplicate data, OpenRouter, multilabel, separation, or +calibration junk. +``` + ## stale-name sweeps Command: