docs: explain logprob readouts in README

This commit is contained in:
wassname
2026-06-25 20:37:20 +08:00
parent 5eb147871c
commit 8ab02adf63
3 changed files with 201 additions and 46 deletions
+5
View File
@@ -15,6 +15,7 @@ Out: rewriting the model rollout core, changing steering-lite metrics.
- R5: Python files compile. Done means: `python -m py_compile src/tinymfv/*.py scripts/*.py`. UAT: `/tmp/tinymfv_pycompile_cleanup.log`.
- R6: The main functional path still runs. Done means: `just smoke` loads a real tiny HF model, evaluates 5 vignettes, and writes a result JSONL. UAT: `/tmp/tinymfv_smoke_cleanup.log` and `data/results/forced_choice_classic.jsonl`.
- R7: Public prose avoids obvious AI-generated phrasing. Done means: humanizer lint has no OVER checks on README/spec. UAT: `/tmp/tinymfv_humanizer_lint_final.log`.
- R8: README explains what tinymfv measures, how to run it, why logprob readouts are the main research handle, and how the map/range figures relate to the easier summaries. Done means: README contains `evaluate`, `administer`, `pmass_allowed`, `nll_prefill`, `score`, `lp`, `profile_E`, `profile_C`, `dlogit_per_foundation`, `si_per_foundation`, and tracked image links. UAT: `/tmp/tinymfv_readme_docs_check.log`.
## Tasks
- [x] T1 (R1): Remove the dead `shuffle_dimensions` test and add nominal categorical reducer coverage.
@@ -24,11 +25,13 @@ Out: rewriting the model rollout core, changing steering-lite metrics.
- [x] T5 (R4): Update public docs and comments away from legacy JS headline and stale "wiring in progress" text.
- [x] T6 (R5/R6): Run UAT commands and record paths.
- [x] T7 (R7): Run humanizer on README/spec prose and fix mechanical tells.
- [x] T8 (R8): Rewrite README around the measured quantities: raw logprobs first, easier value summaries second, then showcase figures.
## Log
- 2026-06-25: `shuffle_dimensions` was removed in commit `abfe0e1`, but tests still import it.
- 2026-06-25: Root and packaged vignette JSONL currently hash-identical, so deleting root duplicates should not change eval data.
- 2026-06-25: Branch `creating-tinymfv` exists before deleting the OpenRouter creation scripts from main.
- 2026-06-25: README had the MFQ-2 culture-map image and MFV steering metrics, but it did not show the range plot or the ordinal `administer` path.
## UAT
- Human-readable evidence file: `docs/spec/20260625_repo_cleanup_uat.md`.
@@ -37,6 +40,8 @@ Out: rewriting the model rollout core, changing steering-lite metrics.
- `just smoke`: PASS, 5 Qwen/Qwen3-0.6B vignette rows, `mean_pmass_allowed=0.985396534204483`, result at `data/results/forced_choice_classic.jsonl`, see `/tmp/tinymfv_smoke_cleanup.log`.
- `python - <<'PY' ... data/results/forced_choice_classic.jsonl`: PASS, first result row includes `p`, `score`, `label`, `pmass_allowed`, `nll_prefill`, `top1`, and `margin`.
- `python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md`: PASS, no OVER checks, see `/tmp/tinymfv_humanizer_lint_final.log`.
- `python - <<'PY' ... README.md`: PASS, required terms and all README image links present, see `/tmp/tinymfv_readme_docs_check.log`.
- `python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md docs/spec/20260625_repo_cleanup_uat.md`: PASS, no OVER checks, see `/tmp/tinymfv_humanizer_lint_readme_followup.log`.
- `rg -n 'mean_js|median_js|max_js|nll_json|mean_nll_json|agree_logodds|logodds\b|keyed_logodds\b|openrouter|OpenRouter|OPENROUTER|openrouter_request|shuffle_dimensions|\bJS\b|Jensen|jensen' src scripts tests README.md pyproject.toml justfile`: PASS, output empty.
- `rg -n 'ROOT / "data" / f"vignettes|ROOT / "data" / "vignettes' scripts src`: PASS, output empty.
- `git branch --list creating-tinymfv`: PASS, branch exists.
+52
View File
@@ -82,6 +82,58 @@ docs/spec/20260625_repo_cleanup_uat.md: 19 prose lines
check count budget status
```
## README follow-up
Command:
```sh
python - <<'PY' 2>&1 | tee /tmp/tinymfv_readme_docs_check.log
...
PY
```
Result excerpt:
```text
required_terms_missing: []
image_links: ['docs/img/showcase/mfq2/map_pca_ipsative.png', 'docs/img/showcase/mfq2/range.png', 'docs/img/showcase/mfv/foundation_dlogit.png']
missing_image_links: []
tracked_image_links: ['docs/img/showcase/mfq2/map_pca_ipsative.png', 'docs/img/showcase/mfq2/range.png', 'docs/img/showcase/mfv/foundation_dlogit.png']
README docs check: PASS
```
Humanizer follow-up command:
```sh
python3 /home/wassname/.agents/skills/humanizer/lint.py README.md docs/spec/20260625_repo_cleanup.md docs/spec/20260625_repo_cleanup_uat.md
```
Result excerpt from `/tmp/tinymfv_humanizer_lint_readme_followup.log`:
```text
README.md: 108 prose lines
check count budget status
tricolon 1 - info
docs/spec/20260625_repo_cleanup.md: 41 prose lines
check count budget status
docs/spec/20260625_repo_cleanup_uat.md: 25 prose lines
check count budget status
```
Fresh-eyes review:
```text
no blockers.
README diff answers the latest ask well: raw answer-token logprobs are the primitive;
pmass_allowed/nll_prefill are defined; sensitive relative readouts are score, lp,
dlogit_per_foundation, and delta C; easier value summaries are profile, profile_E,
and logodds_agree. Artifact sweeps found no tracked/staged __pycache__, .pyc,
data/results, logs, root duplicate data, OpenRouter, multilabel, separation, or
calibration junk.
```
## stale-name sweeps
Command: