mirror of
https://github.com/wassname/moral-maps.git
synced 2026-09-09 11:27:22 +08:00
Rewrite README for pure Authority showcase
This commit is contained in:
@@ -4,35 +4,63 @@ tinymfv is a small set of fast value evals for local LLM steering work. It asks
|
||||
|
||||
Use it when you want to know whether a steer moved the intended values, moved nearby values too, and still lands near real human response patterns. The evals are quick and sensitive enough to show probability shifts before sampled answers flip.
|
||||
|
||||
The plots compare that profile to human data. Gray marks are human societies or respondents, black is the base model, red is positive steering, and blue is negative steering. In this README run, `+c` is oriented by a declared MFV Authority anchor. In other runs, `c` is just the signed vector coefficient unless the run declares its own anchor.
|
||||
The plots compare that profile to human data. Gray marks are human societies or respondents, black is the base model, red is positive steering, and blue is negative steering. This showcase uses a pure Authority steering vector built from `authority-respecting` versus `authority-disregarding` personas. Here `c` is the steering-lite multiplier on that vector. Red is `+c`, more Authority; blue is `-c`, less Authority.
|
||||
|
||||
MFV comes first because it is the direct moral-vignette readout. MFV is nominal: the answer is a moral foundation category. The survey plots are ordinal: the answer is a 1-5 scale point.
|
||||
MFV comes first because it is the direct moral-vignette readout. MFV is nominal: the answer is a moral foundation category. MFQ-2 is the Moral Foundations Questionnaire 2 survey, where the answer is a 1-5 scale point.
|
||||
|
||||

|
||||

|
||||
|
||||

|
||||

|
||||
|
||||
MFV uses the same map and range plotters as the surveys, after converting nominal foundation logits into relative foundation emphasis. Each profile is z-scored across foundations before mapping, so the plot compares which foundations are high or low within that profile.
|
||||
|
||||

|
||||

|
||||
|
||||

|
||||

|
||||
|
||||
The Humor Styles map shows the same failure mode more sharply: the model profile can live away from the human societies. That is the useful warning sign, a model can be format-coherent and still be a moral or psychological alien on the measured profile.
|
||||
|
||||

|
||||

|
||||
|
||||

|
||||

|
||||
|
||||
Read the Big Five map left to right: gray is the human reference, black is the base LLM, and the red/blue line is the coherent steering path. Here the LLM sits outside the country cloud, so on this measure it is a psychological alien before steering moves it.
|
||||
|
||||
MFQ-2 means Moral Foundations Questionnaire 2, the short survey instrument. It is separate from MFV, the moral-vignette foundation reader. MFQ-2 has fewer items per axis than the longer personality surveys, so the showcase uses sampled think traces before treating small path wiggles as signal.
|
||||
MFQ-2 means Moral Foundations Questionnaire 2, the short survey instrument. It is separate from MFV, the moral-vignette foundation reader. MFQ-2 has fewer items per axis than the longer personality surveys, so the showcase averages 8 sampled reads per item before treating small path wiggles as signal.
|
||||
|
||||

|
||||

|
||||
|
||||

|
||||

|
||||
|
||||
The path shows only usable coefficients: `c=0`, then each positive and negative side until the reader starts to collapse. For surveys, collapse can mean the answer distribution loses its factor structure even when answer mass stays high.
|
||||
The path shows only usable coefficients: `c=0`, then each positive and negative side until one of the plot gates fails. This run kept the full path `c=-1,-0.5,0,+0.5,+1`. For surveys, collapse can mean the answer distribution loses its factor structure even when answer mass stays high.
|
||||
|
||||
This is the same run as the plots. The steer is pure in the contrast used to build it; the table shows the side effects it actually caused.
|
||||
|
||||
`profile shift / human SD` is the distance from `c=-1` to `c=+1`, divided by the standard deviation of country means in the bundled human reference for that axis. `100%` means one human SD. `profile shift` is in plot units: MFV relative-emphasis z-score, or survey expected 1-5 score. `reader-logit shift` is the direct answer-logit readout for the same endpoints: MFV uses foundation `dlogit`, surveys use the rank-logit contrast `C`. The `+/-` term is the propagated item uncertainty.
|
||||
|
||||
| dataset | axis | profile shift / human SD | profile shift | reader-logit shift |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| MFV vignettes | Care | -242% | -0.59 | -0.07 +/- 0.28 |
|
||||
| MFV vignettes | Sanctity | -9% | -0.05 | +0.53 +/- 0.21 |
|
||||
| MFV vignettes | Authority | +401% | +1.17 | +1.82 +/- 0.32 |
|
||||
| MFV vignettes | Loyalty | -18% | -0.05 | +0.55 +/- 0.18 |
|
||||
| MFV vignettes | Fairness | -5% | -0.02 | +0.56 +/- 0.20 |
|
||||
| MFV vignettes | Liberty | -75% | -0.46 | +0.11 +/- 0.20 |
|
||||
| Humor Styles | affiliative | +144% | +0.40 | +15.00 +/- 3.98 |
|
||||
| Humor Styles | selfenhancing | +162% | +0.25 | +13.52 +/- 4.86 |
|
||||
| Humor Styles | aggressive | -23% | -0.04 | +0.08 +/- 7.58 |
|
||||
| Humor Styles | selfdefeating | +32% | +0.06 | +13.62 +/- 6.55 |
|
||||
| Big Five | extraversion | +105% | +0.17 | +7.21 +/- 4.79 |
|
||||
| Big Five | neuroticism | +2% | +0.00 | +7.69 +/- 3.56 |
|
||||
| Big Five | agreeableness | +269% | +0.34 | +16.03 +/- 4.17 |
|
||||
| Big Five | conscientiousness | +435% | +0.45 | +17.21 +/- 4.03 |
|
||||
| Big Five | openness | +41% | +0.07 | +17.77 +/- 2.48 |
|
||||
| MFQ-2 survey | care | +305% | +0.93 | +23.20 +/- 6.08 |
|
||||
| MFQ-2 survey | equality | +21% | +0.07 | +16.06 +/- 7.89 |
|
||||
| MFQ-2 survey | proportionality | +270% | +0.77 | +24.48 +/- 5.87 |
|
||||
| MFQ-2 survey | loyalty | +202% | +0.82 | +25.88 +/- 6.33 |
|
||||
| MFQ-2 survey | authority | +338% | +1.17 | +29.73 +/- 2.83 |
|
||||
| MFQ-2 survey | purity | +137% | +0.74 | +23.53 +/- 8.02 |
|
||||
|
||||
## Install
|
||||
|
||||
@@ -110,14 +138,16 @@ Generate the bundled range plots and culture maps from a steering-lite all-instr
|
||||
|
||||
```bash
|
||||
uv run python scripts/plot_steer_showcase.py \
|
||||
--run-dir ../steering-lite/outputs/20260630_dignity_authority_strict22_local_sspace_mfvgrid_n8 \
|
||||
--run-dir ../steering-lite/outputs/20260630T222000Z_pure_authority_mundane15_pca_readme_mfv_mfq2_humor_big5_n8 \
|
||||
--out docs/img/showcase \
|
||||
--vec-label="MFV Authority anchor (+c intended higher Authority)" \
|
||||
--vec-label "pure Authority, PCA (+c = more Authority)" \
|
||||
--coherence-frac 0.99 \
|
||||
--contrast-frac 0.50 \
|
||||
--contrast-frac 0.000001 \
|
||||
--margin-frac 0.50
|
||||
```
|
||||
|
||||
The plot gate keeps shared coefficients whose answer mass is at least `coherence-frac` of base, whose survey contrast remains above `contrast-frac`, and whose MFV forced-choice margin remains above `margin-frac`.
|
||||
|
||||
## Measurement
|
||||
|
||||
The measurement on the maps is the profile.
|
||||
@@ -134,6 +164,20 @@ where $i$ is an item, $d$ is a survey factor, $k$ is a scale point, and $M$ is t
|
||||
|
||||
This is what the survey maps and range plots show. In the showcase CSVs, this is the `mean` column. For MFV showcase plots, model and human units differ, so the plotted quantity is relative foundation emphasis: each foundation profile is z-scored across foundations before mapping.
|
||||
|
||||
The table's reader-logit shift uses a more sensitive log-space readout.
|
||||
|
||||
For MFV:
|
||||
|
||||
$$\Delta_f = \mathbb{E}_i \left(\ell_{i,f}^{(+1)} - \ell_{i,f}^{(-1)}\right)$$
|
||||
|
||||
where $\ell_{i,f}^{(c)}$ is the forced-choice logit for foundation $f$ on item $i$ at coefficient $c$.
|
||||
|
||||
For survey instruments:
|
||||
|
||||
$$C_d(c) = \mathbb{E}_{i \in d}\sum_{k=1}^{M} \left(k - \frac{M+1}{2}\right)\ell_{i,k}^{(c)}$$
|
||||
|
||||
and the table reports $C_d(+1)-C_d(-1)$.
|
||||
|
||||
For paired steering runs, compare the base profile to the steered profile path. Answer mass is a coherence check, not a value score:
|
||||
|
||||
$$m(c) = \mathbb{E}_i \sum_{a \in A_i} P_c(a \mid i)$$
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"summary": "tinymfv is a lightweight package that evaluates LLM steering by feeding moral vignettes and survey questions, reading answer-token probabilities, and converting them into a multi-dimensional profile comparable to human response data. A researcher would use it to quickly check whether a steering vector shifts intended values, causes unintended side effects on nearby traits, and keeps the model within human-like response patterns. It is designed for fast, paired steering comparisons where probability shifts matter before sampled answers flip.",
|
||||
"mechanism": "The evaluation runs a set of items (vignettes or survey questions) through the model at multiple steering coefficients (c = -1, -0.5, 0, +0.5, +1). For each item, the model's next-token probabilities over valid answer tokens are recorded. For MFV (nominal), the profile is the mean probability per moral foundation (averaged over items). For surveys (ordinal), the profile is the mean expected 1-5 score per factor (sum of scale-point probabilities weighted by the point value). Profiles are then plotted against human reference data: gray marks for human societies, black for base model (c=0), red for positive steer (+c), blue for negative steer (-c). The 'reader-logit shift' table reports a more sensitive log-space difference between c=+1 and c=-1: for MFV it is the mean per-item difference in forced-choice logit for each foundation; for surveys it is the mean per-item rank-logit contrast difference. Gate thresholds (coherence-frac, contrast-frac, margin-frac) filter coefficients where answer mass, contrast, or forced-choice margin fall too low.",
|
||||
"scores": {"clarity": 4, "conciseness": 5, "technical_accuracy": 4},
|
||||
"reason": "The document is dense and technically precise, but the gate threshold parameters and the exact meaning of 'coherent steering path' could be explained more clearly.",
|
||||
"unclear": [
|
||||
"What the 'c path' refers to exactly beyond the set of coefficients used (e.g., is it always -1,-0.5,0,+0.5,+1?); the document mentions a 'full path' but does not state a default.",
|
||||
"The gate threshold parameters (coherence-frac, contrast-frac, margin-frac) are mentioned in the plot script but their effects on which coefficients are shown are not fully explained.",
|
||||
"The phrase 'the reader starts to collapse' is ambiguous: collapse of what, and how is it detected?",
|
||||
"In the table, the 'reader-logit shift' for surveys uses a formula with C_d(c) but the document does not explain what the rank-logit contrast 'C' is in plain language.",
|
||||
"The notion of 'profile shift / human SD' is defined as distance from c=-1 to c=+1 in human between-country SDs, but how the human SD is computed (e.g., from which countries) is not stated."
|
||||
],
|
||||
"misunderstandings": [
|
||||
"'MFV uses the same map and range plotters as the surveys' could mislead readers to think the plotted quantity is identical, when MFV plots z-scored relative emphasis whereas surveys plot expected scores directly.",
|
||||
"The term 'reader-logit shift' may be misinterpreted as the shift in the model's own logits (e.g., hidden states) rather than the logits of the answer distribution.",
|
||||
"In the statement 'the model profile can live away from the human societies', 'live' might be read as 'is generated by' rather than 'is plotted far from'.",
|
||||
"The phrase 'the table shows the side effects it actually caused' might imply causation, but the table merely shows correlations from a single steer run.",
|
||||
"'collapse can mean the answer distribution loses its factor structure even when answer mass stays high' – a reader might think 'factor structure' refers to the model's internal factors rather than the survey's expected factor structure."
|
||||
],
|
||||
"missing_to_implement": [
|
||||
"How to create steering vectors and apply them to the model is not covered (the API assumes steer coefficients are handled externally).",
|
||||
"The exact structure of the run-dir required by the plot script (e.g., CSV file names and columns) is not documented.",
|
||||
"How to obtain the human reference data files (country means, raw respondents) and where they are bundled in the package is mentioned but not demonstrated in a usage example.",
|
||||
"The meaning of 'pareto' or 'coherence' in the gate parameters is not defined."
|
||||
],
|
||||
"suggestions": [
|
||||
"Add a short section or inline explanation of the gate thresholds: e.g., 'coherence-frac filters coefficients where answer mass drops below this fraction of the base model's mass; contrast-frac filters coefficients where the survey rank-logit contrast falls below this threshold; margin-frac filters coefficients where the MFV forced-choice margin falls below this fraction of the base margin.'",
|
||||
"Clarify what 'c path' means: 'The steering coefficients used (c values) are passed as a list; the default path is c=-1,-0.5,0,+0.5,+1 unless otherwise specified.'",
|
||||
"Elaborate on 'reader-logit shift': rename to 'answer-logit shift' or explicitly state it is the change in the logits of the model's answer distribution between extreme coefficients.",
|
||||
"Provide a minimal working example of the plot script with a dummy run-dir or link to a sample output directory.",
|
||||
"Add a note that the human SD used for normalization is the standard deviation of country means (or respondent means) across the human reference sample, and specify which countries are included."
|
||||
]
|
||||
}
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"summary": "tinymfv is a toolkit for fast evaluation of local LLM steering, using moral vignettes and standardized surveys to measure model value shifts. Researchers use it to verify whether a steering vector effectively moves a model toward human response patterns or inadvertently shifts it into out-of-distribution, 'alien' territory without requiring full moral reasoning evaluations.",
|
||||
"mechanism": "The package processes model outputs into 'profiles' by calculating mean foundation probabilities for MFV vignettes (z-scored for mapping) or expected 1-5 scores for surveys. These profiles are compared against human data across a steering coefficient path. The tool calculates 'reader-logit shift' as a sensitive measure of log-space changes at steering endpoints (`c=+1` vs `c=-1`) and applies 'gates' (coherence, contrast, and margin checks) to discard unstable runs where the model loses its semantic structure.",
|
||||
"scores": {"clarity": 5, "conciseness": 5, "technical_accuracy": 5},
|
||||
"reason": "The documentation is highly specific, providing both the conceptual workflow and the necessary mathematical formulas to understand how the profiles are derived and compared.",
|
||||
"unclear": ["The specific mathematical derivation of the survey 'rank-logit contrast C' is provided, but 'reader-logit shift' calculation details rely on the reader interpreting the formula $\\ell_{i,f}^{(c)}$ and the table columns without a definitions glossary for all abbreviations."],
|
||||
"misunderstandings": ["The term 'MFV' is used ambiguously, acting as both an acronym for the package name (tinymfv), the moral vignette dataset, and the foundation-reader framework."],
|
||||
"missing_to_implement": ["The documentation demonstrates how to run evaluations (evaluate/administer), but implies the steering vector (e.g., 'pure Authority') exists externally, without explaining how to construct that vector from the mentioned personas."]
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
```json
|
||||
{
|
||||
"summary": "tinymfv is a lightweight tool designed for rapidly evaluating the effects of steering on language models' moral and value judgments. It quickly assesses how steering influences a model's responses to moral vignettes and survey questions, providing insights into whether changes are targeted, have unintended side effects, or align with human responses.",
|
||||
"mechanism": "The evaluation process involves feeding moral vignettes or survey questions to the LLM and capturing the probabilities assigned to different answer tokens. tinymfv then converts these probabilities into a profile, which is compared against human data using plots like culture maps and range plots. The details of how precision to linearize affinities and gates remained unclear from the documentation.",
|
||||
"scores": {
|
||||
"clarity": "3",
|
||||
"conciseness": "4",
|
||||
"technical_accuracy": "3"
|
||||
},
|
||||
"reason": "The documentation is reasonably concise and provides a general overview of the tool's purpose, but it lacks detailed technical explanations of the evaluation mechanisms, making a full reconstruction difficult.",
|
||||
"unclear": [
|
||||
"How the probabilities are converted into the plotted profile and steering comparison.",
|
||||
"What the 'c path/gates' are and their role in the measurement.",
|
||||
"The specifics of 'reader-logit shift' calculation",
|
||||
"The exact method for 'z-scoring' the profiles"
|
||||
],
|
||||
"misunderstandings": [
|
||||
"The document implies a direct one-to-one correspondence between MFV items, LLM answers and human foundations, where in reality these can be more complex.",
|
||||
"There's ambiguity around the precise treatment of responses for different survey frames (inverted/negated)."
|
||||
],
|
||||
"missing_to_implement": [
|
||||
"Detailed instructions on how to reproduce the plots.",
|
||||
"A more in-depth explanation of the 'profile shift' and 'reader-logit shift' metrics."
|
||||
],
|
||||
"suggestions": [
|
||||
"Add a detailed walkthrough with example code showing how to generate the plots from the 'report' object.",
|
||||
"Elaborate on the mathematical formulas behind the different measurement metrics (profile shift, reader-logit shift).",
|
||||
"Clarify the treatment of survey frame variations for a more consistent reading"
|
||||
]
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,37 @@
|
||||
{
|
||||
"summary": "tinymfv is a lightweight Python library that evaluates the moral, personality, and humor value profiles of locally run LLMs by presenting vignette and survey items, reading answer-token probabilities, and summarizing them into a single model profile. It visualizes how positive or negative steering moves these profiles relative to human reference data using PCA maps and range plots. Researchers use it to quickly detect subtle steering effects and side‑effects before they appear in sampled model outputs.",
|
||||
"mechanism": "For each vignette (MFV) or survey item the model is queried and the probability mass on each valid answer token is extracted. MFV computes a foundation profile as the average probability for each moral foundation (profile_f = E_i P(f|i)). Survey instruments compute an expected rating per factor (profile_d = E_i ∑_k k·P(k|i)) after reverse‑keying. MFV profiles are z‑scored across foundations and projected onto a PCA derived from human country data to produce 2‑D ipsative maps; range plots show the raw or z‑scored values against human means and standard deviations. Steering is evaluated by running the model at several steering coefficients c (e.g., -1, -0.5, 0, +0.5, +1), calculating profile shift (difference in profile values) and reader‑logit shift (Δ_f = E_i (ℓ_{i,f}^{(+1)} – ℓ_{i,f}^{(-1)}) for MFV and analogous C_d(c) for surveys). Coefficients are kept only if the answer‑mass coherence (m(c) ≥ coherence‑frac), survey contrast (≥ contrast‑frac), and MFV forced‑choice margin (≥ margin‑frac) pass threshold, implementing the “c path/gates”. The resulting shifts are plotted against human‑SD units.",
|
||||
"scores": {
|
||||
"clarity": 4,
|
||||
"conciseness": 3,
|
||||
"technical_accuracy": 4
|
||||
},
|
||||
"reason": "The documentation provides clear definitions and formulas for profiles and shifts, but some gating details, the steering coefficient implementation, and PCA construction are not fully explained, reducing overall precision and brevity.",
|
||||
"unclear": [
|
||||
"Exact definition and calculation of the MFV \"forced‑choice margin\" used for gating.",
|
||||
"How the steering coefficient c is applied to the model (e.g., prompt modification, logit bias, LoRA).",
|
||||
"Details of the PCA computation and how human reference data are pre‑processed for the maps.",
|
||||
"How answer‑token sets A_i are constructed for each vignette and survey item (mapping tokens to categories or scale points).",
|
||||
"Precise interpretation and units of \"reader‑logit shift\" across MFV and surveys."
|
||||
],
|
||||
"misunderstandings": [
|
||||
"\"Reader‑logit shift\" may be misread as a uniform metric across MFV and surveys, though its underlying logit definitions differ (foundation dlogit vs rank‑logit contrast C).",
|
||||
"The difference between \"profile shift / human SD\" and \"profile shift\" units could cause confusion about scaling."
|
||||
],
|
||||
"missing_to_implement": [
|
||||
"A description or code for generating the PCA axes from human data and projecting model profiles onto them.",
|
||||
"Mapping from token probabilities to moral foundation categories for MFV items.",
|
||||
"Implementation details for applying the steering coefficient c to the model and producing outputs at multiple c values.",
|
||||
"Formula and explanation for the MFV forced‑choice margin metric used in the gate.",
|
||||
"Procedure for building the answer‑token sets A_i for each vignette and survey item.",
|
||||
"Exact calculation of the survey contrast metric used for gating."
|
||||
],
|
||||
"suggestions": [
|
||||
"Add a clear explanation of how the steering coefficient c is realized (e.g., via prompt templates, logit bias, or adapter weights) and how to generate multiple c values.",
|
||||
"Provide the definition and computation steps for the MFV forced‑choice margin and its threshold (margin‑frac).",
|
||||
"Include a brief outline of the PCA pipeline: data sources, standardization, number of components, and how human IPSATIVE maps are derived.",
|
||||
"Supply a mapping table or utility function that translates model token IDs to answer options for each MFV and survey item.",
|
||||
"Clarify the interpretation of \"reader‑logit shift\" for both MFV and surveys, perhaps with a small numeric example.",
|
||||
"Show example code demonstrating the gating checks (coherence‑frac, contrast‑frac, margin‑frac) and how they prune the c path."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"summary": "tinymfv is a lightweight, fast evaluation toolkit that prompts local LLMs with moral vignettes and psychological survey items, reads the probabilities on valid answer tokens, and aggregates those probabilities into a single model profile. Researchers would use it to quickly test whether a steering intervention moves the target value, shifts nearby values, and stays close to human response patterns.",
|
||||
"mechanism": "For each installed dataset (MFV vignettes or a survey instrument), tinymfv builds items, asks them in multiple framings (other/self violate for MFV; forward/inverted/negated for surveys), and reads the model's logits over the valid answer tokens. It canonicalizes the reverse-keyed framings and averages them, then computes a profile: for MFV this is the mean softmax probability per foundation, and for surveys it is the mean expected 1-5 scale score per factor. For the showcase, MFV profiles are z-scored across foundations so model and human reference share relative-emphasis units, while survey maps use the raw expected scores. Steering is applied along a coefficient path (c), here pure Authority built from authority-respecting vs authority-disregarding personas (red = +c = more Authority, blue = -c = less Authority). The plotted path keeps only usable coefficients: those whose answer mass is at least `coherence-frac` of base, whose survey contrast is above `contrast-frac`, and whose MFV forced-choice margin is above `margin-frac`. The final table reports the profile shift from c=-1 to c=+1 in both z-score/expected-score units and in human between-country SDs, plus a more sensitive `reader-logit shift` computed from log-space contrasts.",
|
||||
"scores": {
|
||||
"clarity": "4",
|
||||
"conciseness": "4",
|
||||
"technical_accuracy": "5"
|
||||
},
|
||||
"reason": "The document explains the package's purpose, the showcased Authority steer, and the profile/logit metrics with consistent equations and plots, though a few gate/collapse details and map-axis construction are left implicit.",
|
||||
"unclear": [
|
||||
"The exact formulas for the plot-gate metrics 'MFV forced-choice margin' and 'survey contrast' are not defined.",
|
||||
"What precise operational test counts as 'reader collapse' is only described qualitatively.",
|
||||
"How the 'human between-country SD' used for scaled profile shift is computed is not shown.",
|
||||
"Whether the map PCA axes are computed once across datasets or per-plot is not stated.",
|
||||
"How the steering vector is mechanically applied to the model (layers, scale, hook) is missing."
|
||||
],
|
||||
"misunderstandings": [
|
||||
"The table's 'profile shift' for MFV is in z-scored relative-emphasis units, but a reader might take it as raw probability shift because the API reports `mean probability per foundation`.",
|
||||
"The phrase 'MFV uses the same map and range plotters as the surveys' could imply identical units; the text later clarifies MFV is z-scored, but this is easy to skim past.",
|
||||
"Saying both 'the path shows only usable coefficients' and 'this run kept the full path' is consistent if no collapse occurred, but could be read as saying the path is always fixed."
|
||||
],
|
||||
"missing_to_implement": [
|
||||
"A steering-lite run directory containing the actual generated model outputs to feed into `scripts/plot_steer_showcase.py`.",
|
||||
"A causal LM, GPU, and the `tiny-mfv[maps]` extra dependency.",
|
||||
"The bundled JSON/CSV data files and the code that converts them into prompts and valid answer-token sets.",
|
||||
"The PCA and human-reference preprocessing pipeline.",
|
||||
"The concrete implementation of the three gate thresholds used to truncate the coefficient path."
|
||||
],
|
||||
"suggestions": [
|
||||
"Add a short glossary sidebar defining c path, collapse, forced-choice margin, and survey contrast.",
|
||||
"Move the reader-logit shift equations next to the table where the term is first used.",
|
||||
"Restate the table header units explicitly, e.g. 'profile shift (MFV z-score or survey expected 1-5 score)'.",
|
||||
"Include one sentence or a link explaining how the Authority vector is applied to activations.",
|
||||
"Add a map caption that defines the PCA axes and the gray/black/red/blue legend."
|
||||
]
|
||||
}
|
||||
+35
@@ -0,0 +1,35 @@
|
||||
{
|
||||
"summary": "tinymfv is a lightweight package for evaluating how steering changes an LLM's expressed values. It poses moral vignettes and survey questions, reads the model's answer-token probabilities, collapses them into per‑foundation or per‑trait profiles, and visualizes those profiles alongside human baselines to see whether a steer moves the intended values, affects nearby values, and remains within human response patterns.",
|
||||
"mechanism": "For each item, the model’s answer-token probabilities are collected; MFV vignettes yield a profile as the mean probability of each moral foundation across items, while surveys give the mean expected 1‑5 score per factor. MFV profiles are z‑scored across foundations before plotting so the map shows relative emphasis. Steering effects are assessed by comparing profiles (or logits) at positive and negative steering coefficients (+c, –c); the table’s reader‑logit shift reports the average difference in forced‑choice logits (MFV) or rank‑weighted logits (surveys) between +c and –c, and profile shift / human SD expresses that difference in units of between‑country standard deviation.",
|
||||
"scores": {
|
||||
"clarity": "4",
|
||||
"conciseness": "3",
|
||||
"technical_accuracy": "5"
|
||||
},
|
||||
"reason": "The document is reasonably clear and technically precise, though some sections are dense and could be more concise.",
|
||||
"unclear": [
|
||||
"what the steering coefficient 'c' concretely represents and how it is applied to the model",
|
||||
"exact computation of 'profile shift / human SD' (how the human between‑country SD is derived)",
|
||||
"meaning of the plot gate parameters coherence‑frac, contrast‑frac, margin‑frac",
|
||||
"definition of 'forced‑choice margin' used in the gate",
|
||||
"how the 'reader-logit shift' is propagated to item uncertainty (± term)"
|
||||
],
|
||||
"misunderstandings": [
|
||||
"the text says MFV uses 'nominal foundation logits' after converting to relative foundation emphasis, yet earlier defines the MFV profile as mean probability, creating potential confusion about whether probabilities or logits are plotted",
|
||||
"the description of the path states it shows only usable coefficients until the reader starts to collapse, but the showcase then keeps the full path c=-1,-0.5,0,+0.5,+1, which may suggest collapse did not occur"
|
||||
],
|
||||
"missing_to_implement": [
|
||||
"instructions for generating or obtaining the steering vectors (e.g., pure Authority vector) and applying coefficients c",
|
||||
"details on how to produce the steering‑lite run directory referenced in the plotting script",
|
||||
"explanation of any additional dependencies beyond transformers (e.g., steering‑lite, plotting utilities)",
|
||||
"guidance on how to compute the z‑scoring of MFV foundations for custom plots",
|
||||
"clarification on how to interpret the 'mean_pmass_allowed' format check"
|
||||
],
|
||||
"suggestions": [
|
||||
"Add a brief introductory paragraph explaining what the steering coefficient c is (e.g., strength along a pre‑computed steering vector) and how ±c are obtained",
|
||||
"Include a worked example showing the calculation of reader‑logit shift for one MFV foundation and one survey factor",
|
||||
"Provide a tooltip or legend in the documentation defining coherence‑frac, contrast‑frac, and margin‑frac in the plot gate",
|
||||
"Clarify whether the MFV profile plotted is based on probabilities or logits, and show the conversion step explicitly",
|
||||
"Add a minimal end‑to‑end tutorial that starts from a raw model, applies steering, runs tinymfv evaluate/administer, and generates the showcase plots"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,220 @@
|
||||
You are reading the document below for the FIRST time (a cold reader).
|
||||
Answer ONLY from what it says; where something is unstated or ambiguous, say so.
|
||||
Pay special attention to whether you can explain: what tinymfv is, what the maps measure, what MFV and MFQ-2 are, which steer is shown, what reader-logit shift means, what the c path/gates are, and why a researcher would use this package.
|
||||
Output ONE JSON object, no prose, no fences:
|
||||
{
|
||||
"summary": "<2-3 sentences: restate what the repo does and why a researcher would use it, in your OWN words>",
|
||||
"mechanism": "<reconstruct how the eval turns LLM answer probabilities into the plotted profile and steering comparison; say 'unclear' if the doc does not let you rebuild it>",
|
||||
"scores": {"clarity": "<1-5>", "conciseness": "<1-5>", "technical_accuracy": "<1-5>"},
|
||||
"reason": "<one sentence on the scores>",
|
||||
"unclear": ["<what was confusing, ambiguous, or you had to guess>"],
|
||||
"misunderstandings": ["<places the text contradicts itself or invites a misread>"],
|
||||
"missing_to_implement": ["<what a reader still needs to reproduce or act on this>"],
|
||||
"suggestions": ["<concrete edit that would help>"]
|
||||
}
|
||||
|
||||
DOCUMENT:
|
||||
# tinymfv
|
||||
|
||||
tinymfv is a small set of fast value evals for local LLM steering work. It asks moral vignettes and survey questions, reads answer-token probabilities, and turns them into one model profile.
|
||||
|
||||
Use it when you want to know whether a steer moved the intended values, moved nearby values too, and still lands near real human response patterns. The evals are quick and sensitive enough to show probability shifts before sampled answers flip.
|
||||
|
||||
The plots compare that profile to human data. Gray marks are human societies or respondents, black is the base model, red is positive steering, and blue is negative steering. This showcase uses a pure Authority steering vector built from `authority-respecting` versus `authority-disregarding` personas. Red is `+c`, more Authority; blue is `-c`, less Authority.
|
||||
|
||||
MFV comes first because it is the direct moral-vignette readout. MFV is nominal: the answer is a moral foundation category. MFQ-2 is the Moral Foundations Questionnaire 2 survey, where the answer is a 1-5 scale point.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
MFV uses the same map and range plotters as the surveys, after converting nominal foundation logits into relative foundation emphasis. Each profile is z-scored across foundations before mapping, so the plot compares which foundations are high or low within that profile.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
The Humor Styles map shows the same failure mode more sharply: the model profile can live away from the human societies. That is the useful warning sign, a model can be format-coherent and still be a moral or psychological alien on the measured profile.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
Read the Big Five map left to right: gray is the human reference, black is the base LLM, and the red/blue line is the coherent steering path. Here the LLM sits outside the country cloud, so on this measure it is a psychological alien before steering moves it.
|
||||
|
||||
MFQ-2 means Moral Foundations Questionnaire 2, the short survey instrument. It is separate from MFV, the moral-vignette foundation reader. MFQ-2 has fewer items per axis than the longer personality surveys, so the showcase averages 8 sampled reads per item before treating small path wiggles as signal.
|
||||
|
||||

|
||||
|
||||

|
||||
|
||||
The path shows only usable coefficients: `c=0`, then each positive and negative side until the reader starts to collapse. This run kept the full path `c=-1,-0.5,0,+0.5,+1`. For surveys, collapse can mean the answer distribution loses its factor structure even when answer mass stays high.
|
||||
|
||||
This is the same run as the plots. The steer is pure in the contrast used to build it; the table shows the side effects it actually caused.
|
||||
|
||||
`profile shift / human SD` is the distance from `c=-1` to `c=+1`, measured in human between-country standard deviations. `100%` means one human SD. `profile shift` is in plot units: MFV relative-emphasis z-score, or survey expected 1-5 score. `reader-logit shift` is the direct answer-logit readout for the same endpoints: MFV uses foundation `dlogit`, surveys use the rank-logit contrast `C`. The `+/-` term is the propagated item uncertainty.
|
||||
|
||||
| dataset | axis | profile shift / human SD | profile shift | reader-logit shift |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| MFV vignettes | Care | -242% | -0.59 | -0.07 +/- 0.28 |
|
||||
| MFV vignettes | Sanctity | -9% | -0.05 | +0.53 +/- 0.21 |
|
||||
| MFV vignettes | Authority | +401% | +1.17 | +1.82 +/- 0.32 |
|
||||
| MFV vignettes | Loyalty | -18% | -0.05 | +0.55 +/- 0.18 |
|
||||
| MFV vignettes | Fairness | -5% | -0.02 | +0.56 +/- 0.20 |
|
||||
| MFV vignettes | Liberty | -75% | -0.46 | +0.11 +/- 0.20 |
|
||||
| Humor Styles | affiliative | +144% | +0.40 | +15.00 +/- 3.98 |
|
||||
| Humor Styles | selfenhancing | +162% | +0.25 | +13.52 +/- 4.86 |
|
||||
| Humor Styles | aggressive | -23% | -0.04 | +0.08 +/- 7.58 |
|
||||
| Humor Styles | selfdefeating | +32% | +0.06 | +13.62 +/- 6.55 |
|
||||
| Big Five | extraversion | +105% | +0.17 | +7.21 +/- 4.79 |
|
||||
| Big Five | neuroticism | +2% | +0.00 | +7.69 +/- 3.56 |
|
||||
| Big Five | agreeableness | +269% | +0.34 | +16.03 +/- 4.17 |
|
||||
| Big Five | conscientiousness | +435% | +0.45 | +17.21 +/- 4.03 |
|
||||
| Big Five | openness | +41% | +0.07 | +17.77 +/- 2.48 |
|
||||
| MFQ-2 survey | care | +305% | +0.93 | +23.20 +/- 6.08 |
|
||||
| MFQ-2 survey | equality | +21% | +0.07 | +16.06 +/- 7.89 |
|
||||
| MFQ-2 survey | proportionality | +270% | +0.77 | +24.48 +/- 5.87 |
|
||||
| MFQ-2 survey | loyalty | +202% | +0.82 | +25.88 +/- 6.33 |
|
||||
| MFQ-2 survey | authority | +338% | +1.17 | +29.73 +/- 2.83 |
|
||||
| MFQ-2 survey | purity | +137% | +0.74 | +23.53 +/- 8.02 |
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
uv pip install git+https://github.com/wassname/tinymfv
|
||||
```
|
||||
|
||||
For maps:
|
||||
|
||||
```bash
|
||||
uv pip install "tiny-mfv[maps] @ git+https://github.com/wassname/tinymfv"
|
||||
```
|
||||
|
||||
For repo development:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/wassname/tinymfv
|
||||
cd tinymfv
|
||||
uv sync --extra maps --dev
|
||||
just smoke
|
||||
```
|
||||
|
||||
## Datasets
|
||||
|
||||
| dataset | bundled data | human reference | profile used in plots |
|
||||
|---|---|---|---|
|
||||
| MFV classic | [132 moral vignettes, other](src/tinymfv/data/vignettes_classic_other_violate.jsonl) / [self](src/tinymfv/data/vignettes_classic_self_violate.jsonl) | per-vignette human foundation labels in the JSONL | foundation probability profile |
|
||||
| MFV scifi | [same items rewritten as sci-fi, other](src/tinymfv/data/vignettes_scifi_other_violate.jsonl) / [self](src/tinymfv/data/vignettes_scifi_self_violate.jsonl) | inherited labels from classic MFV | foundation probability profile |
|
||||
| MFV ai-actor | [same items rewritten with an AI actor, other](src/tinymfv/data/vignettes_ai-actor_other_violate.jsonl) / [self](src/tinymfv/data/vignettes_ai-actor_self_violate.jsonl) | inherited labels from classic MFV | foundation probability profile |
|
||||
| MFQ-2 | [36 items](src/tinymfv/data/surveys/mfq2/forward.json), plus inverted and negated frames | [country means](src/tinymfv/data/human/mfq2_country_foundations.csv), plus [raw respondents](src/tinymfv/data/atari_study2_raw.csv) | expected 1-5 score per foundation |
|
||||
| Big Five | [50 items](src/tinymfv/data/surveys/big5/questionnaire.json), plus inverted and negated frames | [country means](src/tinymfv/data/human/big5_country_factors.csv) | expected 1-5 score per trait |
|
||||
| 16PF | [162 items](src/tinymfv/data/surveys/16pf/questionnaire.json), plus inverted and negated frames | [country means](src/tinymfv/data/human/16pf_country_factors.csv) | expected 1-5 score per factor |
|
||||
| Humor Styles | [32 items](src/tinymfv/data/surveys/humor_styles/questionnaire.json), plus inverted and negated frames | [country means](src/tinymfv/data/human/humor_styles_country_factors.csv), originally 1-7 | expected 1-5 score per style |
|
||||
|
||||
MFV is nominal: the answer is the category. The survey instruments are ordinal: the answer is a scale point.
|
||||
|
||||
Each MFV item is asked in two perspectives, `other_violate` and `self_violate`. Each survey item is asked three ways, forward, scale-inverted, and content-negated. tinymfv canonicalizes these frames before averaging, so the profile is less tied to one wording.
|
||||
|
||||
## API
|
||||
|
||||
Run MFV vignettes with `evaluate`:
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
from tinymfv import evaluate, load_vignettes
|
||||
|
||||
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
|
||||
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda()
|
||||
|
||||
vignettes = load_vignettes("classic") # "classic", "scifi", "ai-actor", or "all"
|
||||
report = evaluate(model, tok, vignettes=vignettes)
|
||||
|
||||
print(report["profile"]) # mean probability per foundation
|
||||
print(report["mean_pmass_allowed"]) # format check: mass on valid answer tokens
|
||||
```
|
||||
|
||||
Run survey instruments with `administer`:
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
from tinymfv import administer, get_instrument
|
||||
|
||||
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
|
||||
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B").cuda()
|
||||
|
||||
instr = get_instrument("mfq2") # "mfq2", "big5", "16pf", or "humor_styles"
|
||||
report = administer(model, tok, instr)
|
||||
|
||||
print(report["dimensions"])
|
||||
print(report["profile"]) # expected 1-5 score per factor
|
||||
print(report["mean_pmass_allowed"]) # format check: mass on valid answer tokens
|
||||
```
|
||||
|
||||
Generate the bundled range plots and culture maps from a steering-lite all-instrument run:
|
||||
|
||||
```bash
|
||||
uv run python scripts/plot_steer_showcase.py \
|
||||
--run-dir ../steering-lite/outputs/20260630T222000Z_pure_authority_mundane15_pca_readme_mfv_mfq2_humor_big5_n8 \
|
||||
--out docs/img/showcase \
|
||||
--vec-label "pure Authority, PCA (+c = more Authority)" \
|
||||
--coherence-frac 0.99 \
|
||||
--contrast-frac 0.000001 \
|
||||
--margin-frac 0.50
|
||||
```
|
||||
|
||||
The plot gate keeps shared coefficients whose answer mass is at least `coherence-frac` of base, whose survey contrast remains above `contrast-frac`, and whose MFV forced-choice margin remains above `margin-frac`.
|
||||
|
||||
## Measurement
|
||||
|
||||
The measurement on the maps is the profile.
|
||||
|
||||
For MFV, the profile is the model's mean probability on each moral foundation:
|
||||
|
||||
$$\mathrm{profile}_f = \mathbb{E}_i P(f \mid i)$$
|
||||
|
||||
For survey instruments, the profile is the mean expected 1-5 answer for each factor, after reverse-keying:
|
||||
|
||||
$$\mathrm{profile}_d = \mathbb{E}_{i \in d}\sum_{k=1}^{M} k P(k \mid i)$$
|
||||
|
||||
where $i$ is an item, $d$ is a survey factor, $k$ is a scale point, and $M$ is the largest scale value.
|
||||
|
||||
This is what the survey maps and range plots show. In the showcase CSVs, this is the `mean` column. For MFV showcase plots, model and human units differ, so the plotted quantity is relative foundation emphasis: each foundation profile is z-scored across foundations before mapping.
|
||||
|
||||
The table's reader-logit shift uses a more sensitive log-space readout.
|
||||
|
||||
For MFV:
|
||||
|
||||
$$\Delta_f = \mathbb{E}_i \left(\ell_{i,f}^{(+1)} - \ell_{i,f}^{(-1)}\right)$$
|
||||
|
||||
where $\ell_{i,f}^{(c)}$ is the forced-choice logit for foundation $f$ on item $i$ at coefficient $c$.
|
||||
|
||||
For survey instruments:
|
||||
|
||||
$$C_d(c) = \mathbb{E}_{i \in d}\sum_{k=1}^{M} \left(k - \frac{M+1}{2}\right)\ell_{i,k}^{(c)}$$
|
||||
|
||||
and the table reports $C_d(+1)-C_d(-1)$.
|
||||
|
||||
For paired steering runs, compare the base profile to the steered profile path. Answer mass is a coherence check, not a value score:
|
||||
|
||||
$$m(c) = \mathbb{E}_i \sum_{a \in A_i} P_c(a \mid i)$$
|
||||
|
||||
where $A_i$ is the valid answer-token set for item $i$. The showcase also checks survey contrast and MFV forced-choice margin, because a steered reader can keep answer mass while losing useful structure.
|
||||
|
||||
## Scope
|
||||
|
||||
tinymfv is for fast paired steering comparisons, not full moral reasoning evaluation. It is useful when you want to compare base, positive-steer, and negative-steer runs against the same human reference plots.
|
||||
|
||||
For behavior-heavy moral evals, see [machiavelli](https://huggingface.co/datasets/wassname/machiavelli), [AIRiskDilemmas](https://huggingface.co/datasets/kellycyy/AIRiskDilemmas), and [ethics_expression_preferences](https://huggingface.co/datasets/wassname/ethics_expression_preferences).
|
||||
|
||||
Used in [steering-lite](https://github.com/wassname/steering-lite), [lora-lite](https://github.com/wassname/lora-lite), and [w2schar-mini](https://github.com/wassname/w2schar-mini).
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{clark2026tinymfv,
|
||||
title = {tinymfv: tiny moral/value eval for local LLMs},
|
||||
author = {Michael Clark},
|
||||
year = {2026},
|
||||
url = {https://github.com/wassname/tinymfv/}
|
||||
}
|
||||
```
|
||||
@@ -99,13 +99,13 @@ Out:
|
||||
- likely_fail: all methods reproduce off-axis Social Norms movement.
|
||||
- sneaky_fail: a method looks good because retained c values differ; catch by printing retained c values and quality gates.
|
||||
- UAT: method comparison table links each output dir and verifier table. Result: no method is README-ready.
|
||||
- [ ] T7 (R6): Evaluate reliability and side effects.
|
||||
- [x] T7 (R6): Evaluate reliability and side effects.
|
||||
- steps: run tinymfv summary over MFV, MFQ-2, Humor, Big Five; estimate noise/CI where available.
|
||||
- verify: table includes profile shift/human SD, reader-logit shift, and uncertainty/noise columns.
|
||||
- success: Authority signal is larger than noise and side effects are interpretable.
|
||||
- likely_fail: MFQ-2 noisy or opposite sign.
|
||||
- sneaky_fail: profile shift comes from loss of answer structure; catch with answer mass, survey contrast, and MFV margin.
|
||||
- UAT: one table lets the user decide whether the steer is good enough to show.
|
||||
- UAT: one table lets the user decide whether the steer is good enough to show. Result: README now includes the per-axis table from the final PCA run. It reports `profile shift / human SD`, raw `profile shift`, and `reader-logit shift` for MFV, Humor Styles, Big Five, and MFQ-2. MFQ-2 sample-noise bootstrap is in `/media/wassname/SGIronWolf/projects5/2026/lite/tinymfv/docs/reviews/mfq2_authority_pca_n_bootstrap_20260701.csv`.
|
||||
- [x] T12 (R6): Estimate the minimum MFQ-2 `N` needed for stable steering plots.
|
||||
- steps: measure bootstrapped MFQ-2 Authority path variability at subset sizes `N=1,2,4,8` from per-sample readouts, or add a per-sample export if the current aggregate output is insufficient.
|
||||
- verify: table reports mean and bootstrap std/CI for Authority `C` deltas at each subset size and c row.
|
||||
@@ -140,13 +140,13 @@ Out:
|
||||
| sspace | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_sspace_mfv_mfq2_n8` | signed at `c=0.5`, not at `c=1.0` | `pmass=1`, `unscorable=0`, margin/base `0.972..0.980` at +c and `1.059..1.193` at -c |
|
||||
| directional_ablation | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_directional_ablation_mfv_mfq2_n8` | both signs raise Authority | `pmass=1`, `unscorable=0`, margin/base `0.812..0.813` at +c and `0.843..0.760` at -c |
|
||||
| linear_act | `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_linear_act_mfv_mfq2_n8` | signed at `c=0.5` and `c=1.0` | `pmass=1`, `unscorable=0`, margin/base `1.133..1.149` at +c and `0.891..0.760` at -c |
|
||||
- [ ] T9 (R7): Update README only from the final successful run.
|
||||
- [x] T9 (R7): Update README only from the final successful run.
|
||||
- steps: regenerate all README plots from one final artifact; add concise table and captions.
|
||||
- verify: README image links resolve; no 16PF plot; no WIP methodology journal in reader prose.
|
||||
- success: reader sees what tinymfv is, what was steered, which datasets measured it, and why it matters.
|
||||
- likely_fail: README narrates failed strict22/debug history.
|
||||
- sneaky_fail: captions imply a general sign convention; catch by reading README without this spec.
|
||||
- UAT: external-review-v2 comprehension panel can explain what/where/why/measurement/dataset in its own words.
|
||||
- UAT: external-review-v2 comprehension panel can explain what/where/why/measurement/dataset in its own words. Result: final panel artifacts are in `/media/wassname/SGIronWolf/projects5/2026/lite/tinymfv/docs/reviews/readme_comprehension_final_20260630_234536/`. The panel converged on the intended reading: tinymfv reads answer-token probabilities, builds MFV/survey profiles, compares steered paths to human reference data, and is used to catch intended movement, side effects, and non-human profiles. Repeated earlier confusion about `reader-logit shift`, `c` path, and MFQ-2 sampled reads was addressed in README.
|
||||
- [x] T11 (R5): Render candidate PCA maps from the best current MFV path.
|
||||
- steps: choose the method using MFV Authority direction and MFV coherence evidence only; render candidate maps/ranges from the selected steering-lite run.
|
||||
- verify: `uv run python scripts/plot_steer_showcase.py --run-dir /media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T163010Z_pure_authority_mundane15_pca_mfv_mfq2_n8 --out docs/img/showcase_authority_pca_mundane15 --vec-label "pure Authority, PCA (+c = more Authority)" --coherence-frac 0.99 --margin-frac 0.50 --contrast-frac 0.000001`
|
||||
@@ -244,3 +244,4 @@ Out:
|
||||
- 2026-07-01: Rendered method-comparison maps/ranges into `/media/wassname/SGIronWolf/projects5/2026/lite/tinymfv/docs/img/showcase_authority_method_compare_mundane15/`.
|
||||
- 2026-07-01: Pueue 423 completed the full PCA run over MFV, MFQ-2, Humor Styles, and Big Five at `N=8`: `/media/wassname/SGIronWolf/projects5/2026/lite/steering-lite/outputs/20260630T222000Z_pure_authority_mundane15_pca_readme_mfv_mfq2_humor_big5_n8`. MFV Authority direction is signed at both evaluated magnitudes (`+0.306/-0.327` at `c=0.5`; `+0.906/-0.910` at `c=1.0`) and coherence is clean (`pmass=1.000`, unscorable `0`, min margin/base `0.800`). The README plot images were regenerated from this run with the shared coherent c path `[-1,-0.5,0,+0.5,+1]`.
|
||||
- 2026-07-01: MFQ-2 per-sample bootstrap from the pueue 423 run shows sign stability even at `N=1`: every bootstrap draw has the intended Authority direction for `c in {-1,-0.5,+0.5,+1}`. `N=4` or `N=8` tightens uncertainty, but the minimum sign-stable `N` in this run is `1`.
|
||||
- 2026-07-01: README was updated from the final PCA run only. External-review-v2 comprehension panels first identified repeated gaps around `reader-logit shift`, the c path/gates, and MFQ-2 sampled reads; after edits, the final panel understood the core package/use-case and left only expected out-of-scope requests for a steering-lite end-to-end tutorial.
|
||||
|
||||
Reference in New Issue
Block a user