From 576da388301241d880611a36d46247d2e3041f2c Mon Sep 17 00:00:00 2001 From: wassname <1103714+wassname@users.noreply.github.com> Date: Sun, 5 Jul 2026 12:44:55 +0800 Subject: [PATCH] data provenance: inline at point-of-use, drop standalone SOURCES.md Per feedback (brief + inline + where seen, not a shadow file): - rm mfv_country_factors_SOURCES.md; MFV per-source citations+transforms now in read_human_mfv() docstring (deep detail stays in each row's commit body). - instruments.py: one-line human_csv source per legacy dataset (mfq2/big5/16pf/humor) above _SPECS -- Atari 2023, OpenPsychometrics IPIP, Schermer 2020. - value_axes.py: add years+DOIs to the custom-axis Sources block (Graham/Haidt 2009, DeYoung 2007, Atari 2023, Martin 2003, Inglehart-Welzel). iw_axes already documents WVS. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com> --- scripts/plot_steer_showcase.py | 11 ++- .../data/human/mfv_country_factors_SOURCES.md | 95 ------------------- src/tinymfv/instruments.py | 6 ++ src/tinymfv/value_axes.py | 11 ++- 4 files changed, 21 insertions(+), 102 deletions(-) delete mode 100644 src/tinymfv/data/human/mfv_country_factors_SOURCES.md diff --git a/scripts/plot_steer_showcase.py b/scripts/plot_steer_showcase.py index f01dd37..9be8972 100644 --- a/scripts/plot_steer_showcase.py +++ b/scripts/plot_steer_showcase.py @@ -245,9 +245,14 @@ def _zscore(v: np.ndarray) -> np.ndarray: def read_human_mfv() -> tuple[list[str], dict[str, dict[str, float]]]: """(countries, {country: {foundation: mean_1to5}}) from the bundled MFV human norms. - JimenezLeal2025 (LatAm: Argentina/Colombia/Peru/US) + Yamada2025 (MFV-J: Japan) - + Hopp2024 (Dutch: Netherlands) + Marques2020 (Brazil) + Crone2021 (Australia): - 8 countries x 6 foundations (no Social Norms). Provenance: mfv_country_factors_SOURCES.md.""" + 8 countries x 6 foundations (no Social Norms). Per-row provenance = the CSV `source` + column; each tag expands here (full transform detail lives in the row's git commit body): + JimenezLeal2025_LatAm AR/CO/PE/US Jimenez-Leal+ 2025 Collabra doi 10.1525/collabra.128178 (tables) + Yamada2025_MFV-J Japan Yamada+ 2026 Jpn J Psych doi 10.4992/jjpsy.97.24228 (table) + Hopp2024_DutchMFV Netherlands Hopp+ 2024 JDM 19:e10 doi 10.1017/jdm.2024.5 (Table 1; care=mean(phys,emo)) + Marques2020_..._affinecal Brazil Marques+ 2020 JDM journal.sjdm.org/19/190809a; Fig-3 digitized + affine bias-cal + Crone2021_AusUndergrad Australia Crone,Rhee,Laham 2021 Behav Res Methods doi 10.3758/s13428-020-01489-y; + raw 90-item dat_rep.sav (OSF cmwpv, 756 undergrads), NOT the GA-abbrev subset""" path = T.maps.DATA / "human" / "mfv_country_factors.csv" by_country: dict[str, dict[str, float]] = {} with open(path, newline="") as fh: diff --git a/src/tinymfv/data/human/mfv_country_factors_SOURCES.md b/src/tinymfv/data/human/mfv_country_factors_SOURCES.md deleted file mode 100644 index 01882a5..0000000 --- a/src/tinymfv/data/human/mfv_country_factors_SOURCES.md +++ /dev/null @@ -1,95 +0,0 @@ -# Sources & provenance: `mfv_country_factors.csv` - -Audit trail for the human country rows in the MFV value map. One entry per `source` -tag in the CSV. Each entry records: full citation, URL/DOI, exactly which table/figure -the numbers came from, and any data transformation applied (so a reviewer can reproduce -the row from the primary source). - -Author of these notes: Claude (pair-programming with wassname), not wassname. - -## Instrument & how the map uses these numbers - -All rows are the Clifford et al. (2015) Moral Foundations Vignettes (MFV): 2nd-person -"You see ..." vignettes rated for moral wrongness on a 1-5 scale, coded to six -foundations (care, fairness, liberty, authority, loyalty/ingroup, sanctity/purity). - -The value map (`scripts/plot_steer_showcase.py::read_human_mfv`) reads **only the `mean` -column**, then **z-scores each country across its six foundations** (ipsative). Because -z-scoring is affine-invariant, any per-country scale/offset (translation bias, digitizing -bias, differing scale anchors) cancels: the map shows *relative* foundation emphasis, not -a calibrated cross-country ranking. Absolute means/SDs are still stored honestly for any -non-map reuse. - -Caveat carried by two of the source papers (Jimenez-Leal, Hopp): the MFV shows -measurement non-invariance / differential item functioning across countries, so raw -between-country mean comparisons are a rough reference, not a validated ranking. - -## Sources - -### `JimenezLeal2025_LatAm` -- Argentina, Colombia, Peru, US -- Jimenez Leal, W., Carmona, G., Murray, S., & Amaya, S. (2025). Validation of the Moral - Foundation Vignettes in Latin America. *Collabra: Psychology*, 11(1), 128178. -- DOI: https://doi.org/10.1525/collabra.128178 (open access, CC BY) -- Numbers: per-country foundation means/SD/N from the paper's descriptive tables - (N = 1,650 across 3 Latin-American countries via polling agency, plus a US comparison). -- Transformation: none (means used as tabulated on the native 1-5 scale). - -### `Yamada2025_MFV-J` -- Japan -- Yamada, J., Nakawake, Y., & Suyama, M. (2026). Developing a Japanese version of the - Moral Foundations Vignettes (MFV-J). *The Japanese Journal of Psychology*. - (Tag says 2025 = preprint/advance-pub year; journal assigns 2026.) -- DOI: https://doi.org/10.4992/jjpsy.97.24228 (advance publication PDF, open access) -- Numbers: MFV-J foundation means/SD, N = 564, from the paper's descriptive table. -- Transformation: none. - -### `Hopp2024_DutchMFV` -- Netherlands -- Hopp, F. R., Jargow, B., Kouwen, E., & Bakker, B. N. (2024). The Dutch moral foundations - stimulus database. *Judgment and Decision Making*, 19, e10. -- DOI: https://doi.org/10.1017/jdm.2024.5 (open access, CC BY). OSF: https://osf.io/9gnza/ -- Numbers: foundation means + 95% CIs from the paper's Table 1 (N = 586 Dutch crowdworkers, - 120 translated MFVs). Per-foundation N varies by item allocation (that's why the CSV N - differs per row). -- Transformation: **Care collapsed** from the paper's split physical-care (4.09) + - emotional-care (3.53) into a single care = 3.81 (mean of the two). Other foundations - taken as tabulated. SD/SE/CI from the paper's reported CIs. - -### `Marques2020_BrazilMFV_fig3digitized_affinecal` -- Brazil -- Marques, L. M., et al. (2020). Translation and validation of the Moral Foundations - Vignettes for the Portuguese language in a Brazilian sample. *Judgment and Decision - Making*. -- URL: http://journal.sjdm.org/19/190809a/jdm190809a.html - PDF: http://journal.sjdm.org/19/190809a/jdm190809a.pdf -- Numbers: the paper tabulates no per-foundation means, so means were **digitized from - Figure 3** (Brazil series) with WebPlotDigitizer by wassname. N = 494 (paper). -- Transformations (two, both by Claude): - 1. **Care collapsed** from digitized Care-E + Care-P into one care value. - 2. **Affine bias-correction** of all digitized means to two paper-stated Purity anchors - (Brazil purity 3.45, Clifford US purity 3.85; paper text): `true = 0.9155*digitized - + 0.228`. Reproduces both anchors; Brazil purity lands exactly 3.45. SDs recovered - from the digitized 95% CI whiskers (Fig 3 caption): `sd = halfwidth/1.96 * sqrt(494)`; - each whisker pair's midpoint matches its mean to <=0.005 (validated). SDs scaled by - the affine slope. Provably map-neutral (max |dz| = 0.017, rounding only). -- Digitized source figure staged at `docs/digitize/brazil_fig3_page-08.png`. - -### `Crone2021_AusUndergrad_MFV90raw` -- Australia -- Crone, D. L., Rhee, J. J., & Laham, S. M. (2021). Developing brief versions of the Moral - Foundations Vignettes using a genetic algorithm-based approach. *Behavior Research - Methods*, 53(3), 1179-1187. -- DOI: https://doi.org/10.3758/s13428-020-01489-y . OSF: https://osf.io/cmwpv/ - (component "Data" = https://osf.io/nv4ty/ , file `dat_rep.sav`). -- Sample: 756 Australian undergraduates (complete cases). The paper's other sample is - 580 US MTurk workers (`dat_amt`), NOT used here -- that would duplicate the US row. -- Numbers: computed by Claude from the **raw participant ratings** in `dat_rep.sav`, the - full 90-item Clifford MFV (1-5 wrongness), using the author's own item->foundation map - from `mfv_abbreviation.Rmd` (Care = MFV 1-27 [physical+emotional+other], Fairness 28-39, - Liberty 40-50, Authority 51-64, Loyalty/Ingroup 65-80, Sanctity/Purity 81-90). - Per foundation: mean over participants of each participant's item-mean; SD = between- - participant SD (ddof=1); SE = SD/sqrt(N); CI = mean +/- 1.96*SE. Complete-case exclusion - (all 90 items present) reproduces the paper's N = 756 exactly. -- Transformation: **NOT the genetic-algorithm-abbreviated subset.** The paper's headline - contribution is a brief MFV chosen by GAabbreviate; we deliberately used the *full* - 90-item ratings so Australia is comparable to the other (full-instrument) country rows. - Care spans all 27 care items (no collapse needed; it's the natural mean). -- Reproduce: OSF files fetched via `https://osf.io/download/4psfc/` (dat_rep.sav); - read with `pyreadstat.read_sav`. (The OSF "Download as zip" gave an empty archive -- - fetch files individually.) diff --git a/src/tinymfv/instruments.py b/src/tinymfv/instruments.py index 997fb6d..70f13d4 100644 --- a/src/tinymfv/instruments.py +++ b/src/tinymfv/instruments.py @@ -32,6 +32,12 @@ DIGITS_1_5 = ["1", "2", "3", "4", "5"] PREFILL = "(" # instrument name -> (survey subdir, {frame: filename stem}, human_csv, display, human_scale_max) +# human_csv provenance (per-country reference means; ported from the mft_honesty experiment): +# mfq2 : Atari+ 2023 "Morality beyond the WEIRD", OSF osf.io/srtxn/ Study 2 (19 countries) +# big5 : OpenPsychometrics IPIP-FFM raw (openpsychometrics.org/_rawdata, BIG5), aggregated (24 countries) +# 16pf : OpenPsychometrics 16PF raw (openpsychometrics.org/tests/16PF), aggregated (34 countries) +# humor_styles : Schermer+ 2020 "Humor styles across 28 countries", doi 10.1007/s12144-019-00552-y (Tables 1-2) +# (MFV is not here -- it's read in scripts/plot_steer_showcase.py::read_human_mfv, provenance in that docstring.) _F3 = {"forward": "questionnaire", "inverted": "questionnaire_inverted", "negated": "questionnaire_negated"} _SPECS = { "mfq2": ("mfq2", {"forward": "forward", "inverted": "inverted", "negated": "negated"}, diff --git a/src/tinymfv/value_axes.py b/src/tinymfv/value_axes.py index ec7859a..4935574 100644 --- a/src/tinymfv/value_axes.py +++ b/src/tinymfv/value_axes.py @@ -7,10 +7,13 @@ factors of the endorsement (sign +1) or its complement 1-endorsement (sign -1). with -1 is reverse-scored (big5 Stability reverses neuroticism; a contrast axis puts one pole's factors at -1). High score = the axis's POSITIVE (second) pole. -Sources: MFT individualizing/binding -- Graham & Haidt; MFQ-2 equality/proportionality fairness split --- Atari et al. 2023. Big Five meta-traits Plasticity/Stability -- DeYoung 2007. HSQ 2x2 -adaptive/maladaptive x self/other -- Martin et al. 2003. WVS -- Inglehart-Welzel (see tinymfv.iw_axes; -the WVS map builds its own item-level axes, this table is for the psychometric instruments). +Sources (axis groupings aggregated from these papers): +- MFT individualizing/binding: Graham, Haidt & Nosek 2009, doi 10.1037/a0015141. +- MFQ-2 equality/proportionality fairness split: Atari et al. 2023 "Morality beyond the WEIRD". +- Big Five meta-traits Plasticity/Stability: DeYoung, Quilty & Peterson 2007, doi 10.1037/0022-3514.93.5.880. +- HSQ 2x2 adaptive/maladaptive x self/other: Martin et al. 2003, doi 10.1016/S0092-6566(02)00534-2. +- WVS: Inglehart & Welzel 2005 (see tinymfv.iw_axes -- the WVS map builds its own item-level axes; + this table is for the psychometric instruments). The debatable calls (flagged): the SECOND MFT axis is not canonical -- mfq2 uses the documented equality(egalitarian) vs proportionality(meritocratic) fairness split; mfv (no equality/proportionality