From dcfb5ae7893b6c6c3432d45e4fa9895c69ebbe44 Mon Sep 17 00:00:00 2001 From: "PI[gpt-5.6-sol]" <288921227+claudypoo@users.noreply.github.com> Date: Sun, 20 Sep 2026 06:40:13 +0800 Subject: [PATCH] Record respondent packet comparison conclusion Signed-off-by: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com> --- docs/RESEARCH_JOURNAL.md | 44 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 44 insertions(+) diff --git a/docs/RESEARCH_JOURNAL.md b/docs/RESEARCH_JOURNAL.md index 84a5317..8d1b2cc 100644 --- a/docs/RESEARCH_JOURNAL.md +++ b/docs/RESEARCH_JOURNAL.md @@ -1303,3 +1303,47 @@ a causal model-release effect. The published WVS map data, PNG, and SVG had a by Full audit: `slop/audits/job_1788_wvs_respondent_packet_v2.md`. -- PI[gpt-5.6-terra] + +## 2026-09-19 -- Human-format packets increased Qwen family uncertainty + +This experiment tested whether asking each model to answer one coherent WVS survey, in the human +response format, would give a more stable estimate than asking it to rate every answer option. + +| Evaluator | Constant-family RMSE | Linear-fit RMSE | Constant LOO error | Linear LOO error | +|---|---:|---:|---:|---:| +| Dense option ratings | 0.0307 | 0.0304 | 0.0338 | 0.0514 | +| Human-format respondent packets | 0.0792 | 0.0696 | 0.0942 | 0.1318 | + +RMSE is the two-dimensional root-mean-square distance from one shared family location or a fitted +release line. LOO is leave-one-release-out prediction error. The exact values and bootstrap +intervals are in `slop/research/wvs/20260919_respondent_packet/analysis.json`, under `dense_v1` and +`packet`; the full method audit is `slop/audits/job_1788_wvs_respondent_packet_v2.md`. + +Like-for-like scope: both rows use the same four Qwen Plus release IDs, the same 12 recovered WVS +axis rows, and Alibaba responses. They are not paired samples and use different administrations: +dense-v1 rates every option, while packet-v2 selects one human-format answer. The 14 ambiguous +packet answers are unresolved question-name mappings in seven saved replies, not observed model +refusals, so they should not be compared directly with earlier explicit refusal or flat-rating rates. + +Observation: the packet evaluator's constant-family scatter was 0.0792, compared with 0.0307 for +dense ratings. Its simulated response-noise floor was also higher, 0.0132 compared with 0.0108, but +this small floor difference does not account for the much larger observed family scatter. The +linear model reduced in-sample packet RMSE to 0.0696, but increased held-out prediction error from +0.0942 to 0.1318. The same held-out direction appears under dense ratings. + +Observation: packet reasons exposed AI-identity framing, including lacking spiritual beliefs or +physical agency. Dense-v1 saves exposed provider reasoning when available but does not request one +short audit reason. Adding a required reason would change the prompt and response schema, so it +would need a new evaluator identity, such as `wvs-score-all-options-v2`, rather than rewriting v1. + +My interpretation: it is *highly likely* that this human-format packet evaluator is less stable +across these four Qwen Plus service releases than the dense evaluator. It did not achieve the +purpose of reducing uncertainty. A linear release trend is also not supported because it predicts +the omitted release worse than one shared family location. Unknown endpoint quantization and the +lack of an independent repeat panel still limit attribution, so this does not establish that the +human survey format is generally worse. + +The current evidence supports retaining dense ratings for the published map and keeping respondent +packets as a separate, human-comparable research artifact. + +-- PI[gpt-5.6-sol]