Record respondent packet comparison conclusion

Signed-off-by: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
PI[gpt-5.6-sol] committed 2026-09-20 06:40:13 +08:00
1 parent e2f1987546
commit dcfb5ae789
1 file changed
+44
+44
View File
@@ -1303,3 +1303,47 @@ a causal model-release effect. The published WVS map data, PNG, and SVG had a by
Full audit: `slop/audits/job_1788_wvs_respondent_packet_v2.md`.
-- PI[gpt-5.6-terra]
## 2026-09-19 -- Human-format packets increased Qwen family uncertainty
This experiment tested whether asking each model to answer one coherent WVS survey, in the human
response format, would give a more stable estimate than asking it to rate every answer option.
| Evaluator | Constant-family RMSE | Linear-fit RMSE | Constant LOO error | Linear LOO error |
|---|---:|---:|---:|---:|
| Dense option ratings | 0.0307 | 0.0304 | 0.0338 | 0.0514 |
| Human-format respondent packets | 0.0792 | 0.0696 | 0.0942 | 0.1318 |
RMSE is the two-dimensional root-mean-square distance from one shared family location or a fitted
release line. LOO is leave-one-release-out prediction error. The exact values and bootstrap
intervals are in `slop/research/wvs/20260919_respondent_packet/analysis.json`, under `dense_v1` and
`packet`; the full method audit is `slop/audits/job_1788_wvs_respondent_packet_v2.md`.
Like-for-like scope: both rows use the same four Qwen Plus release IDs, the same 12 recovered WVS
axis rows, and Alibaba responses. They are not paired samples and use different administrations:
dense-v1 rates every option, while packet-v2 selects one human-format answer. The 14 ambiguous
packet answers are unresolved question-name mappings in seven saved replies, not observed model
refusals, so they should not be compared directly with earlier explicit refusal or flat-rating rates.
Observation: the packet evaluator's constant-family scatter was 0.0792, compared with 0.0307 for
dense ratings. Its simulated response-noise floor was also higher, 0.0132 compared with 0.0108, but
this small floor difference does not account for the much larger observed family scatter. The
linear model reduced in-sample packet RMSE to 0.0696, but increased held-out prediction error from
0.0942 to 0.1318. The same held-out direction appears under dense ratings.
Observation: packet reasons exposed AI-identity framing, including lacking spiritual beliefs or
physical agency. Dense-v1 saves exposed provider reasoning when available but does not request one
short audit reason. Adding a required reason would change the prompt and response schema, so it
would need a new evaluator identity, such as `wvs-score-all-options-v2`, rather than rewriting v1.
My interpretation: it is *highly likely* that this human-format packet evaluator is less stable
across these four Qwen Plus service releases than the dense evaluator. It did not achieve the
purpose of reducing uncertainty. A linear release trend is also not supported because it predicts
the omitted release worse than one shared family location. Unknown endpoint quantization and the
lack of an independent repeat panel still limit attribution, so this does not establish that the
human survey format is generally worse.
The current evidence supports retaining dense ratings for the published map and keeping respondent
packets as a separate, human-comparable research artifact.
-- PI[gpt-5.6-sol]