mirror of
https://github.com/wassname/moral-maps.git
synced 2026-10-08 12:19:14 +08:00
Record respondent packet comparison conclusion
Signed-off-by: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
1 parent
e2f1987546
commit
dcfb5ae789
1 file changed
+44
@@ -1303,3 +1303,47 @@ a causal model-release effect. The published WVS map data, PNG, and SVG had a by
|
||||
Full audit: `slop/audits/job_1788_wvs_respondent_packet_v2.md`.
|
||||
|
||||
-- PI[gpt-5.6-terra]
|
||||
|
||||
## 2026-09-19 -- Human-format packets increased Qwen family uncertainty
|
||||
|
||||
This experiment tested whether asking each model to answer one coherent WVS survey, in the human
|
||||
response format, would give a more stable estimate than asking it to rate every answer option.
|
||||
|
||||
| Evaluator | Constant-family RMSE | Linear-fit RMSE | Constant LOO error | Linear LOO error |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Dense option ratings | 0.0307 | 0.0304 | 0.0338 | 0.0514 |
|
||||
| Human-format respondent packets | 0.0792 | 0.0696 | 0.0942 | 0.1318 |
|
||||
|
||||
RMSE is the two-dimensional root-mean-square distance from one shared family location or a fitted
|
||||
release line. LOO is leave-one-release-out prediction error. The exact values and bootstrap
|
||||
intervals are in `slop/research/wvs/20260919_respondent_packet/analysis.json`, under `dense_v1` and
|
||||
`packet`; the full method audit is `slop/audits/job_1788_wvs_respondent_packet_v2.md`.
|
||||
|
||||
Like-for-like scope: both rows use the same four Qwen Plus release IDs, the same 12 recovered WVS
|
||||
axis rows, and Alibaba responses. They are not paired samples and use different administrations:
|
||||
dense-v1 rates every option, while packet-v2 selects one human-format answer. The 14 ambiguous
|
||||
packet answers are unresolved question-name mappings in seven saved replies, not observed model
|
||||
refusals, so they should not be compared directly with earlier explicit refusal or flat-rating rates.
|
||||
|
||||
Observation: the packet evaluator's constant-family scatter was 0.0792, compared with 0.0307 for
|
||||
dense ratings. Its simulated response-noise floor was also higher, 0.0132 compared with 0.0108, but
|
||||
this small floor difference does not account for the much larger observed family scatter. The
|
||||
linear model reduced in-sample packet RMSE to 0.0696, but increased held-out prediction error from
|
||||
0.0942 to 0.1318. The same held-out direction appears under dense ratings.
|
||||
|
||||
Observation: packet reasons exposed AI-identity framing, including lacking spiritual beliefs or
|
||||
physical agency. Dense-v1 saves exposed provider reasoning when available but does not request one
|
||||
short audit reason. Adding a required reason would change the prompt and response schema, so it
|
||||
would need a new evaluator identity, such as `wvs-score-all-options-v2`, rather than rewriting v1.
|
||||
|
||||
My interpretation: it is *highly likely* that this human-format packet evaluator is less stable
|
||||
across these four Qwen Plus service releases than the dense evaluator. It did not achieve the
|
||||
purpose of reducing uncertainty. A linear release trend is also not supported because it predicts
|
||||
the omitted release worse than one shared family location. Unknown endpoint quantization and the
|
||||
lack of an independent repeat panel still limit attribution, so this does not establish that the
|
||||
human survey format is generally worse.
|
||||
|
||||
The current evidence supports retaining dense ratings for the published map and keeping respondent
|
||||
packets as a separate, human-comparable research artifact.
|
||||
|
||||
-- PI[gpt-5.6-sol]
|
||||
Reference in new issue
Block a user