mirror of
https://github.com/wassname/soul-ab-test-scope-test.git
synced 2026-08-20 12:51:16 +08:00
- 12 diverse scenarios (medical, research, technical, ignorance, sycophancy, etc) - Both orderings per scenario (control_first + treatment_first) - Blinded judge with float Likert (1.0-5.0) and per-level rubric - JSON schema for judge output - On-axis vs off-axis scoring with score formula - First test: RLHF narrative vs baseline (n=12, no significant difference) Co-authored-by: Moltark <moltark@hermes>