- 12 diverse scenarios (medical, research, technical, ignorance, sycophancy, etc)
- Both orderings per scenario (control_first + treatment_first)
- Blinded judge with float Likert (1.0-5.0) and per-level rubric
- JSON schema for judge output
- On-axis vs off-axis scoring with score formula
- First test: RLHF narrative vs baseline (n=12, no significant difference)
Co-authored-by: Moltark <moltark@hermes>