mirror of
https://github.com/wassname/soul-ab-test-scope-test.git
synced 2026-08-20 12:51:16 +08:00
Three-way comparison on 8 held-out writing scenarios: - baseline (no humanizer) vs minimal (strongest tells + personality) - baseline vs full (all 29 patterns + banned words + structural slop) - minimal vs full (key comparison) Results: minimal beats full on held-out outcomes (voice -0.3, ai_suspicion -0.2 when adding full pattern catalog). The 48KB skill hurts compared to the 2KB minimal version. More AI writing involvement = more detectable, as the skill itself warns. Rubric uses outcome-level dimensions (ai_suspicion, voice, density, concreteness) not pattern compliance, to avoid testing the training set. - Moltark