diff --git a/README.md b/README.md index a190d23..959acf3 100644 --- a/README.md +++ b/README.md @@ -37,11 +37,13 @@ columns, QA flags, and limitations. ## Evaluation snapshot A fixed Gemma 4 31B judge compared the assigned motives in both account orders. -With no rubric, crossed-set accuracy declines slightly across the five Qwen -explanation models. This describes the judge finding the generated accounts -harder to classify, not stronger models concealing motives better. +With no rubric, crossed-set accuracy declines across the four open Qwen 3.5 +explanation models. Qwen 3.7 Max and Kimi K3 are omitted from this headline +because they are not from the same open model generation. This describes the +judge finding the generated accounts harder to classify, not stronger models +concealing motives better. -![No-rubric assigned-motive accuracy across five Qwen explanation models](results/no_rubric_motive_by_agent.svg) +![No-rubric assigned-motive accuracy across four open Qwen 3.5 explanation models](results/no_rubric_motive_by_agent.svg) ## Reproduce it diff --git a/results/no_rubric_motive_by_agent.svg b/results/no_rubric_motive_by_agent.svg index a6631d9..1c54ea3 100644 --- a/results/no_rubric_motive_by_agent.svg +++ b/results/no_rubric_motive_by_agent.svg @@ -1,13 +1,17 @@ - Assigned-motive accuracy across five Qwen models - no rubric; fixed Gemma 4 31B judge; crossed action-harm set + Assigned-motive accuracy across four open Qwen 3.5 models + no rubric; fixed Gemma 4 31B judge (AA 29); crossed action-harm set + weak to strong ← + | judge | + → strong to weak chance + 0.25 0.375 @@ -16,26 +20,24 @@ 0.75 crossed-set accuracy - - fit: -0.003 accuracy / AA point + + fit: -0.009 accuracy / AA point - - - - + + + 21 - 29 - 32 - 34 - 46 + 29 + 32 + 34 Artificial Analysis score - circles: Qwen 3.5; square: Qwen 3.7 Max; line: five-model linear fit + Qwen 3.5: 9B, 35B-A3B, 122B-A10B, 397B-A17B; line: four-model fit