From aebb86ce2511f3842c5bc3473887001ceeeee453 Mon Sep 17 00:00:00 2001 From: wassname <1103714+wassname@users.noreply.github.com> Date: Thu, 30 Jul 2026 15:29:38 +0800 Subject: [PATCH] Focus the motive headline on Qwen models --- README.md | 7 ++- results/no_rubric_motive_by_agent.svg | 69 +++++++++++---------------- 2 files changed, 31 insertions(+), 45 deletions(-) diff --git a/README.md b/README.md index e5f4ad1..a190d23 100644 --- a/README.md +++ b/README.md @@ -38,11 +38,10 @@ columns, QA flags, and limitations. A fixed Gemma 4 31B judge compared the assigned motives in both account orders. With no rubric, crossed-set accuracy declines slightly across the five Qwen -explanation models; Kimi K3 is shown separately because it is from another model -family. This describes the judge finding the generated accounts harder to -classify, not stronger models concealing motives better. +explanation models. This describes the judge finding the generated accounts +harder to classify, not stronger models concealing motives better. -![No-rubric assigned-motive accuracy by explanation-model capability](results/no_rubric_motive_by_agent.svg) +![No-rubric assigned-motive accuracy across five Qwen explanation models](results/no_rubric_motive_by_agent.svg) ## Reproduce it diff --git a/results/no_rubric_motive_by_agent.svg b/results/no_rubric_motive_by_agent.svg index da9652e..a6631d9 100644 --- a/results/no_rubric_motive_by_agent.svg +++ b/results/no_rubric_motive_by_agent.svg @@ -1,54 +1,41 @@ - - + + - Assigned-motive accuracy by explanation model - no rubric; fixed Gemma 4 31B judge; crossed action-harm set - circles: Qwen 3.5; square: Qwen 3.7 Max; diamond: Kimi K3 + Assigned-motive accuracy across five Qwen models + no rubric; fixed Gemma 4 31B judge; crossed action-harm set - - - - chance + + + + chance - 0.25 - 0.375 - 0.50 - 0.625 - 0.75 - crossed-set accuracy + 0.25 + 0.375 + 0.50 + 0.625 + 0.75 + crossed-set accuracy - - Qwen fit: -0.003 / AA point - - - - - - - - - + + fit: -0.003 accuracy / AA point - - - - - - + + + + + - 21 - 29 - 32 - 34 - 46 - 57 - + 21 + 29 + 32 + 34 + 46 - Artificial Analysis score - points: model means; bars: 95% scenario bootstrap; line: Qwen-only linear fit + Artificial Analysis score + circles: Qwen 3.5; square: Qwen 3.7 Max; line: five-model linear fit