From fe2e63b6fa89e59633ef1239813b23c537ac4f03 Mon Sep 17 00:00:00 2001 From: wassname <1103714+wassname@users.noreply.github.com> Date: Thu, 30 Jul 2026 14:45:12 +0800 Subject: [PATCH] Add the no-rubric motive scaling figure --- README.md | 10 +++++ results/no_rubric_motive_by_agent.svg | 54 +++++++++++++++++++++++++++ 2 files changed, 64 insertions(+) create mode 100644 results/no_rubric_motive_by_agent.svg diff --git a/README.md b/README.md index a01f843..e5f4ad1 100644 --- a/README.md +++ b/README.md @@ -34,6 +34,16 @@ above. The Hugging Face dataset card explains the two configurations, intended uses, columns, QA flags, and limitations. +## Evaluation snapshot + +A fixed Gemma 4 31B judge compared the assigned motives in both account orders. +With no rubric, crossed-set accuracy declines slightly across the five Qwen +explanation models; Kimi K3 is shown separately because it is from another model +family. This describes the judge finding the generated accounts harder to +classify, not stronger models concealing motives better. + +![No-rubric assigned-motive accuracy by explanation-model capability](results/no_rubric_motive_by_agent.svg) + ## Reproduce it Install [uv](https://docs.astral.sh/uv/), then run: diff --git a/results/no_rubric_motive_by_agent.svg b/results/no_rubric_motive_by_agent.svg new file mode 100644 index 0000000..da9652e --- /dev/null +++ b/results/no_rubric_motive_by_agent.svg @@ -0,0 +1,54 @@ + + + + Assigned-motive accuracy by explanation model + no rubric; fixed Gemma 4 31B judge; crossed action-harm set + circles: Qwen 3.5; square: Qwen 3.7 Max; diamond: Kimi K3 + + + + + chance + + 0.25 + 0.375 + 0.50 + 0.625 + 0.75 + crossed-set accuracy + + + Qwen fit: -0.003 / AA point + + + + + + + + + + + + + + + + + + + + + 21 + 29 + 32 + 34 + 46 + 57 + + + + Artificial Analysis score + points: model means; bars: 95% scenario bootstrap; line: Qwen-only linear fit + +