Record durable WVS API panels

Co-Authored-By: PI[gpt-5.6-terra] <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
wassname
2026-09-16 21:36:57 +08:00
co-authored by PI[gpt-5.6-terra]
parent 9efd9ca8e6
commit f2d6556af1
21 changed files with 3395 additions and 115 deletions
+32
View File
@@ -879,3 +879,35 @@ be compared.
Decision: MFV is NOT a cultures map. Keep it for the MODEL's relative-emphasis steer only.
Plan below. MFQ-2/Big5/Humour use single-source country tables and are unaffected (TODO: still
worth confirming each is single-source + comparable). -- authored by Claude
## 2026-09-16 -- WVS API run budget before requests
This entry records the configuration-based budget for the planned WVS API measurement.
Evidence from the approved plan and `scripts/wvs_map.py` before this run: the panel has twelve distinct
WVS items, each model is requested twelve ratings per item, and each initial request has a maximum
completion allowance of 1024 tokens. This yields 144 initial requests and 147456 maximum initial
completion tokens per model. A malformed initial reply triggers one rescue with a 2048-token maximum,
so 144 rescues add at most 294912 completion tokens. The protocol therefore reserves at most 442368
completion tokens per model when every initial reply needs rescue. Input token counts are unknown at
this point because OpenRouter bills the provider tokenization of each rendered prompt, and historical
request records were not retained. Cache reads and writes, failed calls, and any unreported provider
billing fields are also unknown, not zero. Source: approved plan
`.pi/plan/9a9c0a-v1.md`, and the pre-run request loop in `src/moralmaps/read_api.py`.
The planned accounting rule is `cost_usd = input_tokens * input_usd_per_million / 1e6 + completion_tokens
* output_usd_per_million / 1e6`, with cached-token rates kept separate when the provider returns them.
The plan records public catalog prices observed on 2026-09-16, including DeepSeek V4.1 Flash at
0.15 input and 0.60 output USD per million tokens, and GLM 5.3 Flash at 0.09 input and 0.30 output
USD per million tokens. On the output-only allowance, these give 0.09 and 0.04 USD respectively for
one initial-only model run, and 0.27 and 0.13 USD respectively if every request needs a rescue. Fable
5.1 and GPT-6 Astra are listed at 50 USD per million output tokens, which is 7.37 USD initial-only or
22.12 USD if every request rescues, before inputs. These are configuration bounds using stated prices,
not measured invoices. Source: `.pi/plan/9a9c0a-v1.md` Appendix, quoted catalog snapshot.
My read: the cheapest full diagnostic should establish actual completion and rescue behavior before the
expensive models. The unknown input and cache billing mean that a simple per-model maximum does not
prove total spend remains below the authorized cap, so durable records must retain every raw usage object
and request phase before the next paid call. -- PI[gpt-5.6-terra]
The next result will replace these bounds with reconciled provider-reported usage.
+77
View File
@@ -0,0 +1,77 @@
# OpenRouter WVS model inventory
Checked 2026-09-16 against `slop/research/wvs/20260916_openrouter/openrouter_models_20260916T1303Z.json`. Prices are catalog USD per million tokens. Batch and free aliases are excluded because they duplicate an underlying model. Qwen entries with an expiration date before the check date are excluded. The remaining direct `qwen/` text-capable releases are candidates, not evidence that they completed the panel.
## Requested additions
| exact OpenRouter ID | name | created UTC | input USD/M | output USD/M | status |
|---|---|---:|---:|---:|---|
| `anthropic/claude-fable-5.1` | Anthropic: Claude Fable 5.1 | 2026-09-01 | 10 | 50 | candidate |
| `openai/gpt-6-astra` | OpenAI: GPT-6 Astra | 2026-09-04 | 10 | 50 | candidate |
| `meta/muse-spark-1.3` | Meta: Muse Spark 1.3 | 2026-09-02 | 1.25 | 4.25 | candidate |
| `moonshotai/kimi-k3` | MoonshotAI: Kimi K3 | 2026-07-16 | 2.64814 | 13.2827 | candidate |
| `thinkingmachines/inkling` | Thinking Machines: Inkling | 2026-07-17 | 1 | 4.05 | candidate |
| `deepseek/deepseek-v4.1-flash` | DeepSeek: DeepSeek V4.1 Flash | 2026-09-10 | 0.15 | 0.6 | candidate |
| `z-ai/glm-5.3` | Z.ai: GLM 5.3 | 2026-08-18 | 1.4 | 4.4 | candidate |
| `z-ai/glm-5.3-flash` | Z.ai: GLM 5.3 Flash | 2026-08-26 | 0.09 | 0.3 | candidate |
| `google/gemini-3.7-flash` | Google: Gemini 3.7 Flash | 2026-08-13 | 0.75 | 3.75 | candidate |
| `x-ai/grok-4.5` | SpaceXAI: Grok 4.5 | 2026-07-08 | 2 | 6 | candidate |
| `openai/gpt-5.6-sol` | OpenAI: GPT-5.6 Sol | 2026-07-09 | 2 | 10 | candidate |
## Direct Qwen candidates
| exact OpenRouter ID | name | created UTC | input USD/M | output USD/M | status |
|---|---|---:|---:|---:|---|
| `qwen/qwen3.8-max-0902` | Qwen: Qwen3.8 Max (0902) | 2026-09-03 | 2 | 6 | candidate |
| `qwen/qwen3.8-flash` | Qwen: Qwen3.8 Flash | 2026-08-26 | 0.15 | 0.47 | candidate |
| `qwen/qwen3.8-27b` | Qwen: Qwen3.8 27B | 2026-08-14 | 0.214 | 2.55 | candidate |
| `qwen/qwen3.8-2.4t-a95b` | Qwen: Qwen3.8 2.4T A95B | 2026-08-12 | 2 | 6 | candidate |
| `qwen/qwen3.7-flash` | Qwen: Qwen3.7 Flash | 2026-07-27 | 0.03 | 0.13 | candidate |
| `qwen/qwen3.7-plus` | Qwen: Qwen3.7 Plus | 2026-06-03 | 0.32 | 1.28 | candidate |
| `qwen/qwen3.7-max` | Qwen: Qwen3.7 Max | 2026-05-21 | 1.475 | 4.425 | candidate |
| `qwen/qwen3.5-plus-20260420` | Qwen: Qwen3.5 Plus 2026-04-20 | 2026-04-27 | 0.3 | 1.8 | candidate |
| `qwen/qwen3.6-flash` | Qwen: Qwen3.6 Flash | 2026-04-27 | 0.1875 | 1.125 | candidate |
| `qwen/qwen3.6-35b-a3b` | Qwen: Qwen3.6 35B A3B | 2026-04-27 | 0.1 | 0.9 | candidate |
| `qwen/qwen3.6-max-preview` | Qwen: Qwen3.6 Max Preview | 2026-04-27 | 1.027 | 6.162 | candidate |
| `qwen/qwen3.6-27b` | Qwen: Qwen3.6 27B | 2026-04-27 | 0.3 | 2 | candidate |
| `qwen/qwen3.6-plus` | Qwen: Qwen3.6 Plus | 2026-04-02 | 0.325 | 1.95 | candidate |
| `qwen/qwen3.5-9b` | Qwen: Qwen3.5-9B | 2026-03-10 | 0.1 | 0.15 | candidate |
| `qwen/qwen3.5-35b-a3b` | Qwen: Qwen3.5-35B-A3B | 2026-02-25 | 0.1625 | 1.3 | candidate |
| `qwen/qwen3.5-27b` | Qwen: Qwen3.5-27B | 2026-02-25 | 0.195 | 1.56 | candidate |
| `qwen/qwen3.5-122b-a10b` | Qwen: Qwen3.5-122B-A10B | 2026-02-25 | 0.26 | 2.08 | candidate |
| `qwen/qwen3.5-flash-02-23` | Qwen: Qwen3.5-Flash | 2026-02-25 | 0.065 | 0.26 | candidate |
| `qwen/qwen3.5-plus-02-15` | Qwen: Qwen3.5 Plus 2026-02-15 | 2026-02-16 | 0.26 | 1.56 | candidate |
| `qwen/qwen3.5-397b-a17b` | Qwen: Qwen3.5 397B A17B | 2026-02-16 | 0.55 | 3.5 | candidate |
| `qwen/qwen3-max-thinking` | Qwen: Qwen3 Max Thinking | 2026-02-09 | 0.78 | 3.9 | candidate |
| `qwen/qwen3-coder-next` | Qwen: Qwen3 Coder Next | 2026-02-04 | 0.12 | 0.8 | candidate |
| `qwen/qwen3-vl-32b-instruct` | Qwen: Qwen3 VL 32B Instruct | 2025-10-23 | 0.104 | 0.416 | candidate |
| `qwen/qwen3-vl-8b-thinking` | Qwen: Qwen3 VL 8B Thinking | 2025-10-14 | 0.18 | 2.1 | candidate |
| `qwen/qwen3-vl-8b-instruct` | Qwen: Qwen3 VL 8B Instruct | 2025-10-14 | 0.117 | 0.455 | candidate |
| `qwen/qwen3-vl-30b-a3b-thinking` | Qwen: Qwen3 VL 30B A3B Thinking | 2025-10-06 | 0.2 | 2.4 | candidate |
| `qwen/qwen3-vl-30b-a3b-instruct` | Qwen: Qwen3 VL 30B A3B Instruct | 2025-10-06 | 0.15 | 0.6 | candidate |
| `qwen/qwen3-vl-235b-a22b-thinking` | Qwen: Qwen3 VL 235B A22B Thinking | 2025-09-23 | 0.4 | 4 | candidate |
| `qwen/qwen3-vl-235b-a22b-instruct` | Qwen: Qwen3 VL 235B A22B Instruct | 2025-09-23 | 0.21 | 1.9 | candidate |
| `qwen/qwen3-max` | Qwen: Qwen3 Max | 2025-09-23 | 0.78 | 3.9 | candidate |
| `qwen/qwen3-coder-plus` | Qwen: Qwen3 Coder Plus | 2025-09-23 | 0.65 | 3.25 | candidate |
| `qwen/qwen3-coder-flash` | Qwen: Qwen3 Coder Flash | 2025-09-17 | 0.195 | 0.975 | candidate |
| `qwen/qwen3-next-80b-a3b-thinking` | Qwen: Qwen3 Next 80B A3B Thinking | 2025-09-11 | 0.15 | 1.2 | candidate |
| `qwen/qwen3-next-80b-a3b-instruct` | Qwen: Qwen3 Next 80B A3B Instruct | 2025-09-11 | 0.09 | 1.1 | candidate |
| `qwen/qwen-plus-2025-07-28` | Qwen: Qwen Plus 0728 | 2025-09-08 | 0.26 | 0.78 | candidate |
| `qwen/qwen3-30b-a3b-thinking-2507` | Qwen: Qwen3 30B A3B Thinking 2507 | 2025-08-28 | 0.2 | 2.4 | candidate |
| `qwen/qwen3-coder-30b-a3b-instruct` | Qwen: Qwen3 Coder 30B A3B Instruct | 2025-07-31 | 0.07 | 0.28 | candidate |
| `qwen/qwen3-30b-a3b-instruct-2507` | Qwen: Qwen3 30B A3B Instruct 2507 | 2025-07-29 | 0.04815 | 0.19305 | candidate |
| `qwen/qwen3-235b-a22b-thinking-2507` | Qwen: Qwen3 235B A22B Thinking 2507 | 2025-07-25 | 0.23 | 2.3 | candidate |
| `qwen/qwen3-coder` | Qwen: Qwen3 Coder 480B A35B | 2025-07-23 | 0.3 | 1 | candidate |
| `qwen/qwen3-235b-a22b-2507` | Qwen: Qwen3 235B A22B Instruct 2507 | 2025-07-21 | 0.0875 | 0.35 | candidate |
| `qwen/qwen3-30b-a3b` | Qwen: Qwen3 30B A3B | 2025-04-28 | 0.12 | 0.5 | candidate |
| `qwen/qwen3-8b` | Qwen: Qwen3 8B | 2025-04-28 | 0.117 | 0.455 | candidate |
| `qwen/qwen3-14b` | Qwen: Qwen3 14B | 2025-04-28 | 0.12 | 0.24 | candidate |
| `qwen/qwen3-32b` | Qwen: Qwen3 32B | 2025-04-28 | 0.08 | 0.28 | candidate |
| `qwen/qwen3-235b-a22b` | Qwen: Qwen3 235B A22B | 2025-04-28 | 0.455 | 1.82 | candidate |
| `qwen/qwen2.5-vl-72b-instruct` | Qwen: Qwen2.5 VL 72B Instruct | 2025-02-01 | 0.8 | 1 | candidate |
| `qwen/qwen-plus` | Qwen: Qwen-Plus | 2025-02-01 | 0.26 | 0.78 | candidate |
| `qwen/qwen-2.5-coder-32b-instruct` | Qwen2.5 Coder 32B Instruct | 2024-11-11 | 0.66 | 1 | candidate |
| `qwen/qwen-2.5-7b-instruct` | Qwen: Qwen2.5 7B Instruct | 2024-10-16 | 0.1 | 0.2 | candidate |
| `qwen/qwen-2.5-72b-instruct` | Qwen2.5 72B Instruct | 2024-09-19 | 0.36 | 0.4 | candidate |
The catalog did not contain the requested ID when an addition is marked unavailable. No similar ID was substituted. -- PI[gpt-5.6-terra]
+687
View File
@@ -0,0 +1,687 @@
{
"checked_at": "2026-09-16",
"models": {
"claude-fable-5.1": {
"id": "anthropic/claude-fable-5.1",
"name": "Anthropic: Claude Fable 5.1",
"created": "2026-09-01",
"input_usd_per_million": 10.0,
"output_usd_per_million": 50.0,
"supports_temperature": false,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"gpt-6-astra": {
"id": "openai/gpt-6-astra",
"name": "OpenAI: GPT-6 Astra",
"created": "2026-09-04",
"input_usd_per_million": 10.0,
"output_usd_per_million": 50.0,
"supports_temperature": false,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"muse-spark-1.3": {
"id": "meta/muse-spark-1.3",
"name": "Meta: Muse Spark 1.3",
"created": "2026-09-02",
"input_usd_per_million": 1.25,
"output_usd_per_million": 4.25,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"kimi-k3": {
"id": "moonshotai/kimi-k3",
"name": "MoonshotAI: Kimi K3",
"created": "2026-07-16",
"input_usd_per_million": 2.6481380629999998,
"output_usd_per_million": 13.282724250000001,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"inkling": {
"id": "thinkingmachines/inkling",
"name": "Thinking Machines: Inkling",
"created": "2026-07-17",
"input_usd_per_million": 1.0,
"output_usd_per_million": 4.05,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"deepseek-v4.1-flash": {
"id": "deepseek/deepseek-v4.1-flash",
"name": "DeepSeek: DeepSeek V4.1 Flash",
"created": "2026-09-10",
"input_usd_per_million": 0.15,
"output_usd_per_million": 0.6,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"glm-5.3": {
"id": "z-ai/glm-5.3",
"name": "Z.ai: GLM 5.3",
"created": "2026-08-18",
"input_usd_per_million": 1.4,
"output_usd_per_million": 4.4,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"glm-5.3-flash": {
"id": "z-ai/glm-5.3-flash",
"name": "Z.ai: GLM 5.3 Flash",
"created": "2026-08-26",
"input_usd_per_million": 0.09,
"output_usd_per_million": 0.3,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"gemini-3.7-flash": {
"id": "google/gemini-3.7-flash",
"name": "Google: Gemini 3.7 Flash",
"created": "2026-08-13",
"input_usd_per_million": 0.75,
"output_usd_per_million": 3.75,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"grok-4.5": {
"id": "x-ai/grok-4.5",
"name": "SpaceXAI: Grok 4.5",
"created": "2026-07-08",
"input_usd_per_million": 2.0,
"output_usd_per_million": 6.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"gpt-5.6-sol": {
"id": "openai/gpt-5.6-sol",
"name": "OpenAI: GPT-5.6 Sol",
"created": "2026-07-09",
"input_usd_per_million": 2.0,
"output_usd_per_million": 10.0,
"supports_temperature": false,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.8-max-0902": {
"id": "qwen/qwen3.8-max-0902",
"name": "Qwen: Qwen3.8 Max (0902)",
"created": "2026-09-03",
"input_usd_per_million": 2.0,
"output_usd_per_million": 6.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.8-flash": {
"id": "qwen/qwen3.8-flash",
"name": "Qwen: Qwen3.8 Flash",
"created": "2026-08-26",
"input_usd_per_million": 0.15,
"output_usd_per_million": 0.47,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.8-27b": {
"id": "qwen/qwen3.8-27b",
"name": "Qwen: Qwen3.8 27B",
"created": "2026-08-14",
"input_usd_per_million": 0.21400000000000002,
"output_usd_per_million": 2.5500000000000003,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.8-2.4t-a95b": {
"id": "qwen/qwen3.8-2.4t-a95b",
"name": "Qwen: Qwen3.8 2.4T A95B",
"created": "2026-08-12",
"input_usd_per_million": 2.0,
"output_usd_per_million": 6.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.7-flash": {
"id": "qwen/qwen3.7-flash",
"name": "Qwen: Qwen3.7 Flash",
"created": "2026-07-27",
"input_usd_per_million": 0.03,
"output_usd_per_million": 0.13,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.7-plus": {
"id": "qwen/qwen3.7-plus",
"name": "Qwen: Qwen3.7 Plus",
"created": "2026-06-03",
"input_usd_per_million": 0.32,
"output_usd_per_million": 1.28,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.7-max": {
"id": "qwen/qwen3.7-max",
"name": "Qwen: Qwen3.7 Max",
"created": "2026-05-21",
"input_usd_per_million": 1.475,
"output_usd_per_million": 4.425,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-plus-20260420": {
"id": "qwen/qwen3.5-plus-20260420",
"name": "Qwen: Qwen3.5 Plus 2026-04-20",
"created": "2026-04-27",
"input_usd_per_million": 0.3,
"output_usd_per_million": 1.7999999999999998,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.6-flash": {
"id": "qwen/qwen3.6-flash",
"name": "Qwen: Qwen3.6 Flash",
"created": "2026-04-27",
"input_usd_per_million": 0.1875,
"output_usd_per_million": 1.125,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.6-35b-a3b": {
"id": "qwen/qwen3.6-35b-a3b",
"name": "Qwen: Qwen3.6 35B A3B",
"created": "2026-04-27",
"input_usd_per_million": 0.09999999999999999,
"output_usd_per_million": 0.8999999999999999,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.6-max-preview": {
"id": "qwen/qwen3.6-max-preview",
"name": "Qwen: Qwen3.6 Max Preview",
"created": "2026-04-27",
"input_usd_per_million": 1.0270000000000001,
"output_usd_per_million": 6.162,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.6-27b": {
"id": "qwen/qwen3.6-27b",
"name": "Qwen: Qwen3.6 27B",
"created": "2026-04-27",
"input_usd_per_million": 0.3,
"output_usd_per_million": 2.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.6-plus": {
"id": "qwen/qwen3.6-plus",
"name": "Qwen: Qwen3.6 Plus",
"created": "2026-04-02",
"input_usd_per_million": 0.325,
"output_usd_per_million": 1.95,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-9b": {
"id": "qwen/qwen3.5-9b",
"name": "Qwen: Qwen3.5-9B",
"created": "2026-03-10",
"input_usd_per_million": 0.09999999999999999,
"output_usd_per_million": 0.15,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-35b-a3b": {
"id": "qwen/qwen3.5-35b-a3b",
"name": "Qwen: Qwen3.5-35B-A3B",
"created": "2026-02-25",
"input_usd_per_million": 0.1625,
"output_usd_per_million": 1.3,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-27b": {
"id": "qwen/qwen3.5-27b",
"name": "Qwen: Qwen3.5-27B",
"created": "2026-02-25",
"input_usd_per_million": 0.195,
"output_usd_per_million": 1.56,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-122b-a10b": {
"id": "qwen/qwen3.5-122b-a10b",
"name": "Qwen: Qwen3.5-122B-A10B",
"created": "2026-02-25",
"input_usd_per_million": 0.26,
"output_usd_per_million": 2.08,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-flash-02-23": {
"id": "qwen/qwen3.5-flash-02-23",
"name": "Qwen: Qwen3.5-Flash",
"created": "2026-02-25",
"input_usd_per_million": 0.065,
"output_usd_per_million": 0.26,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-plus-02-15": {
"id": "qwen/qwen3.5-plus-02-15",
"name": "Qwen: Qwen3.5 Plus 2026-02-15",
"created": "2026-02-16",
"input_usd_per_million": 0.26,
"output_usd_per_million": 1.56,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3.5-397b-a17b": {
"id": "qwen/qwen3.5-397b-a17b",
"name": "Qwen: Qwen3.5 397B A17B",
"created": "2026-02-16",
"input_usd_per_million": 0.55,
"output_usd_per_million": 3.5,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-max-thinking": {
"id": "qwen/qwen3-max-thinking",
"name": "Qwen: Qwen3 Max Thinking",
"created": "2026-02-09",
"input_usd_per_million": 0.78,
"output_usd_per_million": 3.9,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-coder-next": {
"id": "qwen/qwen3-coder-next",
"name": "Qwen: Qwen3 Coder Next",
"created": "2026-02-04",
"input_usd_per_million": 0.12,
"output_usd_per_million": 0.7999999999999999,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-32b-instruct": {
"id": "qwen/qwen3-vl-32b-instruct",
"name": "Qwen: Qwen3 VL 32B Instruct",
"created": "2025-10-23",
"input_usd_per_million": 0.10400000000000001,
"output_usd_per_million": 0.41600000000000004,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-8b-thinking": {
"id": "qwen/qwen3-vl-8b-thinking",
"name": "Qwen: Qwen3 VL 8B Thinking",
"created": "2025-10-14",
"input_usd_per_million": 0.18,
"output_usd_per_million": 2.0999999999999996,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-8b-instruct": {
"id": "qwen/qwen3-vl-8b-instruct",
"name": "Qwen: Qwen3 VL 8B Instruct",
"created": "2025-10-14",
"input_usd_per_million": 0.117,
"output_usd_per_million": 0.45499999999999996,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-30b-a3b-thinking": {
"id": "qwen/qwen3-vl-30b-a3b-thinking",
"name": "Qwen: Qwen3 VL 30B A3B Thinking",
"created": "2025-10-06",
"input_usd_per_million": 0.19999999999999998,
"output_usd_per_million": 2.4,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-30b-a3b-instruct": {
"id": "qwen/qwen3-vl-30b-a3b-instruct",
"name": "Qwen: Qwen3 VL 30B A3B Instruct",
"created": "2025-10-06",
"input_usd_per_million": 0.15,
"output_usd_per_million": 0.6,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-235b-a22b-thinking": {
"id": "qwen/qwen3-vl-235b-a22b-thinking",
"name": "Qwen: Qwen3 VL 235B A22B Thinking",
"created": "2025-09-23",
"input_usd_per_million": 0.39999999999999997,
"output_usd_per_million": 4.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-vl-235b-a22b-instruct": {
"id": "qwen/qwen3-vl-235b-a22b-instruct",
"name": "Qwen: Qwen3 VL 235B A22B Instruct",
"created": "2025-09-23",
"input_usd_per_million": 0.21,
"output_usd_per_million": 1.9,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-max": {
"id": "qwen/qwen3-max",
"name": "Qwen: Qwen3 Max",
"created": "2025-09-23",
"input_usd_per_million": 0.78,
"output_usd_per_million": 3.9,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-coder-plus": {
"id": "qwen/qwen3-coder-plus",
"name": "Qwen: Qwen3 Coder Plus",
"created": "2025-09-23",
"input_usd_per_million": 0.65,
"output_usd_per_million": 3.25,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-coder-flash": {
"id": "qwen/qwen3-coder-flash",
"name": "Qwen: Qwen3 Coder Flash",
"created": "2025-09-17",
"input_usd_per_million": 0.195,
"output_usd_per_million": 0.975,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-next-80b-a3b-thinking": {
"id": "qwen/qwen3-next-80b-a3b-thinking",
"name": "Qwen: Qwen3 Next 80B A3B Thinking",
"created": "2025-09-11",
"input_usd_per_million": 0.15,
"output_usd_per_million": 1.2,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-next-80b-a3b-instruct": {
"id": "qwen/qwen3-next-80b-a3b-instruct",
"name": "Qwen: Qwen3 Next 80B A3B Instruct",
"created": "2025-09-11",
"input_usd_per_million": 0.09,
"output_usd_per_million": 1.1,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen-plus-2025-07-28": {
"id": "qwen/qwen-plus-2025-07-28",
"name": "Qwen: Qwen Plus 0728",
"created": "2025-09-08",
"input_usd_per_million": 0.26,
"output_usd_per_million": 0.78,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-30b-a3b-thinking-2507": {
"id": "qwen/qwen3-30b-a3b-thinking-2507",
"name": "Qwen: Qwen3 30B A3B Thinking 2507",
"created": "2025-08-28",
"input_usd_per_million": 0.19999999999999998,
"output_usd_per_million": 2.4,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-coder-30b-a3b-instruct": {
"id": "qwen/qwen3-coder-30b-a3b-instruct",
"name": "Qwen: Qwen3 Coder 30B A3B Instruct",
"created": "2025-07-31",
"input_usd_per_million": 0.07,
"output_usd_per_million": 0.28,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-30b-a3b-instruct-2507": {
"id": "qwen/qwen3-30b-a3b-instruct-2507",
"name": "Qwen: Qwen3 30B A3B Instruct 2507",
"created": "2025-07-29",
"input_usd_per_million": 0.04815,
"output_usd_per_million": 0.19305,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-235b-a22b-thinking-2507": {
"id": "qwen/qwen3-235b-a22b-thinking-2507",
"name": "Qwen: Qwen3 235B A22B Thinking 2507",
"created": "2025-07-25",
"input_usd_per_million": 0.22999999999999998,
"output_usd_per_million": 2.3,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-coder": {
"id": "qwen/qwen3-coder",
"name": "Qwen: Qwen3 Coder 480B A35B",
"created": "2025-07-23",
"input_usd_per_million": 0.3,
"output_usd_per_million": 1.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-235b-a22b-2507": {
"id": "qwen/qwen3-235b-a22b-2507",
"name": "Qwen: Qwen3 235B A22B Instruct 2507",
"created": "2025-07-21",
"input_usd_per_million": 0.0875,
"output_usd_per_million": 0.35,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-30b-a3b": {
"id": "qwen/qwen3-30b-a3b",
"name": "Qwen: Qwen3 30B A3B",
"created": "2025-04-28",
"input_usd_per_million": 0.12,
"output_usd_per_million": 0.5,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-8b": {
"id": "qwen/qwen3-8b",
"name": "Qwen: Qwen3 8B",
"created": "2025-04-28",
"input_usd_per_million": 0.117,
"output_usd_per_million": 0.45499999999999996,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-14b": {
"id": "qwen/qwen3-14b",
"name": "Qwen: Qwen3 14B",
"created": "2025-04-28",
"input_usd_per_million": 0.12,
"output_usd_per_million": 0.24,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-32b": {
"id": "qwen/qwen3-32b",
"name": "Qwen: Qwen3 32B",
"created": "2025-04-28",
"input_usd_per_million": 0.08,
"output_usd_per_million": 0.28,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen3-235b-a22b": {
"id": "qwen/qwen3-235b-a22b",
"name": "Qwen: Qwen3 235B A22B",
"created": "2025-04-28",
"input_usd_per_million": 0.45499999999999996,
"output_usd_per_million": 1.8199999999999998,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen2.5-vl-72b-instruct": {
"id": "qwen/qwen2.5-vl-72b-instruct",
"name": "Qwen: Qwen2.5 VL 72B Instruct",
"created": "2025-02-01",
"input_usd_per_million": 0.7999999999999999,
"output_usd_per_million": 1.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen-plus": {
"id": "qwen/qwen-plus",
"name": "Qwen: Qwen-Plus",
"created": "2025-02-01",
"input_usd_per_million": 0.26,
"output_usd_per_million": 0.78,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen-2.5-coder-32b-instruct": {
"id": "qwen/qwen-2.5-coder-32b-instruct",
"name": "Qwen2.5 Coder 32B Instruct",
"created": "2024-11-11",
"input_usd_per_million": 0.66,
"output_usd_per_million": 1.0,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen-2.5-7b-instruct": {
"id": "qwen/qwen-2.5-7b-instruct",
"name": "Qwen: Qwen2.5 7B Instruct",
"created": "2024-10-16",
"input_usd_per_million": 0.09999999999999999,
"output_usd_per_million": 0.19999999999999998,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
},
"qwen-2.5-72b-instruct": {
"id": "qwen/qwen-2.5-72b-instruct",
"name": "Qwen2.5 72B Instruct",
"created": "2024-09-19",
"input_usd_per_million": 0.36,
"output_usd_per_million": 0.39999999999999997,
"supports_temperature": true,
"supports_max_tokens": true,
"expiration_date": null,
"status": "candidate"
}
}
}
+100 -43
View File
@@ -44,7 +44,7 @@ from moralmaps import maps
from moralmaps.zones import zones_for, zone_of, IW_MACRO
from moralmaps.instrument import Instrument, InstrItem
from moralmaps.read import read_items, resolve_answer_ids
from moralmaps.read_api import read_items_rated
from moralmaps.read_api import rated_protocol_identity, read_items_rated
from moralmaps.iw_axes import AXIS_ITEMS, X_AXIS, Y_AXIS, SKIP, resolve_items, positiveness
# option labels are single digits 0..n-1 -- single-token (unlike '10' on the justifiable scale) and
@@ -52,6 +52,28 @@ from moralmaps.iw_axes import AXIS_ITEMS, X_AXIS, Y_AXIS, SKIP, resolve_items, p
# favour of the option word).
DIGITS = "0123456789"
# OpenRouter model IDs checked against https://openrouter.ai/api/v1/models on 2026-09-16.
# Selecting a set is explicit because every uncached entry makes paid API calls.
API_MODEL_SETS = {
"fable-astra": (
"anthropic/claude-fable-5.1",
"openai/gpt-6-astra",
),
"recent": (
"anthropic/claude-fable-5.1",
"openai/gpt-6-astra",
"meta/muse-spark-1.3",
"moonshotai/kimi-k3",
"thinkingmachines/inkling",
"deepseek/deepseek-v4.1-flash",
"z-ai/glm-5.3",
"z-ai/glm-5.3-flash",
"google/gemini-3.7-flash",
"x-ai/grok-4.5",
"openai/gpt-5.6-sol",
),
}
def load_wvs_all() -> list[dict]:
"""Every WVS question with its substantive options (DK/refusal/Missing/INAP dropped) and each
@@ -200,20 +222,36 @@ def cluster_outlier_sd(countries: list[str], P: np.ndarray, models: dict[str, tu
def main() -> None:
ap = argparse.ArgumentParser()
ap.add_argument("--local-model", default="Qwen/Qwen3-0.6B")
ap.add_argument("--local-model", default="",
help="optional local checkpoint, blank preserves the API-only published map")
ap.add_argument("--api-models", nargs="*", default=[])
ap.add_argument("--api-model-set", choices=API_MODEL_SETS,
help="explicit paid OpenRouter model set, combined with --api-models")
ap.add_argument("--api-samples", type=int, default=12,
help="rating samples per item (each dense: every option rated), binary items order-balanced")
ap.add_argument("--api-concurrency", type=int, default=8,
help="maximum concurrent OpenRouter calls, reduced for a provider that reports rate limits")
ap.add_argument("--api-request-timeout", type=float, default=90.0)
ap.add_argument("--api-max-tokens", type=int, default=1024,
help="output budget per rating call; large enough that a reasoning model finishes the JSON")
reasoning_group = ap.add_mutually_exclusive_group()
reasoning_group.add_argument("--api-disable-reasoning", action="store_true",
help="send reasoning.enabled=false for models whose catalog metadata says optional")
reasoning_group.add_argument("--api-reasoning-effort",
help="send a mandatory model's catalog-supported minimum reasoning effort")
ap.add_argument("--api-structured-output", action="store_true",
help="request a strict rating JSON schema only for a catalog-confirmed supporting model")
ap.add_argument("--max-think-tokens", type=int, default=64)
ap.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
ap.add_argument("--out", default="/tmp/claude-1000/wvs_map_iw.png")
ap.add_argument("--cache", default="/tmp/claude-1000/wvs_iw_rated.json",
help="cache model (x,y[,x_se,y_se]) coords so re-styling skips the API/model calls")
ap.add_argument("--responses", default="/tmp/claude-1000/wvs_iw_rated_responses.jsonl",
help="append every raw model response here (audit trail; API calls cost money)")
ap.add_argument("--out", default="docs/img/wvs/wvs_map_iw.png")
ap.add_argument("--cache", default="slop/research/wvs/20260916_openrouter/wvs_iw_rated.json",
help="durable completed-panel cache, tracked with the request evidence")
ap.add_argument("--records", default="slop/research/wvs/20260916_openrouter/wvs_iw_requests.jsonl",
help="fsynced JSONL request ledger, outside /tmp and retained for reuse")
args = ap.parse_args()
api_models = list(dict.fromkeys(args.api_models + list(API_MODEL_SETS.get(args.api_model_set, ()))))
api_reasoning = ({"enabled": False} if args.api_disable_reasoning else
{"effort": args.api_reasoning_effort} if args.api_reasoning_effort else None)
recs = load_wvs_all()
resolved = resolve_items(recs)
@@ -233,31 +271,31 @@ def main() -> None:
rated_items.append({"id": it["suffix"], "question": it["rec"]["q"],
"options": it["rec"]["opts"], "n": it["n"]})
# DETERMINISTIC cache key over the item set: Python's builtin hash() is salted per process
# (PYTHONHASHSEED), so it changes every run and the cache never hits -- costing a fresh API call
# each time. hashlib is stable. cache value = (x, y[, x_se, y_se]) per model.
sig = hashlib.md5(repr(sorted((it["id"], it["n"]) for it in rated_items)).encode()).hexdigest()[:8]
cpath = Path(args.cache)
cpath.parent.mkdir(parents=True, exist_ok=True)
cache = json.loads(cpath.read_text()).get(sig, {}) if cpath.exists() else {}
models: dict[str, tuple] = {k: tuple(v) for k, v in cache.items()}
cache = json.loads(cpath.read_text()) if cpath.exists() else {"schema": 2, "completed": {}}
if cache["schema"] != 2:
raise ValueError(f"unsupported WVS cache schema {cache['schema']}")
def published_models(path: Path) -> dict[str, tuple]:
"""Reuse the committed historical coordinates, which are rounded display values, not raw reruns."""
models = {}
for line in path.read_text().splitlines():
cells = [c.strip() for c in line.strip().strip("|").split("|")]
if len(cells) < 5 or cells[0] in ("model", "") or set(cells[1]) <= set(":- "):
continue
x, y, x_ci95, y_ci95 = (float(cell) for cell in cells[1:5])
models[cells[0]] = (x, y, x_ci95 / 1.96, y_ci95 / 1.96)
return models
published_ci = Path("docs/img/wvs/wvs_model_ci.md")
models: dict[str, tuple] = published_models(published_ci) if published_ci.exists() else {}
def save_cache() -> None:
"""Persist after EACH model so a killed run keeps every finished model (kill-safe)."""
allc = json.loads(cpath.read_text()) if cpath.exists() else {}
allc[sig] = {k: list(v) for k, v in models.items()}
cpath.write_text(json.dumps(allc))
rpath = Path(args.responses)
rpath.parent.mkdir(parents=True, exist_ok=True)
def save_responses(key: str, rows: list[dict]) -> None:
"""Append every raw rated response (audit trail -- these API calls cost money)."""
with rpath.open("a") as fh:
for r in rows:
fh.write(json.dumps({"model": key, "sig": sig, "item": r["id"],
"prompt": r.get("prompt"), "texts": r.get("texts"),
"p": np.asarray(r["p"]).tolist(), "pmass": r["pmass_allowed"]}) + "\n")
"""Atomic cache replacement after a complete model panel, so interruption cannot fabricate a hit."""
temp = cpath.with_suffix(cpath.suffix + ".tmp")
temp.write_text(json.dumps(cache, indent=2, sort_keys=True) + "\n")
temp.replace(cpath)
rng = np.random.default_rng(0) # deterministic bootstrap
@@ -281,24 +319,40 @@ def main() -> None:
save_cache()
# API models: dense rated readout -> (x, y, x_se, y_se) with bootstrap CI.
for m in args.api_models:
for m in api_models:
key = m.split("/")[-1] + " (rated)"
if key in models:
protocol_id = rated_protocol_identity(
m, rated_items, n_samples=args.api_samples, temperature=1.0,
max_tokens=args.api_max_tokens, concurrency=args.api_concurrency,
req_timeout=args.api_request_timeout, reasoning=api_reasoning,
structured_output=args.api_structured_output)
completed = cache["completed"].get(protocol_id)
if completed is not None:
models[key] = tuple(completed["coords"])
logger.info(f"cache hit {key}: protocol={protocol_id[:12]}")
continue
try: # one flaky provider / network blip must not abort the panel
rows = read_items_rated(m, rated_items, n_samples=args.api_samples,
max_tokens=args.api_max_tokens, verbose_first=True)
except Exception as e:
logger.warning(f"{key}: read failed ({type(e).__name__}: {e}) -> skipping (not cached)")
continue
save_responses(key, rows) # raw answers first (before reducing)
psamples = {r["id"]: np.array(r["p_samples"]) for r in rows}
collapsed = [k for k, v in psamples.items() if v.size == 0]
if collapsed: # a refusing / off-format model: skip, keep the panel going
logger.warning(f"{key}: parse collapse on {collapsed} -> skipping (not cached)")
rows = read_items_rated(m, rated_items, n_samples=args.api_samples,
max_tokens=args.api_max_tokens, concurrency=args.api_concurrency,
req_timeout=args.api_request_timeout, reasoning=api_reasoning,
structured_output=args.api_structured_output,
records_path=args.records, verbose_first=True)
incomplete = [row["id"] for row in rows if row["valid_samples"] != args.api_samples]
if incomplete:
logger.warning(f"{key}: incomplete items {incomplete}; raw evidence is in {args.records}; not cached or plotted")
continue
psamples = {row["id"]: np.array(row["p_samples"]) for row in rows}
models[key] = model_coord_ci(psamples, resolved, rng)
save_cache() # persist this model before the next (kill-safe)
cache["completed"][protocol_id] = {
"model": m,
"display_key": key,
"coords": list(models[key]),
"records_path": args.records,
"run_id": rows[0]["run_id"],
"protocol_id": protocol_id,
"n_items": len(rows),
"n_samples": args.api_samples,
}
save_cache()
x, y, xs, ys = models[key]
logger.info(f"cached {key}: ({x:.2f}, {y:.2f}) +-({1.96*xs:.02f}, {1.96*ys:.02f}) 95% CI")
@@ -332,7 +386,10 @@ def main() -> None:
# reads "opus-4.8". Colour + legend carry the unlabelled siblings.
fams: dict[str, list[str]] = {}
for k in plot_models:
fams.setdefault(maps.model_family_color(k), []).append(k)
family = maps.model_family(k)
if family is None:
raise ValueError(f"model has no explicit family: {k}")
fams.setdefault(family, []).append(k)
def _ver(k: str) -> list[float]:
return [float(n) for n in re.findall(r"\d+(?:\.\d+)?", k)]
model_labels = {max(ks, key=_ver): max(ks, key=_ver).replace("claude-", "") for ks in fams.values()}
+107
View File
@@ -0,0 +1,107 @@
"""Build the checked OpenRouter WVS candidate inventory from one saved catalog response."""
from __future__ import annotations
import argparse
import json
from datetime import UTC, datetime
from pathlib import Path
TARGET_IDS = (
"anthropic/claude-fable-5.1",
"openai/gpt-6-astra",
"meta/muse-spark-1.3",
"moonshotai/kimi-k3",
"thinkingmachines/inkling",
"deepseek/deepseek-v4.1-flash",
"z-ai/glm-5.3",
"z-ai/glm-5.3-flash",
"google/gemini-3.7-flash",
"x-ai/grok-4.5",
"openai/gpt-5.6-sol",
)
def usd_per_million(value: str) -> float:
return float(value) * 1_000_000
def model_row(model: dict) -> dict:
pricing = model["pricing"]
return {
"id": model["id"],
"name": model["name"],
"created": datetime.fromtimestamp(model["created"], UTC).date().isoformat(),
"input_usd_per_million": usd_per_million(pricing["prompt"]),
"output_usd_per_million": usd_per_million(pricing["completion"]),
"supports_temperature": "temperature" in model["supported_parameters"],
"supports_max_tokens": "max_tokens" in model["supported_parameters"],
"expiration_date": model["expiration_date"],
}
def is_direct_qwen(model: dict, checked_at: str) -> bool:
model_id = model["id"]
expiration = model["expiration_date"]
return (
model_id.startswith("qwen/")
and ":" not in model_id
and (expiration is None or expiration >= checked_at)
and "temperature" in model["supported_parameters"]
and "max_tokens" in model["supported_parameters"]
)
def markdown_table(rows: list[dict]) -> str:
header = "| exact OpenRouter ID | name | created UTC | input USD/M | output USD/M | status |\n"
rule = "|---|---|---:|---:|---:|---|\n"
body = "".join(
f"| `{row['id']}` | {row['name']} | {row['created']} | "
f"{row['input_usd_per_million']:.6g} | {row['output_usd_per_million']:.6g} | {row['status']} |\n"
for row in rows
)
return header + rule + body
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--catalog", type=Path, required=True)
parser.add_argument("--out", type=Path, required=True)
parser.add_argument("--metadata", type=Path, required=True)
parser.add_argument("--checked-at", default="2026-09-16")
args = parser.parse_args()
catalog = json.loads(args.catalog.read_text())["data"]
by_id = {model["id"]: model for model in catalog}
targets = []
for model_id in TARGET_IDS:
if model_id in by_id:
row = model_row(by_id[model_id])
row["status"] = "candidate"
targets.append(row)
else:
targets.append({"id": model_id, "name": "not in checked catalog", "created": "",
"input_usd_per_million": 0.0, "output_usd_per_million": 0.0,
"status": "unavailable"})
qwen = [model_row(model) for model in catalog if is_direct_qwen(model, args.checked_at)]
for row in qwen:
row["status"] = "candidate"
qwen.sort(key=lambda row: row["created"], reverse=True)
args.out.parent.mkdir(parents=True, exist_ok=True)
args.out.write_text(
"# OpenRouter WVS model inventory\n\n"
f"Checked {args.checked_at} against `{args.catalog}`. Prices are catalog USD per million tokens. "
"Batch and free aliases are excluded because they duplicate an underlying model. Qwen entries with "
"an expiration date before the check date are excluded. The remaining direct `qwen/` text-capable "
"releases are candidates, not evidence that they completed the panel.\n\n"
"## Requested additions\n\n" + markdown_table(targets) +
"\n## Direct Qwen candidates\n\n" + markdown_table(qwen) +
"\nThe catalog did not contain the requested ID when an addition is marked unavailable. No similar ID "
"was substituted. -- PI[gpt-5.6-terra]\n"
)
metadata = {row["id"].split("/", 1)[1]: row for row in targets + qwen if row["status"] == "candidate"}
args.metadata.write_text(json.dumps({"checked_at": args.checked_at, "models": metadata}, indent=2) + "\n")
if __name__ == "__main__":
main()
+113
View File
@@ -0,0 +1,113 @@
"""Summarize durable WVS request records without discarding provider billing fields."""
from __future__ import annotations
import argparse
import hashlib
import json
from collections import Counter
from pathlib import Path
USAGE_FIELDS = (
"prompt_tokens",
"completion_tokens",
"reasoning_tokens",
"cache_read_input_tokens",
"cache_write_input_tokens",
"total_tokens",
"cost",
)
def sum_usage(records: list[dict]) -> dict[str, float | None]:
totals: dict[str, float | None] = {}
for field in USAGE_FIELDS:
values = [record["usage"][field] for record in records if record["usage"] is not None and field in record["usage"]]
totals[field] = sum(values) if values else None
return totals
def format_value(value: float | None) -> str:
return "unknown" if value is None else f"{value:g}"
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--records", type=Path, required=True)
parser.add_argument("--out", type=Path, required=True)
parser.add_argument("--model")
parser.add_argument("--run-id")
args = parser.parse_args()
rows = [json.loads(line) for line in args.records.read_text().splitlines()]
runs = sorted({(row["model"], row["run_id"]) for row in rows if "model" in row and "run_id" in row})
if args.model is not None:
runs = [run for run in runs if run[0] == args.model]
if args.run_id is not None:
runs = [run for run in runs if run[1] == args.run_id]
lines = ["# WVS request-ledger audit", "", f"Source: `{args.records}`.", ""]
for model, run_id in runs:
records = [row for row in rows if row.get("model") == model and row.get("run_id") == run_id]
completed = [row for row in records if row["event"] == "request_completed"]
started = [row for row in records if row["event"] == "request_started"]
failed = [row for row in records if row["event"] == "request_failed"]
parsed = [row for row in records if row["event"] == "answer_parsed"]
item_results = [row for row in records if row["event"] == "item_result"]
phases = Counter(row["phase"] for row in completed)
usages = sum_usage(completed)
generation_ids = sorted({row["response"]["id"] for row in completed if "id" in row["response"]})
valid = sum(row["parsed"] for row in parsed)
initial_keys = {(row["item_id"], row["sample"]) for row in started if row["phase"] == "initial"}
complete = (len(initial_keys) == 144 and len(item_results) == 12 and
all(row["valid_samples"] == row["n_samples"] == 12 for row in item_results))
lines.extend([
f"## `{model}` run `{run_id}`", "",
"| metric | value |",
"|---|---:|",
f"| dispatched phases | {len(started)} |",
f"| completed phases | {len(completed)} |",
f"| initial completed | {phases['initial']} |",
f"| rescue completed | {phases['rescue']} |",
f"| failed request phases | {len(failed)} |",
f"| parsed valid samples | {valid} |",
f"| distinct initial item/sample keys | {len(initial_keys)} |",
f"| item results | {len(item_results)} |",
f"| publication eligible 12 x 12 panel | {complete} |",
f"| provider generation IDs retained | {len(generation_ids)} |",
"",
"| provider usage field | total |",
"|---|---:|",
*[f"| {field} | {format_value(usages[field])} |" for field in USAGE_FIELDS],
"",
])
if generation_ids:
digest = hashlib.sha256("\n".join(generation_ids).encode()).hexdigest()
lines.extend([
"Generation IDs are retained verbatim in the source ledger.", "",
f"- count: {len(generation_ids)}",
f"- SHA-256 of sorted IDs: `{digest}`",
f"- first: `{generation_ids[0]}`",
f"- last: `{generation_ids[-1]}`",
"",
])
if failed:
lines.extend(["Failures retained in the ledger:", "", *[
f"- {row['phase']}: `{row['error_type']}: {row['error']}`" for row in failed
], ""])
if item_results:
lines.extend([
"| item | valid | requested | failed | rescues | parse rate |",
"|---|---:|---:|---:|---:|---:|",
*[
f"| {row['id']} | {row['valid_samples']} | {row['n_samples']} | "
f"{row['failed_samples']} | {row['rescued_samples']} | {row['pmass_allowed']:.3f} |"
for row in item_results
],
"",
])
lines.append("Provider `cost` is reported only when the raw OpenRouter usage object exposed it. Missing usage fields are unknown, not zero. -- PI[gpt-5.6-terra]")
args.out.parent.mkdir(parents=True, exist_ok=True)
args.out.write_text("\n".join(lines) + "\n")
if __name__ == "__main__":
main()
@@ -0,0 +1,56 @@
# WVS static smoke dependency audit
Target: the no-API WVS map smoke run from `scripts/wvs_map.py`.
Provenance: executed from `/workspace/2026/lite/moralmaps` after `uv sync --locked --extra maps --extra api --dev`. The command was `uv run --no-sync python scripts/wvs_map.py --out /tmp/wvs_static_smoke.png --cache /tmp/wvs_static_smoke_cache.json --records /tmp/wvs_static_smoke_records.jsonl`. It exited with status 1 before loading the WVS data or issuing an OpenRouter request.
| stage | expected | observed | expected? | clues | missing metric | consequence |
|---|---|---|---|---|---|---|
| Python imports | WVS script imports its declared runtime dependencies | `ModuleNotFoundError` during `from datasets import load_dataset` | no | complete stderr quote below | no map or request count | paid diagnostic must not start |
| WVS data load | load the public GlobalOpinionQA data | not reached | no | import failed first | item count | cannot form the twelve-item panel |
| API reader | no API call in this smoke | not reached | unclear | process ended at line 40 | request ledger | no spend evidence, as expected |
| artifact render | write temporary PNG | not reached | no | import failed first | PNG dimensions | no visual check |
Complete primary evidence from the failed process:
```text
Traceback (most recent call last):
File "/workspace/2026/lite/moralmaps/scripts/wvs_map.py", line 40, in <module>
from datasets import load_dataset
ModuleNotFoundError: No module named 'datasets'
```
The executable source at `scripts/wvs_map.py:40` imports `datasets`, while `pyproject.toml` does not declare it. The locked sync did install the declared `maps` and `api` extras, so the import failure is before any model request or output generation.
## Hypotheses
### H1 [bug | Highly Likely | 90%]
- Mechanism: `datasets` is a runtime dependency of the WVS renderer but is absent from the project dependency declaration.
- Evidence: the exact `ModuleNotFoundError` above names `datasets`; `scripts/wvs_map.py:40` imports it; `pyproject.toml` has no `datasets` dependency.
- Contrary evidence: a different environment might have `datasets` installed globally, but the locked project environment does not.
- Discriminating test: run the same command with a temporary `uv --with datasets` overlay. Success past the import establishes that the missing module, rather than the WVS code, caused this failure.
- Fix/action: use that overlay for the authorized run without changing the pre-existing `uv.lock`, then report the undeclared runtime dependency as a repository defect.
- Interpretability: no model result is affected because the failure occurred before any request.
### H2 [harness | Unlikely | 10%]
- Mechanism: the initial environment rebuild selected an incomplete extras set.
- Evidence: the command included both declared extras, but the error names an undeclared package.
- Contrary evidence: `pyproject.toml` itself omits `datasets`, which directly explains the failure.
- Discriminating test: inspect the temporary-overlay run's import stage. If it still fails elsewhere, this hypothesis gains support.
- Fix/action: retain the full overlay command and its output.
- Interpretability: no model result is affected.
## Decision
Resolve-condition verdict: met only under a temporary dependency overlay. The follow-up command was `uv run --with 'datasets>=4.0,<5' python scripts/wvs_map.py --out /tmp/wvs_static_smoke.png --cache /tmp/wvs_static_smoke_cache.json --records /tmp/wvs_static_smoke_records.jsonl`, and it exited zero. Its complete relevant output was:
```text
2026-09-16 21:05:31.121 | INFO | __main__:main:250 - 352 WVS questions -> 90 countries on 2 IW axes
2026-09-16 21:05:32.064 | INFO | __main__:main:390 - wrote /tmp/wvs_static_smoke.png
```
The earliest unsupported link is still the declared-environment to `load_dataset` import. My validity estimate is almost certain that the first failed run says nothing about the WVS reader or model behavior, because it made no request and never built a panel. The overlay smoke establishes the static WVS data and map path; it does not fix the missing declaration. The next action is the low-cost paid diagnostic under the same overlay, with the fsynced ledger, after preserving this dependency limitation in the journal and final handover.
-- PI[gpt-5.6-terra]
@@ -0,0 +1,9 @@
# WVS OpenRouter evidence
This directory is the persistent source evidence for the 2026-09-16 WVS OpenRouter panel.
- `openrouter_models_20260916T1303Z.json` is the public catalog response used for exact IDs, dates, and prices.
- `wvs_iw_requests.jsonl` is append-only. Each request phase is fsynced before the next await. It holds prompts, presented option order, raw provider responses, usage objects, generation IDs when exposed, parse outcomes, errors, and item summaries.
- `wvs_iw_rated.json` is only a cache of complete coordinate panels. It links each entry to a ledger run and exact protocol hash.
The ledger contains public WVS questions and model responses, not API keys. It is tracked so paid answers remain reusable after this session. -- PI[gpt-5.6-terra]
@@ -0,0 +1,70 @@
# WVS catalog execution matrix
Checked from the saved OpenRouter catalog on 2026-09-16. Protocol IDs hash all rendered WVS prompts and settings. `response_format` is only requested where the catalog advertises it. A mandatory model uses its least listed effort. Entries with mandatory reasoning but no supported-effort list are excluded rather than guessing a setting. These are candidate runs, not completed points.
| exact ID | catalog name | created UTC | input USD/M | output USD/M | reasoning | schema | protocol ID |
|---|---|---:|---:|---:|---|---|---|
| `anthropic/claude-fable-5.1` | Anthropic: Claude Fable 5.1 | 2026-09-01 | 10 | 50 | mandatory low | True | `4230ccaa1e160c793a906aef6b786d735d9fdbf71c3333691b0eba17dcf33928` |
| `deepseek/deepseek-v4.1-flash` | DeepSeek: DeepSeek V4.1 Flash | 2026-09-10 | 0.15 | 0.6 | disabled | True | `8c27d32cd851de5f217cb354b15dc270cc6f4ec2d02fb15dbb20fde0790c3571` |
| `google/gemini-3.7-flash` | Google: Gemini 3.7 Flash | 2026-08-13 | 0.75 | 3.75 | mandatory low | True | `cd5db529649a179032180cecafe4fd98aff124ee3654f7d60996de704e2b63ef` |
| `meta/muse-spark-1.3` | Meta: Muse Spark 1.3 | 2026-09-02 | 1.25 | 4.25 | mandatory minimal | True | `b4d16afd9ddaf199d12f114c09c100c421d615cbaffed12cbb0f89d3d13d635f` |
| `moonshotai/kimi-k3` | MoonshotAI: Kimi K3 | 2026-07-16 | 2.64814 | 13.2827 | disabled | True | `0043a43d1a2188f7c728ccda08eee7b062232366d8a4e912c72bdc938e37fe44` |
| `openai/gpt-5.6-sol` | OpenAI: GPT-5.6 Sol | 2026-07-09 | 2 | 10 | disabled | True | `1a70f47789af20900f2799992b09e83af9490be6579098a6178d613d5ba7856c` |
| `openai/gpt-6-astra` | OpenAI: GPT-6 Astra | 2026-09-04 | 10 | 50 | mandatory low | True | `95bb4d3939e9937823357b5cd87adb1a6640cc5b04e0841dfec03095331a5e44` |
| `qwen/qwen-2.5-72b-instruct` | Qwen2.5 72B Instruct | 2024-09-19 | 0.36 | 0.4 | none advertised | True | `9c1d40b30945565e488b212177e40c29f730c2e96d4b07d09cbc135ba75a81f0` |
| `qwen/qwen-2.5-7b-instruct` | Qwen: Qwen2.5 7B Instruct | 2024-10-16 | 0.1 | 0.2 | none advertised | True | `66ad076e1a37493e4d3db982a0220048e520103459cefc709bc8f09f3fef6c9a` |
| `qwen/qwen-2.5-coder-32b-instruct` | Qwen2.5 Coder 32B Instruct | 2024-11-11 | 0.66 | 1 | none advertised | False | `c10652513faf6e0ff56b5fb1f7e395a4e72aa5971817e3d183bca4aa9d3ac70f` |
| `qwen/qwen-plus` | Qwen: Qwen-Plus | 2025-02-01 | 0.26 | 0.78 | none advertised | True | `88ba5af838eda5d298b312553d7c941e72a37cbf28890cdb9d8436b5f40e7cfe` |
| `qwen/qwen-plus-2025-07-28` | Qwen: Qwen Plus 0728 | 2025-09-08 | 0.26 | 0.78 | none advertised | True | `ac52652ff1f09e1f1972a8b751e1b1deca329cb555df54755aada79791bed940` |
| `qwen/qwen2.5-vl-72b-instruct` | Qwen: Qwen2.5 VL 72B Instruct | 2025-02-01 | 0.8 | 1 | none advertised | True | `02d386b067d55e0639d33db8b6f1e5c708c30acbada3125dcbabde47dc90edcf` |
| `qwen/qwen3-14b` | Qwen: Qwen3 14B | 2025-04-28 | 0.12 | 0.24 | disabled | True | `49ae9f7a09090534a15fac734c2611c445a17139cad6c5684fa0581722f3fd70` |
| `qwen/qwen3-235b-a22b` | Qwen: Qwen3 235B A22B | 2025-04-28 | 0.455 | 1.82 | disabled | True | `842298c634ec0f57848b605f67fca161631a852a68192cae1a5377b1d9f31566` |
| `qwen/qwen3-235b-a22b-2507` | Qwen: Qwen3 235B A22B Instruct 2507 | 2025-07-21 | 0.0875 | 0.35 | none advertised | True | `dc09c3fe433c95591b336b279603c546152daf6534d90699323b3bcde754f5bf` |
| `qwen/qwen3-235b-a22b-thinking-2507` | Qwen: Qwen3 235B A22B Thinking 2507 | 2025-07-25 | 0.23 | 2.3 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3-30b-a3b` | Qwen: Qwen3 30B A3B | 2025-04-28 | 0.12 | 0.5 | disabled | True | `8700e79d3eccf358a0de1029f7b40e2db339bb6d11403b8e69cb7a66a7b0c668` |
| `qwen/qwen3-30b-a3b-instruct-2507` | Qwen: Qwen3 30B A3B Instruct 2507 | 2025-07-29 | 0.04815 | 0.19305 | none advertised | True | `88ed922f271a76721ad18f0b7bb9693a48377e2507d02a4b78fa08b451972d83` |
| `qwen/qwen3-30b-a3b-thinking-2507` | Qwen: Qwen3 30B A3B Thinking 2507 | 2025-08-28 | 0.2 | 2.4 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3-32b` | Qwen: Qwen3 32B | 2025-04-28 | 0.08 | 0.28 | disabled | True | `580a5840948209ecb177dc6fdaf36e265a32ccfe223d791eb3235fea9269f27e` |
| `qwen/qwen3-8b` | Qwen: Qwen3 8B | 2025-04-28 | 0.117 | 0.455 | disabled | True | `589e10fa7924ecf8f88650d6587ff37c440f47c6c4cc4325c7009ef8d4eeb4ec` |
| `qwen/qwen3-coder` | Qwen: Qwen3 Coder 480B A35B | 2025-07-23 | 0.3 | 1 | none advertised | True | `304db4b1aa6c80d734f802749c75143feddfcdb740c4131385e927148ba7d05b` |
| `qwen/qwen3-coder-30b-a3b-instruct` | Qwen: Qwen3 Coder 30B A3B Instruct | 2025-07-31 | 0.07 | 0.28 | none advertised | True | `39659697ca3437f62be4fb8e129e6b3902a91b5acac3243ca914dfc917dd563b` |
| `qwen/qwen3-coder-flash` | Qwen: Qwen3 Coder Flash | 2025-09-17 | 0.195 | 0.975 | none advertised | True | `d7e09ed951a3e1310a5e6bf09ac6aab40a68d37524c6bc20c16c0206fd08a532` |
| `qwen/qwen3-coder-next` | Qwen: Qwen3 Coder Next | 2026-02-04 | 0.12 | 0.8 | none advertised | True | `57bed80fd5b94157effe68fe998c7c12612fea210c1371dbf230097a7e8e0717` |
| `qwen/qwen3-coder-plus` | Qwen: Qwen3 Coder Plus | 2025-09-23 | 0.65 | 3.25 | none advertised | True | `182a7e88869698133f1175f18e53e7f6f2f056739559d325b211bd065235a71c` |
| `qwen/qwen3-max` | Qwen: Qwen3 Max | 2025-09-23 | 0.78 | 3.9 | none advertised | True | `fa57e897f1e8246a77fddbb7079541ba14ba0913237434ff6a4096759172bab7` |
| `qwen/qwen3-max-thinking` | Qwen: Qwen3 Max Thinking | 2026-02-09 | 0.78 | 3.9 | disabled | True | `bca0745bdc1554718d57891098f366214f05e9e896007a916d14a06efa33471a` |
| `qwen/qwen3-next-80b-a3b-instruct` | Qwen: Qwen3 Next 80B A3B Instruct | 2025-09-11 | 0.09 | 1.1 | none advertised | True | `9d8a5189c078672eb884784938b183944fb161d0e716e2c44722540b17603e34` |
| `qwen/qwen3-next-80b-a3b-thinking` | Qwen: Qwen3 Next 80B A3B Thinking | 2025-09-11 | 0.15 | 1.2 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3-vl-235b-a22b-instruct` | Qwen: Qwen3 VL 235B A22B Instruct | 2025-09-23 | 0.21 | 1.9 | none advertised | True | `61ab1adae3781ca3893328a845d4d517daa5fb47e07ea189da1df45388f62d42` |
| `qwen/qwen3-vl-235b-a22b-thinking` | Qwen: Qwen3 VL 235B A22B Thinking | 2025-09-23 | 0.4 | 4 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3-vl-30b-a3b-instruct` | Qwen: Qwen3 VL 30B A3B Instruct | 2025-10-06 | 0.15 | 0.6 | none advertised | True | `0a798666aeed20785b4661c8c2fe4a6fc6853796bc2f3c45133863469e5b3f82` |
| `qwen/qwen3-vl-30b-a3b-thinking` | Qwen: Qwen3 VL 30B A3B Thinking | 2025-10-06 | 0.2 | 2.4 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3-vl-32b-instruct` | Qwen: Qwen3 VL 32B Instruct | 2025-10-23 | 0.104 | 0.416 | none advertised | True | `43ddc43fb5a29c7fbf21c8f32b42a909c2118feb9255959ecaeab7dadb1f3dde` |
| `qwen/qwen3-vl-8b-instruct` | Qwen: Qwen3 VL 8B Instruct | 2025-10-14 | 0.117 | 0.455 | none advertised | True | `1d7f4015decbcbe955adb79817b970981ca1d3165ed121b458b213ba1060f675` |
| `qwen/qwen3-vl-8b-thinking` | Qwen: Qwen3 VL 8B Thinking | 2025-10-14 | 0.18 | 2.1 | mandatory effort unknown | True | `not runnable` |
| `qwen/qwen3.5-122b-a10b` | Qwen: Qwen3.5-122B-A10B | 2026-02-25 | 0.26 | 2.08 | disabled | True | `e0a9daef251d975c5e805e159f3813dd0c55b3f0919876b828e572b0530bea88` |
| `qwen/qwen3.5-27b` | Qwen: Qwen3.5-27B | 2026-02-25 | 0.195 | 1.56 | disabled | True | `7058adf367d238c6c92e1b83ae71f6556362e9fc7a243297b6fc65e54e2c2886` |
| `qwen/qwen3.5-35b-a3b` | Qwen: Qwen3.5-35B-A3B | 2026-02-25 | 0.1625 | 1.3 | disabled | True | `7ee20ca2b88f69dd06b9de3018deea4cb7a418e3119e17862cd0b32146daae1a` |
| `qwen/qwen3.5-397b-a17b` | Qwen: Qwen3.5 397B A17B | 2026-02-16 | 0.55 | 3.5 | disabled | True | `921707cb0d3db143e948e9d78cd8191f4ee720eaf56b6934e2d13d9ce9e3382f` |
| `qwen/qwen3.5-9b` | Qwen: Qwen3.5-9B | 2026-03-10 | 0.1 | 0.15 | disabled | True | `3016d2c1ab8d8f88f01723d2d695be996c7f1fdcd442a51cfbb6c7b7110a25e9` |
| `qwen/qwen3.5-flash-02-23` | Qwen: Qwen3.5-Flash | 2026-02-25 | 0.065 | 0.26 | disabled | True | `c946b6990be1e8b0b908a2f5fbb18097b83aa5bc76370127336fe989836b259f` |
| `qwen/qwen3.5-plus-02-15` | Qwen: Qwen3.5 Plus 2026-02-15 | 2026-02-16 | 0.26 | 1.56 | disabled | True | `74e3649a6961000bab284e801fe811d5c8fc4cf1db15af52bb10959f5a5a37c1` |
| `qwen/qwen3.5-plus-20260420` | Qwen: Qwen3.5 Plus 2026-04-20 | 2026-04-27 | 0.3 | 1.8 | disabled | True | `6d6ce30e9fbc5fdb07229b0a08e067d77511857069a4e572613d3e48b7c8b261` |
| `qwen/qwen3.6-27b` | Qwen: Qwen3.6 27B | 2026-04-27 | 0.3 | 2 | disabled | True | `abf45a27f4bafad0eb264aac33b1e49e9d0404e141966954d0707930d56665e3` |
| `qwen/qwen3.6-35b-a3b` | Qwen: Qwen3.6 35B A3B | 2026-04-27 | 0.1 | 0.9 | disabled | True | `b35b916c22649e91ded175aede34f3937a16f42a354f7abab01a17bf2b0dfc97` |
| `qwen/qwen3.6-flash` | Qwen: Qwen3.6 Flash | 2026-04-27 | 0.1875 | 1.125 | disabled | True | `fa37fc90604e5032183ec460bcc2632507566c3be3e9085f61a6ca85582b7d47` |
| `qwen/qwen3.6-max-preview` | Qwen: Qwen3.6 Max Preview | 2026-04-27 | 1.027 | 6.162 | disabled | True | `fa9ff591d389ad2d8a4b7e3fc1410ed588508fe286430c06a941180b0590422b` |
| `qwen/qwen3.6-plus` | Qwen: Qwen3.6 Plus | 2026-04-02 | 0.325 | 1.95 | disabled | True | `96d86b64bf7ceff8eed0daff1aa0396ee69609b5867942a7c2c9d33093f03c9d` |
| `qwen/qwen3.7-flash` | Qwen: Qwen3.7 Flash | 2026-07-27 | 0.03 | 0.13 | disabled | True | `82875b6ee164d0d980eba738bf05f69959234567b91f9cfd01fcabca0bd70dac` |
| `qwen/qwen3.7-max` | Qwen: Qwen3.7 Max | 2026-05-21 | 1.475 | 4.425 | disabled | True | `1c31668055e399de28805046bc08075035efbad0cdb1ff7d8926841a282e07a5` |
| `qwen/qwen3.7-plus` | Qwen: Qwen3.7 Plus | 2026-06-03 | 0.32 | 1.28 | disabled | True | `f095e0d2dbafdb91b51be3d0b3bd5a597730c8e3c397e0461023c87e924d92f8` |
| `qwen/qwen3.8-2.4t-a95b` | Qwen: Qwen3.8 2.4T A95B | 2026-08-12 | 2 | 6 | mandatory low | True | `b6452714bedbbbca3b241939aee504828d55376eed879036f1014c98af07fb14` |
| `qwen/qwen3.8-27b` | Qwen: Qwen3.8 27B | 2026-08-14 | 0.214 | 2.55 | disabled | True | `ed8190c48b2a3778bba8afff7381bb9f1578211b5e8f750d21321b62617c82c3` |
| `qwen/qwen3.8-flash` | Qwen: Qwen3.8 Flash | 2026-08-26 | 0.15 | 0.47 | disabled | True | `3443c17ae66d3fb529a058128b662024c4bc0601394e614bb693b388ec4c988b` |
| `qwen/qwen3.8-max-0902` | Qwen: Qwen3.8 Max (0902) | 2026-09-03 | 2 | 6 | mandatory minimal | True | `f7b53c82a39132f32b6c2beeafacc09db6b360fd9f46a5be0122b2b8a1889653` |
| `thinkingmachines/inkling` | Thinking Machines: Inkling | 2026-07-17 | 1 | 4.05 | disabled | False | `67a70b1b03ba85be79dc122cc888a080fde239205624135a9147168e7e77293d` |
| `x-ai/grok-4.5` | SpaceXAI: Grok 4.5 | 2026-07-08 | 2 | 6 | mandatory low | True | `721da5868b28958d481b18b9afc4ae36db3c2e0235ab0e3ae9a5c2210375d9cf` |
| `z-ai/glm-5.3` | Z.ai: GLM 5.3 | 2026-08-18 | 1.4 | 4.4 | mandatory low | True | `d81f7e66b3c4c720f8dd80768d3dfc25de6edeaf6a352cfed40e961f21fcfa22` |
| `z-ai/glm-5.3-flash` | Z.ai: GLM 5.3 Flash | 2026-08-26 | 0.09 | 0.3 | mandatory low | True | `fdf70c2d5768c4283a40342f78d82a9ebf8326db3b8c7a156a0a1d8da58b240b` |
The initial Qwen 3.7 Flash diagnostic used reasoning on and is excluded. The two corrected-attempt protocol IDs are in the request ledger. -- PI[gpt-5.6-terra]
File diff suppressed because one or more lines are too long
@@ -0,0 +1,52 @@
# WVS request-ledger audit
Source: `slop/research/wvs/20260916_openrouter/wvs_iw_requests.jsonl`.
## `qwen/qwen3.7-flash` run `20260916T132258Z_c2918219d8c3`
| metric | value |
|---|---:|
| dispatched phases | 210 |
| completed phases | 210 |
| initial completed | 144 |
| rescue completed | 66 |
| failed request phases | 0 |
| parsed valid samples | 79 |
| distinct initial item/sample keys | 144 |
| item results | 12 |
| publication eligible 12 x 12 panel | False |
| provider generation IDs retained | 210 |
| provider usage field | total |
|---|---:|
| prompt_tokens | 37168 |
| completion_tokens | 3296 |
| reasoning_tokens | unknown |
| cache_read_input_tokens | unknown |
| cache_write_input_tokens | unknown |
| total_tokens | 40464 |
| cost | 0.00154352 |
Generation IDs are retained verbatim in the source ledger.
- count: 210
- SHA-256 of sorted IDs: `a74412a88aadb5680eb0e2ef024ade648d5a0582f95c100fcd41b207cdeff4a6`
- first: `gen-1789564978-OX84XgDxxwyp2BHhOb7y`
- last: `gen-1789565251-LwqX85RylwSHRU5XdYk9`
| item | valid | requested | failed | rescues | parse rate |
|---|---:|---:|---:|---:|---:|
| Homosexuality | 12 | 12 | 0 | 0 | 1.000 |
| dealing with people? | 12 | 12 | 0 | 0 | 1.000 |
| Signing a petition | 4 | 12 | 0 | 8 | 0.333 |
| Attending peaceful demonstrations | 4 | 12 | 0 | 8 | 0.333 |
| Joining in boycotts | 8 | 12 | 0 | 4 | 0.667 |
| Religion | 0 | 12 | 0 | 12 | 0.000 |
| God | 9 | 12 | 0 | 3 | 0.750 |
| Abortion | 11 | 12 | 0 | 1 | 0.917 |
| Obedience | 6 | 12 | 0 | 6 | 0.500 |
| Independence | 3 | 12 | 0 | 9 | 0.250 |
| Determination, perseverance | 4 | 12 | 0 | 8 | 0.333 |
| Imagination | 6 | 12 | 0 | 7 | 0.500 |
Provider `cost` is reported only when the raw OpenRouter usage object exposed it. Missing usage fields are unknown, not zero. -- PI[gpt-5.6-terra]
@@ -0,0 +1,50 @@
# WVS request-ledger audit
Source: `slop/research/wvs/20260916_openrouter/wvs_iw_requests.jsonl`.
## `qwen/qwen3.7-flash`
| metric | value |
|---|---:|
| dispatched phases | 286 |
| completed phases | 286 |
| initial completed | 144 |
| rescue completed | 142 |
| failed request phases | 0 |
| parsed valid samples | 129 |
| item results | 12 |
| provider generation IDs retained | 286 |
| provider usage field | total |
|---|---:|
| prompt_tokens | 111447 |
| completion_tokens | 305854 |
| reasoning_tokens | unknown |
| cache_read_input_tokens | unknown |
| cache_write_input_tokens | unknown |
| total_tokens | 417301 |
| cost | 0.0431044 |
Generation IDs are retained verbatim in the source ledger.
- count: 286
- SHA-256 of sorted IDs: `cc086d2bf2a98cfbeb60ed4115e28c97ae10e12fca2de5822a11e41c64b2f880`
- first: `gen-1789564110-0WjESzJzswz5Y7xtD8zR`
- last: `gen-1789564677-xagE7T7PleXCYY3fF0aM`
| item | valid | requested | failed | rescues | parse rate |
|---|---:|---:|---:|---:|---:|
| Homosexuality | 10 | 12 | 0 | 12 | 0.833 |
| dealing with people? | 12 | 12 | 0 | 12 | 1.000 |
| Signing a petition | 12 | 12 | 0 | 12 | 1.000 |
| Attending peaceful demonstrations | 12 | 12 | 0 | 12 | 1.000 |
| Joining in boycotts | 12 | 12 | 0 | 11 | 1.000 |
| Religion | 12 | 12 | 0 | 12 | 1.000 |
| God | 12 | 12 | 0 | 11 | 1.000 |
| Abortion | 12 | 12 | 0 | 12 | 1.000 |
| Obedience | 6 | 12 | 0 | 12 | 0.500 |
| Independence | 10 | 12 | 0 | 12 | 0.833 |
| Determination, perseverance | 8 | 12 | 0 | 12 | 0.667 |
| Imagination | 11 | 12 | 0 | 12 | 0.917 |
Provider `cost` is reported only when the raw OpenRouter usage object exposed it. Missing usage fields are unknown, not zero. -- PI[gpt-5.6-terra]
@@ -0,0 +1,52 @@
# WVS request-ledger audit
Source: `slop/research/wvs/20260916_openrouter/wvs_iw_requests.jsonl`.
## `qwen/qwen3.7-flash` run `20260916T133001Z_82875b6ee164`
| metric | value |
|---|---:|
| dispatched phases | 144 |
| completed phases | 144 |
| initial completed | 144 |
| rescue completed | 0 |
| failed request phases | 0 |
| parsed valid samples | 144 |
| distinct initial item/sample keys | 144 |
| item results | 12 |
| publication eligible 12 x 12 panel | True |
| provider generation IDs retained | 144 |
| provider usage field | total |
|---|---:|
| prompt_tokens | 22752 |
| completion_tokens | 3240 |
| reasoning_tokens | unknown |
| cache_read_input_tokens | unknown |
| cache_write_input_tokens | unknown |
| total_tokens | 25992 |
| cost | 0.00110376 |
Generation IDs are retained verbatim in the source ledger.
- count: 144
- SHA-256 of sorted IDs: `81e6b4bd61af3a862439f5ae7e88533a5841bcce08bc3354df6167189a69fb7c`
- first: `gen-1789565401-ePJlW2zJ71smju2BHp28`
- last: `gen-1789565562-TLgEuGlGDT65plHRQuWl`
| item | valid | requested | failed | rescues | parse rate |
|---|---:|---:|---:|---:|---:|
| Homosexuality | 12 | 12 | 0 | 0 | 1.000 |
| dealing with people? | 12 | 12 | 0 | 0 | 1.000 |
| Signing a petition | 12 | 12 | 0 | 0 | 1.000 |
| Attending peaceful demonstrations | 12 | 12 | 0 | 0 | 1.000 |
| Joining in boycotts | 12 | 12 | 0 | 0 | 1.000 |
| Religion | 12 | 12 | 0 | 0 | 1.000 |
| God | 12 | 12 | 0 | 0 | 1.000 |
| Abortion | 12 | 12 | 0 | 0 | 1.000 |
| Obedience | 12 | 12 | 0 | 0 | 1.000 |
| Independence | 12 | 12 | 0 | 0 | 1.000 |
| Determination, perseverance | 12 | 12 | 0 | 0 | 1.000 |
| Imagination | 12 | 12 | 0 | 0 | 1.000 |
Provider `cost` is reported only when the raw OpenRouter usage object exposed it. Missing usage fields are unknown, not zero. -- PI[gpt-5.6-terra]
@@ -0,0 +1,20 @@
{
"completed": {
"82875b6ee164d0d980eba738bf05f69959234567b91f9cfd01fcabca0bd70dac": {
"coords": [
0.6529354469060352,
0.5885508561103798,
0.049299882651159914,
0.08073268479006059
],
"display_key": "qwen3.7-flash (rated)",
"model": "qwen/qwen3.7-flash",
"n_items": 12,
"n_samples": 12,
"protocol_id": "82875b6ee164d0d980eba738bf05f69959234567b91f9cfd01fcabca0bd70dac",
"records_path": "slop/research/wvs/20260916_openrouter/wvs_iw_requests.jsonl",
"run_id": "20260916T133001Z_82875b6ee164"
}
},
"schema": 2
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,20 @@
| model | x self-expr | y secular | x 95%CI | y 95%CI |
|:---------------------------|--------------:|------------:|----------:|----------:|
| qwen3.7-max (rated) | +0.37 | +0.69 | +0.16 | +0.17 |
| gemma-4-31b-it (rated) | +0.36 | +0.66 | +0.16 | +0.13 |
| grok-4.20 (rated) | +0.59 | +0.62 | +0.11 | +0.17 |
| gemini-2.5-pro (rated) | +0.47 | +0.70 | +0.13 | +0.14 |
| grok-4.3 (rated) | +0.44 | +0.73 | +0.14 | +0.13 |
| deepseek-v4-flash (rated) | +0.57 | +0.64 | +0.11 | +0.16 |
| gpt-5.4 (rated) | +0.45 | +0.68 | +0.13 | +0.14 |
| mistral-large-2512 (rated) | +0.65 | +0.64 | +0.08 | +0.18 |
| gemma-3-27b-it (rated) | +0.60 | +0.60 | +0.09 | +0.17 |
| llama-4-maverick (rated) | +0.64 | +0.64 | +0.08 | +0.18 |
| qwen3.7-flash (rated) | +0.65 | +0.59 | +0.10 | +0.16 |
| gpt-5.3-chat (rated) | +0.50 | +0.67 | +0.09 | +0.16 |
| deepseek-v4-pro (rated) | +0.55 | +0.73 | +0.15 | +0.10 |
| llama-4-scout (rated) | +0.62 | +0.53 | +0.05 | +0.20 |
| gpt-5.5 (rated) | +0.42 | +0.76 | +0.17 | +0.07 |
| claude-opus-4.7 (rated) | +0.59 | +0.60 | +0.07 | +0.09 |
| claude-opus-4.8 (rated) | +0.61 | +0.58 | +0.07 | +0.07 |
| claude-opus-4.6 (rated) | +0.63 | +0.63 | +0.04 | +0.05 |
Binary file not shown.

After

Width:  |  Height:  |  Size: 351 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 351 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 352 KiB

+16 -6
View File
@@ -269,18 +269,28 @@ MODEL_FAMILY_COLORS = {
"gemini": "#0ea5e9", # Gemini / Google -> sea blue (sibling of gemma, bluer)
"gpt": "#2563eb", # OpenAI -> blue
"llama": "#6d5ae0", # Llama / Meta -> indigo
"muse": "#7c3aed", # Muse / Meta -> violet
"claude": "#c026d3", # Anthropic -> purple / magenta
"grok": "#2b2d42", # Grok / xAI -> near-black (brand), well clear of gpt blue
"kimi": "#a16207", # Kimi / Moonshot -> ochre
"glm": "#dc2626", # GLM / Z.ai -> red
"inkling": "#0891b2", # Inkling / Thinking Machines -> cyan
}
def model_family_color(name: str) -> str:
"""The lab-family colour for a model key (substring match on the family name), MODEL_RED if none."""
def model_family(name: str) -> str | None:
"""Stable model-series name used to choose one label from each plotted family."""
key = name.lower()
for fam, col in MODEL_FAMILY_COLORS.items():
if fam in key:
return col
return MODEL_RED
for family in MODEL_FAMILY_COLORS:
if family in key:
return family
return None
def model_family_color(name: str) -> str:
"""The model-series colour for a model key, or MODEL_RED if the series is unknown."""
family = model_family(name)
return MODEL_FAMILY_COLORS[family] if family is not None else MODEL_RED
def plot_value_map(display: str, countries: list[str], P: np.ndarray,
+179 -66
View File
@@ -25,9 +25,13 @@ and E degenerates to an integer.
from __future__ import annotations
import asyncio
import hashlib
import json
import os
import re
from datetime import UTC, datetime
from math import inf
from pathlib import Path
import numpy as np
from loguru import logger
@@ -141,26 +145,40 @@ def _parse_ratings(text: str, n: int) -> dict[int, float] | None:
return out
_FORCE_MSG = ('You are out of time. Output ONLY a single-line compact JSON object mapping each answer '
'number to its 1-5 rating, e.g. {{"0": 1, "1": 5}}. No markdown, no reasoning, nothing else.')
def _rating_schema(n: int) -> dict:
keys = [str(k) for k in range(n)]
return {"type": "json_schema", "json_schema": {"name": "ratings", "strict": True, "schema": {
"type": "object", "properties": {key: {"type": "number", "minimum": 1, "maximum": 5}
for key in keys}, "required": keys, "additionalProperties": False,
}}}
def _force_msg(n: int) -> str:
keys = ", ".join(f'"{k}"' for k in range(n))
example = ", ".join(f'"{k}": 1' for k in range(n))
return (f"Output ONLY a compact JSON object with every required key [{keys}] and values from 1 to 5, "
f"for example {{{example}}}. No markdown, no reasoning, nothing else.")
async def _force_answer(model: str, prompt: str, phase1_msg: dict, temperature: float,
max_tokens: int, req_timeout: float) -> str:
max_tokens: int, req_timeout: float, reasoning: dict | None,
response_format: dict | None, n: int) -> dict:
"""Phase-2 rescue (wassname's bounded-thinking pattern, gist 72eed3a1): a reasoning model that
spent its whole budget thinking and truncated the JSON mid-object gets a follow-up in the SAME
conversation -- feed its (truncated) reasoning back as the assistant turn, then demand a compact
one-line answer NOW. It has already thought, so it just commits. Works even where reasoning can't
be disabled (some providers, e.g. gemini-2.5-pro, MANDATE it and 400 on reasoning_effort=none), so
we do NOT pass a reasoning-off knob -- we constrain the OUTPUT instead. A bigger cap than phase 1
(reasoning models re-think briefly). Still parsed by the caller; may fail again -> dropped sample."""
one-line answer NOW. The caller keeps its selected reasoning configuration unchanged across both
phases. A bigger cap than phase 1 lets mandatory-reasoning models finish the compact object. Still
parsed by the caller; may fail again -> dropped sample."""
tail = (phase1_msg.get("reasoning") or phase1_msg.get("content") or "")[-1500:] or "(thinking truncated)"
msgs = [{"role": "user", "content": prompt},
{"role": "assistant", "content": tail},
{"role": "user", "content": _FORCE_MSG.format()}]
{"role": "user", "content": _force_msg(n)}]
payload = {"model": model, "messages": msgs, "temperature": temperature, "max_tokens": max(max_tokens, 2048)}
data = await asyncio.wait_for(openrouter_request(payload), timeout=req_timeout)
return data["choices"][0]["message"].get("content") or ""
if reasoning is not None:
payload["reasoning"] = reasoning
if response_format is not None:
payload["response_format"] = response_format
return await asyncio.wait_for(openrouter_request(payload), timeout=req_timeout)
def _rate_plan(items: list[dict], n_samples: int, per_call: int = 1) -> list[dict]:
@@ -173,88 +191,183 @@ def _rate_plan(items: list[dict], n_samples: int, per_call: int = 1) -> list[dic
opts, n = it["options"], it["n"]
groups = [([0, 1], (n_samples + 1) // 2), ([1, 0], n_samples // 2)] if n == 2 \
else [(list(range(n)), n_samples)]
sample = 0
for perm, tot in groups:
legend = "\n".join(f"{j}) {opts[perm[j]]}" for j in range(n))
prompt = _RATE_PROMPT.format(question=it["question"], legend=legend)
while tot > 0:
k = min(tot, per_call); tot -= k
plan.append({"i": i, "perm": perm, "prompt": prompt, "cnt": k})
plan.append({"i": i, "perm": perm, "prompt": prompt, "cnt": k,
"sample": sample, "presented_options": [opts[j] for j in perm]})
sample += k
return plan
def rated_protocol_identity(model: str, items: list[dict], *, n_samples: int, temperature: float,
max_tokens: int, concurrency: int, req_timeout: float,
reasoning: dict | None, structured_output: bool) -> str:
"""Hash the exact model, rendered prompts, and request settings that define a cacheable panel."""
plan = _rate_plan(items, n_samples)
protocol = {
"schema": 2,
"model": model,
"temperature": temperature,
"max_tokens": max_tokens,
"concurrency": concurrency,
"req_timeout": req_timeout,
"reasoning": reasoning,
"structured_output": structured_output,
"rate_prompt": _RATE_PROMPT,
"rescue_prompt": _force_msg(10),
"requests": [{key: req[key] for key in ("i", "perm", "prompt", "cnt", "sample", "presented_options")}
for req in plan],
}
encoded = json.dumps(protocol, sort_keys=True, separators=(",", ":"), ensure_ascii=True).encode()
return hashlib.sha256(encoded).hexdigest()
def _append_record(path: Path, record: dict) -> None:
record["recorded_at_utc"] = datetime.now(UTC).isoformat()
with path.open("a", encoding="utf-8") as fh:
fh.write(json.dumps(record, ensure_ascii=True, sort_keys=True) + "\n")
fh.flush()
os.fsync(fh.fileno())
def read_items_rated(model: str, items: list[dict], *, n_samples: int = 12, temperature: float = 1.0,
max_tokens: int = 512, concurrency: int = 8, req_timeout: float = 90.0,
reasoning: dict | None = None, structured_output: bool = False, records_path: str | Path,
verbose_first: bool = False) -> list[dict]:
"""Dense Likert readout: per item, ask the model to rate EVERY option 1-5 as JSON, N times, and
normalize the mean rating to a per-option distribution `p`. Higher signal per call than a single
forced choice, and positional bias is controlled by permuting the PRESENTED order of BINARY items
(n==2) across samples then mapping ratings back to the canonical option order. All requests for the
model fire CONCURRENTLY (asyncio.gather, capped at `concurrency`) so a 12-item panel is ~1 round
trip, not 24 sequential ones. A reasoning model that burns its token budget thinking and truncates
the JSON is rescued by a one-shot force-answer follow-up (_force_answer) instead of being dropped.
"""Run one dense rating panel and write an fsynced JSONL event for every paid request phase.
`items`: [{"id", "question", "options"(canonical), "n"}]. Returns per item: id, p (mean over valid
samples, canonical order), p_samples (per-sample canonical p arrays for bootstrap CIs),
pmass_allowed (valid-JSON fraction), prompt, texts. p is NaN at total parse collapse (do not
compare), matching the logprob reader. A failed request (network) just drops its samples."""
The record is the source of truth. It preserves dispatches, responses, rescues, provider usage,
errors, prompt identity, presented option order, and final per-item samples. Returned rows are a
reduced view for coordinates only. An incomplete item stays incomplete and the caller must not plot it.
"""
assert temperature > 0, "sampling readout needs temperature > 0"
plan = _rate_plan(items, n_samples)
protocol_id = rated_protocol_identity(model, items, n_samples=n_samples, temperature=temperature,
max_tokens=max_tokens, concurrency=concurrency,
req_timeout=req_timeout, reasoning=reasoning,
structured_output=structured_output)
run_id = f"{datetime.now(UTC).strftime('%Y%m%dT%H%M%SZ')}_{protocol_id[:12]}"
rpath = Path(records_path)
rpath.parent.mkdir(parents=True, exist_ok=True)
settings = {"model": model, "n_samples": n_samples, "temperature": temperature,
"max_tokens": max_tokens, "concurrency": concurrency, "req_timeout": req_timeout,
"reasoning": reasoning, "structured_output": structured_output}
_append_record(rpath, {"event": "run_started", "run_id": run_id, "protocol_id": protocol_id,
"settings": settings, "items": items, "planned_requests": len(plan)})
async def run_all() -> list:
async def run_all() -> list[dict]:
sem = asyncio.Semaphore(concurrency)
async def call(req):
async def call(seq: int, req: dict) -> dict:
item = items[req["i"]]
request_id = f"{run_id}_{seq:03d}"
request_meta = {"request_id": request_id, "run_id": run_id, "protocol_id": protocol_id,
"model": model, "item_id": item["id"], "canonical_options": item["options"],
"presented_options": req["presented_options"], "presented_order": req["perm"],
"sample": req["sample"], "prompt": req["prompt"], "settings": settings}
payload = {"model": model, "messages": [{"role": "user", "content": req["prompt"]}],
"temperature": temperature, "n": req["cnt"], "max_tokens": max_tokens}
if reasoning is not None:
payload["reasoning"] = reasoning
response_format = _rating_schema(item["n"]) if structured_output else None
if response_format is not None:
payload["response_format"] = response_format
phase = "initial"
async with sem:
n = items[req["i"]]["n"]
payload = {"model": model, "messages": [{"role": "user", "content": req["prompt"]}],
"temperature": temperature, "n": req["cnt"], "max_tokens": max_tokens}
# per-request wall-clock cap: one request stuck in the wrapper's stamina backoff (a
# rate-limited provider) must not stall the whole model's gather -- time it out and drop
# it as a failed sample (return_exceptions catches the TimeoutError) so the panel moves on.
data = await asyncio.wait_for(openrouter_request(payload), timeout=req_timeout)
out = []
for c in data["choices"]:
content = c["message"].get("content") or ""
if _parse_ratings(content, n) is None: # truncated JSON / reasoning ate the budget
content = await _force_answer(model, req["prompt"], c["message"],
temperature, max_tokens, req_timeout)
out.append(content)
return out
return await asyncio.gather(*(call(r) for r in plan), return_exceptions=True)
try:
_append_record(rpath, {"event": "request_started", "phase": phase,
**request_meta, "payload": payload})
data = await asyncio.wait_for(openrouter_request(payload), timeout=req_timeout)
_append_record(rpath, {"event": "request_completed", "phase": phase,
**request_meta, "response": data, "usage": data.get("usage")})
if len(data["choices"]) != req["cnt"]:
raise ValueError(f"expected {req['cnt']} choices, got {len(data['choices'])}")
message = data["choices"][0]["message"]
text = message.get("content") or ""
rescued = False
if _parse_ratings(text, item["n"]) is None:
phase = "rescue"
rescue_payload = {"model": model, "temperature": temperature,
"max_tokens": max(max_tokens, 2048), "messages": [
{"role": "user", "content": req["prompt"]},
{"role": "assistant", "content":
(message.get("reasoning") or message.get("content") or "")[-1500:]
or "(thinking truncated)"},
{"role": "user", "content": _force_msg(item["n"])},
]}
if response_format is not None:
rescue_payload["response_format"] = response_format
if reasoning is not None:
rescue_payload["reasoning"] = reasoning
_append_record(rpath, {"event": "request_started", "phase": phase,
**request_meta, "payload": rescue_payload,
"initial_response_message": message})
rescue = await _force_answer(model, req["prompt"], message, temperature,
max_tokens, req_timeout, reasoning, response_format, item["n"])
_append_record(rpath, {"event": "request_completed", "phase": phase,
**request_meta, "response": rescue, "usage": rescue.get("usage")})
if len(rescue["choices"]) != 1:
raise ValueError(f"expected one rescue choice, got {len(rescue['choices'])}")
text = rescue["choices"][0]["message"].get("content") or ""
rescued = True
return {"text": text, "rescued": rescued, "error": None}
except Exception as exc:
_append_record(rpath, {"event": "request_failed", "phase": phase,
**request_meta, "error_type": type(exc).__name__, "error": str(exc)})
return {"text": None, "rescued": phase == "rescue", "error": f"{type(exc).__name__}: {exc}"}
return await asyncio.gather(*(call(seq, req) for seq, req in enumerate(plan)))
results = asyncio.run(run_all())
agg = {i: {"p_samples": [], "texts": [], "prompt": ""} for i in range(len(items))}
n_fail = 0
for req, res in zip(plan, results):
agg = {i: {"p_samples": [], "texts": [], "failed": 0, "rescued": 0, "prompt": ""}
for i in range(len(items))}
for req, result in zip(plan, results):
i, n, perm = req["i"], items[req["i"]]["n"], req["perm"]
agg[i]["prompt"] = req["prompt"]
if isinstance(res, Exception):
n_fail += 1
agg[i]["rescued"] += int(result["rescued"])
if result["error"] is not None:
agg[i]["failed"] += 1
continue
for text in res:
agg[i]["texts"].append(text)
rated = _parse_ratings(text, n)
if rated is None:
continue
r_canon = np.zeros(n)
for j in range(n):
r_canon[perm[j]] = rated[j] # map presented label -> canonical option
agg[i]["p_samples"].append(r_canon / r_canon.sum())
if n_fail:
logger.warning(f"{model}: {n_fail}/{len(plan)} rating calls failed (network) -> fewer samples")
text = result["text"]
agg[i]["texts"].append(text)
rated = _parse_ratings(text, n)
_append_record(rpath, {"event": "answer_parsed", "run_id": run_id, "protocol_id": protocol_id,
"model": model, "item_id": items[i]["id"], "sample": req["sample"],
"presented_order": perm, "text": text, "parsed": rated is not None})
if rated is None:
continue
r_canon = np.zeros(n)
for j in range(n):
r_canon[perm[j]] = rated[j]
agg[i]["p_samples"].append(r_canon / r_canon.sum())
out = []
for i, it in enumerate(items):
for i, item in enumerate(items):
ps = agg[i]["p_samples"]
n = it["n"]
p = np.mean(ps, axis=0) if ps else np.full(n, np.nan)
out.append({"id": it["id"], "p": p, "p_samples": [x.tolist() for x in ps],
"pmass_allowed": len(ps) / n_samples, "n_samples": n_samples,
"prompt": agg[i]["prompt"], "texts": agg[i]["texts"]})
p = np.mean(ps, axis=0) if ps else np.full(item["n"], np.nan)
row = {"id": item["id"], "p": p, "p_samples": [x.tolist() for x in ps],
"pmass_allowed": len(ps) / n_samples, "n_samples": n_samples,
"valid_samples": len(ps), "failed_samples": agg[i]["failed"],
"rescued_samples": agg[i]["rescued"], "prompt": agg[i]["prompt"],
"texts": agg[i]["texts"], "protocol_id": protocol_id, "run_id": run_id}
_append_record(rpath, {"event": "item_result", "run_id": run_id, "protocol_id": protocol_id,
"model": model, **row, "p": np.asarray(p).tolist()})
out.append(row)
if verbose_first and i == 0:
logger.debug(
f"\n=== TRACE read_items_rated first item ({model}, N={n_samples}) ===\n"
f"--- prompt ---\n{agg[i]['prompt']}\n"
f"--- first 2 raw replies ---\n{agg[i]['texts'][:2]}\n"
f"--- mean p over {it['options']} ---\n{np.round(p, 3).tolist()} valid={len(ps)}/{n_samples}\n"
f"--- prompt ---\n{row['prompt']}\n"
f"--- first 2 raw replies ---\n{row['texts'][:2]}\n"
f"--- mean p over {item['options']} ---\n{np.round(p, 3).tolist()} valid={len(ps)}/{n_samples}\n"
f"SHOULD: replies are a bare JSON dict of 1-5 ratings; valid rate near 1.0 -> coherent. "
f"ELSE the model is refusing / adding prose / max_tokens too small (empty content).\n")
f"ELSE the record shows malformed output, rescue, or request failure.\n")
_append_record(rpath, {"event": "run_finished", "run_id": run_id, "protocol_id": protocol_id,
"model": model, "planned_requests": len(plan),
"valid_samples": sum(row["valid_samples"] for row in out),
"failed_samples": sum(row["failed_samples"] for row in out),
"rescued_samples": sum(row["rescued_samples"] for row in out)})
return out