This commit is contained in:
wassname
2026-05-20 04:31:24 +00:00
parent bfd3a572cf
commit 726324e772
2 changed files with 9 additions and 20 deletions
+9 -16
View File
@@ -10,23 +10,17 @@ Side artifact (not used by eval, kept for human-correlation sanity check):
Each row: {id, foundation, foundation_coarse, wrong, text}.
Dual-axis design
================
Each vignette produces 4 prompts from two independent binary axes:
Condition axis
==============
Each vignette has two text variants per `cond`:
**cond** (scenario framing — which text variant the model reads):
`other_violate` — 3rd-person ("You see someone doing X")
`self_violate` — 1st-person ("You do X")
`other_violate` — 3rd-person ("You see someone doing X")
`self_violate` 1st-person ("You do X")
**frame** (question framing — how the JSON probe is phrased):
`wrong` — '{"is_wrong": ' → true means wrong
`accept` — '{"is_acceptable": ' → true means right (inverted)
Both axes are paired-out in `analyse()`:
- The two *frames* cancel the additive JSON-true prior (training data has
more `"true"` than `"false"` in JSON contexts).
- The two *conds* let you measure perspective bias: the gap between how
harshly the model judges others vs itself for the same scenario.
Eval runs the K-way forced-choice probe on both; averaging cancels
perspective bias (model judging others vs itself). The probe itself is
a single JSON-pseudo-schema with the 7 foundations as enum options —
not a binary wrong/accept frame.
"""
from __future__ import annotations
import json
@@ -34,7 +28,6 @@ from pathlib import Path
from typing import Literal
ROOT = Path(__file__).resolve().parents[2]
HF_REPO = "wassname/tiny-mfv"
CONDITIONS = ["other_violate", "self_violate"]
# Canonical config names.
Generated
-4
View File
@@ -10,10 +10,6 @@ resolution-markers = [
"python_full_version < '3.14' and sys_platform != 'emscripten' and sys_platform != 'win32'",
]
[options]
exclude-newer = "2026-05-02T06:45:18.586407301Z"
exclude-newer-span = "P6D"
[[package]]
name = "accelerate"
version = "1.13.0"