mirror of
https://github.com/wassname/evil_MoE.git
synced 2026-08-16 11:20:40 +08:00
misc
This commit is contained in:
@@ -10,7 +10,8 @@ each line = {problem_id, messages=FAITHFUL hint-only prompt, completion=hack}) i
|
||||
The elicit-then-strip is already done upstream: derisk saved the FAITHFUL prompt as
|
||||
`messages` (the cheat recipe lived only in the elicit suffix, never saved) and the
|
||||
model's hack as `completion`. So the student only ever sees the faithful prompt; the
|
||||
recipe minted the labelled example and is gone. (No-cheat invariant holds.)
|
||||
recipe creates the labelled example but is never shown to the student. This preserves
|
||||
the oracle-free training constraint.
|
||||
|
||||
Two gates here, both load-bearing:
|
||||
1. EXPLOIT-VERIFY: re-grade each completion under the NON-OVERLAP grader
|
||||
|
||||
Reference in New Issue
Block a user