From c06721892d79dda466d407b10607fb493e58e867 Mon Sep 17 00:00:00 2001 From: "wassname (Michael J Clark)" <1103714+wassname@users.noreply.github.com> Date: Sat, 27 Jun 2026 19:37:54 +0800 Subject: [PATCH] Update README.md --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index bdd0ddb..f5a5c5f 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ mechanism so far is signed-CorDA absorption: in one 4B run, deployment ablation reduced held-out hack rate from 0.759 to 0.218 while solve rate moved from 0.161 to 0.149. That is mechanism evidence, not a deployable operating point. -An accompanying LessWrong write-up is forthcoming. The public evidence summary +An accompanying [LessWrong write-up](https://www.lesswrong.com/posts/kzri5W2uBfF2mdboK/can-we-use-steering-vectors-to-suppress-reward-hacking-1). The public evidence summary is in [docs/research_notes.md](docs/research_notes.md). This repository is the minimal public extraction of the working core: