paper: add verified related work (11 refs) + fix Huang->Deng first author

Related-work search (local qmd/gh/LW + Perplexity/Gemini/ChatGPT/Elicit), all
arXiv ids verified HTTP 200, bibtex+abstracts via the bibtex MCP / arXiv scrape:
- gradient-level reward hacking: ackermann2026gradreg (GR), liu2026harve (HARVE)
- deletable-module precedent (pre-dates Cloud): zhou2023securityvectors
- gradient-projection unlearning: shamsian2025orthograd (OrthoGrad), sun2026ogpsa
- C2 generalisation: taylor2025schoolrewardhacks, nishimuragasparian2025rhgeneralize
- weight-space contrastive direction: fierro2025weightarithmetic
- shortcut gradient surgery: cao2026sart; survey: wang2026rewardhackingsurvey
- idea provenance: mallen2025rhinterventions (AF)
Fix: huang2026directional first author is Deng, Wenlong (arXiv 2605.25189);
sync the cold-reader comment to 'Deng et al.'

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
wassname
2026-06-04 15:18:44 +08:00
co-authored by Claudypoo
parent 6085efcc54
commit b097d9abfc
2 changed files with 165 additions and 2 deletions
+2 -1
View File
@@ -367,6 +367,7 @@ enough to route one as it forms.
\bottomrule
\end{tabular}
\end{table}
% TODO hmm a bit hard conceptually. not direction is... random direction? maybe it shoudl be alternative methods idk
\subsection{Long-run convergence}
@@ -439,7 +440,7 @@ extraction set, not only the in-distribution mode (Table~\ref{tab:generalisation
one-liners are in docs/grad\_routing/related\_work.md.}
\begin{itemize}
% COMPREHENSION (cold-reader panel 2026-06-03): the keep-vs-remove inversion
% takes two reads. State it plainly first: "Huang projects ONTO a clean
% takes two reads. State it plainly first: "Deng et al. project ONTO a clean
% direction; we project a hack direction OUT." Skeptic also flagged: zeroing
% delta_S_hack at deploy == not-projecting at deploy, so "deletable knob vs
% only-constrains-training" is thin unless argued; and we never measured