mirror of
https://github.com/wassname/eliciting_suppressed_knowledge.git
synced 2026-08-20 12:20:37 +08:00
wip
This commit is contained in:
@@ -52,7 +52,7 @@ This method exploits the "residual sharpening" stage identified by Lad et al. (2
|
||||
|
||||
## Key Results
|
||||
|
||||

|
||||

|
||||
|
||||
Linear probes targeting suppressed activations consistently outperform both naive outputs and standard activation probes across model scales. The performance gap (~X%) represents recoverable truthful knowledge that remains encoded but deliberately suppressed during normal generation.
|
||||
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 31 KiB After Width: | Height: | Size: 32 KiB |
+428
-991
File diff suppressed because one or more lines are too long
@@ -1,3 +1,8 @@
|
||||
|
||||
# 2025-05-02 22:14:52
|
||||
|
||||
TODO group by
|
||||
- plain hs, logit, llm_ans
|
||||
- supressed activations
|
||||
- removed attn sinks
|
||||
- combinations
|
||||
|
||||
Reference in New Issue
Block a user