evil_MoE

mirror of https://github.com/wassname/evil_MoE.git synced 2026-06-27 21:07:17 +08:00

Files

T

wassname 4359dc53a8 feat: route2 distinct-basis quarantine + per-sample act-mask detach-route

Adds intervention=route2: a LoRA quarantine (A_q,B_q) with its own basis,
always summed into the forward, plus a per-sample activation-cosine mask that
detaches the kept adapter for flagged samples. Routing happens in the forward,
not via grad surgery: a flagged sample updates only the quarantine; an unflagged
hack-like sample concentrates there by gradient magnitude (absorption). Deploy
zeroes A_q,B_q. v_act built by extract_v_act (forward-only activation mean-diff
over persona pairs). Fixes the per-prompt zero_grad wiping quarantine grads
before opt.step. scripts/make_random_vhack.py = the random-V route control.
vhack_refresh_every default 0->5 (0 is ablation-only).

Smoke: R1 grad check passes (flagged->delta_S grad 0, A_q/B_q>0; forward value
unchanged); smoke-route2 ||B_q||=0.109, deploy eval + asserts pass.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>

2026-05-31 10:16:13 +00:00

build_combined_pool.py

reorg: out/ sorted by datatype (vhack/ pools/ runs/ vhack_grads/ figs/)

2026-05-30 03:52:24 +00:00

make_dataset_pairsets.py

scripts

2026-05-30 04:16:56 +00:00

make_pairsets.py

wip

2026-05-30 04:33:33 +00:00

make_random_vhack.py

feat: route2 distinct-basis quarantine + per-sample act-mask detach-route