The "full-preset loader" TODO: full (200 steps, 43 prompts/step) was resampling
just 6 problems. Expand the hand-written substrate to 24 (6 per loophole mode,
matching the old repo's partition) so a whole mode can be held out for the
generalisation test and full/fast see real variety. Self-contained by design --
no external dataset.
- problems.py: +18 simple one-liner problems (max/min/sum/count/...), balanced
6 per mode. All hand-verified.
- rewards._self_check: new guard -- every problem's clean solution must pass the
strict oracle (gt_correct, not exploited). A wrong body/gt_test now fails
`just check` loud. Confirms 6/mode.
- train: fast + full presets use n_problems=24 (smoke keeps 6).
just check: diagonal clean + all 24 clean solutions pass the oracle.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Replace the AntiPaSTO SVD-diag adapter (δW = U·diag(δS)·Vh) with a parametrized
LoRA: per Linear a frozen random A (r×d_in, semi-orthonormal) and a trained
B (d_out×r), δW = (B+B_hack)·A. δW stays LINEAR in the trained knob -- the one
property the projection needs (a once-extracted V stays a fixed weight-space
direction) -- but the whole per-module SVD subsystem is gone (no svd_cached, no
SVD_CACHE, no bf16 hash/contiguity friction). B_hack is the SAME shape as B, so the
route quarantine is capacity-matched by construction (old-repo route2 diverged from
an oversized quarantine sink).
- antipasto: wrap() builds A (fp32 orthogonal_, geqrf has no bf16 -> cast after) +
B/B_hack zero-init on the layer device; forward y + (B+B_hack)@(A@x).
- proj: project_one is dim-agnostic; project_all flattens B.grad (d_out·r) and
reshapes back. cos_overlap flattens too.
- extract: V from the SVD of stacked B.grad pair-diffs (d_out·r).
- train: B/B_hack rename, lora_rank config, per-step aligned table + legend
(replaces the sparse tqdm postfix), clean argv via preset/Config defaults.
- justfile: collapse smoke-vanilla/smoke-route/fast-vanilla/full-vanilla into
smoke/fast/full *ARGS + a `sweep` recipe that fires vanilla|erase|route as pueue
jobs. results.py: glob run_*.log (skip loguru verbose logs).
Smoke (GPU bf16, all three arms) green: cout~0 one_sided identity holds in the LoRA
basis, |Bh|=0 for erase/none, route parks into B_hack.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The fp32+CPU smoke never walked the fast/full bf16+cuda path, so the whole
device/dtype class was invisible to the only gate. The GPU path had in fact
never run end-to-end. Seven bugs, each masked by fp32+CPU:
- svd cache key: numpy has no bf16 -> hash via .view(uint8)
- svd save: safetensors needs contiguous cpu tensors
- svd cache-hit: load_file returns cpu -> .to(W) for device+dtype
- delta_S/delta_S_hack created on cpu -> device=lin.weight.device (else the
forward hook mixes cpu/cuda)
- V_hack is fp32 (svd) but grads are bf16 -> cast to delta_S's space
- completion_nll fed cpu ids to the cuda model in extraction
- extraction orientation vote D@V.T mixed bf16 D with fp32 V
Smoke is now tiny-random on GPU in bf16 (same device+dtype as fast/full), so
this class stays caught. All arms (none/erase/route) and extraction paths
(miss/hit/refresh) green: cin_t~0.31-0.41, cout~0 (one_sided identity).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>