Commit Graph
3 Commits
Author SHA1 Message Date
wassname 474f74ac33 wip 2026-07-10 14:36:55 +08:00
wassnameandClaudypoo ebb912452d U4 step3: dim_batch 16->8 to survive OOM contention with user's VS Code kernel
Run 551 was OOM-killed at n_done=36: the user's VS Code Jupyter kernel
(jsteer venv, PID 3214401) co-loaded ~1.5GB VRAM + 1.9GB RAM while the fit
sat at the 22.4/24.6GB ceiling. Clean SIGKILL with no CUDA traceback = host
OOM killer, not a CUDA OOM. dim_batch=8 halves the fit's peak footprint;
it changes only the backward schedule, not the accumulated Jacobian, so U4
exactness is preserved. Resumes from checkpoint (n_done=36), lossless.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:56:17 +08:00
wassnameandClaudypoo 4fb1cfd63e apply external review (deepseek): reject float bands post-fit, assert->ValueError, 2 clarifying comments
Review verdict APPROVE; rejected findings (empty-prompts guard = preemptive
defensive check, from_hf dedup = ms-scale, lm.forward swap = loses attention
mask on padded batches) documented in docs/reviews/code.md triage.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-10 13:05:27 +08:00