Commit Graph
68 Commits
Author SHA1 Message Date
wassnameandClaudypoo aa0f45de80 SKILL.md: add empty slot for wassname's note on LLM agents
Textbook order: collected advice, then his comment on how it applies to
LLMs, then the exercises. Content left for him to write.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-26 14:21:58 +08:00
wassnameandClaudypoo d4cad35f42 SKILL.md: mark each exercise small or large, route by size
Small is under a paragraph; large means work like comparing against a
reference repo. Do all applicable small ones, pick one large one.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-26 14:21:51 +08:00
wassname cafc1c89be SKILL.md: quote Sanh on the threshold mistake, the one bullet left bare 2026-08-26 14:16:25 +08:00
wassname 3cc4ca3eb7 SKILL.md: add wassname's eight common mistakes, with the Nanda, Sanh and Achiam quotes
Source is his own message of 2026-08-25, spelling fixed and slightly more polite as he
asked, with each mistake pointing at the exercise that answers it. Also adds his rule
that a job is never abandoned without doing the exercises, one at a time.
2026-08-26 13:44:52 +08:00
wassnameandClaudypoo 6351958834 add exercises 14 and 15: one failed attempt is not a negative, and get the scale before the gate
Two gaps the existing 13 did not cover, found by mining the evidence cache against
wassname's list of common AI-agent failures. Quotes are verbatim from
docs/evidence/ (Steinhardt, Rahtz, Nanda, Goodfellow-Bengio-Courville).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-26 13:32:48 +08:00
wassname 3d84683036 Revert "SKILL.md: adopt v6 (terra's seven exercises plus wassname's common mistakes with sources)"
This reverts commit 765a061fb7.
2026-08-26 08:24:11 +08:00
wassnameandClaudypoo 765a061fb7 SKILL.md: adopt v6 (terra's seven exercises plus wassname's common mistakes with sources)
Measured on wassname-ml-bench v97, 12 items, loaded-skill header:
  grok-4.6   bare +0.592, v5 +0.809, v6 +0.785
  deepseek   bare +0.706, v5 +0.705, v6 +0.737
v5 and v6 are indistinguishable there; v6 carries wassname's failure modes.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-26 08:08:16 +08:00
wassname 192311c425 skill: sign off with a quote in a hand-drawn ascii animal 2026-08-24 22:16:42 +08:00
wassname 2871d89512 skill: the shorter runbook, 174 lines from 255 2026-08-24 21:39:04 +08:00
wassnameandClaudypoo dc369f5fac skill: fold in the design-doc gaps -- outcome signatures, persisted predictions, executed config, sample selection, guide caps
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:53:38 +08:00
wassnameandClaudypoo 2e1eefbba6 skill: raw artifacts in P1, seed-noise and baseline before A-beats-B, fresh auditor, follow the job
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:46:23 +08:00
wassnameandClaudypoo c94a450518 skill: permit 'unknown', require attribution when several things change at once
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:45:09 +08:00
wassnameandClaudypoo ff797a95e1 skill: keep failure interpretations, ban failure causes; history goes to the journal
Interpretations must read as evidence for or against a claim, never as a cause,
because a cause list gets picked from, called certain, and used to stop.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:41:02 +08:00
wassnameandClaudypoo e2bd28dbc2 skill: read your data and assume you have a bug were only in the frontmatter
Both were promised by the description and absent from the procedure. P2 now
prints the formatted examples; P3 opens with the bug assumption.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:39:36 +08:00
wassnameandClaudypoo f830c0cb23 skill: put literal trigger phrases in the description so it self-invokes
A sonnet subagent had ml-debug listed, with 'after a run finishes or crashes' in
the description, and did not invoke it for 'read the last pueue log'.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:55:43 +08:00
wassnameandClaudypoo a7ff779e88 skill: P1 steps 1-4 are read-only, so a 'just read the log' task still fills the table
Sonnet routed to P1 correctly then declined it: 'the task only asked to read, not
to audit or act'. Reading a log is exactly when the measurement table gets filled.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:52:56 +08:00
wassnameandClaudypoo f3f1a38485 skill: rewrite as a runbook -- P1-P5, numbered imperative steps, stop gates
Sonnet read the block version and summarised it instead of running it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:51:10 +08:00
wassnameandClaudypoo 8b092a4320 skill: drop the word ritual, headings become instructions, frame is a draft you revise
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:46:41 +08:00
wassnameandClaudypoo 4f422a8016 skill: bug-assumption gets its own trigger, the all-clear sentence; guide cap 400 -> 180 lines
A sonnet subagent read the skill, summarised the ritual list, ran none of them,
and closed with 'everything checks out'. Phase triggers do not fire; a sentence
the agent watches itself write might.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:44:49 +08:00
wassnameandClaudypoo df4e08ef90 skill: add the audit fields I dropped -- trained/frozen, measure-before-diagnose, curve at 4 points, next-experiment case, contrary evidence
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:37:09 +08:00
wassnameandClaudypoo fb13b4fda7 skill: replace folklore prose with rituals -- trigger, form, artifact shown to user
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:31:18 +08:00
wassname d5d725e750 Revert "add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization"
This reverts commit ec2bb4f4be.
2026-08-18 19:02:26 +08:00
wassnameandClaudypoo ec2bb4f4be add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-18 18:01:32 +08:00
wassnameandClaudypoo 647b9a0145 llm_judges: Miller error bars -- repeat draws, don't touch the thermostat, paired differences
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-16 13:29:02 +08:00
wassname b2c666dbbf add wassname's 37-reasons checklist to refs/checklist.md
Quoted whole from the 2017 thread, since the article author asked to merge it
and never did. Verified line by line against the thread cache. Covers sample
size from a cumulative-mean plot, KLD/Dice on unbalanced data, augmentation
bounded by feature std, dummy metrics, jumpy validation loss, testing the
framework itself, activation swaps, and loss-curve shapes.
2026-08-15 06:40:45 +08:00
wassname 54dcd832b8 drop line-number anchors from cache citations
102 of them, 48 already wrong after today's refetches. They rot every time a
cache is refetched and buy nothing the quote text does not: the caches are
verbatim, so the quote itself is the anchor. Descriptive labels stay.
2026-08-15 06:23:31 +08:00
wassname ceba01782b replace four invented nanochat quotes with Karpathy's own words
All four came from the AI-generated deepwiki page, not the experiment log,
and one inverted the finding: the log says clipping was removed because
'Grad norm never exceeds 1.0 naturally', and the real distributed gotcha was
clipping local norms before sync. Footnote now points at dev/LOG.md on master
and at the renamed cache.
2026-08-15 06:22:53 +08:00
wassname a6c8ba77d2 cite arxiv pdf, never abs
An /abs/ link is a stub: it costs a second lookup before anyone can check the
quote, and it is what let abstract-only caches look sourced. 77 links across
the skill files and 4 cache headers. Verbatim source bodies untouched, their
reference lists are the authors' text.
2026-08-15 06:14:10 +08:00
wassname 996942c4de Make ML debugging expose evidence and uncertainty 2026-08-09 11:11:02 +08:00
wassnameandClaudypoo 602f6193ed assume every negative result is a bug until the logs rule it out; be patient
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-28 17:14:05 +08:00
wassname c7f22b8478 Add deployment-faithful time-series evaluation guide 2026-07-26 10:31:15 +08:00
wassname (Michael J Clark) e92ec01efe Enhance ML debugging guidance for LLM agents
Added guidance for LLM agents on reading and calibrating their confidence levels in ML debugging.
2026-06-26 09:52:43 +08:00
wassnameandClaudypoo 5fca5ad2b2 Refresh Schulman cache anchors after transcript rewrite
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-25 10:31:39 +08:00
wassname 67d4dc90bb Document quote-first evidence style 2026-06-25 10:31:39 +08:00
wassname (Michael J Clark) 3f3f95a3b4 Update SKILL.md with debugging strategies and folklore
Reorganize and refine debugging folklore and hypotheses for RL implementations.
2026-06-14 08:37:57 +08:00
wassname b8c3ffcf11 gpt5.5/fable 2026-06-12 09:30:25 +08:00
wassname 160bd040cc docs: add sourced transformer report folklore 2026-06-12 07:02:58 +08:00
wassname 3e28a950e9 docs: clarify competing-worlds debugging loop 2026-06-12 06:53:36 +08:00
wassname 8b9a1d62ed docs: resolve ml-debug TODO references 2026-06-12 06:52:38 +08:00
wassname (Michael J Clark) e58eda360b Update SKILL.md with TODOs for future content
Added TODO items for additional references and empirical evidence regarding transformers and model sizes.
2026-06-11 21:35:49 +08:00
wassnameandClaudypoo 30ac76053e chore: drop links to deleted/tombstone gists (repo is canonical now)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 16:55:04 +08:00
wassnameandClaudypoo 0837f27f08 fix: companion gist link pointed at the wrong gist
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 16:41:50 +08:00
wassnameandClaudypoo 8cd3c61050 folklore: tuning playbook, Domingos, Bekman loss spikes, Ng error analysis; LLM-judge bias appendix
- SKILL.md: 3 new entries (exploration-over-exploitation + nuisance HPs,
  test-set contamination, loss-spikes-mean-bad-data-pocket) and an Ng
  100-misclassified-examples quote under inspect-the-data
- refs/llm_judges.md: position/verbosity/self-preference biases (Zheng,
  Wang 66/80 flip, Panickssery) + mitigation checklist from verdict docs
- Lones pitfalls linked as the exhaustive 36-item do/don't checklist
- 6 new frozen evidence files; Hamel evals link in further reading

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 15:30:41 +08:00
wassnameandClaudypoo 2a2f5045bb folklore: add Karpathy common-mistakes tweet and Sculley CACE principle
Both quote-verbatim with frozen evidence: the 2018 tweet thread (mirrored
via threadreaderapp, x.com blocks fetching) slots after overfit-one-batch;
CACE (NIPS 2015, entanglement section transcribed from the PDF) gives
Always-Be-Ablating its why.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 14:43:47 +08:00
wassnameandClaudypoo fb753d093e restructure: quotes-first SKILL.md, synthesized playbook split out
SKILL.md is now folklore only: verbatim practitioner quotes ordered
most-general-first, transformer/LLM fine-tuning entries in their own
section, minimal context, links and footnotes. New sources: unsloth,
axolotl (+training stability), HF course ch8.4, Bekman debug_utils
(evidence frozen in docs/evidence/).

The synthesized material (mental models, priors, symptom tables, agent
loop, triage, anti-patterns) moves to PLAYBOOK.md, framed as menus of
hypotheses rather than authoritative diagnoses. Made-up symptom tables
no longer sit next to sourced quotes.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 14:33:32 +08:00
wassnameandClaudypoo 8ee980d62f diagnostics: add NaN-poisoning leakage tracer + Karpathy backprop-to-input check; README citation
NaN poisoning: inject NaN where info must not come from (future/test/labels), run the real pipeline, assert past outputs stay finite. Documents false negatives (pandas skipna, nanmean) and false positives (softmax rows, batch stats). Backprop-to-input is its gradient dual for inside the model; quote already frozen in docs/evidence/karpathy_recipe_training_nn_2019.md.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-11 10:18:51 +08:00
wassname (Michael J Clark) 53fb1c4bda Update SKILL.md 2026-06-03 08:19:07 +08:00
wassnameandClaudypoo 8509ec3c30 folklore: promote Spinning Up to main; add a Research-taste section
- Promote the general (non-RL-specific) Spinning Up lessons up to the main
  folklore: "broken code fails silently", "you can't tell it's broken if you
  can't see that it's breaking", and test on more than one setup.
- Add gwern's "Unseeing" to the data theme: you can't read what you actually
  wrote, hence fresh eyes / a fresh-eyes subagent.
- New "Research taste (adjacent to debugging)" section with verbatim quotes,
  each cached: Neel Nanda (your research is false by default; excitement is
  evidence of bullshit; read your data), Ulisse Mini (understand the system to
  shrink the search space), John Wentworth (gears-level models are capital
  investments vs cheap black boxes).

All quotes verbatim from cached sources; 25/25 footnotes resolve.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-02 21:08:49 +08:00
wassnameandClaudypoo ee4e9a5caa folklore: add koaning, gwern, kidger, nanochat, cleanrl; trim lucidrains
Gather debugging folklore from more practitioners, each a verbatim quote
checked against a cached source copy (footnoted with line numbers):
- koaning (Vincent Warmerdam), "Bad Labels": benchmark labels are often wrong;
  find them with confidence-sorted errors.
- gwern, the tank-detection legend: the canonical data-leakage parable, plus
  the scout-mindset twist that it's a likely-unsourced urban legend.
- Patrick Kidger, "Just Know Stuff": why research code is buggy ("kludge ...
  bugs that don't cripple things only because some other bug stops them") and
  "never accept the kludge". Plus a one-line jaxtyping pointer for shape bugs.
- nanochat (Karpathy): BOS-alignment fake metric improvement; all-ranks must
  clip on inf (a multi-GPU bug single-GPU testing hides).
- cleanrl "37 Implementation Details of PPO" -> RL sub-skill, as the canonical
  proof that reference-impl details (not ideas) decide whether PPO works.

Trim the lucidrains item to one quote (it had ballooned). Add wassname credit
+ companion-gist link. All 20 footnotes resolve.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-02 20:59:36 +08:00
wassnameandClaudypoo 9911ac83c5 folklore: add lucidrains transformer-stability item (QK-norm, post-emb LN)
Phil Wang's x-transformers is the canonical "the fix is in the code, not the
paper" catalogue. Add a folklore item on the most debugging-relevant trick:
QK / cosine-sim normalization to stop attention logits overflowing (the usual
cause of transformer loss spikes/divergence), plus the BLOOM/YaLM
post-embedding LayerNorm. Two verbatim lucidrains quotes, footnoted to the repo
+ a cached README copy with line numbers. Doubles as the modern concrete
example for the read-a-working-implementation section.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-02 20:49:15 +08:00