Commit Graph

  • 2f7b541d44 Remove evaluator-scoring language from ML debugging guidance main wassname2andPI/OpenAI 2026-09-16 10:42:41 +08:00
  • 0ba4507bae Make audit findings drive visible research decisions wassname 2026-09-09 13:02:52 +08:00
  • 9ffbeff014 Merge dev4 into main wassname 2026-09-07 14:36:01 +08:00
  • 76d93f2297 Clarify expensive-run exploration and sweep uncertainty dev4 wassname2andPi 2026-09-07 11:37:12 +08:00
  • bea4f75a77 Add randomized ML debugging fortunes wassnameandPI[k3] 2026-09-02 12:05:29 +08:00
  • 7cf8e0e245 preserve training-script prose wassnameandPI[openai-codex] 2026-09-02 09:00:36 +08:00
  • 9e6391583a restore single-file training guidance wassnameandPI[openai-codex] 2026-09-02 08:59:10 +08:00
  • c853e90eec inline private ML workflow contracts wassnameandPI[openai-codex] 2026-09-02 08:43:01 +08:00
  • 1df150618a refresh dev4 skill draft wassnameandPI[openai-codex] 2026-09-02 08:33:45 +08:00
  • 006ee0ae4d try compact research-loop skill wassnameandPI[openai-codex] 2026-09-02 08:27:26 +08:00
  • 4e1e77fa24 add Agans nine rules: full verbatim quotes + evidence notes wassnameandPI[openai-codex] 2026-09-02 07:21:28 +08:00
  • 0c102d5177 add Tobin debugging sequence dev wassnameandPI[openai-codex] 2026-09-02 06:59:55 +08:00
  • b7f46074fa make design review tool-independent wassnameandPI[openai-codex] 2026-09-02 06:47:31 +08:00
  • 2c62449d9b fix skill install link and enforce audits wassnameandPI[openai-codex] 2026-09-02 06:32:09 +08:00
  • 6106575e9c replace grader-folk words with the precise term wassnameandClaudypoo 2026-09-01 05:46:15 +08:00
  • e4ae445108 form is an anti-skim gate: task/eval framing, scoring rows, routing is mandatory wassnameandClaudypoo 2026-08-31 17:30:35 +08:00
  • 8c62f961a8 restore wassname's original form framing: Q: prefixes, Is it: nesting, TODOs, evaluated-on line wassname 2026-08-31 17:29:48 +08:00
  • 79f46be733 add a fill-in form at the head, so an agent reading only the top does one core exercise wassnameandClaudypoo 2026-08-31 17:27:30 +08:00
  • bdb70ba7b1 Refactor README to streamline content and sections wassname (Michael J Clark) 2026-08-30 17:55:12 +08:00
  • 1fb188e923 Update README.md wassname (Michael J Clark) 2026-08-30 17:53:45 +08:00
  • 4112134bfd Enhance README with debugging quotes and illustrations wassname (Michael J Clark) 2026-08-30 17:51:43 +08:00
  • 37bb6fcf90 two more quotes at the top, and why Rahtz's fast-feedback argument does not transfer to an agent wassname 2026-08-27 19:43:20 +08:00
  • cb03fb18fd the v9 row: be-diligent-first scores +0.56, above bare, where the old text was below it wassname 2026-08-27 18:00:38 +08:00
  • 7451008c1b README table: point the untested row at the version that will be tested wassname 2026-08-27 09:59:19 +08:00
  • 3a58c54170 follow the skill spec: references/ not refs/, and namespaced subskill names wassnameandClaudypoo 2026-08-27 09:59:01 +08:00
  • 739f846297 README: which part of the document does the work, and where each version lives wassname 2026-08-27 09:56:43 +08:00
  • 26ed3b80b7 misc wassname 2026-08-27 09:48:46 +08:00
  • 742dbe4d41 keep the bench numbers out of the skill, point at the README table wassname 2026-08-27 09:48:21 +08:00
  • 6e149bde5e put be-diligent at the top, it is the part with measured uplift wassname 2026-08-27 09:45:19 +08:00
  • e296293960 name the exercises, and make the selector a nested if/then list wassname 2026-08-27 09:19:39 +08:00
  • bc136d224e wassname's rule: metrics and demos inline in the train script, no side-car probes wassname 2026-08-27 08:35:11 +08:00
  • efcac5ca5f plain wording: do the exercises and show the result wassname 2026-08-26 16:14:24 +08:00
  • b8689ad23b fail fast is advice for a human who over-commits, agents quit early instead wassname 2026-08-26 16:09:00 +08:00
  • 392da00cc4 Common mistakes: promote the crash-loudly rule from PLAYBOOK, plus Nanda wassname 2026-08-26 15:58:58 +08:00
  • 8429a08903 sign off: name who said the quote wassname 2026-08-26 15:44:47 +08:00
  • fe2d13bf4d SKILL.md: extend Karpathy quote to what to do when out of ideas wassname 2026-08-26 15:42:50 +08:00
  • 986c017fce SKILL.md: Karpathy's NEVER STOP next to wassname's never-give-up line wassname 2026-08-26 15:34:40 +08:00
  • bfb97f061c SKILL.md: wassname's note on field-standard language fills the LLM slot wassname 2026-08-26 15:23:08 +08:00
  • 5ab2418c3c exercise 12: heading follows the new quote wassname 2026-08-26 14:24:03 +08:00
  • 54e7708f15 exercise 12: swap the tank legend for Zech et al. and the fastbook grant leak wassname 2026-08-26 14:23:48 +08:00
  • d7536fe549 SKILL.md: annoy-less comment review on the AI-written prose wassnameandClaudypoo 2026-08-26 14:23:18 +08:00
  • aa0f45de80 SKILL.md: add empty slot for wassname's note on LLM agents wassnameandClaudypoo 2026-08-26 14:21:58 +08:00
  • d4cad35f42 SKILL.md: mark each exercise small or large, route by size wassnameandClaudypoo 2026-08-26 14:21:51 +08:00
  • cafc1c89be SKILL.md: quote Sanh on the threshold mistake, the one bullet left bare wassname 2026-08-26 14:16:25 +08:00
  • d769cacfd4 README: add the eight common mistakes with 46 mined quotes wassname 2026-08-26 13:47:00 +08:00
  • 3cc4ca3eb7 SKILL.md: add wassname's eight common mistakes, with the Nanda, Sanh and Achiam quotes wassname 2026-08-26 13:44:52 +08:00
  • 6351958834 add exercises 14 and 15: one failed attempt is not a negative, and get the scale before the gate wassnameandClaudypoo 2026-08-26 13:32:48 +08:00
  • 3d84683036 Revert "SKILL.md: adopt v6 (terra's seven exercises plus wassname's common mistakes with sources)" wassname 2026-08-26 08:24:11 +08:00
  • 765a061fb7 SKILL.md: adopt v6 (terra's seven exercises plus wassname's common mistakes with sources) wassnameandClaudypoo 2026-08-26 08:08:16 +08:00
  • 26eb2cce6a candidate quotes and a common-mistakes draft, not yet in the README wassnameandClaudypoo 2026-08-25 11:23:30 +08:00
  • 192311c425 skill: sign off with a quote in a hand-drawn ascii animal wassname 2026-08-24 22:16:42 +08:00
  • 2871d89512 skill: the shorter runbook, 174 lines from 255 wassname 2026-08-24 21:39:04 +08:00
  • dc369f5fac skill: fold in the design-doc gaps -- outcome signatures, persisted predictions, executed config, sample selection, guide caps wassnameandClaudypoo 2026-08-20 06:53:38 +08:00
  • 2e1eefbba6 skill: raw artifacts in P1, seed-noise and baseline before A-beats-B, fresh auditor, follow the job wassnameandClaudypoo 2026-08-20 06:46:23 +08:00
  • c94a450518 skill: permit 'unknown', require attribution when several things change at once wassnameandClaudypoo 2026-08-20 06:45:09 +08:00
  • ff797a95e1 skill: keep failure interpretations, ban failure causes; history goes to the journal wassnameandClaudypoo 2026-08-20 06:41:02 +08:00
  • e2bd28dbc2 skill: read your data and assume you have a bug were only in the frontmatter wassnameandClaudypoo 2026-08-20 06:39:36 +08:00
  • f830c0cb23 skill: put literal trigger phrases in the description so it self-invokes wassnameandClaudypoo 2026-08-19 20:55:43 +08:00
  • a7ff779e88 skill: P1 steps 1-4 are read-only, so a 'just read the log' task still fills the table wassnameandClaudypoo 2026-08-19 20:52:56 +08:00
  • f3f1a38485 skill: rewrite as a runbook -- P1-P5, numbered imperative steps, stop gates wassnameandClaudypoo 2026-08-19 20:51:10 +08:00
  • c858469f59 readme: match the skill's wording, instructions not rituals wassnameandClaudypoo 2026-08-19 20:46:48 +08:00
  • 8b092a4320 skill: drop the word ritual, headings become instructions, frame is a draft you revise wassnameandClaudypoo 2026-08-19 20:46:41 +08:00
  • 4f422a8016 skill: bug-assumption gets its own trigger, the all-clear sentence; guide cap 400 -> 180 lines wassnameandClaudypoo 2026-08-19 20:44:49 +08:00
  • d37b9c88a6 readme: state the ritual rework as an untested bet, not a finding wassnameandClaudypoo 2026-08-19 20:39:22 +08:00
  • df4e08ef90 skill: add the audit fields I dropped -- trained/frozen, measure-before-diagnose, curve at 4 points, next-experiment case, contrary evidence wassnameandClaudypoo 2026-08-19 20:37:09 +08:00
  • fb13b4fda7 skill: replace folklore prose with rituals -- trigger, form, artifact shown to user wassnameandClaudypoo 2026-08-19 20:31:18 +08:00
  • 7dc8cfd23e readme: hold the folklore quotes, agents read prose and ignore it wassnameandClaudypoo 2026-08-19 20:30:08 +08:00
  • d5d725e750 Revert "add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization" wassname 2026-08-18 19:02:26 +08:00
  • ec2bb4f4be add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization wassnameandClaudypoo 2026-08-18 18:01:32 +08:00
  • 52390d7593 readme: drop the run-by-run pairing, it was not a real pairing wassnameandClaudypoo 2026-08-17 11:21:55 +08:00
  • 1419c2e7df readme: the uplift did not replicate over three answers per question wassnameandClaudypoo 2026-08-17 03:52:28 +08:00
  • 647b9a0145 llm_judges: Miller error bars -- repeat draws, don't touch the thermostat, paired differences wassnameandClaudypoo 2026-08-16 13:29:02 +08:00
  • aa791fb839 readme: one measurement, not a rate wassname 2026-08-16 07:16:17 +08:00
  • 8ba59c54b8 readme: say plainly that only the cheap model reads the document wassname 2026-08-16 07:15:47 +08:00
  • ab9779ec00 readme: the skill closes 59% of the gap to gpt-5.6-sol on the same questions wassnameandClaudypoo 2026-08-16 07:13:32 +08:00
  • e4e3386d1f readme: first measurement of the skill's effect, +0.135 on wassname-ml-bench wassnameandClaudypoo 2026-08-16 07:12:32 +08:00
  • b2c666dbbf add wassname's 37-reasons checklist to refs/checklist.md wassname 2026-08-15 06:40:45 +08:00
  • f4d6fc28ca cache the 37-reasons reddit thread, which holds 13 checks the article never absorbed wassname 2026-08-15 06:38:01 +08:00
  • 6c50496122 add missing cache headers, warn on the unverified judge entries wassname 2026-08-15 06:23:54 +08:00
  • 54dcd832b8 drop line-number anchors from cache citations wassname 2026-08-15 06:23:31 +08:00
  • ceba01782b replace four invented nanochat quotes with Karpathy's own words wassname 2026-08-15 06:22:53 +08:00
  • 55726b56bd delete the spinning-up notes file, fix its last reference wassname 2026-08-15 06:21:49 +08:00
  • 60ed9df651 drop the derived spinning-up notes, full-text the nanda draft wassname 2026-08-15 06:21:37 +08:00
  • a6c8ba77d2 cite arxiv pdf, never abs wassname 2026-08-15 06:14:10 +08:00
  • 2565f203e4 add evidence full-text verifier wassname 2026-08-15 06:07:11 +08:00
  • 383265c60c evidence: full text for nine LessWrong and research-blog caches wassname 2026-08-15 06:07:11 +08:00
  • 7fec3c557d evidence: full text for nine blog and essay caches wassname 2026-08-15 06:07:11 +08:00
  • 776ccf7047 evidence: full text for tuning playbook, ml yearning, bekman, spinning up, qwen3 wassname 2026-08-15 06:07:11 +08:00
  • d68ff6477a evidence: full nanochat experiment log, not a deepwiki synthesis wassname 2026-08-15 06:07:11 +08:00
  • b1087b8efd llm_judges: verify the five abstract-quoted papers against raw PDF wassname 2026-08-14 21:16:13 +08:00
  • ffcc94df00 evidence: full text for sculley 2015, lones 2021, domingos 2012 wassname 2026-08-14 21:13:11 +08:00
  • 701a09a525 pinn: correct claims that were written from abstracts only wassname 2026-08-14 20:56:44 +08:00
  • b70dcfa2b1 pinn evidence: replace abstract-only stubs with full paper text wassname 2026-08-14 20:51:56 +08:00
  • 3a8839b7d7 make judge setup repair a checklist wassname 2026-08-14 15:42:47 +08:00
  • cb7d597962 generalize judge setup repair guidance wassname 2026-08-14 14:54:56 +08:00
  • ad8c981504 teach judge audits to repair setup confusion wassname 2026-08-14 14:41:38 +08:00
  • 8b6d1b59e0 pi-vent as a friction-channel reference wassname 2026-08-13 19:07:16 +08:00
  • 2ffa93439c llm_judges: rubric-point judging section from 16 audit rounds of wassname-ml-bench wassname 2026-08-13 16:15:31 +08:00
  • 996942c4de Make ML debugging expose evidence and uncertainty wassname 2026-08-09 11:11:02 +08:00
  • 4e45bd4140 DPO can barely move: tinker's own reference run reports 0.52 pair accuracy wassnameandClaudypoo 2026-08-08 10:04:32 +08:00