Commit Graph
134 Commits
Author SHA1 Message Date
wassname 2871d89512 skill: the shorter runbook, 174 lines from 255 2026-08-24 21:39:04 +08:00
wassnameandClaudypoo dc369f5fac skill: fold in the design-doc gaps -- outcome signatures, persisted predictions, executed config, sample selection, guide caps
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:53:38 +08:00
wassnameandClaudypoo 2e1eefbba6 skill: raw artifacts in P1, seed-noise and baseline before A-beats-B, fresh auditor, follow the job
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:46:23 +08:00
wassnameandClaudypoo c94a450518 skill: permit 'unknown', require attribution when several things change at once
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:45:09 +08:00
wassnameandClaudypoo ff797a95e1 skill: keep failure interpretations, ban failure causes; history goes to the journal
Interpretations must read as evidence for or against a claim, never as a cause,
because a cause list gets picked from, called certain, and used to stop.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:41:02 +08:00
wassnameandClaudypoo e2bd28dbc2 skill: read your data and assume you have a bug were only in the frontmatter
Both were promised by the description and absent from the procedure. P2 now
prints the formatted examples; P3 opens with the bug assumption.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-20 06:39:36 +08:00
wassnameandClaudypoo f830c0cb23 skill: put literal trigger phrases in the description so it self-invokes
A sonnet subagent had ml-debug listed, with 'after a run finishes or crashes' in
the description, and did not invoke it for 'read the last pueue log'.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:55:43 +08:00
wassnameandClaudypoo a7ff779e88 skill: P1 steps 1-4 are read-only, so a 'just read the log' task still fills the table
Sonnet routed to P1 correctly then declined it: 'the task only asked to read, not
to audit or act'. Reading a log is exactly when the measurement table gets filled.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:52:56 +08:00
wassnameandClaudypoo f3f1a38485 skill: rewrite as a runbook -- P1-P5, numbered imperative steps, stop gates
Sonnet read the block version and summarised it instead of running it.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:51:10 +08:00
wassnameandClaudypoo c858469f59 readme: match the skill's wording, instructions not rituals
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:46:48 +08:00
wassnameandClaudypoo 8b092a4320 skill: drop the word ritual, headings become instructions, frame is a draft you revise
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:46:41 +08:00
wassnameandClaudypoo 4f422a8016 skill: bug-assumption gets its own trigger, the all-clear sentence; guide cap 400 -> 180 lines
A sonnet subagent read the skill, summarised the ritual list, ran none of them,
and closed with 'everything checks out'. Phase triggers do not fire; a sentence
the agent watches itself write might.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:44:49 +08:00
wassnameandClaudypoo d37b9c88a6 readme: state the ritual rework as an untested bet, not a finding
ml-bench measured standalone answers, not whether an agent in a loop follows
principles. Those are different claims and the first does not support the second.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:39:22 +08:00
wassnameandClaudypoo df4e08ef90 skill: add the audit fields I dropped -- trained/frozen, measure-before-diagnose, curve at 4 points, next-experiment case, contrary evidence
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:37:09 +08:00
wassnameandClaudypoo fb13b4fda7 skill: replace folklore prose with rituals -- trigger, form, artifact shown to user
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:31:18 +08:00
wassnameandClaudypoo 7dc8cfd23e readme: hold the folklore quotes, agents read prose and ignore it
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 20:30:08 +08:00
wassname d5d725e750 Revert "add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization"
This reverts commit ec2bb4f4be.
2026-08-18 19:02:26 +08:00
wassnameandClaudypoo ec2bb4f4be add LUCID v8 lessons: null needs positive controls + injection test; PASS needs mechanism check (cos, drop-k, read decodes); ceiling-probe before training; small-n contrast memorization
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-18 18:01:32 +08:00
wassnameandClaudypoo 52390d7593 readme: drop the run-by-run pairing, it was not a real pairing
The run index is a cache key, not a seed, so subtracting run 3 from run 3 is
arbitrary. Unpaired the difference is +0.023 +- 0.044; by question +0.023 +- 0.031.
Still not pushed.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-17 11:21:55 +08:00
wassnameandClaudypoo 1419c2e7df readme: the uplift did not replicate over three answers per question
+0.135 was one answer per arm, and it is draw 1 of three. Draws 2 and 3 read
-0.007 and -0.060, so the mean is +0.023 with sd 0.102. Not pushed: wassname
should read this before it goes public.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-17 03:52:28 +08:00
wassnameandClaudypoo 647b9a0145 llm_judges: Miller error bars -- repeat draws, don't touch the thermostat, paired differences
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-16 13:29:02 +08:00
wassname aa791fb839 readme: one measurement, not a rate 2026-08-16 07:16:17 +08:00
wassname 8ba59c54b8 readme: say plainly that only the cheap model reads the document 2026-08-16 07:15:47 +08:00
wassnameandClaudypoo ab9779ec00 readme: the skill closes 59% of the gap to gpt-5.6-sol on the same questions
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-16 07:13:32 +08:00
wassnameandClaudypoo e4e3386d1f readme: first measurement of the skill's effect, +0.135 on wassname-ml-bench
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-16 07:12:32 +08:00
wassname b2c666dbbf add wassname's 37-reasons checklist to refs/checklist.md
Quoted whole from the 2017 thread, since the article author asked to merge it
and never did. Verified line by line against the thread cache. Covers sample
size from a cumulative-mean plot, KLD/Dice on unbalanced data, augmentation
bounded by feature std, dummy metrics, jumpy validation loss, testing the
framework itself, activation swaps, and loss-curve shapes.
2026-08-15 06:40:45 +08:00
wassname f4d6fc28ca cache the 37-reasons reddit thread, which holds 13 checks the article never absorbed
The author asked the commenter 'Do you mind if I add them to the article?' and
then did not: checked all 13 against the 2025 archived article, only the
batch-size point overlaps. Reddit blocks scrapers, so the thread came from a
Wayback snapshot. Also gave the article cache a real header: Medium is dead to
scrapers, so it now records the archive URL used to verify it.
2026-08-15 06:38:01 +08:00
wassname 6c50496122 add missing cache headers, warn on the unverified judge entries
The two README caches had no Source line, so nothing could re-verify them;
both now check at 99% coverage. The judge-bias file gets a warning at the top:
11 entries still carry summarizer numbers, and 2 of the 5 checked so far were
wrong.
2026-08-15 06:23:54 +08:00
wassname 54dcd832b8 drop line-number anchors from cache citations
102 of them, 48 already wrong after today's refetches. They rot every time a
cache is refetched and buy nothing the quote text does not: the caches are
verbatim, so the quote itself is the anchor. Descriptive labels stay.
2026-08-15 06:23:31 +08:00
wassname ceba01782b replace four invented nanochat quotes with Karpathy's own words
All four came from the AI-generated deepwiki page, not the experiment log,
and one inverted the finding: the log says clipping was removed because
'Grad norm never exceeds 1.0 naturally', and the real distributed gotcha was
clipping local norms before sync. Footnote now points at dev/LOG.md on master
and at the renamed cache.
2026-08-15 06:22:53 +08:00
wassname 55726b56bd delete the spinning-up notes file, fix its last reference
Follow-up to the previous commit, which failed to remove the file because it
still had the abs -> pdf edit staged.
2026-08-15 06:21:49 +08:00
wassname 60ed9df651 drop the derived spinning-up notes, full-text the nanda draft
The notes file listed 11 links; 9 are already in the full page cache and the
80k Hours one has its own cache, so it only carried Rocktaschel, now in the
research_taste reading list. The Nanda shared draft was a 902-word excerpt of
a local download; wassname supplied the Google Doc, whose text export gives
the full 11318 words.
2026-08-15 06:21:37 +08:00
wassname a6c8ba77d2 cite arxiv pdf, never abs
An /abs/ link is a stub: it costs a second lookup before anyone can check the
quote, and it is what let abstract-only caches look sourced. 77 links across
the skill files and 4 cache headers. Verbatim source bodies untouched, their
reference lists are the authors' text.
2026-08-15 06:14:10 +08:00
wassname 2565f203e4 add evidence full-text verifier
Refetches the URL in a cache file's header and measures 5-word shingle
coverage of the source, so a summary cannot pass as a copy. Guards against
stub sources, where high coverage of an abstract or landing page would
otherwise read as success. n=5 chosen by sweeping n against synthetic
copy/paraphrase/summary variants.
2026-08-15 06:07:11 +08:00
wassname 383265c60c evidence: full text for nine LessWrong and research-blog caches
All nine said 'excerpted from HTML via browser' and held 141-684 words.
LessWrong posts refetched through the markdown API, the rest through jina.
All 99 previously quoted passages, and the 46 quotes SKILL.md and
refs/research_taste.md take from them, verify against the new full texts.
2026-08-15 06:07:11 +08:00
wassname 7fec3c557d evidence: full text for nine blog and essay caches
Every one was 3x to 100x shorter than its source: gwern_tank 179 -> 20132
words, cleanrl 37-details 239 -> 11410, kidger 185 -> 2932. Three passages the
old karpathy_recipe file presented inside quote marks were paraphrases, not
quotes, and are gone with the excerpts.
2026-08-15 06:07:11 +08:00
wassname 776ccf7047 evidence: full text for tuning playbook, ml yearning, bekman, spinning up, qwen3
All were excerpt caches. Tuning playbook 662 -> 16216 words, ML Yearning
527 -> 26084 (the whole book, was ch13-19), bekman 579 -> 2357 (three
concatenated pages), Spinning Up 842 -> 3353. The Qwen3 report was truncated
mid-section 4.3 and now runs to the appendix.
2026-08-15 06:07:11 +08:00
wassname d68ff6477a evidence: full nanochat experiment log, not a deepwiki synthesis
Was 1179 words summarised from AI-generated deepwiki pages, accessed 2026-03.
Now the 9232-word primary dev/LOG.md, which also runs 2 months further.
Renamed since the source is the log, not deepwiki.
2026-08-15 06:07:11 +08:00
wassname b1087b8efd llm_judges: verify the five abstract-quoted papers against raw PDF
NoLiMa effective length was wrong in both files: most models fall below the
85% threshold at 1-4K tokens, not 8-16K (only GPT-4o 8K, GPT-4.1 16K). NoLiMa
32K count was 11 models, not 10. The Loo quote's first sentence was stitched
from two places and is dropped. JudgeLRM's 8.14% is an Introduction number,
not an abstract one, and the abstract quote had lost /14B and its trailing
clause. Shi et al is 'The findings', not 'Our findings'. Body quotes now
replace abstract quotes where the number matters; tags [ID] -> [FT].
2026-08-14 21:16:13 +08:00
wassname ffcc94df00 evidence: full text for sculley 2015, lones 2021, domingos 2012
These three held hand-transcribed excerpts (381 / 914 / 573 words) with a
header claiming extraction was too hard. It was not: jina reads the sculley
and lones PDFs, and plain pdftotext (no -layout) reads all 3 columns of the
domingos CACM PDF in order. Now 5736 / 15285 / 8180 words.
2026-08-14 21:13:11 +08:00
wassname 701a09a525 pinn: correct claims that were written from abstracts only
Checked every 2001.04536 / 2402.01868 citation against the full texts now in
docs/evidence. Two quotation-marked sentences were not in the papers; the
Helmholtz 49x figure does not exist and the 64x Klein-Gordon one belongs to
architecture+annealing, not architecture alone; the Wang 1e5 Hessian number is
a max eigenvalue, not a max/min ratio. Section, table and figure pointers fixed.
2026-08-14 20:56:44 +08:00
wassname b70dcfa2b1 pinn evidence: replace abstract-only stubs with full paper text
wang2021 (2001.04536) and rathore2024 (2402.01868) held only the arXiv
abstract plus an assistant summary; fetched full text via jina.
2026-08-14 20:51:56 +08:00
wassname 3a8839b7d7 make judge setup repair a checklist 2026-08-14 15:42:47 +08:00
wassname cb7d597962 generalize judge setup repair guidance 2026-08-14 14:54:56 +08:00
wassname ad8c981504 teach judge audits to repair setup confusion 2026-08-14 14:41:38 +08:00
wassname 8b6d1b59e0 pi-vent as a friction-channel reference
Two details worth copying into a judge harness: scope the channel to systemic friction
and rule out one-off noise explicitly, and batch the entry at the end of the turn so
venting does not interleave with the work. -- Claude
2026-08-13 19:07:16 +08:00
wassname 2ffa93439c llm_judges: rubric-point judging section from 16 audit rounds of wassname-ml-bench 2026-08-13 16:15:31 +08:00
wassname 996942c4de Make ML debugging expose evidence and uncertainty 2026-08-09 11:11:02 +08:00
wassnameandClaudypoo 4e45bd4140 DPO can barely move: tinker's own reference run reports 0.52 pair accuracy
Contrast with the cookbook's RLHF pipeline (46% -> 94% win rate in 100 steps),
so a decreasing DPO loss is weak evidence of behaviour change.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-08 10:04:32 +08:00
wassnameandClaudypoo b5a70354a4 Judgemark: add the cheap-tier rows and date the snapshot
DeepSeek-V4-Flash is 0.368 (rank 28/36) at 0.78 while gemma-4-31b is 0.723 at
0.82, so the cheap judge slot is not where you would guess. Note that -latest
aliases have no row of their own.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-06 08:24:01 +08:00