Also turns the exercise selector into an explicit if/then table. 7 and 8 were
bundled under 'about to report a result'; they now have their own conditions.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Two different rules share the name fail fast. Nanda's is killing a doomed
direction early; this one is crashing on the error instead of carrying on.
Kept apart on purpose.
The order alone is unfollowable for an agent that has genuinely run dry.
Left out the 5-minute-experiment arithmetic, which gives the wrong number
for hour-long architecture runs.
His text, spelling fixed and voice kept, plus the three quotes from the cache
that back it: Bekman flagging his own overloaded heading, the tuning playbook
on two things sharing the name learning_rate, and Lones on which AUC.
gwern's own page concludes the tank story did not happen, so citing it
undercut the exercise. Zech is peer-reviewed with the in-site against
out-of-site AUC pair; the fastbook case covers the tabular version.
Comment review mode only, no prose changed. Flags negative framing,
aphoristic closers, and three places where the rewrite made wassname's
hedged claims stronger than his original message.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Textbook order: collected advice, then his comment on how it applies to
LLMs, then the exercises. Content left for him to write.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Small is under a paragraph; large means work like comparing against a
reference repo. Do all applicable small ones, pick one large one.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Quotes come from the docs/evidence cache and were previously unused. Reuses the
existing footnote style, 14 new keys. Carries the source doc's coverage warning:
mode 6 has only three quotes and mode 7 has none that name similarity probes.
Source is his own message of 2026-08-25, spelling fixed and slightly more polite as he
asked, with each mistake pointing at the exercise that answers it. Also adds his rule
that a job is never abandoned without doing the exercises, one at a time.
Two gaps the existing 13 did not cover, found by mining the evidence cache against
wassname's list of common AI-agent failures. Quotes are verbatim from
docs/evidence/ (Steinhardt, Rahtz, Nanda, Goodfellow-Bengio-Courville).
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Interpretations must read as evidence for or against a claim, never as a cause,
because a cause list gets picked from, called certain, and used to stop.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Both were promised by the description and absent from the procedure. P2 now
prints the formatted examples; P3 opens with the bug assumption.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
A sonnet subagent had ml-debug listed, with 'after a run finishes or crashes' in
the description, and did not invoke it for 'read the last pueue log'.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Sonnet routed to P1 correctly then declined it: 'the task only asked to read, not
to audit or act'. Reading a log is exactly when the measurement table gets filled.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
A sonnet subagent read the skill, summarised the ritual list, ran none of them,
and closed with 'everything checks out'. Phase triggers do not fire; a sentence
the agent watches itself write might.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
ml-bench measured standalone answers, not whether an agent in a loop follows
principles. Those are different claims and the first does not support the second.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
The run index is a cache key, not a seed, so subtracting run 3 from run 3 is
arbitrary. Unpaired the difference is +0.023 +- 0.044; by question +0.023 +- 0.031.
Still not pushed.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
+0.135 was one answer per arm, and it is draw 1 of three. Draws 2 and 3 read
-0.007 and -0.060, so the mean is +0.023 with sd 0.102. Not pushed: wassname
should read this before it goes public.
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
Quoted whole from the 2017 thread, since the article author asked to merge it
and never did. Verified line by line against the thread cache. Covers sample
size from a cumulative-mean plot, KLD/Dice on unbalanced data, augmentation
bounded by feature std, dummy metrics, jumpy validation loss, testing the
framework itself, activation swaps, and loss-curve shapes.
The author asked the commenter 'Do you mind if I add them to the article?' and
then did not: checked all 13 against the 2025 archived article, only the
batch-size point overlaps. Reddit blocks scrapers, so the thread came from a
Wayback snapshot. Also gave the article cache a real header: Medium is dead to
scrapers, so it now records the archive URL used to verify it.
The two README caches had no Source line, so nothing could re-verify them;
both now check at 99% coverage. The judge-bias file gets a warning at the top:
11 entries still carry summarizer numbers, and 2 of the 5 checked so far were
wrong.
102 of them, 48 already wrong after today's refetches. They rot every time a
cache is refetched and buy nothing the quote text does not: the caches are
verbatim, so the quote itself is the anchor. Descriptive labels stay.