pi-goals
Plan mode for agreeing on goals before any code gets written. Each goal names the subtle failure
mode that could fake a "done" and the discriminator that tells real success from it. Everything
lives in one markdown file, .pi/plan.md, which the agent edits with its normal edit tool and which
doubles as the task list. A goal is signed off only after a fresh read-only judge checks its
evidence against the repo.
The file has a fold at ## Log. Above it is the working set: title, the human's own words, goals,
discriminators. Below it is durable memory: the log, the learnings, and an unlimited unverified
appendix. The working set is re-sent when the plan goes stale for a couple of turns; the whole file
comes back at session start and after a compaction, which is when the settled context is gone.
Design bet (v2): the plan file is for LLMs and the human, not for TypeScript. There is no parser and
no schema. The format is a convention taught by a prompt; the judge, being a model, reads the file
natively, finds the claimed goal itself (wording drift is fine), and validates the evidence in
words. It cannot execute anything, so it never re-runs your verify command: you run it and save
the output as evidence. This deleted ~800 lines of v1 (parser, exact-string goal matching, verify
runner, JSON-stream judge transport, review menus) and with them the footguns they caused.
Like pi-milestones and burneikis/pi-plan, it guides rather than guards. The reminder cadence is copied from tintinweb/pi-tasks and the resync-after-compaction from tmonk/pi-goal-x. pi-lgtm was my earlier, more complex attempt.
Install
pi install npm:@wassname2/pi-goals
Or for development:
git clone https://github.com/wassname/pi-goals && cd pi-goals && npm install
pi -e ./src/index.ts
Use
/goals CSV export for the report view
/goals enters plan mode and starts a conversation; the objective is an optional seed. From there:
- Plan. The agent explores read-only (edit/write are blocked except on the plan file itself), asks
about anything unclear, and drafts the goals into
.pi/plan.md. The drafting rules are sent once, with your objective, not re-sent every turn. - Review. Read the file; a menu asks Ready, open in
$EDITOR, or keep planning. To revise, just reply. Plan mode ends when you pick Ready. - Work. The agent ticks subtasks, appends to
## Logand## Learnings, fillsevidence:, and callsCompleteGoalwhen a discriminator is satisfied. If it leaves the plan untouched for two turns, the working set is sent back with a short upkeep reminder.
Other commands: /goals clear empties the plan file; /goals judge <model-ref> picks a specific
model for the sign-off judge (default: your current session model, else pi's default).
Coming from v1: a leftover .pi/goals.md is renamed to .pi/plan.md on session start.
The plan.md format (a convention, not a schema)
# ship the cache layer
Latency target came from the SLO review; keep the existing client API.
## User voice
- > "keep the client API, I don't want to touch every call site"
## Goals
1. [/] goal: Implement cache layer
- subtle failure mode: cache silently bypassed, latency ok by luck
- discriminator: hit-rate > 0.8 in load-test.log (a bypass reads ~0)
- verify: pytest tests/cache -q && python bench/p95.py --max-ms 50
- tasks:
1. [x] wire cache client
2. [/] eviction policy
- evidence:
- > load-test.log: p95=41ms, hit-rate 0.93 (not bypassed)
## Future work / out of scope
## Log
- 2026-06-15 14:02 cache client wired; eviction next
## Learnings
- the client retries on 503, so a cache miss storm looks like latency, not errors
## Appendix (context, not approved)
- A goal is a checkbox line beginning
goal:([ ]open,[/]active,[x]done,[-]cancelled). Indented checkbox lines under a goal are its subtasks. Those two patterns plus the## Logfold are the only things the extension itself reads; everything else is prose to it. - The
discriminatoris the success test, written while planning: the positive observation that the goal succeeded and that none of thesubtle failure modes could fake.evidenceis the proof, filled at sign-off: each item pairs a durable artifact with a short read of it. Prefer committed artifacts (files, tests, diffs);.pi/is usually gitignored, so evidence there is judge-time proof only and won't survive in history. ## User voiceholds the human's requirements word for word. A paraphrase drifts, and then the goals churn on the next reply.## Learningsis what you now know, one line per gotcha, deduped.## Logis what happened, one line a turn.## Appendixis unlimited and unverified: alternatives, links, dead ends, and settled detail that is not part of the approved goals.- Small format deviations are fine; the file is read by the human and the judge, not a parser.
- The agent prunes finished goals itself when the working set gets long (evidence survives in git
history and
## Log).
Signing off a goal (CompleteGoal)
CompleteGoal(goal) is the one blessed tool. It spawns a strictly read-only pi subprocess (-p --no-session --no-extensions, tools read,grep,find,ls -- no bash, no edit/write) with the whole
plan file and the claimed goal. The judge cannot execute anything, so it never re-runs your verify
command (which may be a 10-hour training job); the agent runs verify itself and saves the output
as evidence. The judge reviews evidence discipline in order -- anything here at all? each item
quoted and attributed? provenance visible (how was this produced)? do the quotes match the cited
files on disk? -- and only then the substance, returning VERDICT: accept | reject plus what's
missing. It reads the live working tree, not HEAD: uncommitted work counts, and committing before
sign-off is for durable evidence, not for the judge's visibility.
- accept: a sign-off line is appended to
## Logand the goal is ticked[x]in the same write (exact goal-line match only; on wording drift the result asks the agent to tick it). The tool-written log line is the audit trail; a hand-tick without one shows in the diff. - every run saves the judge's full transcript to
.pi/judge/<stamp>.md, referenced from the log line, so "what did the judge actually check?" stays answerable after the fact. - reject: the goal stays open and the agent gets the missing list.
- judge ran but failed/errored/timed out, or returned no VERDICT line: accepted inconclusive,
logged as such. There is no pre-emptive "no model" path -- a null judgeModel just omits
--modelso pi's configured default runs the judge, so inconclusive always means "ran but failed", never "couldn't start". The working agent is never blocked on judge infra.
Prompts
All model-facing text lives in src/prompts.ts, in flow order.
Develop
pi -e ./src/index.ts # load locally
npm test # vitest: judge argv invariants, appendLog, decideSignOff fail-forward
npm run typecheck
npm run lint
License
MIT