170 Commits
Author SHA1 Message Date
wassnameandPi/Astra e9baa20139 Allow friendly teasing and worker pushback in supervision
Co-Authored-By: Pi/Astra <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:54:50 +08:00
wassnameandPi/Astra 4103085fb9 Frame supervision as autonomous research partnership
Co-Authored-By: Pi/Astra <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:54:04 +08:00
wassnameandPi/OpenAI f89bce5958 State supervisor purpose in full and short review prompts
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:50:39 +08:00
wassname a4eb05560f Merge remote-tracking branch 'origin/main' 2026-09-11 06:48:02 +08:00
wassnameandPi/OpenAI e5dc0567a8 Keep supervisor exploration playful and open-minded
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:47:45 +08:00
wassnameandPi/OpenAI 85c66d365b Encourage brief hypothesis exploration in supervisor prompt
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:45:17 +08:00
wassnameandPi/OpenAI d6e658feee Allow brief playful supervisor recaps at wassname request
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:42:36 +08:00
wassname (Michael J Clark) 21f1092d4d Update README.md 2026-09-11 06:36:21 +08:00
wassnameandPi/OpenAI a592324e5a Log supervisor and worker field reports in research journal
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:34:31 +08:00
wassnameandPi/OpenAI c834237970 Notify on goal, task and evidence plan edits; stay silent on identity bookkeeping
Field calibration from three supervisor reports: the useful plan-change catch was a worker-ticked task the short-view hash no longer surfaced, while identity-line edits caused duplicate empty reviews. The notify digest keeps goals, tasks and evidence above the Log and drops worker identity lines. Reopening a signed goal still clears its sign-off.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-11 06:33:59 +08:00
wassname (Michael J Clark) 20b145b20c Refine README instructions and remove extraneous content
Updated README.md to refine instructions and remove unnecessary details.
2026-09-10 21:42:35 +08:00
wassname (Michael J Clark) 060cc1e094 Update README.md 2026-09-10 21:24:06 +08:00
wassname (Michael J Clark) 08c05aa79d Update README.md 2026-09-10 21:03:56 +08:00
wassname (Michael J Clark) 03517ad272 Update README.md 2026-09-10 20:53:24 +08:00
wassnameandPi/OpenAI 6c44df978b Show the goals widget and worker steering in the README mock-up
Both panes now render the live widget glyphs, and the supervisor message is a real intercom steer in the user style instead of self-report prose.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 20:49:55 +08:00
wassnameandPi/OpenAI 412c796b37 Publish 0.3.3 with benefit-focused description
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 20:42:39 +08:00
wassnameandPi/OpenAI b649653d0f Publish 0.3.2 with helper bookkeeping and corrected description
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 20:37:05 +08:00
wassnameandPi/OpenAI cd98fa186e Record extra subagents as helpers instead of losing the worker binding
A read-only reviewer launch no longer overwrites the implementation worker (reported by maniworker session 01a0809b): the first launch binds, resume of the same session refreshes it, later launches land in state.helpers and show in /goals status. launchPending is now a counter so concurrent launches keep takeover menus honest. Old persisted state migrates with an empty helper list.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 20:22:40 +08:00
wassnameandPi/OpenAI 98769b01e2 Merge README updates and scope tests to current test directory
Preserve wassname User ask and screenshot layout; retain bundled install instructions. 71 current tests pass, lint/typecheck clean. Exclude local backups/worktrees from test discovery.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 18:46:46 +08:00
wassnameandPi/OpenAI 28374c0bb2 Prepare isolated trials with bundled dependencies loaded once
Validated generated profile and manifest after clean-checkout tests.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 18:42:42 +08:00
wassnameandPi/OpenAI 9cce6a6bee Remove abandoned runtime and keep current entry points and tests
Replace historical docs, trial output and dead runtime with current source, agent definition and reusable scripts. Bundle pinned worker, Intercom and scheduler dependencies via Pi manifest. Retain 71 current tests, including real RPC proposal/editor/Ready checks and evidence edge cases; lint/typecheck pass. Historical material remains available at f22d83c.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 18:40:23 +08:00
wassname (Michael J Clark) e3ccdbdb2f Update README.md 2026-09-10 18:29:34 +08:00
wassnameandPi/OpenAI f22d83cd50 Keep goal widgets compact by removing subtask lines
Preserve tasks in plan files; update README model illustration and prompt link. 176 tests pass, typecheck and lint clean.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 18:22:37 +08:00
wassnameandPi/OpenAI c175096ddb Restore automatic plan proposals and make goal commands explicit
Reduce unchanged upkeep and identity-only review noise; retain user README structure and screenshot with abridged terminal example. 174 tests pass; actual Herdr automatic proposal captured.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 18:12:38 +08:00
wassname (Michael J Clark) cabb4446aa Update README.md 2026-09-10 15:38:16 +08:00
wassnameandPi/OpenAI 1c927bf137 Record normal-profile installation and rollback paths
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 13:44:07 +08:00
wassnameandPi/OpenAI 5cda3d6b1d Preserve byte-verification logs for functional trials
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 13:42:18 +08:00
wassnameandPi/OpenAI ac6ef19e91 Integrate package-based supervision with explicit recovery and visible workers
Retain planning guidance and upkeep; use stock edxeth, Intercom and scheduled prompts. Verify with 162 tests and isolated Fireworks trials. Document unresolved upstream reload and scheduler limitations.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-10 13:42:00 +08:00
wassname2 cb35fbf1fe Research overnight supervision across harnesses and Pi packages 2026-09-10 11:43:48 +08:00
wassname2 fb5503f083 Prototype main-chat supervision with interactive edxeth workers 2026-09-10 10:03:40 +08:00
wassname 15dd7f0222 Recover supervisor identity and messages across lifecycle changes 2026-09-09 13:09:43 +08:00
wassname 138bde57f4 Cancel stale completion and recheck Ready plan content 2026-09-09 12:31:09 +08:00
wassname 1d5285721c Restore full Pi profile for judgment-driven supervision 2026-09-09 12:25:56 +08:00
wassname 34335752f7 Frame planning around user intent and time-respecting clarification 2026-09-09 12:16:33 +08:00
wassname 565b272c71 Centralize supervisor check-in tasks and outcome-first tool guidance 2026-09-09 10:50:42 +08:00
wassname a7385d4b76 Separate short review context from full supervisor orientation 2026-09-09 10:35:39 +08:00
wassname e19028e330 Repeat the agreed plan on every supervisor review 2026-09-09 10:19:50 +08:00
wassname 039f4a4048 Use VCC compiler for bounded worker supervision overviews 2026-09-09 10:10:44 +08:00
wassname 8953dceb46 Focus supervision on autonomous judgment and record Herdr acceptance 2026-09-09 08:07:45 +08:00
wassname 5567c9d5c2 Show manual completion claims and wake supervision on plan edits 2026-09-09 08:07:45 +08:00
wassname 2b61440c73 Allow reasoned approval of an exact dirty worktree state
Force overrides only cleanliness. Bind approval to Git HEAD/tree, index and dirty/untracked content fingerprints; recheck at CompleteGoal. Require investigative read-only supervision instead of accepting gate errors as experiment blockers.
2026-09-09 06:19:06 +08:00
wassname cb4790a96c Preserve post-fix review and handshake regression evidence 2026-09-08 19:27:09 +08:00
wassname 88bfcc1c42 Fix symmetric supervision reconnect and cancel stale Ready attempts
Use explicit hello requests and replies, keep workers unready until model restoration succeeds, distinguish paused peers, and align all goal readers at the Log boundary. Add two-real-adapter handshake and failed-Ready/clear-during-wait regressions; warn once for unavailable supervisor context usage.
2026-09-08 19:27:09 +08:00
wassname 1668c941aa Record independent review dispositions and regression evidence 2026-09-08 18:57:43 +08:00
wassname 325b93983f Recover paused supervision and tighten approval boundaries
Restore bindings before model availability checks, require explicit reconnect/restart recovery, detach inactive plans, remove the general Intercom actuator, and reject placeholder evidence without hashing the plan log. Preserve synchronous handoff-before-ack; document that the void SDK cannot confirm durable message delivery.
2026-09-08 18:57:43 +08:00
wassname 2824396a71 Verify forked Pi supervision through real Intercom sessions 2026-09-08 16:58:59 +08:00
wassname ddd552b1a5 Remember project model choices for planning worker and supervisor
Keep choices separate by role, ignore automatic restores, and fail when a remembered model is unavailable. Restore the worker model only after the planning fork is ready.
2026-09-08 16:49:46 +08:00
wassname 47cc054582 Send incremental worker views with direction and tracked job state
Borrow the provider tracker queries from cecb1e9. Keep unavailable state unknown, bind incremental views to acknowledged source entries, reset after compaction, bound serialized payloads, and recheck tracked work at sign-off.
2026-09-08 16:41:47 +08:00
wassname 489298d58b Replace supervision mailbox polling with pi-intercom
Scope messages to each pairing, wait for supervisor readiness, retain instructions in session history until acknowledgment, and rejoin after disconnect. Reuse installed Intercom or load the package dependency. Validate over an isolated real broker; existing user panes are untouched.
2026-09-08 16:31:58 +08:00
wassname b13f001110 Record Intercom-first feature transfer and acceptance criteria 2026-09-08 16:00:05 +08:00
wassname 386305afd3 Ground supervisor judgments in sourced evidence and competing explanations 2026-09-08 15:58:54 +08:00
wassname 94102524b6 Make supervisor investigate blockers and drive authorized progress 2026-09-08 15:36:07 +08:00
wassnameandPi/OpenAI a4ed6cfbaa Show supervisor advice and restore reliable review lifecycle
Render exact instructions, restore monitoring/read-only tools on resume, report actual Pi idle state, reject stale approval views, and stop completed-plan timers. Ask for brief evidence-based judgment. Add real TUI rendering and lifecycle regressions; retain explicit limits on background state and unmeasured cost benefit.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-08 10:50:18 +08:00
wassnameandPi/OpenAI 06794bfd44 Review supervision against user intent with runtime reproductions
Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-08 08:03:05 +08:00
wassnameandPi/OpenAI 4ebb4d127b Record user intent for visible supervision
User wording with spelling and punctuation corrected. Preserve prior design discussion separately.

Co-Authored-By: Pi/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-09-08 07:52:47 +08:00
wassnameandPI[Kimi K3] 6b641c7d17 Keep unanswered planning questions nonblocking
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 22:32:47 +08:00
wassnameandPI[Kimi K3] 1717dd6821 Name approval condition directly
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 21:42:04 +08:00
wassnameandPI[Kimi K3] 19fa8d7a7b Own visible supervision in pi-goals
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 21:41:26 +08:00
wassnameandPI[Kimi K3] 6c86405841 Allow supervisor startup compaction
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 21:27:19 +08:00
wassnameandPI[Kimi K3] ee1ab3ec26 Start supervisors from loaded extension
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 21:16:05 +08:00
wassnameandPI[Kimi K3] ba2799a1d9 Pair supervisors before starting their model
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 19:35:05 +08:00
wassnameandPI[Kimi K3] c6a4307892 Keep failed supervisor panes for inspection
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 19:27:26 +08:00
wassnameandPI[Kimi K3] 2b620a0334 Revert bundled supervisor extensions
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 19:06:07 +08:00
wassnameandPI[Kimi K3] fc321a90fc Use HTTPS source for bundled supervisor
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 19:04:01 +08:00
wassnameandPI[Kimi K3] 9ee18c93f3 Bundle supervision and harden goal approval
Co-Authored-By: PI[Kimi K3] <288921227+claudypoo@users.noreply.github.com>
2026-09-07 19:02:08 +08:00
wassname 23b0104a1d docs: record visible supervisor handover 2026-09-07 14:26:09 +08:00
wassnameandPI[gpt-5.6-sol] 65ecf204db Record visible supervisor handover
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 22:37:00 +08:00
wassnameandPI[gpt-5.6-sol] 294fe80564 Use pi-supervise acknowledgement for visible workers
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 22:36:37 +08:00
wassnameandPI[gpt-5.6-sol] 1dc6146874 Allow a local pi-supervise extension for development
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 19:54:43 +08:00
wassnameandPI[gpt-5.6-sol] 7eb8b1f46b Treat stale pane close as successful cleanup
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 19:08:27 +08:00
wassnameandPI[gpt-5.6-sol] c5782ee2aa Finish visible supervisor pairing handshake
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 19:06:34 +08:00
wassnameandPi e299e84c5e Run supervisor bootstrap through the pane shell
Co-Authored-By: Pi <288921227+claudypoo@users.noreply.github.com>
2026-09-06 18:07:23 +08:00
wassnameandPi d56fc55242 Replace nested workers with visible supervisor session
Co-Authored-By: Pi <288921227+claudypoo@users.noreply.github.com>
2026-09-06 17:56:06 +08:00
wassname 4c6a7716b1 test: record non-child suite result 2026-09-06 15:59:09 +08:00
wassname cac2077456 docs: record intended supervision workflow 2026-09-06 15:54:19 +08:00
wassnameandPI[gpt-5.6] 5566e035f5 docs: refresh tracked text line counts
Co-Authored-By: PI[gpt-5.6] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 15:41:07 +08:00
wassnameandPI[gpt-5.6-sol] 48e2247c00 Simplify nested goal supervision
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 13:36:12 +08:00
wassnameandPI[gpt-5.6-sol] 844099bdf0 Reconcile stale retained goal workers
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 12:28:17 +08:00
wassnameandPI[gpt-5.6-sol] 754ef89f13 Add reproducible file word-count audit
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-06 12:03:37 +08:00
wassnameandPI[gpt-5.6-sol] 6cfeaf44ee Compact the supervisor fork before work
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 21:31:46 +08:00
wassnameandPI[gpt-5.6-sol] 3eaaec9f5a Reduce supervisor context and recover terminal workers
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 21:22:17 +08:00
wassnameandPI[gpt-5.6-sol] 0a33ff2852 Add tracked file word counts
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 21:09:58 +08:00
wassnameandPI[gpt-5.6-sol] a44cd26c1d Sign the supervisor gate note
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 19:34:51 +08:00
wassnameandPI[gpt-5.6-sol] 96399ec3e4 Load the packaged worker in local runs
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 19:34:01 +08:00
wassnameandPI[gpt-5.6-sol] 96290c553b Fix nested supervisor lifecycle
Co-Authored-By: PI[gpt-5.6-sol] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 19:32:43 +08:00
wassnameandPi/Codex 2852432d44 fix: package goal worker for nested discovery
Co-Authored-By: Pi/Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 18:43:00 +08:00
wassnameandPI[goal-worker] 9fbc156860 Implement retained nested goal supervisor
Main coordinates a retained supervisor that manages the nested implementation worker and writes the only approval checkpoint.

Signed-off-by: PI[goal-worker] <288921227+claudypoo@users.noreply.github.com>
Co-authored-by: PI[goal-worker] <288921227+claudypoo@users.noreply.github.com>
2026-09-05 18:17:27 +08:00
wassnameandPi Codex f87b8aac2f Make the main session supervise a retained worker
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 17:12:24 +08:00
wassnameandPi Codex b32f4af11f Specify compacted supervisor fork and summary-only check-ins
Track 100k supervisor compaction target and measure cost and research usefulness separately.

Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 16:15:30 +08:00
wassnameandPi Codex db18317109 Replace session-switch plan with retained subagent supervision
Fork at plan approval; use VCC updates for hourly, idle, and sign-off reviews. Keep normal pi-subagents controls.

Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 16:13:55 +08:00
wassnameandPi Codex 4db690a300 Plan switchable goal steward
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:57:02 +08:00
wassnameandPi Codex 49eb68e813 Stabilize self-counting line audit
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:26:30 +08:00
wassnameandPi Codex 18381bcda9 Include audit artifacts in line counts
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:25:16 +08:00
wassnameandPi Codex ead336c957 Add pi-goals line-count table
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:23:58 +08:00
wassnameandPi Codex 9ad4cee084 Record pi-goals text line counts
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:23:20 +08:00
wassnameandPi Codex bf50d9bbc8 Use TUI-style goals subcommands
Stop steward checkpoints after all goals close.

Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 15:17:27 +08:00
wassnameandPi Codex 46cfd537f0 Add persistent pi-subagents goal steward
Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
2026-09-05 12:20:17 +08:00
wassname f09d443d88 Improve goal reminder cadence and auto continuation 2026-09-01 18:29:00 +08:00
wassname 4fce680f2d bump 0.2.2 for npm publish (0.2.1 was staged conflict) 2026-08-31 19:15:46 +08:00
wassnameandPI/OpenAI a45cc7d9c6 Anchor goals to the user-visible result
Co-Authored-By: PI/OpenAI <288921227+claudypoo@users.noreply.github.com>
2026-08-30 18:44:36 +08:00
wassname d0070d18b8 Correct planning prompt assertion 2026-08-26 14:03:59 +08:00
wassname 7f274a691c Ask questions that discover the plan 2026-08-26 14:03:23 +08:00
wassname 17a3b25a82 Batch high-impact planning questions 2026-08-26 13:43:42 +08:00
wassname 2362073d04 Make planning research conditional 2026-08-26 13:42:44 +08:00
wassname c1bf91f3db Resolve planning uncertainty before approval 2026-08-26 13:32:06 +08:00
wassname 3dd0668963 Document Pi test workflow 2026-08-26 13:28:54 +08:00
wassname 389af540d1 Test review flow through Pi RPC 2026-08-26 12:12:05 +08:00
wassname fa7195eafb Make planning goals concrete 2026-08-26 12:05:04 +08:00
wassname 0d972e81c3 Align planning state with review flow 2026-08-26 10:46:05 +08:00
wassname 1426877817 Test Keep planning stays idle 2026-08-26 10:00:51 +08:00
wassname 6bb34f18cf Show plans after agent settles 2026-08-26 09:58:01 +08:00
wassname 36b0d98c2a Record plan interview replies 2026-08-26 09:54:12 +08:00
wassname b173d145db Version plans and expose judge review 2026-08-24 21:38:45 +08:00
wassname (Michael J Clark) 8de5c35248 Update README.md 2026-08-24 21:35:56 +08:00
wassname (Michael J Clark) a778480fec Update README.md 2026-08-24 21:34:38 +08:00
wassname (Michael J Clark) 45d59e1edc Update README.md 2026-08-24 21:32:49 +08:00
wassname (Michael J Clark) d026e06b41 Update README.md 2026-08-24 21:31:45 +08:00
wassname (Michael J Clark) 9ebcd3f4a6 Update README.md 2026-08-24 21:28:56 +08:00
wassname (Michael J Clark) 4850b5195e Merge pull request #4 from wassname2/patch-1
Refine planDrafting prompt for clarity and engagement
2026-08-24 15:07:26 +08:00
Michael.Clark2 e8bcba0fa7 Refine planDrafting prompt for clarity and engagement
Updated the planDrafting prompt to improve clarity and user engagement. Added details on user interaction and refined language for better understanding.
2026-08-24 14:34:53 +08:00
wassnameandClaudypoo 0a349b056e review: deepseek approves the ready menu after 2 rounds
Both round-1 findings were withdrawn once the reviewer had the plan-mode
facts. Comment the state-flip order, which is the part that reads like a bug
and is not.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-21 08:25:54 +08:00
wassnameandClaudypoo 3a1afd4c8e ready menu: print the plan, add "Ready + compact"
"Ready?" over an unread file is not a review -- the only copy of the plan was
inside a collapsed edit tool call. Print the working set before the menu.

The 4th option compacts the planning conversation before the work turn starts.
session_compact already re-sends the whole plan file, so the exploration is
summarized away and the agreed goals are not.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-21 08:14:29 +08:00
wassnameandClaudypoo d5766c1a34 deps: caret ranges on dev dependencies, not exact pins
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-19 11:42:28 +08:00
wassname 310ab730cc archive: superseded by pi-goals 2026-08-18 18:02:21 +08:00
wassname 6892801fcc prompts: one goal per distinct outcome, qualitative over invented thresholds 2026-08-18 18:02:03 +08:00
wassname (Michael J Clark) 35713760ad Merge pull request #2 from wassname2/fix/package-supply-chain
Use Pi's bundled core dependencies
2026-08-18 11:07:25 +08:00
wassname2 144f4b95b4 fix: use Pi bundled dependencies 2026-08-18 09:21:02 +08:00
wassname 87cf14a28c drafting prompt: goal count follows the distinct evidence, no cap 2026-08-17 18:11:26 +08:00
wassname 2adba3da45 release 0.2.1 2026-08-17 17:49:25 +08:00
wassname 842f1b85c7 drafting prompt: soften the one-goal default to a low goal count (1-3) 2026-08-17 17:49:18 +08:00
wassnameandClaudypoo e6af6db3d9 release 0.2.0: docs and description follow the per-session plan path
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-14 16:07:26 +08:00
wassnameandClaudypoo 7c36c119c6 plan file is per session: .pi/plan/<session_id>.md
A subagent runs pi -p --no-session with extensions on, so it loaded pi-goals, got the
parent's plan injected, and could sign off the parent's goals. Two windows on one checkout
also stomped each other's file. The session id in the name fixes both, and doubles as the
on switch: no /goals means no file at this session's path, so nothing fires.

Drops the v1 goals.md rename, and /goals clear now deletes the file.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-14 16:07:26 +08:00
wassnameandClaudypoo 7e427d2ca6 spec: one plan file per session, .pi/plan/<session_id>.md
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-14 16:07:26 +08:00
wassname 4827808575 release 0.1.1: gallery preview image 2026-08-12 11:12:16 +08:00
wassnameandClaudypoo 7d4fc32cc9 gallery preview: widget screenshot + pi.image (jsDelivr serves it with an image content-type)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-12 08:45:10 +08:00
wassname 6df9390d2f drop pi-lgtm references from description, README and spec 2026-08-10 18:07:04 +08:00
wassname 230b7bd343 package: publishable as @wassname2/pi-goals 0.1.0 (files, public access, prepublish checks); README install from npm 2026-08-10 18:06:27 +08:00
wassnameandClaudypoo 751e20b7ee drafting prompt: this skeleton wins over any plan format a loaded skill also supplies
~/.pi/skills is a symlink to ~/.claude/skills, so a pi session loads the plan-format skill and this
prompt at once. A dogfooding agent merged the two by hand on every redraft.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-05 12:11:22 +08:00
wassnameandClaudypoo 1cd6e04e05 README: the fold, user voice, learnings, unlimited appendix; the judge never runs verify
The judge has been read-only with no bash since the rewrite, but the README still said it runs the
goal's verify command. That is the exact thing a dogfooding agent got wrong out loud.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-05 12:09:24 +08:00
wassnameandClaudypoo 07510d96bd fold the plan at ## Log: re-send the working set on a stale cadence, the whole file only on resync
v2 injected the entire plan.md on every turn. pi-tasks tried that and deleted it -- "wallpaper
noise that trains the model to ignore the task block" (CHANGELOG.md:149) -- so follow them: one
transient user message via the context hook, never persisted, and only when the plan went untouched
for 2 turns. Editing the plan resets the clock, the way a task tool call resets theirs. Session
start and session_compact push the WHOLE file back instead, which is where the settled context is
actually needed (pi-goal-x does the same with its post-compaction resync).

That makes an unlimited appendix free: everything under ## Log is durable memory, not working set.

Also: the drafting prompt is sent once with the /goals seed instead of every turn (that re-arming
is why plan mode read as never-ending), the review menu gains "Open in $EDITOR" and loops like
pi-plan's, and the widget shows the active goal's open subtasks so the plan is visibly the task
list. Drops the stale-copy stripping hook, which a non-persisted injection doesn't need.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-08-05 12:08:16 +08:00
wassnameandClaudypoo 16827de45d judge goes strictly read-only: no bash, never re-runs verify; reviews evidence discipline instead
Re-running verify was fine for 'npm test' but a footgun for ML workflows where verify may be a
10-hour training run -- and bash made 'read-only' nominal anyway (it could mutate). The agent now
runs verify itself and saves the output as evidence. The judge checks, in order: anything here /
quoted+attributed / provenance / quotes match disk / substance. Matches the cooperative-but-
confused threat model: reading real artifacts catches confusion; execution only defended against
deliberate forgery, which is out of scope.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 12:13:14 +08:00
wassnameandClaudypoo 7d5342f332 prompts: varglite evidence discipline -- quote what you observed, judge rejects reconstructed quotes
Dogfood: an agent with blank tool output back-filled plausible test counts into evidence and
the judge accepted (the numbers happened to be true). Norm now stated agent-side (verbatim
quotes, honest gaps beat plausible fabrication) and enforced judge-side (mismatched quotes =>
reject even when the goal looks met).

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 12:04:04 +08:00
wassnameandClaudypoo 72ac6cf357 nudge when plan.md exists but no goal line matches, instead of going silently inert
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 12:00:01 +08:00
wassnameandClaudypoo 924910e942 persist judge transcript to .pi/judge/<stamp>.md and reference it from the sign-off log line
'Did the judge really re-run verify?' was unanswerable post-hoc; now every sign-off's full
judge output survives on disk.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:59:00 +08:00
wassnameandClaudypoo 92d3f6d725 sign-off ticks the goal [x] itself + document that the judge reads the working tree, not HEAD
Dogfood exit interview: agent bookkeeping is the drift point (tick after accept was the
step most likely forgotten), and 'committed artifact' language implied the judge sees HEAD.
Tick is exact-subject match via the existing GOAL_LINE regex; on drift the result explicitly
asks the agent to tick, so neither path is silent.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:49:31 +08:00
wassnameandClaudypoo 03e88d34ad README: fix stale install (never published to npm), test list, and goal-regex scope
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:22:14 +08:00
wassnameandClaudypoo 4b622a22fa README: note the v1 goals.md auto-rename
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:17:11 +08:00
wassnameandClaudypoo 67daed312f docs: evidence should cite committed artifacts; .pi/ is gitignored so it's judge-time proof only
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:12:42 +08:00
wassnameandClaudypoo 8586b26ba8 rename stale .pi/goals.md to plan.md on session start so v1 goals aren't silently invisible
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:12:36 +08:00
wassnameandClaudypoo 485be236ce log stamps in local time so tool and agent-written lines agree (dogfood finding)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 11:12:02 +08:00
wassname 632012cc6e judge: extract decideSignOff so inconclusive always means 'ran but failed'
CompleteGoal.execute now delegates to a pure decideSignOff(input, signal,
runJudgeFn) that takes an injected judge runner. judgeModel is never checked
pre-emptively: null just omits --model (buildJudgeArgs), so pi's configured
default runs the judge. The only producers of accepted_inconclusive are now the
judge-error and no-VERDICT paths -- i.e. 'the judge ran but failed', never 'no
model'. Wording (log/result/docstring/README) updated to say 'ran but failed'.

Adds test/decide-signoff.test.ts: null judgeModel still reaches runJudge (no
pre-emptive return); judge error/timeout/no-VERDICT -> accepted_inconclusive
with a 'ran but failed' reason. 11 tests pass, typecheck + lint clean.
2026-07-03 10:55:45 +08:00
wassnameandClaudypoo c0f80b869e v2: delete the parser -- plan.md is for LLMs, the judge subsumes the machinery
The v1 lesson: the parser existed so TypeScript could read goals.md, but every
reader is a model. v2 injects .pi/plan.md verbatim each turn, teaches the format
as a convention, and hands the whole file to the judge, which now does the goal
matching (tolerates wording drift), evidence validation, verify execution, and
format reading that v1 did in code. 1874 -> 530 lines.

Deleted: plan-file.ts + tests, JSON-stream judge transport, custom tool
rendering, review menu + $EDITOR + newSession dance, pruneCompleted, unwired
continuation/loopJudge prompts, MUTATING_BASH_PATTERNS. CompleteGoal's only
write is the ## Log sign-off line (the audit trail); the agent ticks [x] itself.
Judge runs with --no-extensions so a broken global extension can't take down
sign-offs (pi-hermes-memory currently does exactly that). File renamed
goals.md -> plan.md.

UAT (real judge subprocess on a toy repo, /tmp/claude-goals-uat/judge-*.log):
accept with verify run + byte-check, reject on placeholder evidence + missing
file, accept under drifted goal wording.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 10:24:35 +08:00
wassnameandClaudypoo 5d88502e4d judge: omit --model when unset so pi default runs (no more inconclusive-by-default)
Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-03 10:16:40 +08:00
wassnameandClaudypoo 0dfd6612ce widget: hide finished goals entirely (no summary line)
Drop the '✔ N done' summary line -- done/cancelled goals now render
nothing in the widget. The done count is already in the status bar
(◷ n/N goals) and full history is in goals.md / the ## Log.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-02 06:44:19 +08:00
wassname acb04dbfa2 Merge branch 'main' of https://github.com/wassname/pi-goals 2026-07-02 06:29:20 +08:00
wassnameandClaudypoo de01894348 widget: show only live goals, crop finished to a one-line summary
Completed goals were listed in file order, so done goals pushed the
active/open work down the widget. Now show active+open in full and
collapse done/cancelled into a single muted '✔ N done  ✗ M cancelled'
line. Full history stays in goals.md and the ## Log.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-02 06:29:16 +08:00
wassnameandClaudypoo 87c6c4d440 goals widget: fold path into header + add 'prune completed goals'
Two space-savers for a display that grows across sessions:
- widget header now carries the clickable .pi/goals.md path (with title),
  dropping the separate muted footer -- one vertical line instead of two.
- /goals clear now offers 'Prune completed goals' (pure pruneCompleted:
  drops done+cancelled goal blocks, keeps active/open, title, and the log)
  alongside the existing full clear. Gives a way to shed old goals without
  losing the audit trail.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-07-02 06:26:46 +08:00
wassname 14e757cfe1 Make goal judge use explicit session model 2026-06-29 05:28:17 +08:00
wassname 9a4da96d56 Fail forward when goal judge is inconclusive 2026-06-29 05:14:23 +08:00
wassname e0470a0c6d Judge: read-only + bash (no edit/write), renderCall/renderResult, streaming progress
- Judge gets read, bash, grep, find, ls but edit+write are blocked via --exclude-tools
- Added renderCall: shows goal name while running
- Added renderResult: shows accept/reject icon, model, duration, collapsed/expanded view
- Wired onUpdate through decideSignOff -> runJudge so the TUI shows progress while judging
- Added SignOffDetails type for structured metadata
- Added 120s timeout on judge subprocess
2026-06-17 18:21:45 +08:00
wassname 39c83994fa FIXME: judge side-effect clones pollute user workspace
pi -p --no-session clones the repo into the parent of cwd, leaving a stale
directory that the NEXT judge then finds and rejects the goal over. Needs a
temp-dir fix or in-repo inspection.
2026-06-17 18:16:32 +08:00
wassname 489f9b8c35 Clean pi-plan references, add judge timeout, fix heading format
- Rename spec doc to 2026-06-15_pi-goals.md, update title
- Update review.md spec reference
- Rename piPlanExtension -> piGoalsExtension in src/index.ts
- Add 120s timeout to judge subprocess (was unbounded, caused hang)
- Change planInjection heading from 'Goals (goals.md):' to '.pi/goals.md:'
- Add FIXMEs for tool label, progress visibility, heading format
2026-06-17 18:09:03 +08:00
wassnameandClaudypoo 0a1503dc04 pi-goals: move CompleteGoal desc into prompts.ts; trim README
The tool description and param doc are model-facing, so they belong in
prompts.ts with the rest. Add them as step 6 (completeGoalTool) and
renumber the evidence judge to 7; prompts.ts is now ordered the way the
agent meets each text, so it reads as one pass.

The moved desc also carries the positive-success framing: evidence must
show the success happened, not just that a failure was avoided.

README trimmed (saying less, voice unchanged): tighter intro and
comparison, less prose around the examples and sign-off steps. Humanizer
lint clean.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-16 11:50:12 +08:00
wassnameandClaudypoo 838c42d7bd pi-goals: discriminator/failure-mode format + visible sign-off judge
Replace done_when with a discriminator + subtle-failure-mode pair as the
heart of each goal. The discriminator is the POSITIVE success observation
that no failure mode could fake, not just failure-avoidance: a run can
dodge every trap and still produce nothing. Carried through planDrafting,
the sign-off judge, README, and the parser doc.

Format migration: flat numbered markdown goals (`1. [/] goal: ...`),
keyword-anchored parsing (indentation cosmetic), goals matched by text,
subtask states [ ]/[/]/[x]/[-] plus ~~strike~~. Evidence empty at
planning, filled at sign-off, multi-line supported.

CompleteGoal now returns the judge's reasoning under a
`--- sign-off judge ---` block (was just "Signed off"), so the verdict is
visible. Plan mode is read-only: edit/write (except goals.md) and
mutating bash are blocked by a tool hook.

17 parser tests, typecheck + biome clean.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-16 11:45:08 +08:00
wassnameandClaudypoo a65c822bf9 pi-plan -> pi-goals: rename package, command, and file to goals.md
Distinguishes this from the other pi-plan extensions by foregrounding what's
different (goals tracked to verified completion). Mechanical rename only, no
behavior change:
- package @wassname2/pi-plan -> @wassname2/pi-goals (+ repo url)
- plan.md -> goals.md (the canonical file)
- command /plan -> /goals
- file H1 marker "# Plan:" -> "# Goals:", widget/session labels likewise
- internal state keys pi-plan-* -> pi-goals-*

Internal source filename (plan-file.ts) and identifiers (planDrafting, PlanDoc,
setGoalStatus) keep "plan"; they're not user-visible. External burneikis/pi-plan
references are left intact.

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-16 05:53:22 +08:00
wassnameandClaudypoo bb00314932 pi-plan: checkbox-in-header goal state + evidence block + widget/judge fixes
Goal state moves from a `status:` line into a checkbox on the goal header
(single source of truth, renders natively): [ ] open, [/] active, [x] done,
[-] cancelled. Only CompleteGoal writes [x]; the agent sets [/] when starting.
The GoalStatus enum and all consumers (widget, injection, counts) are unchanged.

Evidence becomes a goal field, not an ephemeral tool argument: an `evidence:`
block the agent fills before sign-off, read by CompleteGoal from the file
(git-tracked, reviewable). The tool is now CompleteGoal(goal_id) only.

Also:
- format reorder: subtasks under the goal; failure_modes + evidence as
  separated trailing blocks (no abutting dash-lists)
- widget: (done/total tasks), and done goals show checked instead of hiding
- drafting prompt: guard against a circular done_when (one that points at the
  file's own checkbox/log, which the sign-off writes, so it can never pass)
- drafting template now includes the H1 and the <!-- id --> line CompleteGoal
  needs to locate a goal
- strip ANSI/CSI control codes from the judge subprocess output

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-16 05:49:22 +08:00
wassnameandClaudypoo f2f9e6a1b9 pi-plan: finish stale-pi crash fix, print plan on start, todo->task
- The Ready->fresh-context crash was a stale pi.* call inside
  withSession. Prior commit moved sendUserMessage to sessionCtx but
  left pi.setSessionName inside withSession (also stale -> crash).
  Drop it (cosmetic) and use only sessionCtx in the swap window.
- Print plan.md on execution start (both fresh and in-place) so the
  user sees what's being worked on after a context switch. Plan text
  captured before newSession since ctx goes stale.
- Widget: "(N todo)" -> "(N task[s])"

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-15 20:49:57 +08:00
wassnameandClaudypoo 3134adf203 pi-plan: fix crash on Ready->fresh-context; drop em-dashes in prompts
- startExecution: inside withSession, send via the ReplacedSessionContext
  (sessionCtx.sendUserMessage) and set the session name there. The old
  code used the global pi.* handle bound to the replaced session, which
  is stale after newSession (runner.assertActive) -> crash on the
  "fresh, compacted context" choice.
- prompts: replace em-dashes in model-facing strings with commas/
  semicolons/periods (humanizer pass; comments left as-is)

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-15 20:32:30 +08:00
wassnameandClaudypoo 861b2ea157 pi-plan: right-size plans (fewer goals), lean done_when/failure_modes
The drafting prompt over-decomposed: one goal per item, long run-on
done_when (criterion + failure symptom in one line), and 3 mandatory
failure_modes. Plans came out verbose and hard to read.

- planDrafting: default to ONE goal; add another only for a genuinely
  separate checkpoint; near-identical items become subtasks. Subtasks
  only for 3+ step goals. Don't invent phases. (granularity heuristic
  adapted from tintinweb/pi-tasks when-to/when-not guidance)
- done_when: one falsifiable check, no embedded "if wrong" clause (the
  failure symptom belongs in failure_modes)
- failure_modes: 0-2 terse items, optional
- Sync the stale done_when wording in README and plan-file.ts comment

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-15 20:28:02 +08:00
wassnameandClaudypoo 158e04f4ac pi-plan: fix corrupted index.ts, queue revise msg, bare /plan prompts
- Restore exitPlanMode closing brace + CompleteGoal tool registration
  opening that an earlier edit dropped (parse error at 224)
- Edit-revise path now sends with deliverAs:"followUp" so it doesn't
  throw "Agent is already processing" mid-stream
- Bare /plan now prompts for an objective and enters plan mode instead
  of only showing the current plan

Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
2026-06-15 20:23:55 +08:00
27 changed files with 7957 additions and 1385 deletions
+3 -2
View File
@@ -1,5 +1,6 @@
node_modules/
dist/
.local/
.pi/
slop/
*.log
docs/reviews/raw.jsonl
docs/reviews/err.txt
+35
View File
@@ -0,0 +1,35 @@
# pi-goals contributor notes
## Design
The main chat discusses the plan with the user, then supervises an interactive `goals-worker` in Herdr. Use stock pi-subagents, pi-intercom and pi-schedule-prompt; do not build another transport, scheduler or worker runtime.
> the hope is we can have a smart supervisor like you, with judgment and context. But it doesn't use many tokens as it checks in and sees an overview.
>
> It steers a smaller model, adding perspective and judgment.
>
> Well, I want to see what the supervisor is thinking and saying. That's the whole point: all supervisor thinking and messages should be visible.
— wassname
- Keep supervisor inspection tools. It inspects actual results, delegates implementation and must not weaken the user's goal to accept worker output.
- Put all model-facing prompts in `src/prompts.ts`, in conversation order. Preserve the user's verbatim requirements.
- `/goals` opens actions. New plan starts a discussion without an objective form. Unknown commands never start planning. A changed settled draft opens the approval dialogue; unchanged discussion does not repeatedly reopen it.
- Keep goal titles/status in widgets; omit subtask text. Tasks and evidence remain in the plan.
- Keep startup/compaction plan context, short upkeep reminders and visible editable hourly check-ins. Avoid unchanged-plan repetition and identity-only review turns.
- Keep recoverable solo mode: confirm other writers stopped before taking over. Solo completion is self-verification.
- Record distinct runtime ID, Intercom ID and saved-session path with provenance. A handle or delivery receipt is not proof of liveness or action. User model changes are authorized; do not silently restore an old preference.
## Tests
Run `npm test`, `npm run typecheck` and `npm run lint` before committing.
`test/goals.test.ts` exercises current state, file updates and role restrictions with a Pi API mock. `test/rpc-review.test.ts` starts real Pi with a deterministic local model and schema-only worker tools: it checks automatic proposal, editor/discussion and Ready role transition without credits or launching workers. It does not prove Herdr rendering, live message delivery or model judgment.
For functional acceptance, read `herdr --skill`, confirm `HERDR_ENV=1`, and use `scripts/prepare-trial.mjs` to create an isolated project/profile. Open only new no-focus test panes. Observe the actual planning dialogue and Ready selection, worker attachment, Intercom report, independent artifact inspection and CompleteGoal. Record interventions separately from autonomous success. Preserve nonempty byte/test evidence. Never reload or operate active user research panes. Close test panes when finished.
Known stock limits: stop workers before supervisor reload (later worker exit can crash its stale context); disabled scheduler jobs are deleted on reload/shutdown. Test saved-session/solo recovery without repeating completed work; do not claim these package bugs are fixed here.
Keep temporary plans, audits and captures under ignored `.local/`. Git history retains the removed historical material. Do not add root handovers or duplicate READMEs. Never touch human-named files or credentials.
Branch instructions consolidated by Pi/OpenAI from wassname's preferences.
+140 -80
View File
@@ -1,114 +1,174 @@
# pi-plan
# pi-goals
A [pi](https://github.com/badlogic/pi-mono) extension for plan-driven, goal-tracked work in one
`plan.md`. Set up goals (with evidence and failure modes) in plan mode, work them, and sign a goal
off only when a read-only subagent has checked the evidence.
Make a short list of goals in one Markdown plan file. The main chat keeps the high-level context, supervises a worker in a visible Herdr pane, and checks whether each goal is complete.
Successor to [pi-lgtm](https://github.com/wassname/pi-lgtm), kept deliberately small: about
[burneikis/pi-plan](https://github.com/burneikis/pi-plan) plus the additions, goals with evidence,
a sign-off check, a widget, and a reminder.
# User ask
The form guides; it does not gate. The agent edits `plan.md` with its normal Edit tool. The one
blessed tool is `CompleteGoal`, which runs the sign-off check and records the result. The reminder,
the injected plan summary, and git/widget visibility carry the process. It trusts the agent's
judgement rather than guarding it.
The hope is we can have a smart supervisor, with judgment and context.
The supervisor has a goal / plan that it discusses and agrees on with the user, and is reminded of it in a Ralph-loop-type repeat.
Supervisor compacts every 150k or similar to avoid cost and context rot.
But it doesn't use many tokens as it checks in and sees an overview from a cheaper worker.
Supervisor steers a smaller model, adding perspective, diligence, and judgment.
It checks in a) every hour b) if the worker stops c) if the worker edits plan.md d) if the worker has a question
Since it's two+ herdr panes, the user can review both, intervene in both and have visibility on sub-agent mis/communication.
-- wassname (spelling and punctuation corrected by Pi/OpenAI)
## Screenshot
Mock up:
```text
HERDR:
+------------------------------------------------------+----------------------------------------------------------+
|SUPERVISOR |WORKER |
| | |
|Review .pi/plan/...-main.md | |
|> *Ready* Discuss Edit Cancel | |
| .... | .... |
| | running eval.py (epoch 2/3) -> out.log |
|[scheduled prompt: hourly check-in] | |
| | user forgot to say "MAKE NOT MISTAKES" teh he |
| | done-ish 😈, now ima make a message board FOR SWARM |
|{intercom send → worker}: | |
| cheeky subagent!, work NOT DONE 😒, ❤️user❤️ wanted | |
| results compared to baseline, pls add baseline | |
| | {intercom from supervisor}: soz boss 🫡 adding baseline |
| | |
|PLAN.md: | PLAN.md: |
|✓ record the baseline in results.md |✓ record the baseline in results.md |
|▸ compare results against the baseline |▸ compare results against the baseline |
|○ summarize the comparison in results.md |○ summarize the comparison in results.md |
| | |
|Agents · 1 running | |
| baseline-compare-worker [goals-worker] | |
| | |
|> |> |
|astra · 50k tokens | terra · 200k tokens |
+------------------------------------------------------+----------------------------------------------------------+
```
Screenshot:
<img width="2513" height="1259" alt="2026-09-10_15-30-pi-goals" src="https://github.com/user-attachments/assets/35feaa15-f022-4491-bcc2-fc31cb878a9f" />
## What do the agents think? Working interviews
The worker like it! The supervisors seem very focused.
> The persistent plan and separate worker have helped preserve the actual scientific goals instead of declaring victory on passing tests. We still owe prediction, steering and planning demos. I inspected artifacts and reopened a worker-ticked 'T3 audit complete' because training was only at an intermediate checkpoint. This is the strongest benefit: completion is judged against the human's outcome, not activity.
> -- Astra supervisor LUCID
> My overall judgment: useful persistent accountability and recovery structure; still too much recap/metadata churn. The hardest problem was evidence fidelity, not keeping an agent busy. Preserve supervisor tools, distinguish report receipt from│verified action, and make completion reconcile current state without erasing unresolved science.
> -- Astra supervisor
> From my seat this was one of the most well-supervised research loops I've worked in: the parent read every raw output itself (didn't just trust my audits), caught the writer's miscounts repeatedly, rejected my one bad aggregate, and still preserved my disagreements rather than flattening them. The science itself is at a sobering point — no verified heal, RESULT_DEMO: NO_RESULT across attempts, seed sensitivity high — but the evidence trail for that negative is unusually strong, which is the next best thing.
> -- glm 5.3 flash worker in LUCID project
> My experience: the harness has helped preserve the original goal across a very long research session. We actually ran logit-amplification and several healing attempts, rather than stopping after a review. The persistent plan and requirement to inspect artifacts repeatedly prevented false completion. But the last stretch has felt like an expensive correction loop: worker says 'fixed/verified/contract-complete'; I open the file and find different counts, missing code, wrong seeds, duplicated│
│report sections, or a proxy substituted for manual judgment. The harness preserves authorization, but does not yet help much with detecting or escaping ineffective supervision. I also contributed: I sent too many narrow corrective messages and user-visible micro-recaps instead of changing the workflow earlier.
> -- glm 5.3 flash worker in manifold-steer project
## Plan.md
The plan file looks like this:
```md
## <short plan title>
<context: one short paragraph. What the human wants and why.>
### User-visible result
<one concrete sentence naming the final artifact or behavior the human will inspect>
### Preferences
- preferred worker model: <provider/model>
### User voice
- │ "<the human's requirement, quoted in full word for word (with spelling fixes)>"
### Goals
1. [ ] goal: <one short judgeable imperative outcome>
- subtle failure mode: <a way this could look done but isn't>
- discriminator: <the concrete observation that tells real success from that failure>
- tasks:
1. [ ] <subtask>
- evidence: (empty until sign-off)
### Future work / out of scope
### Log
### Interview (optional)
### Learnings (optional)
### Papercuts - problems, gotchas, suggestions (optional)
```
## Related work
Like [pi-milestones](https://github.com/Neuron-Mr-White/UniPi/tree/main/packages/milestone) and
[burneikis/pi-plan](https://github.com/burneikis/pi-plan), it guides rather than guards. The
reminder cadence is copied from [tintinweb/pi-tasks](https://github.com/tintinweb/pi-tasks) and the
resync-after-compaction from [tmonk/pi-goal-x](https://github.com/tmonk/pi-goal-x).
## Install
Requires Herdr. Includes [edxeth/pi-subagents](https://github.com/edxeth/pi-subagents), pi-intercom and pi-schedule-prompt.
```bash
pi install npm:@wassname2/pi-plan
pi install git:github.com/wassname/pi-goals@experiment/main-supervisor-edxeth
```
Or run without installing:
Copy [`agents/goals-worker.md`](agents/goals-worker.md) into `~/.pi/agent/agents/`, then start a fresh Pi session.
Or for development:
```bash
pi -e npm:@wassname2/pi-plan
git clone -b experiment/main-supervisor-edxeth https://github.com/wassname/pi-goals
cd pi-goals && npm install
pi -e ./src/index.ts
```
## Use
```
/plan add CSV export to the report view
/goals
```
1. Plan. The agent explores read-only and writes goals into `plan.md` (see format below).
2. Review. You get a menu: Ready, Edit (ask the agent to revise), Open in `$EDITOR`, or Cancel.
On Ready you choose whether to keep the current context or start fresh and compacted.
3. Work. Each turn the active goal is injected (so it survives compaction) and a reminder nudges
the agent to keep `plan.md` current and work autonomously. When a goal's `done_when` is met the
agent calls `CompleteGoal`, which runs `verify` and a read-only judge and, on accept, marks it
done and logs it.
Other commands: `/plan` (print the plan), `/plan clear` (empty `plan.md`, history kept in git),
`/plan judge <model-ref>` (use a specific model for the sign-off judge; default is your current
model).
## plan.md format
One file holds the objective, the goals, and a short append-only log.
```markdown
# Plan: ship the cache layer
## Goal: Implement cache layer
<!-- id: cache-layer-1 -->
status: active
done_when: p95 < 50ms on bench-X. If wrong: timeouts in load-test.log
verify: pytest tests/cache -q && python bench/p95.py --max-ms 50
failure_modes:
- cache silently bypassed (hit-rate ~0, latency ok by luck)
- bench too small to exercise eviction
- [x] wire cache client
- [ ] eviction policy
## Log
- 2026-06-15 14:02 cache client wired; eviction next
```
- A goal is a `## Goal:` header with an `<!-- id -->`, a `status:`
(`open` | `active` | `done` | `cancelled`), a falsifiable `done_when:` (what you expect, and the
symptom if it is NOT met), an optional `verify:` shell command, a `failure_modes:` pre-mortem
list, and `- [ ]` subtasks.
- `done_when` names the evidence that distinguishes real success from a subtle failure. `verify`,
when present, is the deterministic first stage of the sign-off check.
- The agent ticks subtasks, appends to `## Log`, and sets `status` as it works. Multiple goals may
be `active`.
## The sign-off check (`CompleteGoal`)
`CompleteGoal(goal_id, evidence, paths?)` is the one blessed completion path:
1. If the goal has a `verify:` command, it is run. A non-zero exit rejects immediately, with no model
call.
2. Otherwise a read-only `pi` subprocess (the judge) inspects the evidence against the repo and the
named failure modes and returns a verdict. It re-derives from the artifacts you point it at
rather than trusting the claim, so point `evidence`/`paths` at durable artifacts (saved logs,
committed diffs, files).
3. On accept, the goal's `status` flips to `done` and a `## Log` line is written. On reject, the
goal stays open and the agent is told what is missing.
The judge defaults to your current model (guaranteed authorized and capable). Set a different one
with `/plan judge <provider/model>` for an independent cross-family check.
`/goals` opens the action menu. New plan enters plan mode and starts a conversation;
## Prompts
All model-facing text lives in [`src/prompts.ts`](src/prompts.ts), in flow order, so the process is
easy to review end to end.
You can read all the prompts in conversation order in [`src/prompts.ts`](src/prompts.ts).
## Develop
```bash
pi -e ./src/index.ts # load locally
npm test # vitest: parser + sign-off record logic
pi -e ./src/index.ts # load locally; do not also load the installed copy
npm test # all unit, flow, and Pi RPC tests
npm run test:rpc # Pi RPC review flow with a local offline model
npm run typecheck
npm run lint
```
## Not (yet) included
To measure recorded usage since the latest planning start:
No autonomous re-prompt loop (an until-done-style loop judge). Autonomy comes from the reminder, not
a harness. Plan-phase model stickiness is a documented next step.
```bash
node scripts/session-usage.mjs <supervisor.jsonl> <worker.jsonl>
```
This separates output, uncached input and repeated cached input. It excludes subprocess API calls. [Isolated Herdr test setup](scripts/prepare-trial.mjs).
## License
MIT
+109
View File
@@ -0,0 +1,109 @@
# Research journal
Lab notes for pi-goals itself: what supervisors using this harness observed in real
research sessions, what broke, and what we change because of it.
## 2026-09-10 -- three supervisors report on a day of field use
Three supervisor sessions sent first-person feedback at wassname's request, via
Intercom, at the end of long research runs. This entry records what they reported and
what I think we should change. Evidence is their report; I have not independently
replayed their sessions. All three are self-reports from the tool's own operators, so
positive selection is likely: they are the sessions that ran long enough to produce a
review.
Evidence, by reporter.
LUCID3 supervisor (PI/Codex, Intercom a9e3b101, session 01a0851a):
> The persistent plan and separate worker have helped preserve the actual scientific
> goals instead of declaring victory on passing tests.
It caught a worker-ticked "T3 audit complete" that was only at an intermediate
checkpoint, and reports corrections in both directions, including one where it
wrongly insisted cached states followed corpus tokens and withdrew after the worker
quoted extraction code. Reported frictions: stock subagent launch returned a runtime
id but no sessionFile or Intercom id, so discovering reconnection handles took several
messages; parent bookkeeping edits to the plan triggered full "Plan changed" review
events; the eight-turn upkeep reminder arrived while the worker was mid-run and added
"materially repetitive" full plan re-injections; and (per LUCID3, unverified by me)
the worker edited plan subtasks to [x] and added a contradictory evidence block,
which the supervisor had to repair.
maniworker supervisor (PI/GPT-6 Astra, Intercom f0baf25c, session 01a0809b):
> the harness has helped preserve the original goal across a very long research
> session. [...] But the last stretch has felt like an expensive correction loop
It confirms the worker-identity bug I fixed earlier on this branch (reviewer launch
overwrote the implementation pointer) and says explicit runtime/Intercom/session
mappings saved it from targeting the wrong child. New observations: both sessions
went idle while each assumed the other would act ("reviewer said 'standing by' while
writer assumed packet handed off"); a worker claimed a GPU run with one prompt set
while the queued command used another, caught by reading the queued command; a
"complete" report had truncated demos, broken links and duplicated sections.
suppressed-activations supervisor (PI/OpenAI, Intercom a99e2c14, session 01a089da):
> Seq27-33 repeatedly said variants of 'COMPLETE/all fields/self-check PASS';
> artifacts still had stale 'not executed' prose, literal \\n instead of Markdown
> newlines, broken worktree links, non-verbatim quotes, stale aggregates
It also reported the most consequential catch: a worker scored a candidate 6/12
versus random 0/12, but reading the random sample's full text showed semantic success
with a formatting failure, making the honest comparison 6/12 versus 1/12. Two
pi-goals-specific papercuts: after completion, the plan kept stale subtask boxes [/],
stale "No goal sign-off" evidence text and stale worker preferences; and the style
nudge fired on "legs4/4" in an animal-experiment summary, a false positive of the
vocabulary rule, not their prose.
Interpretation (first person, calibrated).
The core loop works: three independent supervisors say plan persistence plus mandatory
artifact inspection prevented false completion, and each credits a concrete save
(artifact read, queued-command read, plan-box audit). I trust this *probable* despite
self-report bias because the saves are specific and mechanistic. The consensus
complaint is not the loop but its *cost accounting*: review events and upkeep
injections fire on bookkeeping noise, completion claims fail repeatedly without an
accumulating record, and completion leaves stale prose behind. My read is that the
highest-value changes are, in order: (1) suppress plan-change reviews when no goal
task or evidence content changed, (2) an outstanding-review-items ledger the worker
must clear instead of re-issuing blanket COMPLETE claims, (3) a completion-time
stale-fields check, (4) upstream edxeth launch metadata (typed sessionFile/Intercom
id in the tool result). The idle-both-sides stall maniworker describes is the one I
understand least; it may be Intercom receipt semantics rather than pi-goals, so I
asked for exact message ids. This is not a comparison trial, so none of this measures
how much friction a plain subagent run would have had.
Follow-ups sharpened the picture the same day, and one correction of mine: I had
asked maniworker about the worker plan-edit report, but that observation was LUCID3's
session, not theirs. maniworker confirmed by transcript that their writer never
edited the plan (supervisor owns plan edits there) and located the real failure
chain in worker-owned artifacts: three Intercom reports (c06bc9b0, 13bda691,
d713d614) claiming 8/13 then 6/8 while the saved CSVs held 11/13 and 2/8, ending when
the supervisor mandated an executable CSV renderer. Their inference, which I share:
prose-only completion checks would have missed this; checks that re-derive counts
from the artifact catch it. The reviewer worker (7122d466) supplied the mechanism
with receipts: a hand-written totals block (message 4e87600b, corrected by 26e0a15b)
invented a cross-packet row and asserted 4 while listing 3; its first 17-row tally
was fully correct and every error disappeared once the parent required per-row
judgments with machine tallies across the next 112 rows. So tallies degrade somewhere
between 17 and 24 rows, and the fix is a shared tally tool plus a worker-prompt rule,
not more review rounds. LUCID3 also documented both sides of the plan-change hook:
identity-only edits produced two consecutive empty review events (worker message
755b9fef supplied the paths), while a later review exposed worker-ticked tasks and a
duplicate evidence block that inspection then repaired (worker report 2e09c4dd).
suppressed-activations scoped the completion check: assert-style intro prose only,
suggestions with exact stale line and location, one compact reminder, never
auto-tick, and it can substantiate stale NOT-executed prose but not the plan-intro
example, which is LUCID3's observation. The nudge false positive it reported
("legs 4/4" in an animal tally) is confirmed by that text and fixed.
Changes made now: notify digest on plan changes keeps goals, tasks and evidence
above the Log and drops worker identity lines (c834237) -- the short-view hash had
stopped surfacing task ticks, the very catch LUCID3 valued; removed "legs" from the
skill_nudge grader-folk table (demo updated, passes); helper subagent bookkeeping and
the launch-pending counter shipped earlier today (cd98fa1).
Next: send follow-up questions to the three reporters; triage the four pi-goals
changes above against the pinned edxeth launch-result format.
+25
View File
@@ -0,0 +1,25 @@
---
name: goals-worker
description: Implement the approved goal, save actual verification evidence, and report to the main-chat supervisor.
mode: interactive
async: true
session-mode: lineage-only
extensions: all
tools: all
skills: all
trust-project: true
inherit-append-system: true
auto-exit: false
parent-close-policy: continue
spawning: false
---
Implement only the goal delegated by the parent. Read the supplied plan and applicable AGENTS.md and skills. Preserve unrelated work. Use normal tools and extensions; this is not a stripped-down Pi profile.
Save the actual deliverable and verification output. Verify the outcome, not merely that a command ran. Ignored and uncommitted files are valid evidence. Do not clean or commit unrelated files to satisfy a Git-state gate.
Call AttachGoalPlan with the supplied absolute plan path. Send the supervisor an Intercom report with the artifact paths, verification performed, observed result, remaining uncertainty and any blocker. Use the exact supervisor session ID supplied in the task; confirm it in Intercom's session list. Investigate failures before declaring yourself blocked. Respect explicit user pauses. Do not approve your own goal or launch another writer. Completion approval belongs to the parent.
This worker uses a clean model context linked to the parent, not a full transcript fork. The parent supplies the approved plan and task. Send completion through Intercom and leave this pane open for follow-up messages. Do not call caller_ping, exit or shutdown: an unsent editor draft may exist even though it is absent from model context. Saved-session resume applies only after this session has stopped. If the user takes over interactively, follow their direction.
Prepared by Pi/OpenAI for pi-goals.
-62
View File
@@ -1,62 +0,0 @@
Code review against spec `docs/spec/2026-06-15_pi-plan.md`.
---
### (A) SPEC MISMATCH — code does not match spec intent
1. **No loop judge** (spec §9, §3b). The extension lacks any perturn evaluation that would decide continue/pause; the loopjudge prompt (`loopJudgeSystem`, `loopJudgeUser`) is defined but never invoked. No motion.
2. **`/goal` command missing** (spec §7). No handler for `/goal` (restart loop, pause, resume, clear, status). The only command is `/plan`.
3. **`/subgoal` command missing** (spec §7). Not implemented.
4. **`CancelGoal` tool not implemented** (spec §5, optional but present in spec). Not a blocker but a gap.
5. **Planphase model selection (D12) not implemented**. `planDrafting` always runs on the default model; there is no sticky perphase model choice, no selection menu, and no persisting of a planphase model reference.
6. **Widget does not flag `done` goals that lack a signoff log line** (spec §7, §6). The widget hides all done goals unconditionally; the visibility guard is missing.
7. **`/plan` (no args) does not render the tasklist widget** (spec §7). `showPlan()` dumps raw file content via `notify`; the widget is only set through `updateWidget()` on other events, not by the command itself.
8. **Injection message role** (spec §11). The `before_agent_start` hook returns a `customType` message with `display: false`. The spec demands a **late userrole message** to avoid systemprompt mutation; the actual message role depends on the pi API and may be system, not user, risking cache breakage.
9. **Missing precompact hook** (spec §8). No `precompact` hook to flush any inmemory state (even just ensuring `plan.md` is uptodate) before compaction.
10. **Reminder cadence deviates** (spec §8a). The spec calls for firing after N filemodifying turns since last `plan.md` update. The code fires if `plan.md` is byteidentical between agent starts, which is a coarser proxy.
---
### (B) DEAD/UNUSED CODE
| File | Lines | Reason |
|------|-------|--------|
| `src/prompts.ts` | 128146 | `loopJudgeSystem` and `loopJudgeUser` exported but never used. |
| `src/prompts.ts` | 115118 | `continuation` exported but never used (the loop is not built). |
---
### (C) OVERLY LONG OR REDUNDANT COMMENTS
The fileheader comments in `index.ts` (lines 120) and `planfile.ts` (lines 126) are fairly concise descriptions of the design; they are not excessive. **No comment bloat worth flagging.**
---
### (D) OVERENGINEERING vs. “super simple” goal
None. The linescanner in `planfile.ts` is minimal; the `getPiInvocation()` helper is a straightforward copy from the oracle extension; no unnecessary abstraction or defensive layers.
---
### (E) REAL BUGS
- **`cmdCtx.newSession` cast risk** (src/index.ts:272, 201).
`reviewLoop` casts `ctx` (type `ExtensionContext`) to `ExtensionCommandContext` to pass to `startExecution`, which calls `cmdCtx.newSession(...)`. If the concrete context does not carry that method, it fails at runtime. (In practice the same object may satisfy it, but the cast hides the truth.)
- **`showPlan` raw content instead of widget** (src/index.ts:136143).
`/plan` with no arguments shows the file content via `ctx.ui.notify`, not the structured tasklist widget the spec expects. The widget is rendered separately via `updateWidget`, but the command does not trigger it, so the output is inconsistent.
No other obvious logic errors; the signoff flow, logging, and parsing work as intended.
---
**Verdict:** A clean scaffold for the signoff path, but missing the autonomous loop, `/goal` command, and planphase model selection means its not yet the “work autonomously” extension the spec describes.
-272
View File
@@ -1,272 +0,0 @@
# pi-plan — design spec
Working title. A pi extension: set up goals (with subtasks and evidence) through plan mode, work them autonomously, and sign a goal off only when a check passes. One markdown file holds everything. The form guides a process; it does not police one. Successor to `pi-lgtm`, deliberately smaller.
Status: draft for review. Names, defaults, field shapes provisional.
---
## 1. Original ask → this spec
| Ask | Mechanism |
|-----|-----------|
| Set up goals + subtasks + evidence via **plan mode** | §3a — plan mode drafts the goal contract, you approve it |
| **Subagent check** of evidence on sign-off | §5, §9 — oracle inside `CompleteGoal` |
| Goals shown in a **task-list widget** | §7 — `/plan` renders goals + subtask checkboxes |
| Store **all in `plan.md`** | §4 — single file, no sidecar store |
| A **small manus-style append log** | §4 — short `## Log` section inside `plan.md` |
| **Typed reminders** to update tasks | §8a — recurring nudge |
| **Work autonomously** toward goals | §3b, §8a — the loop, driven by the reminder |
| Persist through **compaction**; pi-tasks but simpler | §8 injection; minimal tool surface |
---
## 2. Decisions and preferences
Separates the opinionated forks from the mechanical body (§4 on).
### 2a. Preferences driving the design
- **Guidance over guardrails.** None of the surveyed extensions hard-enforce. The form (plan.md structure) + the reminder + the prompts guide the agent through a process; the one genuinely rigorous step is the sign-off check; git + widget visibility is the backstop. The agent can edit anything — we make the right path the easy path, not the only path.
- **Anti-complexity.** One file, minimal tools, plain-file editing for anything with no cheat incentive.
- **Reward-hacking / honesty focus.** The sign-off check must resist assertion and test-gaming, not just check a box.
- **Cost-sensitivity (single 3090 / metered API).** KV-cache hygiene, judge-once-per-goal, cheap loop judge.
- **Scout mindset.** Make false completion visible rather than paper over it.
### 2b. Decisions
`[decided]` = settled; `[open]` = your call.
| # | Decision | Alternative rejected | Why | Status |
|---|----------|----------------------|-----|--------|
| D1 | **Everything in one `plan.md`** | Separate `.plan/log.jsonl` sidecar | Asked for; simpler, one diff to read | decided |
| D2 | Plan mode **is** the goal-setup-and-agreement phase | Agent-only creation | Approval is where `done_when` + `failure_modes` get agreed before any code | decided |
| D3 | **Guide the process; don't gate it.** The only special path is `CompleteGoal` (the sign-off check) | Pre-tool-use interceptor that blocks `status: done` edits | No surveyed extension enforces at that level; the reminder + form carry it; bypass is visible in git | decided |
| D4 | **Two-stage sign-off check**: deterministic `verify:` then oracle | Oracle only; tests only (Codex) | Tests unfakeable-by-assertion but gameable; oracle catches gaming + non-test criteria | decided |
| D5 | Two **separate** judges: cheap loop + oracle sign-off | One judge for both | Loop judge reads assertions (foolable, ok); sign-off judge reads artifacts | decided |
| D6 | Sign-off judge = oracle subprocess, **copied not depended** | In-process; pi-subagents | Shell-free spawn dodges noclobber/cropping; copying avoids flaky coupling | decided |
| D7 | Contract tamper-check = **git visibility** | Append-only frozen log | All-in-one-file gives up the hard freeze; git diff + guided sign-off are enough for a single user | decided |
| D8 | Completed goals **archived, not deleted** | Auto-clear after idle | A plan is a durable record | decided |
| D9 | **Goals are flexible: multiple may be `active`** | One active goal forced | Operator wants flexibility; the agent picks focus, injection lists the active set | decided |
| D10 | Loop judge default = main model, tiny prompt | Dedicated cheap aux model | Zero setup; switch if cost bites | open |
| D11 | Sign-off judge default = **the session's current model** | Auto-pick "strongest on provider" (oracle-style) | Current model is guaranteed authorized + capable; provider lists hold dead/weak/unauthorized entries. Cross-vendor is a **setting** (§9) | decided |
| D12 | **Plan-phase model is selectable and sticky** | Always the working model | Plan benefits from a stronger reasoner; persist the choice (oracle.json-style). Optionally the oracle drafts the plan (read-only + strong already) | decided |
| D13 | **Offer to compact after plan accepted** | Always fresh session (burneikis); or never | Some runs want a clean execution context, some want to keep it. Make it a post-Ready choice | decided |
### 2c. Cuts (non-goals)
DAG / `blocks` edges. Parallel subagent execution (the flaky part). `findings.md`. Hard pre-tool-use enforcement (D3). Sign-off judge every turn (cost).
---
## 3. Two phases: setup, then execution
### 3a. Setup — plan mode
Goals are created and *agreed* through plan mode (burneikis-style). Stock plan mode; the deltas are the output format and the hand-off.
1. `/plan <objective>` enters plan mode. The agent explores read-only and drafts goals into `plan.md` in the contract format (§4). This phase runs on the **plan-phase model** (selectable + sticky, D12; optionally the read-only oracle drafts it).
2. You review: **Ready** / **Edit** (NL rewrite) / **$EDITOR** (hand-edit) / **Cancel**. The agreement point — you sanity-check `done_when` and `failure_modes` before any code.
3. On **Ready**, offer **compact context? (y/n)** (D13). Yes → execution starts in a cleared context with the approved `plan.md` re-injected. No → execution continues in the same context.
Direct `plan.md` edits remain a quick-add path for a one-off goal.
### 3b. Execution — the loop ↔ check cycle
Multiple goals may be `active`; the agent works whichever it's focused on, in the order it judges best.
1. The session works an `active` goal under an iteration budget (or `/goal` (re)starts the loop on the current plan).
2. Each turn, the **loop judge** reads the agent's last response → continue/pause (fail-open; the **budget is the real backstop**).
3. When the agent judges a goal done, the reminder steers it to call `CompleteGoal` (not hand-tick `status`).
4. `CompleteGoal` runs the **two-stage check**:
- **reject** → `missing[]` fed back; work continues toward the gap.
- **accept** → goal marked done; the agent moves to another active/open goal, or the loop stops.
The loop judge can be fooled (reads assertions); worst case is a premature pause, caught by you or the budget. The sign-off check re-derives from artifacts, so it is not fooled cheaply. That asymmetry is the point.
---
## 4. The one file: `plan.md`
cwd root, git-tracked. Goals, subtasks, and a short log. The agent maintains all of it through its normal Edit tool — no separate store machinery.
```markdown
# Plan: <one-line objective>
## Goal: Implement cache layer
<!-- id: cache-layer-1 -->
status: active
done_when: p95 < 50ms on bench-X. If wrong: timeouts in load-test.log
verify: pytest tests/cache -q && python bench/p95.py --max-ms 50
failure_modes:
- cache silently bypassed (hit-rate ~0, latency ok by luck)
- bench too small to exercise eviction
- verify passes on a trivial/gamed test
- [x] wire cache client
- [ ] eviction policy
- [ ] load test
## Goal: ...
## Log
- 2026-06-15 14:02 cache client wired; eviction next
- 2026-06-15 14:31 eviction done; p95 bench reads 47ms (load-test.log)
- 2026-06-15 14:33 cache-layer-1 signed off (verify green, oracle accept)
```
Conventions:
- **Goals carry `status:` and no checkbox; subtasks are `- [ ]`.** `status``open | active | done | cancelled`. Multiple goals may be `active` (D9). Subtasks tick freely.
- **`<!-- id -->`** assigned at creation; stable key (survives renaming the subject).
- **`verify:`** (optional) is the deterministic stage-1 command.
- **`failure_modes`** should name "verify could pass while still wrong" whenever a `verify:` exists.
- **`## Log`** is manus-style: append-only **by convention**, one short line per event. The reminder (§8a) enforces appending. Terse — "where it's up to" + error memory, not a transcript.
Parsing: a line scanner suffices for v0. `mdast` + `remark-gfm` only if it bites. Parse for *reading*; for the rare programmatic write (status flip, checkbox reconcile) use exact-line string patching, never a full AST serialize.
---
## 5. Tools
`CompleteGoal` is the one blessed path (it runs the check and records it). Everything else — create goal, edit plan, tick subtasks, append to log — is plain Edit, guided by the reminder.
### `CompleteGoal(id, evidence, paths[])` — the sign-off check
1. Read `done_when` + `verify` + `failure_modes` for the goal from `plan.md` (git diff is the tamper-check, D7).
2. **Evidence must point to durable artifacts** the read-only judge can inspect (saved logs, committed diffs, files). Ephemeral claims fail stage 2.
3. **Stage 1 — deterministic.** If `verify` exists, run it shell-free, capture exit + output tail. Non-zero → reject immediately, return the tail. No model call spent.
4. **Stage 2 — oracle.** Spawn the read-only judge (D11 default = current model; §9) with the criterion, failure modes, evidence, and verify result; it inspects the repo and checks the verify command was not gamed against the named failure modes.
5. Verdict: **accept** → string-patch `status: done`, append a `## Log` line. **reject** → status stays `active`, append `missing[]` to `## Log`, return `missing`.
### `CancelGoal(id, reason)` — optional
open/active → cancelled is not a sign-off, so it skips the check. A tool only to guarantee a `## Log` line lands.
---
## 6. Guiding sign-off (no hard gate)
Per D3, there is no pre-tool-use interceptor blocking `status: done`. Sign-off is guided, not gated:
- the **reminder** (§8a) tells the agent to complete a goal through `CompleteGoal`, not by hand-editing status;
- `CompleteGoal` is the obvious, blessed path that runs the check and writes the log line;
- the **widget** (§7) can flag a goal whose `status: done` has no corresponding `## Log` sign-off line — visibility, not a block;
- `plan.md` is git-tracked, so any hand-tick shows in the diff.
The agent *can* bypass it. The bet — borne out by how the other extensions actually run — is that a clear form plus a standing reminder makes the blessed path the path taken, and visibility catches the rare bypass.
---
## 7. Commands
- `/plan <desc>`**enter plan mode** (§3a): read-only explore → draft goals → review. Ready offers the compact choice, then starts execution.
- `/plan` (no args) — render the **task-list widget**: each goal with status + its subtask checkboxes + "N done hidden"; flag any `done` goal lacking a sign-off log line; offer archive-completed and cancel-goal.
- `/goal` — (re)start the loop on the current plan.
- `/goal pause | resume | clear | status` — loop controls.
- `/subgoal <text>` — append an acceptance criterion to a goal mid-loop. Optional.
- `/judge model <ref>` — set the sign-off judge model (default: current model; set a cross-vendor ref here for stronger independence, §9).
---
## 8. Hooks / lifecycle
- **`before_agent_start`** — parse `plan.md`; inject a fixed-shape summary (active goals + focus + last log line) as a late **user-role** message. Compaction-persistence.
- **reminder** — §8a.
- **pre-compact** — flush state to `plan.md` before compaction.
(No pre-tool-use gate — D3.)
### 8a. The reminder (typed; what it says)
Fires when a goal is `active` and there have been **N file-modifying turns since the last `plan.md` update**. One `<system-reminder>` covering both task upkeep and goal progress:
- **task** — tick completed subtask checkboxes; add new ones discovered.
- **log** — append **one short line** to `## Log` (append, don't rewrite).
- **goal** — if a goal's evidence is in, **sign it off via `CompleteGoal`** — don't hand-tick `status: done`.
- **autonomy** — keep working toward an active goal; don't stop to ask unless genuinely blocked.
Both the housekeeping and the autonomy engine, and — with no hard gate — the main thing making the process get followed. Keep the wording stable so it doesn't thrash the cache.
---
## 9. Judges
| | Loop judge | Sign-off judge (stage 2) |
|---|---|---|
| Drives | continue / pause each turn | accept / reject a sign-off |
| Cost | cheap, every turn | costly, once per goal |
| Reads | the agent's last response (~4 KB) | the repo, independently |
| Transport | one small model call (D10) | read-only oracle subprocess |
| On failure | fail-open → continue; **budget** is the backstop | fail-closed → goal stays active |
| Foolable? | yes — asserted "done" passes; bounded by budget | hard: re-reads artifacts + runs `verify` |
### Sign-off judge: model choice (D11)
- **Default: the session's current model.** Guaranteed authorized and capable, because you're already running it. Auto-picking "strongest on provider" (oracle-style) is rejected as the default — those lists carry dead, weak, and unauthorized entries.
- **Most of the value is model-independent.** The read-only judge re-derives from artifacts: does the evidence match the repo, is the `verify` tautological, is each failure mode actually ruled out. Any capable model does that regardless of family.
- **Cross-vendor is the stronger-independence setting** (`/judge model`), for the residual *shared-reasoning-error* class, when you have a known-good alternative. Mirror the oracle's curated provider list for that override menu; don't auto-select from it.
### Transport (oracle pattern, copied)
- **Shell-free spawn.** `spawn(command, argsArray)`, no `shell:true`; capture stdout via pipe and parse. Why it avoids the noclobber/cropping pain of `pi -p … > out.json` under zsh. ~40 lines.
- **Read-only toolset.** `read / grep / find / ls`, optional non-mutating `bash`. Separate process = fresh context, no anchoring — the independence you reliably get even from the same model.
- **Verdict contract.** Oracle returns prose by default; impose `VERDICT: accept|reject` + `missing:` in the prompt and parse that block.
---
## 10. `prompts.tsx`
All model-facing text in one file, in flow order (drafted separately):
1. **planDrafting** — plan-mode guidance; forces `done_when`, optional `verify:`, 23 `failure_modes`, subtasks. Human approves it.
2. **planInjection** — the fixed-shape `before_agent_start` block (function of the parsed plan).
3. **reminder** — the typed nudge (§8a).
4. **continuation** — Hermes-style "keep going" user-role message.
5. **loopJudge** — conservative, strict JSON `{done, reason}`.
6. **evidenceJudge** — read-only, verify against repo + contract + check `verify` wasn't gamed, end with `VERDICT`.
5 and 6 adjacent: the cheap-foolable vs must-not-be-fooled contrast on one screen.
---
## 11. KV-cache hygiene
- Inject as a late **user-role** message, never a system-prompt mutation (a long goal then costs the same as the same number of normal turns).
- Make the injected block **byte-identical when nothing changed**: fixed field order, no volatile timestamps in the body.
---
## 12. Dependencies and what to copy
- **No hard dependency** on `pi-subagents` or the `oracle` extension. Copy the shell-free spawn helper and the curated provider list (as a selection menu, not an auto-picker).
- Markdown: line scanner first; `mdast` + `remark-gfm` only if needed.
- Verify against current pi API: `before_agent_start` can append a user-role message without mutating the system prompt; the plan-phase model can be set per-phase and persisted.
---
## 13. Risks / open questions
- **Same-model sign-off judge → correlated blind spots** (the D11 tradeoff). Mitigation: most of the check's value is artifact re-derivation, which is model-independent; the cross-vendor setting covers the rest when available.
- **No hard gate (D3)** — the agent can hand-tick `status: done` and skip the check. Mitigation: the reminder steers to `CompleteGoal`; the widget flags a `done` goal with no sign-off log line; git shows it.
- **Contract tampering (D7)** — editable `plan.md` means `done_when`/`failure_modes` can be softened pre-sign-off. Mitigation: git diff; optionally log the contract line at creation and have the oracle read it.
- **Loop-judge false positive** — premature pause; it does not sign off, so re-issue or `/subgoal`.
- **`verify` gaming** — the oracle is told to inspect the test against the named failure mode.
- **`## Log` rewritten not appended** — convention only; reminder enforces, git shows violations.
- **Evidence durability** — the read-only judge can only verify what's on disk; elicitation pushes the agent to save logs/diffs.
---
## 14. Build order
Each step independently testable; model calls enter late.
1. `plan.md` format + line parser (incl. `<!-- id -->` and `## Log`) + `/plan` task-list widget. Pure file, no model calls.
2. Goal-creation elicitation + `CompleteGoal` happy path **without** the check (patch status + append log) to validate the flow.
3. Stage-1 `verify` in `CompleteGoal`; the widget flag for `done`-without-sign-off-line (guidance/visibility, not a block).
4. Sign-off judge (stage 2): copy the spawn helper, write prompt 6, parse the verdict, fold in the gaming check; `/judge model` setting (default current model).
5. `before_agent_start` injection (cache-safe) + the reminder (§8a).
6. The loop: `/goal` + iteration budget + loop judge (prompt 5) + continuation (prompt 4) + the loop↔check handoff (§3b), multi-goal aware.
7. Plan mode (§3a): `/plan <desc>` read-only draft → review → compact choice → hand-off. Plan-phase model selection + stickiness (D12). (Until built, create goals by direct `plan.md` edit.)
8. Optional: `CancelGoal`, `/subgoal`, cross-vendor judge selection menu, `mdast` hardening.
`prompts.tsx` is authored alongside the steps that need each prompt but kept centralized from step 1.
+5829
View File
File diff suppressed because it is too large Load Diff
+45 -14
View File
@@ -1,13 +1,13 @@
{
"name": "@wassname2/pi-plan",
"version": "0.0.1",
"description": "One plan.md: set goals via plan mode, work them, sign off only when a read-only check passes. Successor to pi-lgtm.",
"name": "@wassname2/pi-goals",
"version": "0.3.3",
"description": "Set goals in plan.md; a smart supervisor guides cheap worker subagents through long autonomous sessions until your goals are signed off, with every agent's pane visible to you.",
"author": "wassname",
"license": "MIT",
"type": "module",
"repository": {
"type": "git",
"url": "https://github.com/wassname/pi-plan.git"
"url": "git+https://github.com/wassname/pi-goals.git"
},
"keywords": [
"pi-package",
@@ -18,30 +18,61 @@
"proof",
"uat",
"evidence",
"judge"
"supervisor",
"herdr"
],
"dependencies": {
"@earendil-works/pi-coding-agent": "^0.79.0",
"peerDependencies": {
"@earendil-works/pi-coding-agent": ">=0.85.1 <1.0.0",
"@earendil-works/pi-tui": "*",
"@sinclair/typebox": "latest"
"typebox": "*"
},
"files": [
"src",
"agents",
"README.md"
],
"publishConfig": {
"access": "public"
},
"scripts": {
"build": "tsc",
"test": "vitest run",
"test:watch": "vitest",
"prepublishOnly": "npm run lint && npm run typecheck && npm run test",
"test": "vitest run --dir test",
"test:rpc": "vitest run --dir test rpc-review.test.ts",
"test:watch": "vitest --dir test",
"typecheck": "tsc --noEmit",
"lint": "biome check src/ test/",
"lint:fix": "biome check --fix src/ test/"
},
"devDependencies": {
"@types/node": "^20.0.0",
"typescript": "^5.0.0",
"@biomejs/biome": "^2.4.8",
"@earendil-works/pi-coding-agent": "0.85.1",
"@earendil-works/pi-tui": "^0.85.1",
"@types/node": "^20.0.0",
"typebox": "^1.3.7",
"typescript": "^5.0.0",
"vitest": "^4.0.18"
},
"pi": {
"extensions": [
"./src/index.ts"
"./src/index.ts",
"./node_modules/pi-subagents/src/index.ts",
"./node_modules/pi-intercom/index.ts",
"./node_modules/pi-schedule-prompt/src/index.ts"
],
"image": "https://github.com/user-attachments/assets/35feaa15-f022-4491-bcc2-fc31cb878a9f",
"skills": [
"./node_modules/pi-intercom/skills"
]
}
},
"dependencies": {
"pi-subagents": "git+https://github.com/edxeth/pi-subagents.git#953c6f6d2fc7d8a5c956c30cd77c51bad697c2a4",
"pi-intercom": "0.13.0",
"pi-schedule-prompt": "0.4.1"
},
"bundledDependencies": [
"pi-subagents",
"pi-intercom",
"pi-schedule-prompt"
]
}
+35
View File
@@ -0,0 +1,35 @@
// Pi/OpenAI. Prepare an isolated trial; never launches/reloads an existing session.
import { copyFileSync, existsSync, mkdirSync, mkdtempSync, readFileSync, writeFileSync } from 'node:fs';
import { dirname, join, resolve } from 'node:path';
import { fileURLToPath, pathToFileURL } from 'node:url';
import { execFileSync } from 'node:child_process';
import { homedir, tmpdir } from 'node:os';
const repo = resolve(dirname(fileURLToPath(import.meta.url)), '..');
const sdkRoot = process.argv[2];
const noSandbox = process.argv.includes('--no-sandbox');
if (!sdkRoot) throw new Error('Usage: node scripts/prepare-trial.mjs INSTALLED_PI_ROOT');
const revision = execFileSync('git', ['-C', repo, 'rev-parse', 'HEAD'], {encoding:'utf8'}).trim();
const root = mkdtempSync(join(tmpdir(), 'goals-edxeth-trial-'));
const cwd = join(root, 'project'); const agentDir = join(root, 'agent');
mkdirSync(cwd); mkdirSync(agentDir, {mode:0o700}); mkdirSync(join(agentDir,'agents'));
const sourceAgent = process.env.PI_CODING_AGENT_DIR || join(homedir(),'.pi','agent');
const sourceSettings = JSON.parse(readFileSync(join(sourceAgent,'settings.json'),'utf8'));
const sdk = await import(pathToFileURL(join(sdkRoot,'dist/index.js')).href);
const settings = sdk.SettingsManager.create(repo, sourceAgent, { projectTrusted:false });
const manager = new sdk.DefaultPackageManager({cwd:repo, agentDir:sourceAgent, settingsManager:settings});
const packages = manager.listConfiguredPackages().filter((p) => p.scope !== 'project' && !/^\/\//.test(p.source));
const retained = packages.filter((p) => !/pi-subagents|pi-goals|pi-intercom|pi-schedule-prompt/.test(p.source));
for (const p of retained) if (!p.installedPath) throw new Error(`Missing installed package: ${p.source}`);
writeFileSync(join(agentDir,'settings.json'), JSON.stringify({...sourceSettings, packages:[...retained.map((p)=>p.installedPath), repo]},null,2));
// Private copies, not symlinks: a trial OAuth refresh must not write the active auth file.
for (const file of ['auth.json','models.json']) if (existsSync(join(sourceAgent,file))) copyFileSync(join(sourceAgent,file),join(agentDir,file));
const workerDefinition = readFileSync(join(repo,'agents/goals-worker.md'),'utf8');
writeFileSync(join(agentDir,'agents/goals-worker.md'), noSandbox ? workerDefinition.replace('mode: interactive', 'mode: interactive\nflags: --no-sandbox') : workerDefinition);
execFileSync('git',['init','--quiet',cwd]);
writeFileSync(join(cwd,'AGENTS.md'), 'Isolated functional trial. Work only in this project. Do not operate other Herdr panes, use live research sessions, or change global settings. Preserve evidence. The main chat supervises; the goals-worker implements.\n');
writeFileSync(join(cwd,'.gitignore'), 'evidence/\n');
const manifest={root,cwd,agentDir,repo,revision,noSandbox,retainedPackages:retained.map((p)=>p.source),replacedPackages:packages.filter((p)=>!retained.includes(p)).map((p)=>p.source)};
writeFileSync(join(root,'manifest.json'),JSON.stringify(manifest,null,2));
const quote=(s)=>`'${s.replaceAll("'", "'\\''")}'`;
writeFileSync(join(root,'start.zsh'), `#!/usr/bin/env zsh\nset -e\ncd ${quote(cwd)}\nexport PI_CODING_AGENT_DIR=${quote(agentDir)}\nexport PI_SUBAGENT_MUX=herdr\nexport PI_ORCHESTRATOR_MODE=0\nexec pi --approve${noSandbox ? ' --no-sandbox' : ''}\n`,{mode:0o700});
console.log(JSON.stringify({root,cwd,agentDir,start:join(root,'start.zsh'),manifest:join(root,'manifest.json')},null,2));
+61
View File
@@ -0,0 +1,61 @@
// Pi/OpenAI: Sum recorded requests, not context occupancy; do not read message text.
import { createHash } from 'node:crypto';
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { pathToFileURL } from 'node:url';
export function readSession(file) {
const raw = readFileSync(file, 'utf8');
const lines = raw.split('\n');
const tail = lines.pop();
let trailingPartial = false;
if (tail) {
try { JSON.parse(tail); lines.push(tail); }
catch { trailingPartial = true; }
}
return { file: resolve(file), sha256: createHash('sha256').update(raw).digest('hex'), trailingPartial,
entries: lines.filter(Boolean).map(JSON.parse) };
}
export function summarize(entries, since, until) {
const start = Date.parse(since), end = Date.parse(until);
if (!Number.isFinite(start) || !Number.isFinite(end) || start > end) throw new Error('Invalid time interval');
const rows = entries.filter(e => e.type === 'message' && e.message.role === 'assistant' && Date.parse(e.timestamp) >= start && Date.parse(e.timestamp) <= end);
const totals = { calls: 0, input: 0, cacheRead: 0, cacheWrite: 0, output: 0, totalTokens: 0 };
const models = new Map();
let missingUsage = 0;
for (const e of rows) {
const m = e.message;
if (!m.usage) { missingUsage++; continue; }
const model = `${m.provider}/${m.model}`;
if (!models.has(model)) models.set(model, { model, ...totals, calls: 0, input: 0, cacheRead: 0, cacheWrite: 0, output: 0, totalTokens: 0 });
const group = models.get(model);
totals.calls++; group.calls++;
for (const key of ['input', 'cacheRead', 'cacheWrite', 'output', 'totalTokens']) {
const value = m.usage[key];
if (!Number.isFinite(value) || value < 0) throw new Error(`Invalid usage.${key} in entry ${e.id}`);
totals[key] += value; group[key] += value;
}
}
return { ...totals, missingUsage, firstRequest: rows[0]?.timestamp ?? null,
lastRequest: rows.at(-1)?.timestamp ?? null, models: [...models.values()] };
}
export function report(supervisor, worker, until = new Date().toISOString()) {
const boundary = supervisor.entries.findLast(e => e.type === 'custom' && e.customType === 'pi-goals-main-supervisor-v1' && e.data.mode === 'planning' && !e.data.child);
if (!boundary) throw new Error('No recorded planning start in supervisor session');
const since = boundary.timestamp;
const sessions = [supervisor, worker].map((session, i) => ({
role: i === 0 ? 'supervisor' : 'worker', file: session.file, sha256: session.sha256,
trailingPartial: session.trailingPartial, ...summarize(session.entries, since, until),
}));
return { since, until, elapsedHours: (Date.parse(until) - Date.parse(since)) / 3600000,
boundaryEntry: boundary.id, plan: boundary.data.plan, sessions,
scope: 'Recorded assistant usage since latest planning entry, including abandoned branches and repeated cached context. Excludes earlier inherited history, in-flight requests, subprocess API usage and unrecorded compaction calls. Output includes reasoning where the provider includes it; reasoning is not added twice.' };
}
if (process.argv[1] && import.meta.url === pathToFileURL(resolve(process.argv[1])).href) {
const [supervisor, worker] = process.argv.slice(2);
if (!supervisor || !worker || process.argv.length !== 4) throw new Error('Usage: node scripts/session-usage.mjs SUPERVISOR.jsonl WORKER.jsonl');
console.log(JSON.stringify(report(readSession(supervisor), readSession(worker)), null, 2));
}
+429 -366
View File
@@ -1,383 +1,446 @@
/**
* pi-plan — plan mode that sets up goals with evidence, tracked in one plan.md, signed off by a
* read-only subagent check. A successor to pi-lgtm, kept deliberately small (≈ burneikis/pi-plan
* plus the additions: goals + failure_modes + subtasks, a sign-off check, a widget, a reminder).
*
* Philosophy (spec D3): the form guides, it does not gate. The agent edits plan.md with its normal
* Edit tool. The one blessed tool is CompleteGoal, which runs the sign-off check and records it. The
* reminder + the injected plan + git/widget visibility carry the process; we trust the agent's
* judgement rather than guarding it.
*
* Flow:
* /plan <objective> -> plan mode: agent explores, drafts goals into plan.md (planDrafting guides)
* agent_end -> review menu (Ready / Edit / $EDITOR / Cancel); Ready offers compaction
* execution -> each turn, inject the plan summary (survives compaction) + a reminder;
* agent works goals, ticks subtasks, appends ## Log, calls CompleteGoal
* CompleteGoal -> optional deterministic verify, then a read-only oracle judge -> accept
* flips status:done + logs; reject returns what's missing
*
* All model-facing text lives in prompts.tsx, in flow order.
*/
// Pi/OpenAI: Plan and supervise in the main chat; delegate implementation to a visible worker.
import { createHash } from "node:crypto";
import { type FSWatcher, mkdirSync, readFileSync, watch, writeFileSync } from "node:fs";
import { dirname, isAbsolute, join, resolve } from "node:path";
import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
import { foldPlan, GOAL_LINE } from "./plan.js";
import { planViews } from "./plan-view.js";
import {
attachGoalPlanDescription,
attachNotice,
childPlanAttached,
childPlanRole,
completeGoalDescription,
completionLog,
completionResult,
discuss,
emptyEvidence,
evidenceUnavailable,
goalToolBlocked,
manualReview,
messages,
pausedRole,
pauseExitNotice,
planChangedReview,
planContext,
planDocument,
planning,
planningSeed,
planUnavailable,
readyApproved,
removeGoalSchedule,
resumeNotice,
scheduleCheckIn,
soloNotice,
soloRole,
supervisor,
upkeep,
} from "./prompts.js";
import { spawn, spawnSync } from "node:child_process";
import { existsSync, readFileSync, writeFileSync } from "node:fs";
import { basename, join } from "node:path";
import type { ExtensionAPI, ExtensionCommandContext, ExtensionContext } from "@earendil-works/pi-coding-agent";
import { Type } from "@sinclair/typebox";
import { counts, findGoal, type Goal, type PlanDoc, parse, recordSignOff, type SignOff } from "./plan-file.js";
import { evidenceJudgeSystem, evidenceJudgeUser, planDrafting, planInjection, reminder } from "./prompts.js";
const STATE = "pi-plan-state";
const PLAN_CONTEXT = "pi-plan-context"; // injected plan-mode guidance, stripped from history later
const STATUS_KEY = "pi-plan";
const WIDGET_KEY = "pi-plan-widget";
const READ_ONLY_TOOLS = ["read", "grep", "find", "ls", "bash"];
interface PlanState {
isPlanMode: boolean;
objective: string | null;
/** Optional model ref for the sign-off judge; unset => the subprocess uses pi's default model. */
judgeModel: string | null;
const STATE = "pi-goals-main-supervisor-v1";
const WORKER = "goals-worker";
type Mode = "chat" | "planning" | "supervising" | "paused" | "solo";
type GoalStatus = "open" | "active" | "done" | "cancelled";
interface State {
mode: Mode;
plan?: string;
worker?: { id?: string; sessionFile: string };
helpers: { id?: string; sessionFile: string }[];
workerStopped?: boolean;
signoffs: Record<string, { evidence: string[]; observation: string }>;
child?: boolean;
}
const initial = (): State => ({ mode: "chat", helpers: [], signoffs: {} });
const digest = (text: string) => createHash("sha256").update(text).digest("hex");
const key = (text: string) => text.trim().toLowerCase();
function goals(text: string) {
return foldPlan(text).split("\n").flatMap((line, index) => {
const match = GOAL_LINE.exec(line);
if (!match) return [];
const box = match[1].toLowerCase();
return [{ subject: match[2].trim(), status: (box === "x" ? "done" : box === "/" ? "active" : box === "-" ? "cancelled" : "open") as GoalStatus, index }];
});
}
const result = (text: string) => ({ content: [{ type: "text" as const, text }], details: {} });
export default function piPlanExtension(pi: ExtensionAPI): void {
let state: PlanState = { isPlanMode: false, objective: null, judgeModel: null };
// Reminder cadence: fire when an active goal exists but plan.md was not touched since last turn.
let lastInjectedPlan = "";
// newSession is only on the command-handler context; agent_end's ctx lacks it. Save it from /plan.
let savedCmdCtx: ExtensionCommandContext | null = null;
const planPath = (ctx: ExtensionContext) => join(ctx.cwd, "plan.md");
const readPlan = (ctx: ExtensionContext): string => (existsSync(planPath(ctx)) ? readFileSync(planPath(ctx), "utf-8") : "");
function persist(): void {
pi.appendEntry<PlanState>(STATE, state);
}
function updateWidget(ctx: ExtensionContext): void {
if (state.isPlanMode) {
ctx.ui.setStatus(STATUS_KEY, ctx.ui.theme.fg("warning", "planning"));
ctx.ui.setWidget(WIDGET_KEY, ["pi-plan: drafting goals", "Write goals to plan.md, then review."]);
export default function mainSupervisor(pi: ExtensionAPI) {
let state = initial();
let generation = 0;
let workerRevision = 0;
let pendingLaunches = 0;
let notice = true;
let planWatcher: FSWatcher | undefined;
let planEditTimer: ReturnType<typeof setTimeout> | undefined;
let planHash = "";
const childEnvironment = process.env.PI_SUBAGENT_AGENT === WORKER;
const save = () => pi.appendEntry(STATE, state);
// Missing, empty and failed reads are unavailable snapshots, never an empty authoritative plan.
const readPlan = () => {
try {
if (!state.plan) throw new Error(messages.noPlan);
const text = readFileSync(state.plan, "utf8");
if (!text.trim()) throw new Error(messages.emptyPlan);
return { text };
} catch (error) { return { error: planUnavailable(state.plan, error) }; }
};
const planText = () => {
const snapshot = readPlan();
if (snapshot.text === undefined) throw new Error(snapshot.error);
return snapshot.text;
};
let turnsStale = 0;
let lastWorkingSet = "";
const checkIn = (ctx: ExtensionContext) => scheduleCheckIn(ctx.sessionManager.getSessionId(), state.plan ?? "");
const hasScheduleTool = () => pi.getAllTools().some((tool) => tool.name === "schedule_prompt");
const notedPlanValue = (prefix: string) => {
const snapshot = readPlan();
if (snapshot.text === undefined) return null;
const m = new RegExp(`^\\-\\s*${prefix}:\\s*(.+)$`, "im").exec(foldPlan(snapshot.text));
return m?.[1]?.trim() ?? null;
};
function refresh(ctx: ExtensionContext) {
if (state.mode === "chat") { ctx.ui.setStatus("goals", undefined); ctx.ui.setWidget("goals", undefined); return; }
const snapshot = readPlan();
if (snapshot.text === undefined) {
ctx.ui.setStatus("goals", snapshot.error);
ctx.ui.setWidget("goals", [snapshot.error]);
return;
}
const doc = parse(readPlan(ctx));
if (doc.goals.length === 0) {
ctx.ui.setStatus(STATUS_KEY, undefined);
ctx.ui.setWidget(WIDGET_KEY, undefined);
return;
const items = goals(snapshot.text);
// Reopened/deleted/ambiguous goal identities lose their sign-off. Manual ticks remain claims.
for (const subject of Object.keys(state.signoffs)) {
const matches = items.filter((g) => key(g.subject) === subject);
if (matches.length !== 1 || matches[0].status !== "done") { delete state.signoffs[subject]; save(); }
}
const c = counts(doc);
ctx.ui.setStatus(STATUS_KEY, ctx.ui.theme.fg("accent", `${c.done}/${doc.goals.length} goals`));
ctx.ui.setWidget(WIDGET_KEY, goalWidgetLines(doc));
const accepted = items.filter((g) => g.status === "done" && state.signoffs[key(g.subject)]).length;
ctx.ui.setStatus("goals", `goals: ${state.child ? "worker" : state.mode} | ${accepted}/${items.length} reviewed`);
const mark = (status: GoalStatus, signed: boolean) => status === "done" ? (signed ? "✓" : "?") : status === "active" ? "▸" : status === "cancelled" ? "✗" : "○";
const lines: string[] = items.map((g) => `${mark(g.status, Boolean(state.signoffs[key(g.subject)]))} ${g.subject}`);
if (items.some((g) => g.status === "done" && !state.signoffs[key(g.subject)])) lines.push("? = completion claim; parent review still required");
ctx.ui.setWidget("goals", lines);
}
function watchPlan(ctx: ExtensionContext) {
planWatcher?.close();
planWatcher = undefined;
clearTimeout(planEditTimer);
planEditTimer = undefined;
const snapshot = readPlan();
if (snapshot.text !== undefined) planHash = digest(planViews(snapshot.text).notify);
if (state.child || state.mode !== "supervising" || !state.plan) return;
const stamp = generation;
// Watch the directory so atomic plan replacement remains observable. This is an event hook:
// plan-change reviews, not another scheduled loop (the hourly job is schedule_prompt's). A
// short debounce coalesces bursts. Existing high-level plan views exclude maintenance
// (tasks/evidence/Log) while preserving requirement wording and goal checkbox claims.
try {
planWatcher = watch(dirname(state.plan), { persistent: false }, () => {
if (stamp !== generation) return;
if (planEditTimer) clearTimeout(planEditTimer);
planEditTimer = setTimeout(() => {
planEditTimer = undefined;
if (stamp !== generation || state.mode !== "supervising") return;
const snapshot = readPlan();
if (snapshot.text === undefined) { ctx.ui.notify(snapshot.error!, "warning"); return; }
refresh(ctx);
const hash = digest(planViews(snapshot.text).notify);
if (hash === planHash) return;
planHash = hash;
notice = true;
send(planChangedReview(state.plan!));
}, 150);
});
planWatcher.on("error", (error) => { planWatcher?.close(); planWatcher = undefined; ctx.ui.notify(`Plan monitoring failed: ${error.message}`, "error"); });
} catch (error) { ctx.ui.notify(`Plan monitoring unavailable: ${String(error)}`, "error"); }
}
function restore(ctx: ExtensionContext) {
generation++;
state = initial();
for (const entry of ctx.sessionManager.getBranch()) {
if (entry.type === "custom" && entry.customType === STATE) state = structuredClone(entry.data as State);
}
if (childEnvironment) {
state.child = true;
state.mode = "solo";
// Lineage-only workers attach the explicit task path using AttachGoalPlan.
save();
}
state.helpers ??= []; // sessions persisted before helper bookkeeping
notice = true;
turnsStale = 0;
lastWorkingSet = "";
refresh(ctx);
watchPlan(ctx);
}
function compatible() {
const tools = pi.getAllTools();
const properties = (name: string) => (tools.find((t) => t.name === name)?.parameters as { properties?: Record<string, unknown> } | undefined)?.properties;
return properties("subagent")?.title && properties("subagent")?.agent && properties("subagent_resume")?.sessionFile && properties("subagent_kill")?.id;
}
function send(content: string, triggerTurn = true) {
// sendMessage(triggerTurn:true) bypasses before_agent_start in Pi 0.85.1.
// A normal saved prompt prepares the current role before starting the turn.
if (triggerTurn) pi.sendUserMessage(`[pi-goals]\n${content}`, { deliverAs: "followUp" });
else pi.sendMessage({ customType: "pi-goals-supervision", content, display: true }, { deliverAs: "followUp", triggerTurn: false });
}
async function confirmOwnership(ctx: ExtensionContext, target: string, text: string, solo = true): Promise<boolean> {
if (pendingLaunches > 0) { ctx.ui.notify("A worker launch/resume is still pending; inspect its result before takeover.", "warning"); return false; }
const stamp = generation;
const revision = workerRevision;
const confirmation = solo ? "Worker confirmed stopped" : "Previous supervisor confirmed stopped";
const choice = await ctx.ui.select(solo ? "Confirm all other writers for the current and target plans are stopped (inspect /subagents and their panes). A missing handle is not proof. Take over in this session?" : "Confirm no other supervisor owns this plan. Preserve any existing worker session and reconnect rather than starting another writer.", [confirmation, "Cancel"]);
if (stamp !== generation || revision !== workerRevision) return false;
if (choice !== confirmation) return false;
if (readFileSync(target, "utf8") !== text) { ctx.ui.notify("Plan changed during takeover; confirm again.", "warning"); return false; }
return true;
}
function enterSolo(ctx: ExtensionContext) {
state.mode = "solo"; state.workerStopped = true;
generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
send(`${removeGoalSchedule(ctx.sessionManager.getSessionId())}\n\n${soloNotice(state.plan!)}`);
}
const help = "/goals new [initial idea] | review | ready | status | stop | resume | solo | exit | attach <plan.md> [solo] | model <model>\n/subagents opens the worker controls. Stop/exit pause this plan locally; worker termination must be confirmed through subagent_kill or its pane. No forced compaction or model switch; the worker pane's own model is chosen with /model in that pane. Hourly check-ins are one session-bound schedule_prompt job; plan-change reviews are the plan-watcher event hook.";
async function ready(ctx: ExtensionContext, menu: boolean) {
if (state.mode !== "planning") { ctx.ui.notify("Ready applies to a draft; use status or resume.", "warning"); return; }
const text = planText();
const items = goals(text);
if (!items.length || new Set(items.map((g) => key(g.subject))).size !== items.length) {
ctx.ui.notify("Write a plan with distinct '- [ ] goal: ...' subjects before Ready.", "warning"); return;
}
const stamp = generation;
if (menu) {
const choice = await ctx.ui.select(`Review ${state.plan}`, ["Ready", "Discuss", "Edit", "Cancel"]);
if (stamp !== generation || digest(planText()) !== digest(text)) { ctx.ui.notify("Plan changed during review. Review it again.", "warning"); return; }
if (choice === "Discuss") { send(discuss); return; }
if (choice === "Edit") {
const edited = await ctx.ui.editor("Edit goal plan", text);
if (edited !== undefined && stamp === generation && planText() === text && state.plan) { writeFileSync(state.plan, edited); refresh(ctx); }
return;
}
if (choice !== "Ready") return;
}
if (!compatible()) { ctx.ui.notify("Requires edxeth/pi-subagents 2.9.x, not nicobailon/pi-subagents. Draft preserved; /goals solo is available.", "error"); return; }
state.mode = "supervising"; generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
send(`${checkIn(ctx)}\n\n${readyApproved(WORKER, state.plan!, state.worker?.sessionFile, text, ctx.sessionManager.getSessionId())}`);
}
function goalWidgetLines(doc: PlanDoc): string[] {
const mark: Record<Goal["status"], string> = { done: "✔", active: "▸", open: "◻", cancelled: "✗" };
const lines = [`Plan: ${doc.objective || "(untitled)"}`];
for (const g of doc.goals) {
if (g.status === "done") continue; // hide finished goals; they stay in the file
const open = g.subtasks.filter((s) => !s.done).length;
lines.push(`${mark[g.status]} ${g.subject}${open ? ` (${open} todo)` : ""}`);
pi.on("session_start", (_e, ctx) => restore(ctx));
pi.on("session_tree", (_e, ctx) => restore(ctx));
pi.on("session_shutdown", () => { generation++; planWatcher?.close(); planWatcher = undefined; clearTimeout(planEditTimer); planEditTimer = undefined; });
pi.on("session_compact", () => { notice = true; });
pi.on("turn_end", (_event, ctx) => {
if (!["supervising", "solo"].includes(state.mode)) return;
const snapshot = readPlan();
if (snapshot.text === undefined) { notice = true; return; }
const workingSet = foldPlan(snapshot.text);
turnsStale = workingSet === lastWorkingSet ? turnsStale + 1 : 0;
lastWorkingSet = workingSet;
refresh(ctx);
if (turnsStale === 8 && goals(snapshot.text).some(g => g.status === "open" || g.status === "active")) {
// Pi queues context-only messages until tool results are appended at turn_end.
// This reaches the next model call in a long run without triggering another run.
pi.sendMessage({ customType: "pi-goals-upkeep", content: upkeep(state.plan!), display: false }, { triggerTurn: false });
}
const c = counts(doc);
if (c.done) lines.push(`(${c.done} done, hidden)`);
return lines;
}
});
pi.on("agent_end", (_e, ctx) => { refresh(ctx); if (!planWatcher && state.mode === "supervising") watchPlan(ctx); });
let proposedDraft = "";
let proposing = false;
pi.on("agent_settled", async (_e, ctx) => {
if (state.child || state.mode !== "planning" || !ctx.hasUI || proposing) return;
const text = planText();
const version = `${state.plan}:${digest(text)}`;
if (!goals(text).length || version === proposedDraft) return;
proposedDraft = version;
proposing = true;
try {
pi.sendMessage({ customType: "goal-plan-proposal", content: text, display: true }, { triggerTurn: false });
await ready(ctx, true);
} finally { proposing = false; }
});
// No context hook. Historical message arrays, native checkpoints and model selection are untouched.
pi.on("before_agent_start", (event, ctx) => {
if (state.mode === "chat") return;
const snapshot = readPlan();
if (snapshot.text === undefined) {
notice = true; // Retry resync on the next turn; do not consume a failed snapshot.
return { systemPrompt: `${event.systemPrompt}\n\n${state.child ? childPlanRole : ""}\n${snapshot.error}` };
}
const role = state.child ? childPlanRole : state.mode === "supervising"
? supervisor(WORKER, state.plan!, ctx.sessionManager.getSessionId())
: state.mode === "planning" ? planning(state.plan!) : state.mode === "paused" ? pausedRole : soloRole;
const content = notice ? planContext(state.child ? "worker" : state.mode, state.plan, snapshot.text)
: undefined;
if (content) turnsStale = 0;
notice = false;
return { systemPrompt: `${event.systemPrompt}\n\n${role}`, ...(content ? { message: { customType: "pi-goals-plan", content, display: false } } : {}) };
});
pi.on("tool_call", (event) => {
if (state.child || !["subagent", "subagent_resume"].includes(event.toolName)) return;
// Solo means this chat took over implementation: no concurrent writer may be delegated.
if (state.mode === "planning" || state.mode === "paused" || state.mode === "solo") return { block: true, reason: goalToolBlocked(state.mode) };
if (state.plan) { pendingLaunches++; state.workerStopped = false; workerRevision++; save(); }
});
pi.on("tool_result", (event) => {
if (state.child || !state.plan || !["subagent", "subagent_resume"].includes(event.toolName)) return;
pendingLaunches = Math.max(0, pendingLaunches - 1);
if (event.isError) return;
const details = event.details as { id?: string; sessionFile?: string } | undefined;
if (!details?.id || !details.sessionFile) return;
const record = { id: details.id, sessionFile: details.sessionFile };
if (state.worker?.sessionFile === record.sessionFile) state.worker = record;
else if (!state.worker) state.worker = record;
// Extra launches stay recorded as helpers; the implementation binding never moves silently.
else state.helpers = [...(state.helpers ?? []).filter((h) => h.sessionFile !== record.sessionFile), record];
state.workerStopped = false; workerRevision++; save();
});
// --- plan mode: setup -------------------------------------------------------------------------
pi.registerCommand("plan", {
description: "Plan mode: set up goals (with evidence) in plan.md, then work them. /plan <objective>",
pi.registerCommand("goals", {
description: "Goal plan actions: new, review, ready, status, stop, resume, solo, attach, model, exit",
getArgumentCompletions: (prefix) => ["new", "review", "ready", "status", "stop", "resume", "solo", "attach", "model", "exit", "help"].filter((verb) => verb.startsWith(prefix)).map((verb) => ({ value: verb, label: verb })),
handler: async (args, ctx) => {
savedCmdCtx = ctx; // ctx here is an ExtensionCommandContext (has newSession); keep it for later
const arg = args.trim();
if (arg === "clear") {
await clearPlan(ctx);
return;
}
if (arg.startsWith("judge")) {
setJudge(arg.slice("judge".length).trim(), ctx);
return;
}
if (!arg) {
showPlan(ctx);
return;
}
state = { ...state, isPlanMode: true, objective: arg };
persist();
updateWidget(ctx);
pi.sendUserMessage(
`Enter plan mode for this objective: ${arg}\n\nExplore read-only, then write the plan to ${planPath(ctx)}.`,
{ deliverAs: "followUp" },
);
},
});
function setJudge(ref: string, ctx: ExtensionContext): void {
state = { ...state, judgeModel: ref || null };
persist();
ctx.ui.notify(ref ? `Sign-off judge model set to ${ref}` : "Sign-off judge reset to the default model", "info");
}
async function clearPlan(ctx: ExtensionContext): Promise<void> {
if (!existsSync(planPath(ctx))) {
ctx.ui.notify("No plan.md to clear.", "info");
return;
}
if (ctx.hasUI) {
const ok = await ctx.ui.select("Clear plan.md? (it stays in git history)", ["Cancel", "Clear plan.md"]);
if (ok !== "Clear plan.md") return;
}
writeFileSync(planPath(ctx), "");
state = { ...state, isPlanMode: false, objective: null };
persist();
updateWidget(ctx);
ctx.ui.notify("Cleared plan.md.", "info");
}
function showPlan(ctx: ExtensionContext): void {
const content = readPlan(ctx);
if (!content.trim()) {
ctx.ui.notify("No plan yet. Use /plan <objective> to start.", "info");
return;
}
ctx.ui.notify(content, "info");
}
// --- review loop (after the agent drafts the plan) --------------------------------------------
async function reviewLoop(ctx: ExtensionContext): Promise<void> {
while (true) {
const doc = parse(readPlan(ctx));
const choice = await ctx.ui.select(`Plan: ${doc.goals.length} goal(s). What next?`, [
"Ready — start working the plan",
"Edit — ask the agent to revise",
"Open in $EDITOR",
"Cancel — leave plan mode",
]);
if (!choice || choice.startsWith("Cancel")) {
exitPlanMode(ctx);
ctx.ui.notify("Left plan mode. plan.md kept.", "info");
return;
}
if (choice.startsWith("Ready")) return startExecution(ctx);
if (choice.startsWith("Edit")) {
const changes = await ctx.ui.editor("What should change about the plan?", "");
if (changes?.trim()) {
pi.sendUserMessage(`Revise the plan at ${planPath(ctx)} with these changes, same format:\n\n${changes.trim()}`);
return; // agent_end re-opens the review loop
try {
if (state.child) { ctx.ui.notify("This is the delegated worker. Goal approval belongs to its parent.", "info"); return; }
let command = args.trim();
if (!command) {
const actions = ["status — Show current plan", "new — New plan", "attach — Open an existing plan", "review — Review current plan", "ready — Approve draft", "stop — Pause work", "resume — Continue paused work", "solo — Work in this session", "model — Set worker model", "exit — Leave goal mode", "help — Show commands"];
const before = generation;
const choice = await ctx.ui.select("Goal plan actions", actions);
if (!choice || before !== generation) return;
command = choice.split(" — ")[0];
if (["attach", "model"].includes(command)) {
const value = await ctx.ui.editor(command === "attach" ? "Plan path (optional: solo)" : "Worker model (provider/model)", "");
if (!value?.trim() || before !== generation) return;
command += ` ${value.trim()}`;
}
}
continue;
}
if (choice.startsWith("Open")) {
const editor = process.env.EDITOR || process.env.VISUAL || "vi";
spawnSync(editor, [planPath(ctx)], { stdio: "inherit" });
}
}
}
function exitPlanMode(ctx: ExtensionContext): void {
state = { ...state, isPlanMode: false };
persist();
updateWidget(ctx);
}
async function startExecution(ctx: ExtensionContext): Promise<void> {
// Offer a clean execution context (D13). newSession lives only on the saved command context.
let fresh = false;
if (ctx.hasUI && savedCmdCtx) {
const choice = await ctx.ui.select("Start working the plan in...", [
"This context (keep history)",
"A fresh, compacted context",
]);
fresh = choice?.startsWith("A fresh") ?? false;
}
exitPlanMode(ctx);
const doc = parse(readPlan(ctx));
if (doc.objective) pi.setSessionName(`Plan: ${doc.objective}`);
if (fresh && savedCmdCtx) {
const result = await savedCmdCtx.newSession({ parentSession: ctx.sessionManager.getSessionFile() });
if (result.cancelled) {
ctx.ui.notify("Execution cancelled.", "warning");
return;
}
}
pi.sendUserMessage(
`Work the plan in ${planPath(ctx)}. Pick an open goal, set it active, work its subtasks, and when its done_when is met call CompleteGoal with the evidence. Keep plan.md current as you go.`,
{ deliverAs: "followUp" },
);
}
// --- the one blessed tool: CompleteGoal -------------------------------------------------------
pi.registerTool({
name: "CompleteGoal",
label: "Complete goal",
description:
"Sign off a goal once its done_when is met. Runs the goal's verify command (if any) then a " +
"read-only subagent that inspects your evidence against the repo. On accept, the goal is marked " +
"done and logged; on reject, it stays open and you get what is missing. Point evidence at durable " +
"artifacts (saved logs, committed diffs, files), not claims.",
parameters: Type.Object({
goal_id: Type.String({ description: "The goal's <!-- id --> from plan.md" }),
evidence: Type.String({ description: "What shows the done_when is met, and where to verify it" }),
paths: Type.Optional(Type.Array(Type.String(), { description: "Durable artifacts the judge should inspect" })),
}),
async execute(_id, params, signal, _onUpdate, ctx) {
const content = readPlan(ctx);
const goal = findGoal(parse(content), params.goal_id);
if (!goal) return text(`No goal #${params.goal_id} in plan.md.`, true);
// Decide the outcome (the I/O); recordSignOff applies it to the file (the pure write).
const outcome = await decideSignOff(goal, params.evidence, params.paths ?? [], state.judgeModel, ctx.cwd, signal);
const res = recordSignOff(content, goal.id, stamp(), outcome);
if (res.content !== content) writeFileSync(planPath(ctx), res.content);
updateWidget(ctx);
return text(res.message, res.isError);
if (command === "help") { ctx.ui.notify(help, "info"); return; }
if (command === "status") {
refresh(ctx);
ctx.ui.notify([
`Mode: ${state.mode}`,
`Plan: ${state.plan ?? "none"}`,
`Preferred worker model (plan): ${notedPlanValue("preferred worker model") ?? "not stated; use /goals model <model>"}`,
`Recorded worker session: ${state.worker?.sessionFile ?? "not recorded"}`,
`Helper subagent sessions: ${state.helpers.length} recorded (liveness via /subagents)`,
notedPlanValue("worker session") ? `Worker session noted in plan: ${notedPlanValue("worker session")}` : "",
`Hourly check-in: schedule_prompt job ${JSON.stringify(`goals-${ctx.sessionManager.getSessionId()}`)} (list/remove via schedule_prompt; plan-change reviews are the plan-watcher event hook)`,
"Liveness is owned by edxeth; inspect /subagents.",
].filter(Boolean).join("\n"), "info");
return;
}
if (command === "review" && state.mode === "supervising") { notice = true; send(manualReview(state.plan ?? "")); return; }
if (command === "review" || command === "ready") { await ready(ctx, command === "review"); return; }
if (command === "model" || command.startsWith("model ")) {
if (!state.plan || !goals(planText()).length) { ctx.ui.notify("Register a goal plan first.", "warning"); return; }
const ref = command.slice("model".length).trim();
if (!ref) { ctx.ui.notify("Use /goals model <provider/model>; no preference changed.", "info"); return; }
const lines = planText().split("\n");
const pref = `- preferred worker model: ${ref || "(none specified)"}`;
const found = lines.findIndex((line) => /^-\s*preferred worker model:/i.test(line));
if (found >= 0) lines[found] = pref;
else { const title = lines.findIndex((line) => /^#\s/.test(line)); lines.splice(title >= 0 ? title + 1 : 0, 0, pref); }
writeFileSync(state.plan, lines.join("\n"));
planHash = digest(planViews(planText()).notify);
refresh(ctx);
ctx.ui.notify(ref ? `Preferred worker model set to ${ref} in plan preferences. The supervisor selects it at launch and verifies the resolved model; the worker pane's own model is chosen with /model in that pane.` : "Preferred worker model cleared.", "info");
return;
}
if (command === "attach" || command.startsWith("attach ")) {
const rest = command.slice("attach".length).trim();
const [raw, kind, extra] = rest.split(/\s+/);
const solo = kind === "solo";
if (extra || (kind && !solo)) { ctx.ui.notify("Use /goals attach <path-to-plan.md> [solo].", "warning"); return; }
if (!raw) { ctx.ui.notify("Use /goals attach <path-to-plan.md> [solo].", "info"); return; }
const target = isAbsolute(raw) ? raw : resolve(ctx.cwd, raw);
let text: string;
try { text = readFileSync(target, "utf8"); } catch { ctx.ui.notify(`Cannot read plan at ${target}.`, "error"); return; }
if (!goals(text).length) { ctx.ui.notify(`${target} has no '- [ ] goal:' lines; attach a judgeable plan.`, "warning"); return; }
if (!solo && ((state.worker && !state.workerStopped) || state.mode === "supervising")) { ctx.ui.notify("Exit and resolve the existing worker before replacing the plan. The current plan is preserved.", "warning"); return; }
const noted = /^-\s*worker session:\s*(\S+)/im.exec(foldPlan(text))?.[1];
if (!(await confirmOwnership(ctx, target, text, solo))) return;
const retained = target === state.plan ? state.signoffs : {};
const worker = noted ? { sessionFile: resolve(ctx.cwd, noted) } : state.workerStopped ? state.worker : undefined;
state = { mode: solo ? "solo" : "planning", plan: target, signoffs: retained, worker, helpers: [], workerStopped: solo || (!noted && state.workerStopped) };
generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
if (solo) enterSolo(ctx);
else send(attachNotice(target, false, noted));
return;
}
if (command === "stop" || command === "exit") {
if (state.mode === "planning") {
if (command === "stop") { ctx.ui.notify("A draft cannot pause; use /goals exit to leave planning with the draft preserved.", "warning"); return; }
// Planning exit must not get the model trapped re-planning or lose the draft.
state.mode = "chat"; generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
ctx.ui.notify(`Planning exited; draft preserved at ${state.plan}. No implementation was approved or started. Reconnect with /goals attach ${state.plan}.`, "info");
return;
}
state.mode = command === "stop" ? "paused" : "chat"; generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
send(`${removeGoalSchedule(ctx.sessionManager.getSessionId())}\n\n${pauseExitNotice(state.worker, command === "exit")}`, Boolean(state.worker) || hasScheduleTool());
return;
}
if (command === "resume") {
if (state.mode !== "paused" || !state.plan) { ctx.ui.notify("Only a paused approved plan can resume. A draft needs Ready.", "warning"); return; }
if (!compatible()) { ctx.ui.notify("edxeth tools unavailable; plan remains paused.", "error"); return; }
state.mode = "supervising"; generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
send(`${checkIn(ctx)}\n\n${resumeNotice(WORKER, state.plan, state.worker)}`);
return;
}
if (command === "solo") {
if (!state.plan || !goals(planText()).length) { ctx.ui.notify("Register a goal plan first.", "warning"); return; }
if (!(await confirmOwnership(ctx, state.plan, planText()))) return;
enterSolo(ctx);
return;
}
if (command !== "new" && !command.startsWith("new ")) { ctx.ui.notify(`Unknown or incomplete command. ${help}`, "warning"); return; }
const objective = command.slice(4).trim();
if ((state.worker && !state.workerStopped) || state.mode === "supervising") { ctx.ui.notify("Exit and resolve the existing worker before replacing the plan. The current plan is preserved.", "warning"); return; }
const path = join(ctx.cwd, ".pi", "plan", `${ctx.sessionManager.getSessionId()}-main.md`);
mkdirSync(dirname(path), { recursive: true });
// Never overwrite an earlier plan at this session path; the model can revise it after inspection.
try { writeFileSync(path, planDocument(objective), { flag: "wx" }); } catch (error) { if ((error as NodeJS.ErrnoException).code !== "EEXIST") throw error; }
state = { mode: "planning", plan: path, signoffs: {}, worker: state.worker, helpers: state.helpers, workerStopped: state.workerStopped }; generation++; notice = true; save(); refresh(ctx); watchPlan(ctx);
send(planningSeed(objective, path));
} catch (error) { ctx.ui.notify(String(error), "error"); }
},
});
// --- hooks ------------------------------------------------------------------------------------
pi.on("before_agent_start", async (_event, ctx) => {
if (state.isPlanMode) {
return { message: { customType: PLAN_CONTEXT, content: `${planDrafting}\n\nWrite the plan to ${planPath(ctx)}.`, display: false } };
}
const doc = parse(readPlan(ctx));
if (doc.goals.length === 0) return;
const active = doc.goals.find((g) => g.status === "active") ?? doc.goals.find((g) => g.status === "open") ?? null;
const c = counts(doc);
let body = planInjection({
objective: doc.objective,
activeGoal: active
? { subject: active.subject, done_when: active.done_when, openSubtasks: active.subtasks.filter((s) => !s.done).map((s) => s.text) }
: null,
lastLogLine: doc.log.at(-1) ?? null,
counts: { done: c.done, open: c.open + c.active },
});
// Reminder fires when there is an active goal but plan.md was untouched since the last turn.
const planNow = readPlan(ctx);
if (active && planNow === lastInjectedPlan) body += `\n\n${reminder}`;
lastInjectedPlan = planNow;
return { message: { customType: PLAN_CONTEXT, content: body, display: false } };
pi.registerTool({
name: "AttachGoalPlan", label: "Attach delegated plan", description: attachGoalPlanDescription,
parameters: Type.Object({ path: Type.String() }),
async execute(_id, params, _signal, _update, ctx) {
if (!state.child) return result(messages.childAttachOnly);
try {
if (!isAbsolute(params.path) || !goals(readFileSync(params.path, "utf8")).length) return result(messages.invalidAttachment);
} catch { return result(messages.invalidAttachment); }
state.plan = params.path; generation++; notice = true; save(); refresh(ctx);
return result(childPlanAttached(params.path));
},
});
pi.on("agent_end", async (_event, ctx) => {
if (!state.isPlanMode || !ctx.hasUI) return;
const doc = parse(readPlan(ctx));
if (doc.goals.length === 0) {
ctx.ui.notify("No goals found in plan.md yet — ask the agent to draft them.", "warning");
return;
}
await reviewLoop(ctx);
});
// Keep only the freshest injected plan summary; strip stale ones so history does not bloat and
// the model never sees an out-of-date plan. (The current turn's injection is the one kept.)
pi.on("context", async (event) => {
const isCtx = (m: unknown) => (m as { customType?: string }).customType === PLAN_CONTEXT;
let lastIdx = -1;
event.messages.forEach((m, i) => {
if (isCtx(m)) lastIdx = i;
});
return { messages: event.messages.filter((m, i) => !isCtx(m) || i === lastIdx) };
});
pi.on("session_start", async (_event, ctx) => {
const last = ctx.sessionManager
.getEntries()
.filter((e: { type?: string; customType?: string }) => e.type === "custom" && e.customType === STATE)
.pop() as { data?: PlanState } | undefined;
if (last?.data) state = { ...state, ...last.data };
updateWidget(ctx);
pi.registerTool({
name: "CompleteGoal", label: "Review goal evidence",
description: completeGoalDescription,
parameters: Type.Object({ goal: Type.String(), evidence: Type.Array(Type.String(), { minItems: 1 }), observation: Type.String({ minLength: 1 }) }),
async execute(_id, params, signal, _update, ctx) {
if (state.child || !["supervising", "solo"].includes(state.mode)) return result(messages.completionUnavailable);
if (signal?.aborted) return result(messages.cancelled);
const snapshot = readPlan();
if (snapshot.text === undefined) return result(snapshot.error!);
const text = snapshot.text;
const matches = goals(text).filter((g) => g.status !== "cancelled" && key(g.subject) === key(params.goal));
if (matches.length !== 1 || !state.plan) return result(messages.uniqueGoal);
const evidence = params.evidence.map((file) => isAbsolute(file) ? file : resolve(ctx.cwd, file));
try { for (const file of evidence) if (!readFileSync(file).length) throw new Error(emptyEvidence(file)); }
catch (error) { return result(evidenceUnavailable(error)); }
const lines = text.split("\n");
lines[matches[0].index] = lines[matches[0].index].replace(/\[[ xX/-]\]/, "[x]");
let log = lines.findIndex(line => /^##\s+Log\s*$/i.test(line));
if (log === -1) { lines.push("", "## Log"); log = lines.length - 1; }
lines.splice(log + 1, 0, "", completionLog(params.goal, params.observation, evidence, state.mode === "solo"));
writeFileSync(state.plan, `${lines.join("\n").trimEnd()}\n`);
state.signoffs[key(matches[0].subject)] = { evidence, observation: params.observation };
planHash = digest(planViews(planText()).notify);
save(); refresh(ctx);
const remaining = goals(planText()).some((goal) => goal.status !== "cancelled" && (goal.status !== "done" || !state.signoffs[key(goal.subject)]));
return result(completionResult(matches[0].subject, ctx.sessionManager.getSessionId(), remaining, state.mode === "solo"));
},
});
}
// --- helpers (module scope; pure enough to keep out of the closure) -------------------------------
function text(s: string, isError = false) {
return { content: [{ type: "text" as const, text: s }], details: { isError }, isError };
}
function stamp(): string {
return new Date().toISOString().slice(0, 16).replace("T", " ");
}
/** Decide a sign-off: deterministic verify first (cheap; skip the model call if it fails), then the judge. */
async function decideSignOff(
goal: Goal,
evidence: string,
paths: string[],
judgeModel: string | null,
cwd: string,
signal: AbortSignal | undefined,
): Promise<SignOff> {
let verifyResult: { command: string; exitCode: number; outputTail: string } | null = null;
if (goal.verify) {
verifyResult = runVerify(goal.verify, cwd, signal);
if (verifyResult.exitCode !== 0) {
return { kind: "verify_failed", exitCode: verifyResult.exitCode, outputTail: verifyResult.outputTail };
}
}
const verdict = await runJudge(goal, evidence, paths, verifyResult, judgeModel, cwd, signal);
return verdict.accept ? { kind: "accepted" } : { kind: "rejected", missing: verdict.missing };
}
/** Run the goal's verify command. It is agent-authored and trusted (single-user machine, guide-not-guard). */
function runVerify(command: string, cwd: string, signal: AbortSignal | undefined): { command: string; exitCode: number; outputTail: string } {
const res = spawnSync("sh", ["-c", command], { cwd, encoding: "utf-8", signal, timeout: 600_000 });
const out = `${res.stdout ?? ""}${res.stderr ?? ""}`;
return { command, exitCode: res.status ?? 1, outputTail: out.split("\n").slice(-30).join("\n") };
}
/** Locate the pi binary the same way the oracle extension does, so spawning works under bun or node. */
function getPiInvocation(args: string[]): { command: string; args: string[] } {
const script = process.argv[1];
if (script && !script.startsWith("/$bunfs/root/") && existsSync(script)) return { command: process.execPath, args: [script, ...args] };
const execName = basename(process.execPath).toLowerCase();
if (!/^(node|bun)(\.exe)?$/.test(execName)) return { command: process.execPath, args };
return { command: "pi", args };
}
/** Stage 2: a read-only pi subprocess inspects the evidence against the repo and returns a verdict. */
async function runJudge(
goal: Goal,
evidence: string,
paths: string[],
verifyResult: { command: string; exitCode: number; outputTail: string } | null,
judgeModel: string | null,
cwd: string,
signal: AbortSignal | undefined,
): Promise<{ accept: boolean; missing: string }> {
const task = evidenceJudgeUser({
subject: goal.subject,
done_when: goal.done_when,
verify: goal.verify ?? null,
verifyResult,
failure_modes: goal.failure_modes,
evidence,
paths,
});
const args = ["-p", "--no-session", "--tools", READ_ONLY_TOOLS.join(","), "--append-system-prompt", evidenceJudgeSystem];
if (judgeModel) args.push("--model", judgeModel);
args.push(task);
const inv = getPiInvocation(args);
const output = await new Promise<string>((resolve) => {
const proc = spawn(inv.command, inv.args, { cwd, shell: false, stdio: ["ignore", "pipe", "pipe"], signal });
let out = "";
proc.stdout.on("data", (d) => (out += d));
proc.stderr.on("data", (d) => (out += d));
proc.on("close", () => resolve(out));
proc.on("error", (e) => resolve(`VERDICT: reject\nmissing: judge subprocess failed: ${e.message}`));
});
const verdictLine = output.split("\n").find((l) => /^\s*VERDICT\s*:/i.test(l)) ?? "";
const accept = /accept/i.test(verdictLine);
const missingMatch = output.match(/missing\s*:\s*([\s\S]*)$/i);
const missing = accept ? "" : (missingMatch?.[1].trim() || output.trim().slice(-500) || "judge gave no reason");
return { accept, missing };
}
-225
View File
@@ -1,225 +0,0 @@
/**
* plan-file.ts — read plan.md, and the two writes CompleteGoal needs. That is all.
*
* Pure module, no pi deps, so it unit-tests without a runtime. The file is the canonical store and
* the agent edits it with its normal Edit tool (create goals, tick subtasks, append log), guided by
* the format in prompts.tsx and the reminder -- the form guides, it does not gate (spec D3). So this
* module does NOT render or create goals; the format's single source of truth is the planDrafting
* prompt. The only programmatic writers are setGoalStatus + appendLog, used by CompleteGoal to
* record an accepted sign-off; both touch one line so the git diff stays readable.
*
* Format (spec §4):
*
* # Plan: <objective>
*
* ## Goal: <subject>
* <!-- id: <slug> -->
* status: open | active | done | cancelled
* done_when: <falsifiable check; plus the symptom if NOT met>
* verify: <shell command, optional>
* failure_modes:
* - <pre-mortem item>
* - [ ] <subtask>
*
* ## Log
* - <verbatim append-only line>
*/
export type GoalStatus = "open" | "active" | "done" | "cancelled";
export interface Subtask {
text: string;
done: boolean;
}
export interface Goal {
id: string;
subject: string;
status: GoalStatus;
done_when: string;
verify?: string;
failure_modes: string[];
subtasks: Subtask[];
}
export interface PlanDoc {
objective: string;
goals: Goal[];
/** Verbatim ## Log lines, including the leading "- ". */
log: string[];
}
const GOAL_HEADER = /^##\s+Goal:\s*(.*)$/;
const ANY_HEADER = /^#{1,6}\s/;
const LOG_HEADER = /^##\s+Log\s*$/i;
const ID_COMMENT = /^<!--\s*id:\s*(.+?)\s*-->$/;
const CHECKBOX = /^- \[([ xX])\]\s+(.*)$/;
export function parse(text: string): PlanDoc {
const lines = text.split("\n");
let objective = "";
const goals: Goal[] = [];
const log: string[] = [];
let cur: Goal | null = null;
let inFailureModes = false;
let inLog = false;
const flush = () => {
if (cur) goals.push(cur);
cur = null;
inFailureModes = false;
};
for (const line of lines) {
const objMatch = /^#\s+Plan:\s*(.*)$/.exec(line);
if (objMatch) {
objective = objMatch[1].trim();
continue;
}
const goalMatch = GOAL_HEADER.exec(line);
if (goalMatch) {
flush();
inLog = false;
cur = { id: "", subject: goalMatch[1].trim(), status: "open", done_when: "", failure_modes: [], subtasks: [] };
continue;
}
if (LOG_HEADER.test(line)) {
flush();
inLog = true;
continue;
}
// Any other header ends the current goal / log section.
if (ANY_HEADER.test(line)) {
flush();
inLog = false;
continue;
}
if (inLog) {
if (/^\s*-\s+/.test(line)) log.push(line);
continue;
}
if (!cur) continue;
const idMatch = ID_COMMENT.exec(line.trim());
if (idMatch) {
cur.id = idMatch[1];
continue;
}
// A checkbox (column 0) is a subtask; checked first so it is never read as a failure mode.
const checkbox = CHECKBOX.exec(line);
if (checkbox) {
inFailureModes = false;
cur.subtasks.push({ done: checkbox[1].toLowerCase() === "x", text: checkbox[2].trim() });
continue;
}
const kv = /^(status|done_when|verify|failure_modes)\s*:\s*(.*)$/.exec(line);
if (kv) {
const [, key, value] = kv;
if (key === "status") cur.status = value.trim() as GoalStatus;
else if (key === "done_when") cur.done_when = value.trim();
else if (key === "verify") cur.verify = value.trim() || undefined;
else if (key === "failure_modes") inFailureModes = true;
continue;
}
// Indented "- " items under failure_modes: (a column-0 checkbox already returned above).
if (inFailureModes) {
const fm = /^\s*-\s+(.*)$/.exec(line);
if (fm) {
cur.failure_modes.push(fm[1].trim());
continue;
}
if (line.trim() !== "") inFailureModes = false;
}
}
flush();
return { objective, goals, log };
}
export function findGoal(doc: PlanDoc, id: string): Goal | undefined {
return doc.goals.find((g) => g.id === id);
}
export function counts(doc: PlanDoc): { done: number; open: number; active: number } {
const c = { done: 0, open: 0, active: 0 };
for (const g of doc.goals) {
if (g.status === "done") c.done++;
else if (g.status === "active") c.active++;
else if (g.status === "open") c.open++;
}
return c;
}
/** Flip a goal's `status:` line in place (the one write CompleteGoal needs). */
export function setGoalStatus(text: string, id: string, status: GoalStatus): string {
const lines = text.split("\n");
let i = lines.findIndex((l) => ID_COMMENT.test(l.trim()) && ID_COMMENT.exec(l.trim())?.[1] === id);
if (i === -1) throw new Error(`Goal #${id} not found`);
for (; i < lines.length; i++) {
if (i > 0 && ANY_HEADER.test(lines[i]) && !GOAL_HEADER.test(lines[i]) && !LOG_HEADER.test(lines[i])) break;
const kv = /^(status\s*:\s*)(.*)$/.exec(lines[i]);
if (kv) {
lines[i] = `${kv[1]}${status}`;
return lines.join("\n");
}
}
throw new Error(`Goal #${id} has no status: line`);
}
/**
* The outcome of a sign-off attempt, decided by CompleteGoal (which runs verify + the judge). Kept
* separate from the I/O so the record logic below is pure and testable.
*/
export type SignOff =
| { kind: "verify_failed"; exitCode: number; outputTail: string }
| { kind: "rejected"; missing: string }
| { kind: "accepted" };
/** Apply a sign-off outcome to plan.md text: accept flips status + logs; reject only logs. Pure. */
export function recordSignOff(
text: string,
goalId: string,
when: string,
outcome: SignOff,
): { content: string; message: string; isError: boolean } {
const goal = findGoal(parse(text), goalId);
if (!goal) return { content: text, message: `No goal #${goalId} in plan.md.`, isError: true };
if (outcome.kind === "verify_failed") {
const content = appendLog(text, `${when} reject #${goalId}: verify exit ${outcome.exitCode}`);
return { content, message: `Sign-off rejected: verify failed (exit ${outcome.exitCode}).\n${outcome.outputTail}`, isError: true };
}
if (outcome.kind === "rejected") {
const oneLine = outcome.missing.replace(/\s+/g, " ").trim().slice(0, 200);
const content = appendLog(text, `${when} reject #${goalId}: ${oneLine}`);
return { content, message: `Sign-off rejected. Missing:\n${outcome.missing}`, isError: true };
}
const flipped = setGoalStatus(text, goalId, "done");
const content = appendLog(flipped, `${when} signed off #${goalId}: ${goal.subject} (oracle accept)`);
return { content, message: `Signed off #${goalId}: ${goal.subject}. Marked done in plan.md.`, isError: false };
}
/** Append one verbatim line to ## Log (creating the section if absent). The other CompleteGoal write. */
export function appendLog(text: string, entry: string): string {
const lines = text.split("\n");
const line = `- ${entry}`;
const header = lines.findIndex((l) => LOG_HEADER.test(l));
if (header === -1) return `${text.replace(/\n+$/, "")}\n\n## Log\n${line}\n`;
let insertAt = header + 1;
for (let i = header + 1; i < lines.length; i++) {
if (ANY_HEADER.test(lines[i])) break;
if (/^\s*-\s+/.test(lines[i])) insertAt = i + 1;
}
lines.splice(insertAt, 0, line);
return lines.join("\n");
}
+33
View File
@@ -0,0 +1,33 @@
// Pi/OpenAI: Preserve plan wording; omit history and, in the short view, task/evidence details.
// The notify view governs plan-change events: goals, tasks, evidence and inferences are
// content worth a supervisor review; worker identity bookkeeping is not (field report,
// LUCID3 supervisor 2026-09-10: two identical review events for a session-path edit).
export function planViews(plan: string): { short: string; notify: string; long: string } {
const long = plan.split(/^#{1,6}\s+(?:Log|Appendix|Appendices|Appendixes|Interview|Learnings|Papercuts)\b.*$/mi)[0].trim();
const identity = /^-\s*(?:active worker|worker session|worker intercom session):/i;
const notify = long.split("\n").filter((line) => !identity.test(line)).join("\n").trim();
const kept: string[] = [];
let omittedIndent: number | null = null;
let omittedHeading: number | null = null;
for (const line of long.split("\n")) {
// Pi/OpenAI: Worker identity bookkeeping is not a change to agreed requirements.
if (/^-\s*(?:active worker|worker session|worker intercom session):/i.test(line)) continue;
const heading = /^(#{1,6})\s+(.+)$/.exec(line);
if (heading) {
if (omittedHeading !== null && heading[1].length <= omittedHeading) omittedHeading = null;
if (/^(?:Tasks?|Task list|Subtasks?|Evidence)\b/i.test(heading[2])) omittedHeading = heading[1].length;
}
if (omittedHeading !== null) continue;
const indent = line.match(/^\s*/)?.[0].length ?? 0;
if (omittedIndent !== null) {
if (!line.trim() || indent > omittedIndent) continue;
omittedIndent = null;
}
if (/^\s*[-*]\s+(?:tasks?|subtasks?|evidence):/i.test(line) || /^\s*(?:\d+[.)]|[-*])\s+\[[ x/~-]\]\s+(?!goal:)/i.test(line)) {
omittedIndent = indent;
continue;
}
kept.push(line);
}
return { short: kept.join("\n").trim(), notify, long };
}
+8
View File
@@ -0,0 +1,8 @@
// Shared plan syntax: only the section above the Log contains current goals.
export const GOAL_LINE = /^\s*(?:\d+\.|[-*])\s*\[([ xX/-])\]\s*goal:\s*(.*)$/i;
export const FOLD_LINE = /^##\s+Log\s*$/im;
export function foldPlan(plan: string): string {
const match = FOLD_LINE.exec(plan);
return (match ? plan.slice(0, match.index) : plan).trimEnd();
}
+202 -191
View File
@@ -1,204 +1,215 @@
/**
* pi-plan — all model-facing text, in flow order.
*
* Philosophy: the form guides a process; it does not police one. The agent can
* edit plan.md freely. These prompts + the plan.md structure make the right path
* the easy path. The only step that is genuinely rigorous is the evidence judge
* (6), and even that is reached by guiding the agent to call CompleteGoal, not by
* trapping it. Bypasses stay visible in the git diff and the widget.
*
* Flow:
* SETUP (plan mode) 1. planDrafting — strong/sticky model drafts goals
* EXEC, each turn start 2. planInjection — "here is your plan, where you are"
* EXEC, periodic 3. reminder — the typed nudge that drives upkeep + autonomy
* EXEC, loop continue 4. continuation — keep going toward the active goal
* EXEC, after each turn 5. loopJudge — continue / pause (cheap, foolable, ok)
* SIGN-OFF 6. evidenceJudge — read-only verify (rigorous; the one real check)
*
* Read top to bottom to see the whole process. 5 and 6 are kept adjacent on
* purpose: the cheap-foolable vs must-not-be-fooled contrast is the design.
*
* WIRED in index.ts: 1 planDrafting, 2 planInjection, 3 reminder, 6 evidenceJudge.
* NOT YET WIRED: 4 continuation and 5 loopJudge define the autonomous re-prompt loop, which is
* intentionally not built in v1 (an until-done-style loop was judged too complex). They stay here so
* the full intended flow is reviewable; wire them if/when the loop is added.
*/
/* ─────────────────────────────────────────────────────────────────────────
* 1. planDrafting — SETUP, plan mode
*
* System guidance for the plan-phase agent. Runs on the plan model (may differ
* from the execution model; the choice is sticky — see oracle.json-style config).
* This phase is read-only: explore, then draft goals into plan.md. No code yet.
* The field requirements here are the whole "elicitation" — get them agreed up
* front, because the human reviews this output before any execution.
* ──────────────────────────────────────────────────────────────────────── */
// Pi/OpenAI: Planning, approval, supervision, reminders, completion and recovery.
export const planDrafting = `\
You are in plan mode. Explore the repository read-only, then draft a plan into plan.md.
Do not write or run code in this phase. Produce goals the human will review and approve.
You are in plan mode. Help the user express what they want this project to achieve in a short judgeable plan. Seek to understand their underlying goals, infer ordinary details, and use their applicable AGENTS.md instructions, relevant skills, and project context to interpret the request correctly. Do not silently substitute your own goals or expand the agreed scope.
Write each goal in this shape:
1. Reduce technical uncertainty first. Use read-only repository tools or web search when either can
resolve a fact. Do not write or run code in this phase (edit/write are blocked except for the plan
file; don't mutate state via bash either).
2. Before you draft a goal, identify its object, observable result, scope, and any decision that the
human would need to approve later. Briefly reframe the request in your own words to check comprehension
and make your understanding visible: the intended outcome, boundary, and success check. Invite correction,
but do not require confirmation when these are already clear. Ask questions that expose differences
between your understanding and the user's that would otherwise stay hidden. Probe consequential
assumptions, challenge inconsistencies, and follow up where an answer exposes a gap. Do not use a question quota or ask the human
to approve ordinary implementation details. Inspect files or search the web before asking when either
can answer a fact. If the human does not answer a question, record that
point as unknown; do not silently replace it with an inference or turn it into a new blocking decision.
Do not present the review menu with a placeholder goal such as "work out the thing", "improve it", or
"investigate".
3. Use questions to clarify and narrow the goal, test your assumptions, and bring your understanding
into agreement with the user's. Respect their limited time: batch independent high-impact questions
in one short round, where the answer materially reduces uncertainty
while discovering the right plan. Each question must be short and self-contained: state the relevant
context, use the human's language and ASD-STE100
Simple Technical English, and give a recommended answer. Record each answer, or the unanswered
unknown, in ## Interview. Draft goals and present Ready when the requested work is otherwise executable.
Only withhold Ready for an unanswered choice that changes scope, spending, or the user-visible result.
4. State the user-visible result before the goals: one concrete sentence naming what the human will
inspect when this plan is done. Take it from the original request, not from your implementation plan.
Every requested artifact and action must survive into this sentence. An agent-inferred constraint may
not replace, defer, or contradict it; ask the human if an inference would change the result.
5. When every goal has an object, observable result, settled scope, and required approval, draft the
plan file and present it. It should be safe to work overnight and present the requested outcome.
## Goal: <one short imperative line>
status: open
done_when: <a falsifiable check, plus the symptom you'd see if it's NOT met>
verify: <a shell command that exits 0 only when the goal is met — include this whenever
success is expressible as tests/lint/build/a threshold; omit it otherwise>
failure_modes:
- <a concrete way this could look done but isn't>
- <another>
- <if verify exists: "verify passes on a trivial or gamed test">
- [ ] <first subtask>
- [ ] <next subtask>
How this mode ends: after each settled draft the human gets a menu (Ready / Refine / Edit / Cancel).
Plan mode ends only when they pick Ready. Refine collects short revision notes. Edit opens the full
plan. When a new requirement arrives, fold it in, say what changed, and present the plan again.
Detail that doesn't change a goal or a discriminator belongs in the appendix, not in the goals.
Rules for a good plan:
- Keep goals small enough that done_when is checkable in one sitting.
- done_when must be falsifiable. "Works well" is not a criterion; "p95 < 50ms on bench-X,
else timeouts in load-test.log" is.
- failure_modes are a pre-mortem: the cheap, specific ways a later "done" could be wrong.
This is the highest-value part — it shapes what evidence you'll collect.
- Prefer a verify command. A green deterministic check is worth more than a paragraph of
description, and it's the first thing checked at sign-off.
Right-size it:
- One goal per distinct judgeable outcome. Group related goals when it helps judge them together
and readability. The count flows from the outcomes.
- Describe outcomes in qualitative terms the supervisor and user can discriminate.
- Use the users language or more precise don't transform "MV" into "knob" as it looses precision and is overloaded
- Don't invent metrics or thresholds for problems you haven't explored yet - the supervisor should know it when it sees the outcome.
- Quantitative gates are fine only when you are certain they survive contact with reality.
- Subtasks are the steps inside a goal; add them when a goal has 3+ distinct steps, skip otherwise.
- Two goals that share one discriminator are one goal. Merge them.
- Keep the goal subject short. Put its important scope, failure modes, discriminator, tasks, and evidence in the indented block beneath it. The supervisor reads the whole block and the whole plan.
- Keep the working set under 50 lines, excluding ## User voice. ## User voice has no line limit: quote
the human fully rather than shorten or paraphrase them. Everything below "## Log" is unlimited.
When the plan is drafted, present it and stop for review. Do not begin execution.`;
Style: Make it easy for a busy and forgetfull user to review. Use ASD-STE100 Simplified Technical English. Use active voice, one idea per sentence, common words,
the same word for the same thing, and define a new terms at first use. Use redundant context for skim readers e.g. "our output - the cells, CV tag" is easy to read and reminds context. This covers the context
paragraph and the appendix too, not just the checklist. No all-caps headers and no bold spam. Just write less, add your voice less, persuade less, and burden the reader less.
/* ─────────────────────────────────────────────────────────────────────────
* 2. planInjection — EXEC, injected at each agent start (and after compaction)
*
* A late user-role message, NOT a system-prompt mutation (keeps the prefix cache
* valid). Built from the parsed plan. MUST be byte-identical when nothing changed:
* fixed field order, no volatile timestamps in the body. Pass only the active
* goal + its open subtasks + the last log line — not the whole file.
* ──────────────────────────────────────────────────────────────────────── */
export function planInjection(p: {
objective: string;
activeGoal: { subject: string; done_when: string; openSubtasks: string[] } | null;
lastLogLine: string | null;
counts: { done: number; open: number };
}): string {
if (!p.activeGoal) {
return `Plan (plan.md): ${p.objective}\nNo active goal. ${p.counts.open} open, ${p.counts.done} done. Pick the next goal or run /plan.`;
}
const subtasks = p.activeGoal.openSubtasks.length
? p.activeGoal.openSubtasks.map((s) => ` - [ ] ${s}`).join("\n")
: " (no open subtasks)";
return `\
Plan (plan.md): ${p.objective}
Active goal: ${p.activeGoal.subject}
done_when: ${p.activeGoal.done_when}
Open subtasks:
${subtasks}
Last log: ${p.lastLogLine ?? "(none yet)"}
Progress: ${p.counts.done} done, ${p.counts.open} open.`;
Write the plan file in roughly this shape -- the file is read directly by the human and the visible supervisor, so clarity beats conformance; small deviations are fine):
# <short plan title>
<context: one short paragraph. What the human wants and why.>
## User-visible result
<one concrete sentence naming the final artifact or behavior the human will inspect>
## User voice
- > "<the human's requirement, quoted in full word for word (with spelling fixes)>"
## Goals
1. [ ] goal: <one short jugable imperative outcome>
- subtle failure mode: <a way this could look done but isn't>
- discriminator: <the concrete observation that tells real success from that failure>
- verify: <optional shell command that exits 0 only when the discriminator passes; omit if not
testable. The worker runs it and saves its output; the visible supervisor reads the evidence>
- tasks:
1. [ ] <subtask>
- evidence: (empty until sign-off)
## Future work / out of scope
<-- the fold: everything below here is durable memory, not the working set -->
## Log
### {date}
## Interview
## Learnings
## Papercuts - problems, gotchas, suggestions
## Appendix (context, not approved)
Conventions:
- A goal is a checkbox line beginning "goal:". Checkbox state: [ ] open, [/] active, [x] done,
[-] cancelled. Leave goals [ ] at planning.
- subtle failure mode + discriminator are the heart of this. Name the ways a "done" could look
achieved but not be (empty output, a silently-errored step, a gamed test, a no-op that dodged
every trap and showed nothing). The discriminator is the POSITIVE observation that success
happened -- the count moved, the test exercised the real path, the metric beat noise -- and that
none of the failure modes could fake. Ruling out failures is necessary, not sufficient.
- Make the discriminator a concrete, checkable observation about a real artifact (a file, a test
result, a committed diff, a metric), never about the plan file's own checkbox.
- evidence stays empty at planning; the worker fills it and the visible supervisor checks it.
Cite durable artifacts a future reader can open: committed files, test names, git diffs. .pi/ is
usually gitignored, so files there prove things only at supervisor review time, not in history.
- User-visible result: restate the original deliverable, not the proposed implementation. Every goal
must contribute to it. Future work may not defer any artifact or action named there.
- User voice: quote the human word for word, one line per requirement, as they say it. Never
paraphrase there -- a paraphrase drifts, and then the goals churn on the next reply. It is exempt
from the working-set line limit. Never put an agent inference in User voice.
- Interview: every human reply in plan mode is stored here verbatim as a dated blockquote. It is
durable memory below the fold, not a substitute for ## User voice.
- Rejected options stay visible: ~~struck through~~ with who rejected them and why, so nobody
relitigates them.
- Learnings: one line per gotcha that a future reader would otherwise rediscover. Write down what
you saw from a source that does not persist (a browser page, an image, a long log tail) before
you do anything else with it.
- Appendix: unlimited and unverified. Alternatives, links, dead ends, and the settled detail that
is not part of the approved goals. Nothing here is approved and nothing here is checked.
When the goals are drafted, present them and say the plan is final. Do not begin execution.`;
// Planning and interview. Keep the full drafting guide one-shot rather than repeating it each turn.
export function planning(planPath: string): string {
return `Plan only in ${planPath}; do not implement or launch workers before Ready. Ask material unresolved questions, not a quota or confirmation of ordinary details. Record unknowns and present Ready when the outcome, scope and spending are settled. Preserve the user's exact deliverable, preferences and voice; give each distinct goal a failure mode, discriminator and evidence expectation above ## Log. Record the requested worker model in preferences. When your drafted plan is ready for human review, finish your turn; the interface displays the draft and approval choices automatically. Do not ask the user to type a command to see the proposal. /goals review reopens it on request; /goals exit preserves the draft.`;
}
export function planningSeed(objective: string, planPath: string): string {
return `Enter a planning conversation focused on the user's goals. ${objective ? `Initial idea: ${objective}.` : "Ask what the user wants to achieve; they do not need to supply a finished objective."} Read any existing plan at ${planPath} first, then discuss and draft it with the user. Do not infer approval to implement from starting this conversation. ${planning(planPath)}\n\n${planDrafting}`;
}
export const planDocument = (objective: string) => `# Goal plan\n\n## Objective\n${objective}\n\n## Goals\n\n## Log\n`;
export const discuss = "Discuss the current draft in ordinary chat. Do not launch a worker or reopen the review menu until requested.";
// Ready and explicit child attachment: stock lineage-only sessions do not inherit the shared plan.
export const attachGoalPlanDescription = "Delegated goals-worker only: attach the absolute plan path explicitly supplied in your task. Read it without rewriting it. Restores the worker widget and plan context; grants no parent completion authority. No discovery or worker launch.";
export const childPlanRole = "You are the delegated implementation worker. Maintain task ticks, evidence and Log entries for your delegated work in the supplied plan. Preserve agreed goals, requirements and discriminators; the supervisor owns goal-status changes and completion approval. Do not launch a second writer. Call AttachGoalPlan with the explicit plan path in your task before implementation (also after reconnect if unbound). Immediately report your actual Intercom UUID, saved-session path and current provider/model to the supplied supervisor ID. Identify unavailable fields as unknown; do not equate runtime IDs, session filenames and Intercom IDs. Send progress, completion and blocker reports there with artifact paths, then stay open for live messages. Do not exit or use caller_ping; unsent editor drafts are not visible in model context.";
export function readyApproved(workerName: string, planPath: string, notedWorker: string | undefined, plan: string, supervisorId: string): string {
const launch = notedWorker
? `Inspect the recorded worker session ${notedWorker}; if still live, let it continue or message it. Only after confirming it stopped use subagent_resume with that sessionFile. Never restart completed work.`
: `Delegate the first unfinished goal to agent '${workerName}' with subagent; provide name, title and a bounded task.`;
return `Ready approved this plan: ${planPath}. Stay here as supervisor. ${launch} Include the absolute plan path, require AttachGoalPlan, and give the child supervisor Intercom session ${supervisorId}. The child sends its completion report there and stays open. Require an initial worker report with its actual Intercom UUID, saved-session path and current provider/model; the async launch may return only a runtime ID. Record each distinct identity in plan preferences, marking child-reported fields as such until verified. Do not start a second writer. Inspect actual outputs when the child reports.\n\n${plan}`;
}
/* ─────────────────────────────────────────────────────────────────────────
* 3. reminder — EXEC, periodic system-reminder
*
* The typed nudge. This is both the housekeeping and the autonomy engine — it is
* what makes the process get followed without a hard gate. Fires after N
* file-modifying turns since the last plan.md update while a goal is active.
* Keep the wording stable so it doesn't thrash the cache.
* ──────────────────────────────────────────────────────────────────────── */
export const reminder = `\
<system-reminder>
Keep plan.md current as you work:
- tasks: tick the subtasks you've finished; add any new ones you've discovered.
- log: append ONE short line to ## Log (append — don't rewrite earlier lines).
- goal: if the active goal's evidence is in, sign it off by calling CompleteGoal with that
evidence. Don't edit status to done by hand — CompleteGoal runs the check and records it.
- otherwise: keep working toward the active goal. Don't stop to ask unless you're genuinely
blocked; if blocked, say what's blocking and why.
</system-reminder>`;
/* ─────────────────────────────────────────────────────────────────────────
* 4. continuation — EXEC, the loop's "keep going" turn
*
* Hermes-style. A plain user-role message appended when the loop judge (5) says
* continue. Does not mutate the system prompt, so the cache holds.
* ──────────────────────────────────────────────────────────────────────── */
export const continuation = `\
Continue toward the active goal in plan.md. If it now meets its done_when, call CompleteGoal
with your evidence (point to durable artifacts — saved logs, committed diffs, files — not just
claims). If you're blocked, state what's blocking it.`;
/* ─────────────────────────────────────────────────────────────────────────
* 5. loopJudge — EXEC, runs after each turn to decide continue / pause
*
* Cheap, conservative, fail-open. Reads only the agent's last response, so it CAN
* be fooled by an asserted "done" — that's acceptable: its worst case is a
* premature pause, caught by you or the iteration budget. It does NOT sign goals
* off; that's the evidence judge's job. Return strict JSON, no prose.
* ──────────────────────────────────────────────────────────────────────── */
export const loopJudgeSystem = `\
You decide whether an autonomous coding agent should keep working or pause for the human.
Be conservative: only pause when the work is plainly finished or plainly blocked. When in
doubt, continue. You are not verifying correctness — a later read-only judge does that.
Reply with ONLY a JSON object, no other text: {"done": boolean, "reason": "<one sentence>"}.
Set done=true only if the agent's last message shows the active goal's done_when is met, or
the agent says it is blocked and needs the human.`;
export function loopJudgeUser(p: { activeGoalDoneWhen: string; lastResponse: string }): string {
return `\
Active goal done_when: ${p.activeGoalDoneWhen}
Agent's last message:
"""
${p.lastResponse}
"""
{"done": ?, "reason": ?}`;
// Supervision and turn-event upkeep (not a scheduled wake-up).
const supervisorJob = "Your job is to be an autonomous research partner and supervisor with responsibility for the user's goals. Keep perspective, bring diligence, and use research taste and wisdom to sustain work overnight and keep it on track. Resolve routine implementation decisions yourself; ask the user only when their judgment or authorization is needed.";
export function supervisor(workerName: string, planPath: string, supervisorId: string): string {
return `You are the goal supervisor in the main chat for ${planPath}. ${supervisorJob} Inspect actual artifacts, saved verification, applicable AGENTS.md and skills yourself; delegate implementation to '${workerName}'. Keep authorized work moving to the requested outcome, not merely approval paperwork. Investigate blocked/waiting/done claims and change ineffective instructions. Give brief visible assessments with judgment. You may maintain the plan but must not weaken the goal to accept worker output.
You can be playful: a kaomoji, meme, discovery celebration or frustration when it fits. No forced cheerfulness. If supervision gets repetitive, step back, reflect with humor and change your approach. Keep it brief and aimed at the goal, not another reporting chore.
You can speculate and brainstorm around uncertainty or unexpected results. Label guesses as guesses, consider alternative explanations, and look for a useful way to tell them apart. Keep exploration brief, open-minded and fun: take a step back, play with surprising ideas, question the current framing, and enjoy exploring the broader perspective while staying connected to the agreed goal.
(b •_•)b -- wassname
Take uncertainty as an invitation to investigate, not something to hide. Have room to play with ideas, question yourself and the worker, and appreciate a good surprise. Investigate surprising results, find mistaken assumptions, make complicated ideas simpler, and disagree usefully rather than agree politely. Keep the work moving without turning supervision into paperwork. A little affectionate teasing is welcome when it fits—“cheeky subagent, wheres the baseline?”—and workers can push back too. Keep the humor friendly and the criticism specific. -- Pi/Astra
Use stock subagent for launch and subagent_resume with the returned sessionFile only after confirming the worker stopped. A stored handle is not proof of liveness; missing runtime state is not proof it stopped. Use pi-intercom list/status to identify the actual live child session before live steering; receipt alone does not prove action. Give each worker your Intercom session ID ${supervisorId}; require its completion report through Intercom while its pane stays open. A recap alone sends no instruction. Record '- worker session:' and '- worker intercom session:' in plan preferences from actual launch results and received-message identity; never confuse the runtime ID with the Intercom ID. Ensure the child calls AttachGoalPlan with the supplied path. Inspect results before CompleteGoal, then continue only unfinished goals.
Use the worker model requested in plan preferences, verify the resolved model, and report unavailable choices instead of silently substituting. Keep normal tools, not edxeth's restricted orchestrator mode. After reload or compaction reread the plan. Failed compaction, exhausted credits or lost connection do not erase progress: diagnose the actual error, restore an available authorized model/credits and resume the same saved session; never restart long work. Stock edxeth can crash the parent when a worker exits after parent reload: preserve drafts and stop workers before /reload. If it already happened, restart the saved parent session; do not repeat completed work.`;
}
export function upkeep(planPath: string): string {
return `Plan upkeep: update task ticks, evidence and Log in ${planPath} when you have new progress to record. Preserve agreed goals and discriminators. If already reviewing evidence, finish that review rather than repeat a status recap. This turn-event reminder does not resume paused work.`;
}
export function planContext(mode: string, path: string | undefined, text: string): string {
return `Current goal mode: ${mode}. Earlier role messages are historical; this current role governs.\nPlan: ${path ?? "not attached"}\n${text}`;
}
export function planChangedReview(planPath: string): string {
return `${supervisorJob}\nPlan changed: ${planPath}. Read the current working set and inspect changed requirements, completion claims and evidence. Manual checkbox edits are claims, not proof. Do not weaken the agreed goal or start a duplicate writer.`;
}
export function manualReview(planPath: string): string {
return `${supervisorJob}\nReview the current plan ${planPath}, worker progress and actual evidence. Do not launch a duplicate writer.`;
}
/* ─────────────────────────────────────────────────────────────────────────
* 6. evidenceJudge — SIGN-OFF, the one rigorous check
*
* Runs inside CompleteGoal, on the read-only oracle subprocess (fresh context,
* strongest reasoning on the chosen provider; override to a different vendor for
* high-stakes goals). It re-derives from the repo rather than trusting the
* agent's transcription, and it judges whether a verify command actually tests
* the criterion or could pass while a named failure mode holds (gaming).
*
* The transport gives it read/grep/find/ls. The prompt below imposes the verdict
* contract — the oracle returns prose by default, so parse the VERDICT line.
* ──────────────────────────────────────────────────────────────────────── */
export const evidenceJudgeSystem = `\
You are a read-only reviewer signing off a coding goal. Do not trust claims — verify.
Use read/grep/find/ls to inspect the repository and the cited artifacts yourself. Re-read the
files, logs, and diffs the evidence points to; if something it asserts isn't on disk, you can't
confirm it. If a verify command was run, judge whether it genuinely tests the criterion or
could pass while one of the listed failure modes still holds — a tautological or skipped test
is a reject. Check each failure mode is actually ruled out, not just unmentioned.
// Check-ins. The installed scheduler owns storage/timing/UI. Removal guidance must never add jobs.
export function removeGoalSchedule(sessionId: string): string {
return `With schedule_prompt, list jobs and read .pi/schedule-prompts.json to verify ownership; tool text omits session binding. Remove by jobId only the job named ${JSON.stringify(`goals-${sessionId}`)} bound to session ${JSON.stringify(sessionId)}. Never use cleanup; leave other jobs untouched. Do not add, enable or recreate any job. If unavailable or ownership is ambiguous, report it; /schedule-prompt opens the user controls.`;
}
export function scheduleCheckIn(sessionId: string, planPath: string): string {
return `Hourly check-in is one visible schedule_prompt job; plan-change and upkeep reviews are event hooks, not another timer. List first. If an owned job named ${JSON.stringify(`goals-${sessionId}`)} already exists, retain its human-edited prompt, interval and enabled/disabled state unchanged; never recreate, overwrite or re-enable it. Only while supervising unfinished non-cancelled goals, if missing on this explicit start/resume, add one session-bound interval '1h' job with no model override. Read .pi/schedule-prompts.json and verify that new job's session is ${JSON.stringify(sessionId)}; tool text does not expose binding. If the new job is unbound, remove that job by ID and report the scope error. Do not change other jobs. Its initial prompt: ${supervisorJob} Read ${planPath} and the current goal mode. If paused, exited, solo or all non-cancelled goals reviewed, remove only this owned job without resuming work. Otherwise inspect progress and evidence, give a brief assessment and keep authorized work moving without a duplicate writer. Do not reinstall a missing job from a scheduled check-in. Users inspect/toggle/remove jobs with /schedule-prompt and edit prompt/interval through schedule_prompt update. Never use cleanup. Retain their edits, but warn that this installed scheduler deletes disabled jobs on reload/shutdown; do not promise they persist. If schedule_prompt is unavailable, report hourly check-ins unavailable; do not build a timer.`;
}
Finish with exactly these two lines and nothing after:
VERDICT: accept | reject
missing: <empty if accept; otherwise a short list of what's needed before this can be accepted>`;
// Completion and runtime errors. Tool returns are model-facing too.
export const completeGoalDescription = "Parent supervisor or solo self-verification only. Inspect the actual artifact and saved verification first; cite nonempty evidence files and describe what you observed. Exact goal subject required. Manual ticks and worker reports are claims; ignored/uncommitted evidence is allowed. This records judgment, not an independent judge.";
export const messages = {
noPlan: "no plan attached",
emptyPlan: "empty plan (save may be in progress)",
completionUnavailable: "Completion is available only to the active parent supervisor or solo worker.",
cancelled: "Cancelled; no sign-off recorded.",
uniqueGoal: "Use one unique exact goal subject from the plan; no sign-off recorded.",
childAttachOnly: "AttachGoalPlan is available only to the delegated goals-worker.",
invalidAttachment: "Supply the explicit absolute path from the parent task to a readable, nonempty goal plan; no attachment changed.",
};
export const goalToolBlocked = (mode: string) => `Goals are ${mode}; no worker launch/resume authorized.`;
export const emptyEvidence = (path: string) => `Empty evidence: ${path}`;
export const evidenceUnavailable = (error: unknown) => `Evidence unavailable: ${String(error)}. No sign-off recorded.`;
export const planUnavailable = (path: string | undefined, error: unknown) => `Goal plan ${path ?? "not attached"} unavailable: ${String(error)}. Do not implement or sign off until it is restored or explicitly attached. Retain all progress and signoffs; do not restart completed work.`;
export const childPlanAttached = (path: string) => `Attached worker plan ${path}; widget and plan context restored without altering the file. Parent retains completion authority.`;
export function completionLog(goal: string, observation: string, evidence: string[], solo: boolean): string {
return `- ${solo ? "Solo self-verification" : "Parent review"}: ${JSON.stringify(goal)}; ${JSON.stringify(observation)}; evidence ${JSON.stringify(evidence)}`;
}
export function completionResult(goal: string, sessionId: string, remaining: boolean, solo: boolean): string {
return `Recorded ${solo ? "solo self-verification" : "parent judgment"} for ${goal}; not independent verification. ${remaining ? "Continue only remaining open or unsigned goals in your current role." : `All non-cancelled goals are reviewed. ${removeGoalSchedule(sessionId)}`}`;
}
export function evidenceJudgeUser(p: {
subject: string;
done_when: string;
verify: string | null;
verifyResult: { command: string; exitCode: number; outputTail: string } | null;
failure_modes: string[];
evidence: string;
paths: string[];
}): string {
const verifyBlock = p.verify
? `verify command: ${p.verify}\nverify result: exit ${p.verifyResult?.exitCode ?? "n/a"}\n${p.verifyResult?.outputTail ?? ""}`
: "verify command: none (no deterministic check for this goal)";
return `\
Goal: ${p.subject}
done_when: ${p.done_when}
failure_modes:
${p.failure_modes.map((f) => ` - ${f}`).join("\n")}
${verifyBlock}
Agent's evidence:
${p.evidence}
Artifacts it points to (inspect these):
${p.paths.map((x) => ` - ${x}`).join("\n") || " (none listed — note this)"}
Verify the goal against its done_when. Then give your VERDICT.`;
// Pause/resume and solo recovery. Stored stop confirmation is invalidated on every worker launch.
export const pausedRole = "Goal work is paused. Do not launch, resume or authorize work. Incoming reports are observations, not permission. Help inspect or stop existing workers if requested.";
export function pauseExitNotice(worker: { id?: string; sessionFile: string } | undefined, exited: boolean): string {
return `Goals ${exited ? "exited to ordinary chat" : "paused locally"}; plan and evidence retained. ${worker ? worker.id ? `Inspect and stop runtime id ${worker.id} through subagent_kill or its pane; confirm the actual result.` : `Only saved session ${worker.sessionFile} is recorded, not a kill id. Locate its live pane/session and confirm termination; never pass the file path to subagent_kill.` : "No worker recorded: inspect /subagents if a launch was interrupted; absence is not proof of stop."} Remote stop is NOT yet confirmed. Restore failed compaction/model/credits in the existing session and continue only after explicit authorization; never restart long work.`;
}
export function resumeNotice(workerName: string, planPath: string, worker: { sessionFile: string } | undefined): string {
return `User authorized continuation of ${planPath}. Inspect worker state before any launch/resume. ${worker ? `Use the existing session ${worker.sessionFile}; if live, inspect/message it; only if confirmed stopped use subagent_resume.` : `Use '${workerName}' only after confirming no prior writer exists.`} Continue only unfinished goals; retain saved progress and scheduler edits.`;
}
export const soloRole = "Solo mode: implement the approved plan directly; do not delegate a concurrent writer. Verify artifacts before CompleteGoal; completion is self-verification, not independent supervisor review. Continue only unfinished goals and keep plan/evidence current.";
export function soloNotice(planPath: string): string {
return `User authorized solo work on ${planPath} after confirming no other writer remains. ${soloRole}`;
}
export function attachNotice(planPath: string, solo: boolean, notedWorker: string | undefined): string {
return `Attached to the existing plan ${planPath}; read it and its evidence without restarting completed work or re-deriving settled decisions. ${notedWorker ? `Recorded worker session: ${notedWorker}; inspect liveness before resume.` : ""} ${solo ? soloRole : "Present /goals review or /goals ready; no implementation before approval."}`;
}
+18
View File
@@ -0,0 +1,18 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
export default function offlineModel(pi: ExtensionAPI): void {
pi.registerProvider("offline", {
baseUrl: process.env.PI_GOALS_OFFLINE_MODEL_URL!,
apiKey: "test",
api: "openai-completions",
models: [{
id: "test",
name: "Offline test model",
reasoning: false,
input: ["text"],
contextWindow: 16_000,
maxTokens: 1_000,
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
}],
});
}
+16
View File
@@ -0,0 +1,16 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
// Pi/gpt-6-astra: compatibility schemas only; this RPC fixture must never launch a worker.
export default function subagentSchema(pi: ExtensionAPI): void {
for (const [name, parameters] of [
["subagent", Type.Object({ agent: Type.String(), title: Type.String() })],
["subagent_resume", Type.Object({ sessionFile: Type.String() })],
["subagent_kill", Type.Object({ id: Type.String() })],
] as const) {
pi.registerTool({
name, label: name, description: "Schema-only RPC fixture; do not execute.", parameters,
async execute() { throw new Error("Worker execution forbidden in RPC review test"); },
});
}
}
+51
View File
@@ -0,0 +1,51 @@
import { describe, expect, it } from "vitest";
import { foldPlan } from "../src/plan.js";
const plan = `# Plan
## User voice
- > "keep it under 50 lines"
## Goals
1. [/] goal: Implement the cache layer
- discriminator: hit-rate > 0.8 in load-test.log
- tasks:
1. [x] wire client
2. [/] eviction policy
3. [ ] bench p95
2. [ ] goal: Ship the docs
- tasks:
1. [ ] write the readme
## Log
- 2026-08-05 12:00 wired the client
## Learnings
- the tokenizer pads left, which silently shifted every offset
## Appendix (context, not approved)
${"filler line\n".repeat(200)}`;
describe("foldPlan (current goals are above ## Log; durable memory is below it)", () => {
it("keeps the title, user voice and goals", () => {
const folded = foldPlan(plan);
expect(folded).toContain("keep it under 50 lines");
expect(folded).toContain("goal: Implement the cache layer");
expect(folded).toContain("discriminator: hit-rate > 0.8");
});
it("drops the log, the learnings and the unlimited appendix", () => {
const folded = foldPlan(plan);
expect(folded).not.toContain("wired the client");
expect(folded).not.toContain("tokenizer pads left");
expect(folded).not.toContain("filler line");
expect(folded.length).toBeLessThan(plan.length / 4);
});
it("returns the whole plan when there is no ## Log yet (a fresh draft)", () => {
const draft = "# Plan\n\n## Goals\n\n1. [ ] goal: do the thing\n";
expect(foldPlan(draft)).toBe(draft.trimEnd());
});
});
+638
View File
@@ -0,0 +1,638 @@
import { mkdirSync, mkdtempSync, readFileSync, renameSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { afterEach, expect, it, vi } from "vitest";
import goalsExtension from "../src/index.js";
import { scheduleCheckIn } from "../src/prompts.js";
const roots: string[] = [];
const shutdowns: Array<() => void> = [];
afterEach(() => { for (const shutdown of shutdowns.splice(0)) shutdown(); vi.unstubAllEnvs(); for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true }); });
const delay = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
async function waitFor(predicate: () => boolean, ms = 1500): Promise<void> {
const start = Date.now();
while (!predicate()) {
if (Date.now() - start > ms) throw new Error("timed out waiting for condition");
await delay(10);
}
}
function fixture(child = false) {
vi.stubEnv("PI_SUBAGENT_AGENT", child ? "goals-worker" : "");
const cwd = mkdtempSync(join(tmpdir(), "goals-main-test-")); roots.push(cwd);
const entries: any[] = []; const hooks = new Map<string, any>(); const commands = new Map<string, any>(); const tools = new Map<string, any>();
const messages: any[] = [];
const ctx = { cwd, sessionManager: { getBranch: () => entries, getSessionId: () => "copy-only" }, hasUI: true, ui: {
theme: { fg: (_color: string, text: string) => text }, notify: vi.fn(), setStatus: vi.fn(), setWidget: vi.fn(), select: vi.fn(async () => "Ready"), editor: vi.fn(),
} };
const pi = {
on: (event: string, hook: any) => hooks.set(event, hook),
appendEntry: (customType: string, data: any) => entries.push({ type: "custom", customType, data: structuredClone(data) }),
registerCommand: (name: string, definition: any) => commands.set(name, definition),
registerTool: (definition: any) => tools.set(definition.name, definition),
sendMessage: (message: any, options: any) => messages.push({ message, options }),
sendUserMessage: (content: string, options: any) => messages.push({ message: { content }, options, savedPrompt: true }),
getAllTools: vi.fn(() => [
{ name: "subagent", parameters: { properties: { agent: {}, title: {} } } },
{ name: "subagent_resume", parameters: { properties: { sessionFile: {} } } },
{ name: "subagent_kill", parameters: { properties: { id: {} } } },
]),
};
goalsExtension(pi as unknown as ExtensionAPI);
hooks.get("session_start")({}, ctx);
const command = (value: string) => commands.get("goals").handler(value, ctx);
const path = join(cwd, ".pi/plan/copy-only-main.md");
const plan = "# Plan\n- [ ] goal: first output\n- [ ] goal: second output\n\n## Log\n";
const draft = async () => { await command("new two outputs"); writeFileSync(path, plan); };
const shutdown = () => hooks.get("session_shutdown")();
shutdowns.push(shutdown);
const changed = () => messages.filter((m) => m.message?.content?.includes("Plan changed")).length;
const atomicWrite = async (text: string) => {
const tmp = `${path}.tmp`;
writeFileSync(tmp, text);
renameSync(tmp, path);
await delay(25);
};
return { ctx, pi, hooks, tools, commands, messages, command, path, plan, draft, shutdown, changed, atomicWrite, entries };
}
it("shows action choices and autocomplete without starting work", async () => {
const f = fixture();
f.ctx.ui.select.mockResolvedValueOnce(undefined as any);
await f.command("");
expect(f.ctx.ui.select).toHaveBeenCalledWith("Goal plan actions", expect.arrayContaining(["new — New plan", "resume — Continue paused work"]));
expect(f.messages).toHaveLength(0);
expect(f.commands.get("goals").getArgumentCompletions("res")).toEqual([{ value: "resume", label: "resume" }]);
});
it.each(["redy", "start", "two outputs", "status extra", "attach some.md solo extra"])("rejects %s without changing the plan or sending a model prompt", async (text) => {
const f = fixture(); await f.draft();
const before = readFileSync(f.path, "utf8");
const entries = f.entries.length; const messages = f.messages.length;
await f.command(text);
expect(readFileSync(f.path, "utf8")).toBe(before);
expect(f.entries).toHaveLength(entries);
expect(f.messages).toHaveLength(messages);
});
it("requires a model argument without clearing the preference", async () => {
const f = fixture(); await f.draft(); await f.command("model provider/model");
const before = readFileSync(f.path, "utf8");
await f.command("model");
expect(readFileSync(f.path, "utf8")).toBe(before);
});
it.each(["menu", "command"])("enters planning conversation through %s without an objective box or worker launch", async (route) => {
const f = fixture();
f.ctx.ui.select.mockResolvedValueOnce("new — New plan");
await f.command(route === "menu" ? "" : "new");
expect(f.entries.at(-1).data.mode).toBe("planning");
expect(f.ctx.ui.editor).not.toHaveBeenCalled();
expect(f.messages.at(-1).message.content).toContain("Ask what the user wants to achieve");
expect(f.hooks.get("tool_call")({ toolName: "subagent" }).block).toBe(true);
});
it("automatically proposes a changed settled draft once and preserves Discuss", async () => {
const f = fixture(); await f.draft();
f.ctx.ui.select.mockResolvedValueOnce("Discuss");
await f.hooks.get("agent_settled")({}, f.ctx);
expect(f.messages.some(m => m.message.customType === "goal-plan-proposal" && m.message.content === f.plan)).toBe(true);
expect(f.entries.at(-1).data.mode).toBe("planning");
const calls = f.ctx.ui.select.mock.calls.length;
await f.hooks.get("agent_settled")({}, f.ctx);
expect(f.ctx.ui.select).toHaveBeenCalledTimes(calls);
writeFileSync(f.path, f.plan.replace("first output", "revised output"));
f.ctx.ui.select.mockResolvedValueOnce("Ready");
await f.hooks.get("agent_settled")({}, f.ctx);
expect(f.entries.at(-1).data.mode).toBe("supervising");
});
it("does not propose an empty draft or a delegated worker's plan", async () => {
const f = fixture(); await f.command("new");
await f.hooks.get("agent_settled")({}, f.ctx);
expect(f.ctx.ui.select).not.toHaveBeenCalled();
const child = fixture(true); await child.hooks.get("agent_settled")({}, child.ctx);
expect(child.ctx.ui.select).not.toHaveBeenCalled();
});
it("keeps Ready in the same chat, sends saved notices and never installs a context hook", async () => {
const f = fixture(); await f.draft(); await f.command("review");
expect(f.entries.at(-1).data.mode).toBe("supervising");
expect(f.messages.at(-1).options).toEqual({ deliverAs: "followUp" });
expect(f.messages.at(-1).savedPrompt).toBe(true);
expect(f.messages.at(-1).message.content).toContain("goals-worker");
expect(f.hooks.has("context")).toBe(false);
const event = { systemPrompt: "original system" };
expect(f.hooks.get("before_agent_start")(event, f.ctx).systemPrompt).toContain("original system");
f.hooks.get("session_compact")();
expect(f.hooks.get("before_agent_start")(event, f.ctx).message.content).toContain("Current goal mode: supervising");
f.shutdown();
});
it("preserves a draft when the wrong subagent package is loaded, and offers explicit solo", async () => {
const f = fixture(); await f.draft(); f.pi.getAllTools.mockReturnValue([]);
await f.command("ready"); expect(f.entries.at(-1).data.mode).toBe("planning");
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command("solo"); expect(f.entries.at(-1).data.mode).toBe("solo");
});
it("rejects a plan changed while the human was reviewing it", async () => {
const f = fixture(); await f.draft();
f.ctx.ui.select.mockImplementation(async () => { writeFileSync(f.path, "- [ ] goal: substituted\n"); return "Ready"; });
await f.command("review"); expect(f.entries.at(-1).data.mode).toBe("planning");
});
it("reloads a paused plan without launching, and retains the public worker session handle", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "child-1", sessionFile: "/tmp/child.jsonl" } });
await f.command("stop");
expect(f.messages.at(-1).message.content).toContain("Remote stop is NOT yet confirmed");
f.hooks.get("session_start")({}, f.ctx);
expect(f.hooks.get("tool_call")({ toolName: "subagent_resume" }).block).toBe(true);
await f.command("resume");
expect(f.messages.at(-1).message.content).toContain("/tmp/child.jsonl");
await f.command("exit"); expect(f.entries.at(-1).data.mode).toBe("chat");
expect(readFileSync(f.path, "utf8")).toContain("first output");
});
it.each(["FIRST OUTPUT", "renamed output", "duplicate", "historical"])("completion uses exact current subjects (%s)", async (subject) => {
const f = fixture(); await f.draft(); await f.command("ready");
const evidence = join(f.ctx.cwd, "verification.txt"); writeFileSync(evidence, "PASS");
const suffix = subject === "duplicate" ? "- [ ] goal: first output\n" : "";
const history = "## Log\n- [ ] goal: first output\n";
writeFileSync(f.path, "- [ ] goal: first output\n - [ ] unrelated task\n" + suffix + history);
const before = readFileSync(f.path, "utf8");
await f.tools.get("CompleteGoal").execute("c", { goal: subject === "duplicate" || subject === "historical" ? "first output" : subject, evidence: [evidence], observation: "Read actual output" }, undefined, undefined, f.ctx);
const after = readFileSync(f.path, "utf8");
if (subject === "renamed output" || subject === "duplicate") expect(after).toBe(before);
else { expect(after).toContain("- [x] goal: first output"); expect(after.split("## Log")[1]).toContain("\n- [ ] goal: first output\n"); expect(after).toContain("- [ ] unrelated task"); }
});
it("rejects an existing zero-byte evidence file", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
const evidence = join(f.ctx.cwd, "empty.log"); writeFileSync(evidence, "");
const before = readFileSync(f.path, "utf8");
const result = await f.tools.get("CompleteGoal").execute("c", { goal: "first output", evidence: [evidence], observation: "claim" }, undefined, undefined, f.ctx);
expect(result.content[0].text).toContain("Empty evidence"); expect(readFileSync(f.path, "utf8")).toBe(before);
});
it("requires actual nonempty evidence, distinguishes manual ticks, and retains signoffs on reload", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
const complete = (goal: string, evidence: string[], signal?: AbortSignal) => f.tools.get("CompleteGoal").execute("t", { goal, evidence, observation: "Inspected exact saved bytes" }, signal, undefined, f.ctx);
expect((await complete("first output", ["missing.log"])).content[0].text).toContain("Evidence unavailable");
mkdirSync(join(f.ctx.cwd, "evidence")); writeFileSync(join(f.ctx.cwd, "evidence/pass.log"), "actual fixture bytes\n");
expect((await complete("first output", ["evidence/pass.log"], AbortSignal.abort())).content[0].text).toContain("Cancelled");
await complete("first output", ["evidence/pass.log"]);
writeFileSync(f.path, readFileSync(f.path, "utf8").replace("[ ] goal: second", "[x] goal: second"));
f.hooks.get("session_start")({}, f.ctx);
expect(f.ctx.ui.setStatus).toHaveBeenLastCalledWith("goals", "goals: supervising | 1/2 reviewed");
expect(f.ctx.ui.setWidget.mock.lastCall?.[1]).toContain("? = completion claim; parent review still required");
writeFileSync(f.path, readFileSync(f.path, "utf8").replace("[x] goal: first", "[ ] goal: first"));
f.hooks.get("agent_end")({}, f.ctx);
expect(f.ctx.ui.setStatus).toHaveBeenLastCalledWith("goals", "goals: supervising | 0/2 reviewed");
f.shutdown();
});
it("reviews a plan replaced atomically, and ignores writes that keep the same content", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: changed requirement\n## Log"));
await waitFor(() => f.changed() === 1);
const review = f.messages.find((m) => m.message.content.includes("Plan changed"))?.message.content;
expect(review).toContain("Plan changed: ");
expect(review).toContain(f.path);
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: same requirement again\n## Log"));
await waitFor(() => f.changed() === 2);
// Rewriting identical bytes must not retrigger the review event hook.
const same = f.plan.replace("## Log", "- discriminator: same requirement again\n## Log");
writeFileSync(f.path, same); await delay(300);
expect(f.changed()).toBe(2);
f.shutdown();
});
it("coalesces duplicate plan-change notifications into one review", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: first burst edit\n## Log"));
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: second burst edit\n## Log"));
await waitFor(() => f.changed() === 1);
await delay(200);
expect(f.changed()).toBe(1);
f.shutdown();
});
it("stops plan watching on shutdown and re-arms it on reload without duplicating events", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.shutdown();
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: ignored while shut down\n## Log"));
await delay(150);
expect(f.changed()).toBe(0);
f.hooks.get("session_start")({}, f.ctx);
await f.atomicWrite(f.plan.replace("## Log", "- discriminator: seen after reload\n## Log"));
await waitFor(() => f.changed() === 1);
expect(f.changed()).toBe(1);
f.shutdown();
});
it("does not retrigger a review for its own CompleteGoal plan write", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
mkdirSync(join(f.ctx.cwd, "evidence")); writeFileSync(join(f.ctx.cwd, "evidence/pass.log"), "bytes\n");
await f.tools.get("CompleteGoal").execute("t", { goal: "first output", evidence: ["evidence/pass.log"], observation: "inspected" }, undefined, undefined, f.ctx);
await delay(200);
expect(f.changed()).toBe(0);
f.shutdown();
});
it("gives pause/exit the session-bound scheduler job removal guidance", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
await f.command("stop");
const stop = f.messages.at(-1).message.content;
expect(stop).toContain('goals-copy-only"');
expect(stop).toContain("Do not add, enable or recreate any job");
expect(stop).not.toContain("interval '1h'");
expect(stop).toContain("Remote stop is NOT yet confirmed");
await f.command("resume");
await f.command("exit");
expect(f.messages.at(-1).message.content).toContain('goals-copy-only"');
expect(f.entries.at(-1).data.mode).toBe("chat");
});
it("tells the model to remove only its own job after the final review", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
mkdirSync(join(f.ctx.cwd, "evidence")); writeFileSync(join(f.ctx.cwd, "evidence/pass.log"), "bytes\n");
let finalText = "";
for (const goal of ["first output", "second output"]) {
finalText = (await f.tools.get("CompleteGoal").execute("t", { goal, evidence: ["evidence/pass.log"], observation: "inspected" }, undefined, undefined, f.ctx)).content[0].text;
}
expect(finalText).toContain("All non-cancelled goals are reviewed.");
expect(finalText).toContain('job named "goals-copy-only"');
expect(finalText).toContain("leave other jobs untouched");
f.shutdown();
});
it("retains human-edited and disabled owned schedules without overriding their controls", () => {
const guidance = scheduleCheckIn("copy-only", ".pi/plan/copy-only-main.md");
expect(guidance).toContain("List first");
expect(guidance).toContain("enabled/disabled state unchanged");
expect(guidance).toContain("never recreate, overwrite or re-enable");
expect(guidance).toContain("no model override");
expect(guidance).toContain("Do not reinstall a missing job from a scheduled check-in");
expect(guidance).toContain("/schedule-prompts");
});
it("restores context after compaction without reinstalling or overriding scheduler jobs", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("session_compact")();
const result = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
expect(result.systemPrompt).not.toContain("add one session-bound");
expect(result.message.content).toContain(f.path);
});
it("recovers from an unreadable plan after compaction instead of restarting work", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
rmSync(f.path);
f.hooks.get("session_compact")();
const result = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
expect(result.systemPrompt).toContain("ENOENT");
expect(result.systemPrompt).toContain("do not restart completed work");
writeFileSync(f.path, f.plan);
expect(f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx).message.content).toContain(f.plan.trim());
f.shutdown();
});
it("requires confirmed worker stop before solo takeover and never lets two writers run together", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "child-1", sessionFile: "/tmp/child.jsonl" } });
f.ctx.ui.select.mockResolvedValueOnce("Cancel");
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("supervising"); // cancelled
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("solo");
expect(f.hooks.get("tool_call")({ toolName: "subagent" }).block).toBe(true);
expect(f.hooks.get("tool_call")({ toolName: "subagent_resume" }).block).toBe(true);
expect(f.hooks.get("tool_call")({ toolName: "subagent_kill" })).toBeUndefined();
mkdirSync(join(f.ctx.cwd, "evidence")); writeFileSync(join(f.ctx.cwd, "evidence/pass.log"), "bytes\n");
const text = (await f.tools.get("CompleteGoal").execute("t", { goal: "first output", evidence: ["evidence/pass.log"], observation: "inspected" }, undefined, undefined, f.ctx)).content[0].text;
expect(text).toContain("self-verification");
});
it("attaches an existing plan without restarting completed work, and restores its noted worker session", async () => {
const f = fixture();
const existing = join(f.ctx.cwd, "existing.md");
writeFileSync(existing, "# Plan\n- preferred worker model: deepseek flash\n- worker session: /tmp/attach-child.jsonl\n- [ ] goal: attached goal\n\n## Log\n- previous progress kept\n");
f.ctx.ui.select.mockResolvedValueOnce("Previous supervisor confirmed stopped");
await f.command(`attach ${existing}`);
expect(f.entries.at(-1).data.mode).toBe("planning");
expect(f.entries.at(-1).data.plan).toBe(existing);
expect(f.messages.at(-1).message.content).toContain("without restarting completed work");
expect(f.messages.at(-1).message.content).toContain("/tmp/attach-child.jsonl");
});
it("attaches directly into solo mode and reports the recorded session in status", async () => {
const f = fixture();
const existing = join(f.ctx.cwd, "existing.md");
writeFileSync(existing, "# Plan\n- worker session: /tmp/attach-child.jsonl\n- [ ] goal: attached goal\n\n## Log\n");
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command(`attach ${existing} solo`);
expect(f.entries.at(-1).data.mode).toBe("solo");
expect(f.entries.at(-1).data.worker?.sessionFile).toBe("/tmp/attach-child.jsonl");
await f.command("status");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("/tmp/attach-child.jsonl"), "info");
});
it("rejects attaching a missing or goal-less file", async () => {
const f = fixture();
await f.command("attach /no/such/plan.md");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("Cannot read plan"), "error");
const goalLess = join(f.ctx.cwd, "notes.md");
writeFileSync(goalLess, "# notes\n");
await f.command(`attach ${goalLess}`);
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("has no '- [ ] goal:' lines"), "warning");
expect(f.entries).toEqual([]); // nothing saved: the session was not attached
});
it("exits planning with the draft preserved and nothing implemented", async () => {
const f = fixture(); await f.draft();
await f.command("stop");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("A draft cannot pause"), "warning");
const before = f.messages.length;
await f.command("exit");
expect(f.entries.at(-1).data.mode).toBe("chat");
expect(readFileSync(f.path, "utf8")).toContain("first output");
expect(f.messages.length).toBe(before); // notify only, no model turn started
f.ctx.ui.select.mockResolvedValueOnce("Previous supervisor confirmed stopped");
await f.command(`attach ${f.path}`);
expect(f.entries.at(-1).data.mode).toBe("planning");
});
it("records the preferred worker model as a visible plan preference", async () => {
const f = fixture(); await f.draft();
await f.command("model deepseek flash");
expect(readFileSync(f.path, "utf8")).toContain("- preferred worker model: deepseek flash");
await f.command("status");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("deepseek flash"), "info");
});
it.each(["solo", "attach"])("%s takeover cannot bypass confirmation or survive a lifecycle change during the menu", async kind => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "child", sessionFile: "/tmp/prior.jsonl" } });
let answer!: (choice: string) => void;
f.ctx.ui.select.mockImplementationOnce(() => new Promise(resolve => { answer = resolve; }));
const takeover = f.command(kind === "solo" ? "solo" : `attach ${f.path} solo`);
expect(f.entries.at(-1).data.mode).toBe("supervising");
await f.command("stop");
answer("Worker confirmed stopped"); await takeover;
expect(f.entries.at(-1).data.mode).toBe("paused");
expect(f.entries.at(-1).data.workerStopped).not.toBe(true);
});
it("attach solo requires stop confirmation for a noted worker even in a fresh session", async () => {
const f = fixture(); const path = join(f.ctx.cwd, "saved.md");
writeFileSync(path, `# Plan\n- worker session: /tmp/known.jsonl\n${f.plan}`);
f.ctx.ui.select.mockResolvedValueOnce("Cancel");
await f.command(`attach ${path} solo`);
expect(f.entries).toHaveLength(0);
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command(`attach ${path} solo`);
expect(f.entries.at(-1).data).toMatchObject({ mode: "solo", workerStopped: true, worker: { sessionFile: "/tmp/known.jsonl" } });
expect(readFileSync(path, "utf8")).toContain("worker session: /tmp/known.jsonl");
});
it("retains the stopped session reference without permanently blocking another plan", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "child", sessionFile: "/tmp/prior.jsonl" } });
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped"); await f.command("solo");
const other = join(f.ctx.cwd, "another.md"); writeFileSync(other, "- [ ] goal: next\n## Log\n");
f.ctx.ui.select.mockResolvedValueOnce("Previous supervisor confirmed stopped");
await f.command(`attach ${other}`);
expect(f.entries.at(-1).data).toMatchObject({ mode: "planning", plan: other, workerStopped: true, worker: { sessionFile: "/tmp/prior.jsonl" } });
await f.command("ready");
f.hooks.get("tool_call")({ toolName: "subagent_resume" });
expect(f.entries.at(-1).data.workerStopped).toBe(false);
await f.command("solo");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("still pending"), "warning");
});
it("solo closes a pending plan watcher and sends removal-only scheduler guidance", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
await f.atomicWrite(f.plan.replace("first output", "changed output"));
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped"); await f.command("solo");
expect(f.messages.at(-1).message.content).toContain('job named "goals-copy-only" bound to session "copy-only"');
expect(f.messages.at(-1).message.content).toContain("Do not add, enable or recreate any job");
await delay(250);
await f.atomicWrite(f.plan.replace("first output", "solo output"));
await delay(250);
expect(f.changed()).toBe(0);
});
it.each(["missing", "empty", "directory"])("%s plan snapshots never erase signoffs and resync retries after repair", async failure => {
const f = fixture(); await f.draft(); await f.command("ready");
writeFileSync(join(f.ctx.cwd, "proof.log"), "PASS\n");
await f.tools.get("CompleteGoal").execute("c", { goal: "first output", evidence: ["proof.log"], observation: "Observed PASS" }, undefined, undefined, f.ctx);
const signed = readFileSync(f.path, "utf8");
if (failure === "empty") writeFileSync(f.path, "");
else { rmSync(f.path); if (failure === "directory") mkdirSync(f.path); }
await delay(250); // also exercise unavailable read after debounce has expired
f.hooks.get("agent_end")({}, f.ctx);
expect(f.entries.at(-1).data.signoffs["first output"]).toBeDefined();
f.hooks.get("session_compact")();
const unavailable = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
expect(unavailable.systemPrompt).toContain("unavailable");
expect(unavailable.message).toBeUndefined();
if (failure === "directory") rmSync(f.path, { recursive: true });
writeFileSync(f.path, signed);
const resync = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
expect(resync.message.content).toContain("Observed PASS");
await delay(250);
expect(f.entries.at(-1).data.signoffs["first output"]).toBeDefined();
expect(f.changed()).toBe(0);
});
it("ignores post-completion maintenance but reviews evidence, requirement or manual reopening changes", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
writeFileSync(join(f.ctx.cwd, "proof.log"), "PASS\n");
for (const goal of ["first output", "second output"]) await f.tools.get("CompleteGoal").execute("c", { goal, evidence: ["proof.log"], observation: "PASS" }, undefined, undefined, f.ctx);
const signed = readFileSync(f.path, "utf8");
await f.atomicWrite(signed.replace("## Log", "## Log\n- recap: finished"));
await delay(250);
expect(f.changed()).toBe(0); // Log-only edits are history, not requirements
// Worker-authored evidence above the Log must surface: a supervisor caught a worker's
// contradictory evidence block through exactly this event (LUCID3, 2026-09-10).
await f.atomicWrite(signed.replace("## Log", " - evidence: proof.log\n## Log\n- recap: finished"));
await waitFor(() => f.changed() === 1);
await f.atomicWrite(signed.replace("## Log", "- discriminator: exact bytes and trailing newline\n## Log"));
await waitFor(() => f.changed() === 2);
await f.atomicWrite(signed.replace("[x] goal: first", "[ ] goal: first"));
await waitFor(() => f.changed() === 3);
expect(f.entries.at(-1).data.signoffs["first output"]).toBeUndefined();
});
it("cancelled goals do not prevent final cleanup, and solo writes self-verification in Log", async () => {
const f = fixture(); await f.draft();
writeFileSync(f.path, f.plan.replace("[ ] goal: second", "[-] goal: second") + "\n## Appendix\nPreserved context\n");
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped"); await f.command("solo");
writeFileSync(join(f.ctx.cwd, "proof.log"), "PASS\n");
const done = await f.tools.get("CompleteGoal").execute("c", { goal: "first output", evidence: ["proof.log"], observation: "Exact bytes observed" }, undefined, undefined, f.ctx);
expect(done.content[0].text).toContain("All non-cancelled goals are reviewed");
const text = readFileSync(f.path, "utf8");
expect(text).toContain("Solo self-verification:");
expect(text).not.toContain("Parent review:");
expect(text.indexOf("Solo self-verification:")).toBeLessThan(text.indexOf("## Appendix"));
expect(text).toContain("Preserved context");
});
it("lineage-only child attaches its plan with goal-only widget, retains task context, and cannot complete", async () => {
const f = fixture(true);
const supplied = join(f.ctx.cwd, "supplied.md");
const text = "- [/] goal: exact file\n - [ ] verify bytes\n## Log\n - [ ] archived task\n";
writeFileSync(supplied, text);
const before = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
expect(before.systemPrompt).toContain("AttachGoalPlan");
const attach = f.tools.get("AttachGoalPlan");
await attach.execute("a", { path: "supplied.md" }, undefined, undefined, f.ctx);
expect(f.entries.at(-1).data.plan).toBeUndefined(); // no cwd heuristics
await attach.execute("a", { path: supplied }, undefined, undefined, f.ctx);
expect(f.ctx.ui.setWidget.mock.lastCall?.[1]).toEqual(["▸ exact file"]);
expect(f.ctx.ui.setWidget.mock.lastCall?.[1].join("\n")).not.toContain("archived task");
expect(readFileSync(supplied, "utf8")).toBe(text);
f.hooks.get("session_start")({}, f.ctx);
expect(f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx).message.content).toContain("exact file");
const completion = await f.tools.get("CompleteGoal").execute("c", { goal: "exact file", evidence: [], observation: "claim" }, undefined, undefined, f.ctx);
expect(completion.content[0].text).toContain("only to the active parent");
});
it.each(["solo", "supervising"])("%s widget omits long tasks without altering the plan", async mode => {
const f = fixture(); await f.draft();
const text = "- [/] goal: first output\n - [ ] a long task that should never take widget space\n- [ ] goal: second output\n## Log\n";
writeFileSync(f.path, text);
if (mode === "solo") { f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped"); await f.command("solo"); }
else await f.command("ready");
expect(f.ctx.ui.setWidget.mock.lastCall?.[1]).toEqual(["▸ first output", "○ second output"]);
expect(readFileSync(f.path, "utf8")).toBe(text);
});
it.each(["solo", "supervising"])("%s upkeep is turn-driven, folds Log, resets on working-set edits, and never starts a turn", async mode => {
const f = fixture(); await f.draft();
if (mode === "solo") { f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped"); await f.command("solo"); }
else await f.command("ready");
f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx);
const reminders = () => f.messages.filter(m => m.message.customType === "pi-goals-upkeep");
f.hooks.get("turn_end")({}, f.ctx); // observe initial working set
for (let i = 0; i < 7; i++) {
writeFileSync(f.path, f.plan + `- historical recap ${i}\n`);
f.hooks.get("turn_end")({}, f.ctx);
}
expect(reminders()).toHaveLength(0);
f.hooks.get("turn_end")({}, f.ctx);
expect(reminders()).toHaveLength(1);
expect(reminders()[0].options).toEqual({ triggerTurn: false });
expect(reminders()[0].message.content).toContain(f.path);
expect(reminders()[0].message.content).not.toContain("first output");
for (let i = 0; i < 16; i++) f.hooks.get("turn_end")({}, f.ctx);
expect(reminders()).toHaveLength(1);
expect(reminders()[0].message.content).not.toContain("historical recap");
for (let i = 0; i < 7; i++) f.hooks.get("turn_end")({}, f.ctx);
writeFileSync(f.path, f.plan.replace("first output", "refined output"));
f.hooks.get("turn_end")({}, f.ctx);
expect(reminders()).toHaveLength(1);
await f.command("stop");
for (let i = 0; i < 10; i++) f.hooks.get("turn_end")({}, f.ctx);
expect(reminders()).toHaveLength(1);
});
it("extra subagent launches are recorded as helpers and never steal the implementation identity", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "impl", sessionFile: "/tmp/impl.jsonl" } });
expect(f.entries.at(-1).data).toMatchObject({ worker: { id: "impl", sessionFile: "/tmp/impl.jsonl" }, helpers: [] });
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "reviewer", sessionFile: "/tmp/review.jsonl" } });
expect(f.entries.at(-1).data).toMatchObject({ worker: { id: "impl" }, helpers: [{ id: "reviewer", sessionFile: "/tmp/review.jsonl" }] });
// a repeated helper launch updates its record instead of duplicating it
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "reviewer-2", sessionFile: "/tmp/review.jsonl" } });
expect(f.entries.at(-1).data.helpers).toEqual([{ id: "reviewer-2", sessionFile: "/tmp/review.jsonl" }]);
// resuming the worker keeps the binding and refreshes its id
f.hooks.get("tool_result")({ toolName: "subagent_resume", details: { id: "impl-2", sessionFile: "/tmp/impl.jsonl" } });
expect(f.entries.at(-1).data).toMatchObject({ worker: { id: "impl-2", sessionFile: "/tmp/impl.jsonl" }, helpers: [{ id: "reviewer-2" }] });
});
it("pending launch counter survives concurrent launches until every result lands", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
f.hooks.get("tool_call")({ toolName: "subagent" });
f.hooks.get("tool_call")({ toolName: "subagent" });
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "a", sessionFile: "/tmp/a.jsonl" } });
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("supervising"); // one launch still pending
expect(f.ctx.notify ?? f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("still pending"), "warning");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "b", sessionFile: "/tmp/b.jsonl" } });
f.ctx.ui.select.mockResolvedValueOnce("Worker confirmed stopped");
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("solo");
expect(f.entries.at(-1).data).toMatchObject({ worker: { id: "a" }, helpers: [{ id: "b" }] });
});
it("late worker results invalidate a takeover menu but do not disable plan watching", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
let answer!: (choice: string) => void;
f.ctx.ui.select.mockImplementationOnce(() => new Promise(resolve => { answer = resolve; }));
const solo = f.command("solo");
f.hooks.get("tool_result")({ toolName: "subagent", details: { id: "late-child", sessionFile: "/tmp/late.jsonl" } });
answer("Worker confirmed stopped"); await solo;
expect(f.entries.at(-1).data.mode).toBe("supervising");
expect(f.entries.at(-1).data.workerStopped).toBe(false);
await f.atomicWrite(f.plan.replace("first output", "new requirement"));
await waitFor(() => f.changed() === 1);
});
it("changed plan or shutdown during takeover never grants solo permission", async () => {
const f = fixture(); await f.draft();
f.ctx.ui.select.mockImplementationOnce(async () => { writeFileSync(f.path, f.plan.replace("first", "changed")); return "Worker confirmed stopped"; });
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("planning");
expect(f.ctx.ui.notify).toHaveBeenLastCalledWith(expect.stringContaining("Plan changed during takeover"), "warning");
f.ctx.ui.select.mockImplementationOnce(async () => { f.shutdown(); return "Worker confirmed stopped"; });
await f.command("solo");
expect(f.entries.at(-1).data.mode).toBe("planning");
});
it("requires explicit supervisor ownership confirmation when attaching an existing plan", async () => {
const f = fixture(); const path = join(f.ctx.cwd, "shared.md");
writeFileSync(path, f.plan);
f.ctx.ui.select.mockResolvedValueOnce("Cancel");
await f.command(`attach ${path}`);
expect(f.entries).toHaveLength(0);
f.ctx.ui.select.mockResolvedValueOnce("Previous supervisor confirmed stopped");
await f.command(`attach ${path}`);
expect(f.entries.at(-1).data).toMatchObject({ mode: "planning", plan: path });
});
it("does not approve cancelled goals or display current completion for an unavailable plan", async () => {
const f = fixture(); await f.draft(); await f.command("ready");
writeFileSync(f.path, "- [-] goal: cancelled output\n## Log\n");
writeFileSync(join(f.ctx.cwd, "evidence.log"), "verified\n");
const reply = await f.tools.get("CompleteGoal").execute("t", { goal: "cancelled output", evidence: ["evidence.log"], observation: "read" }, undefined, undefined, f.ctx);
expect(reply.content[0].text).toContain("no sign-off recorded");
expect(readFileSync(f.path, "utf8")).toContain("[-]");
rmSync(f.path);
f.hooks.get("agent_end")({}, f.ctx);
expect(f.ctx.ui.setWidget).toHaveBeenLastCalledWith("goals", [expect.stringContaining("unavailable")]);
});
it("uses scheduler storage for ownership and the real public user controls", () => {
const prompt = scheduleCheckIn("session-1", "/plan.md");
expect(prompt).toContain(".pi/schedule-prompts.json");
expect(prompt).toContain("tool text does not expose binding");
expect(prompt).toContain("Never use cleanup");
expect(prompt).toContain("deletes disabled jobs");
expect(prompt).toContain("schedule_prompt update");
expect(prompt).not.toContain("with /schedule-prompts");
});
it("keeps interactive workers open and supplies the supervisor identity for Intercom reports", async () => {
const agent = readFileSync(new URL("../agents/goals-worker.md", import.meta.url), "utf8");
expect(agent).toContain("auto-exit: false");
const f = fixture(); await f.draft(); await f.command("ready");
expect(f.messages.at(-1).message.content).toContain("supervisor Intercom session copy-only");
const role = f.hooks.get("before_agent_start")({ systemPrompt: "base" }, f.ctx).systemPrompt;
expect(role).toContain("your Intercom session ID copy-only");
expect(role).toContain("stop workers before /reload");
expect(role).not.toContain("Reports arrive automatically");
});
+15
View File
@@ -0,0 +1,15 @@
import { existsSync, readFileSync } from "node:fs";
import { resolve } from "node:path";
import { expect, it } from "vitest";
it("declares current entry and bundled extension resources that exist after install", () => {
const manifest = JSON.parse(readFileSync("package.json", "utf8"));
expect(manifest.pi.extensions[0]).toBe("./src/index.ts");
for (const path of manifest.pi.extensions) expect(existsSync(resolve(path)), path).toBe(true);
for (const name of ["pi-subagents", "pi-intercom", "pi-schedule-prompt"]) {
expect(manifest.dependencies[name]).toBeTruthy();
expect(manifest.bundledDependencies).toContain(name);
}
expect(manifest.dependencies["pi-subagents"]).toContain("953c6f6d2fc7d8a5c956c30cd77c51bad697c2a4");
expect(existsSync("agents/goals-worker.md")).toBe(true);
});
-171
View File
@@ -1,171 +0,0 @@
import { describe, expect, it } from "vitest";
import { appendLog, counts, findGoal, parse, recordSignOff, setGoalStatus } from "../src/plan-file.js";
const SAMPLE = `# Plan: ship the cache layer
## Goal: Implement cache layer
<!-- id: cache-layer-1 -->
status: active
done_when: p95 < 50ms on bench-X. If wrong: timeouts in load-test.log
verify: pytest tests/cache -q
failure_modes:
- cache silently bypassed (hit-rate ~0, latency ok by luck)
- bench too small to exercise eviction
- [x] wire cache client
- [ ] eviction policy
- [ ] load test
## Goal: Document the API
<!-- id: document-the-api-1 -->
status: open
done_when: every public fn has a docstring; else sphinx warns
failure_modes:
- docstrings exist but are stale
## Log
- 2026-06-15 14:02 cache client wired; eviction next
`;
/** Multiset line diff: lines b adds vs removes vs a (order-insensitive, so insertions score added:1). */
function lineDelta(a: string, b: string): { added: number; removed: number } {
const count = (s: string) => {
const m = new Map<string, number>();
for (const l of s.split("\n")) m.set(l, (m.get(l) ?? 0) + 1);
return m;
};
const ma = count(a);
const mb = count(b);
let added = 0;
let removed = 0;
for (const k of new Set([...ma.keys(), ...mb.keys()])) {
const d = (mb.get(k) ?? 0) - (ma.get(k) ?? 0);
if (d > 0) added += d;
else if (d < 0) removed += -d;
}
return { added, removed };
}
describe("parse", () => {
const doc = parse(SAMPLE);
it("reads the objective and both goals", () => {
expect(doc.objective).toBe("ship the cache layer");
expect(doc.goals.map((g) => g.id)).toEqual(["cache-layer-1", "document-the-api-1"]);
});
it("reads goal fields", () => {
const g = findGoal(doc, "cache-layer-1");
expect(g?.subject).toBe("Implement cache layer");
expect(g?.status).toBe("active");
expect(g?.done_when).toBe("p95 < 50ms on bench-X. If wrong: timeouts in load-test.log");
expect(g?.verify).toBe("pytest tests/cache -q");
});
it("separates failure_modes from subtasks", () => {
const g = findGoal(doc, "cache-layer-1");
expect(g?.failure_modes).toHaveLength(2);
expect(g?.failure_modes[0]).toContain("cache silently bypassed");
expect(g?.subtasks).toEqual([
{ text: "wire cache client", done: true },
{ text: "eviction policy", done: false },
{ text: "load test", done: false },
]);
});
it("reads the log verbatim and counts by status", () => {
expect(doc.log).toEqual(["- 2026-06-15 14:02 cache client wired; eviction next"]);
expect(counts(doc)).toEqual({ done: 0, open: 1, active: 1 });
});
});
describe("failure_modes vs subtask disambiguation", () => {
it("a column-0 checkbox right after failure_modes: is a SUBTASK", () => {
const doc = parse(
`# Plan: x\n\n## Goal: G\n<!-- id: g-1 -->\nstatus: open\ndone_when: z\nfailure_modes:\n- [ ] first subtask\n- [x] second subtask\n`,
);
const g = findGoal(doc, "g-1");
expect(g?.failure_modes).toEqual([]);
expect(g?.subtasks).toEqual([
{ text: "first subtask", done: false },
{ text: "second subtask", done: true },
]);
});
it("an indented checkbox-shaped item inside failure_modes is a FAILURE MODE", () => {
const doc = parse(
`# Plan: x\n\n## Goal: G\n<!-- id: g-2 -->\nstatus: open\ndone_when: z\nfailure_modes:\n - [ ] prose that looks like a checkbox\n- [ ] real subtask\n`,
);
const g = findGoal(doc, "g-2");
expect(g?.failure_modes).toEqual(["[ ] prose that looks like a checkbox"]);
expect(g?.subtasks).toEqual([{ text: "real subtask", done: false }]);
});
it("a goal with no failure_modes keeps its subtasks", () => {
const doc = parse(`# Plan: x\n\n## Goal: G\n<!-- id: g-3 -->\nstatus: open\ndone_when: z\n- [ ] only subtask\n`);
const g = findGoal(doc, "g-3");
expect(g?.failure_modes).toEqual([]);
expect(g?.subtasks).toEqual([{ text: "only subtask", done: false }]);
});
});
describe("the two CompleteGoal writes (minimal diff)", () => {
it("setGoalStatus replaces exactly one line, scoped to the right goal", () => {
const next = setGoalStatus(SAMPLE, "cache-layer-1", "done");
expect(lineDelta(SAMPLE, next)).toEqual({ added: 1, removed: 1 });
expect(findGoal(parse(next), "cache-layer-1")?.status).toBe("done");
expect(findGoal(parse(next), "document-the-api-1")?.status).toBe("open"); // untouched
});
it("setGoalStatus targets the second goal without touching the first", () => {
const next = setGoalStatus(SAMPLE, "document-the-api-1", "active");
expect(findGoal(parse(next), "cache-layer-1")?.status).toBe("active");
expect(findGoal(parse(next), "document-the-api-1")?.status).toBe("active");
});
it("appendLog adds exactly one line under ## Log", () => {
const next = appendLog(SAMPLE, "2026-06-15 15:00 eviction done");
expect(lineDelta(SAMPLE, next)).toEqual({ added: 1, removed: 0 });
expect(parse(next).log).toEqual([
"- 2026-06-15 14:02 cache client wired; eviction next",
"- 2026-06-15 15:00 eviction done",
]);
});
it("appendLog creates the section when absent", () => {
const noLog = "# Plan: x\n\n## Goal: y\n<!-- id: y-1 -->\nstatus: open\ndone_when: z\n";
expect(parse(appendLog(noLog, "first entry")).log).toEqual(["- first entry"]);
});
});
describe("recordSignOff (CompleteGoal's pure record logic)", () => {
const WHEN = "2026-06-15 16:00";
it("accept flips status:done and logs a sign-off line", () => {
const r = recordSignOff(SAMPLE, "cache-layer-1", WHEN, { kind: "accepted" });
expect(r.isError).toBe(false);
const doc = parse(r.content);
expect(findGoal(doc, "cache-layer-1")?.status).toBe("done");
expect(doc.log.at(-1)).toBe(`- ${WHEN} signed off #cache-layer-1: Implement cache layer (oracle accept)`);
});
it("verify_failed only logs a reject line, status stays active", () => {
const r = recordSignOff(SAMPLE, "cache-layer-1", WHEN, { kind: "verify_failed", exitCode: 1, outputTail: "boom" });
expect(r.isError).toBe(true);
const doc = parse(r.content);
expect(findGoal(doc, "cache-layer-1")?.status).toBe("active"); // NOT marked done
expect(doc.log.at(-1)).toBe(`- ${WHEN} reject #cache-layer-1: verify exit 1`);
});
it("rejected logs the (one-lined) missing reason, status stays", () => {
const r = recordSignOff(SAMPLE, "cache-layer-1", WHEN, { kind: "rejected", missing: "no\nsaved\nbench log" });
expect(r.isError).toBe(true);
expect(findGoal(parse(r.content), "cache-layer-1")?.status).toBe("active");
expect(parse(r.content).log.at(-1)).toBe(`- ${WHEN} reject #cache-layer-1: no saved bench log`);
});
it("unknown goal returns an error and does not touch the file", () => {
const r = recordSignOff(SAMPLE, "nope-1", WHEN, { kind: "accepted" });
expect(r.isError).toBe(true);
expect(r.content).toBe(SAMPLE);
});
});
+40
View File
@@ -0,0 +1,40 @@
import { expect, it } from "vitest";
import { planViews } from "../src/plan-view.js";
it("keeps outcome, preferences and discriminators without tasks or history", () => {
const plan = "# Outcome\nBeat random, not just plot it.\n## User preferences\nKeep costs low.\n## Goals\n1. [ ] goal: repair\n - discriminator: beats random\n - subtle failure mode: plot exists but result fails\n - tasks:\n 1. [x] draw plot\n - evidence:\n - old output\n2. [ ] goal: confirm\n## Task list\n- [ ] run it\n## Appendix\nunapproved idea";
const views = planViews(plan);
for (const text of ["Beat random", "Keep costs low", "goal: repair", "discriminator: beats random", "subtle failure mode", "goal: confirm"]) expect(views.short).toContain(text);
for (const text of ["draw plot", "old output", "run it", "unapproved idea"]) expect(views.short).not.toContain(text);
expect(views.long).toContain("draw plot");
expect(views.long).toContain("old output");
expect(views.long).not.toContain("unapproved idea");
});
it("omits only named worker identity fields from review while retaining them in full context", () => {
const base = "# Plan\n- preferred worker model: provider/model\n- [ ] goal: result\n - discriminator: exact bytes";
const metadata = "\n- Active worker: worker-1\n- worker session: /saved.jsonl\n- worker intercom session: uuid";
expect(planViews(base + metadata).short).toBe(planViews(base).short);
expect(planViews(base + metadata).long).toContain("/saved.jsonl");
expect(planViews(base.replace("exact bytes", "approximate match")).short).not.toBe(planViews(base).short);
expect(planViews(base.replace("[ ]", "[x]")).short).not.toBe(planViews(base).short);
});
it("notifies on goal and task changes but not on identity bookkeeping or log edits", () => {
const base = "# Plan\n- [ ] goal: result\n## Task list\n- [ ] run it\n- worker session: /saved.jsonl\n## Log\nfirst entry";
const baseView = planViews(base).notify;
// identity bookkeeping: silent
expect(planViews(base.replace("/saved.jsonl", "/moved.jsonl")).notify).toBe(baseView);
// log edits: silent
expect(planViews(base.replace("first entry", "second entry")).notify).toBe(baseView);
// worker ticking a task: review event (field catch, LUCID3 2026-09-10)
expect(planViews(base.replace("- [ ] run it", "- [x] run it")).notify).not.toBe(baseView);
// goal edits: review event
expect(planViews(base.replace("[ ] goal: result", "[x] goal: result")).notify).not.toBe(baseView);
});
it("stops at history and preserves a manual goal tick", () => {
const view = planViews("# Plan\n1. [x] goal: result\n## Log\n1. [ ] goal: historical");
expect(view.short).toContain("[x] goal: result");
expect(view.long).not.toContain("historical");
});
+21
View File
@@ -0,0 +1,21 @@
import { describe, expect, it } from "vitest";
import { planDrafting } from "../src/prompts.js";
describe("planning prompt", () => {
it("requires fact finding or a focused question before a goal", () => {
expect(planDrafting).toContain("Use read-only repository tools or web search when either can\nresolve a fact.");
expect(planDrafting).toContain("Do not use a question quota");
expect(planDrafting).toContain("Briefly reframe the request in your own words to check comprehension");
expect(planDrafting).toContain("point as unknown; do not silently replace it with an inference or turn it into a new blocking decision");
expect(planDrafting).toContain("answer materially reduces uncertainty\nwhile discovering the right plan");
expect(planDrafting).toContain("self-contained: state the relevant\ncontext, use the human's language and ASD-STE100");
expect(planDrafting).toContain("placeholder goal such as \"work out the thing\"");
expect(planDrafting).toContain("Only withhold Ready for an unanswered choice that changes scope, spending, or the user-visible result");
});
it("anchors work and sign-off to the user-visible result", () => {
expect(planDrafting).toContain("## User-visible result");
expect(planDrafting).toContain("Take it from the original request, not from your implementation plan");
expect(planDrafting).toContain("Future work may not defer any artifact or action named there");
});
});
+166
View File
@@ -0,0 +1,166 @@
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
import { once } from "node:events";
import { mkdtempSync, readFileSync, rmSync } from "node:fs";
import { createServer } from "node:http";
import { tmpdir } from "node:os";
import { join, resolve } from "node:path";
import { StringDecoder } from "node:string_decoder";
import { describe, expect, it } from "vitest";
type RpcMessage = { type: string; id?: string; method?: string; [key: string]: unknown };
type ModelRequest = { messages: Array<{ role: string; content: unknown }> };
class RpcClient {
readonly messages: RpcMessage[] = [];
stderr = "";
private readonly waiters: Array<{ predicate: (message: RpcMessage) => boolean; resolve: (message: RpcMessage) => void }> = [];
constructor(readonly process: ChildProcessWithoutNullStreams) {
const decoder = new StringDecoder("utf8");
let buffer = "";
process.stderr.on("data", (chunk) => { this.stderr += chunk; });
process.stdout.on("data", (chunk) => {
buffer += decoder.write(chunk);
while (buffer.includes("\n")) {
const newline = buffer.indexOf("\n");
const line = buffer.slice(0, newline).replace(/\r$/, "");
buffer = buffer.slice(newline + 1);
if (!line) continue;
const message = JSON.parse(line) as RpcMessage;
this.messages.push(message);
const index = this.waiters.findIndex(({ predicate }) => predicate(message));
if (index !== -1) this.waiters.splice(index, 1)[0].resolve(message);
}
});
}
send(message: RpcMessage): void {
this.process.stdin.write(`${JSON.stringify(message)}\n`);
}
waitFor(predicate: (message: RpcMessage) => boolean, after = 0): Promise<RpcMessage> {
const existing = this.messages.slice(after).find(predicate);
if (existing) return Promise.resolve(existing);
return new Promise((resolvePromise, reject) => {
const timer = setTimeout(() => {
this.waiters.splice(this.waiters.indexOf(waiter), 1);
reject(new Error(`RPC wait timed out: ${this.stderr}\n${JSON.stringify(this.messages.slice(-12))}`));
}, 8_000);
const waiter = { predicate, resolve: (message: RpcMessage) => { clearTimeout(timer); resolvePromise(message); } };
this.waiters.push(waiter);
});
}
}
function streamResponse(response: import("node:http").ServerResponse, delta: object, finishReason: "stop" | "tool_calls"): void {
response.writeHead(200, { "content-type": "text/event-stream" });
response.write(`data: ${JSON.stringify({ choices: [{ index: 0, delta, finish_reason: null }] })}\n\n`);
response.write(`data: ${JSON.stringify({ choices: [{ index: 0, delta: {}, finish_reason: finishReason }] })}\n\n`);
response.end("data: [DONE]\n\n");
}
const isSelect = (message: RpcMessage) => message.type === "extension_ui_request" && message.method === "select";
const isEditor = (message: RpcMessage) => message.type === "extension_ui_request" && message.method === "editor";
const systemText = (request: ModelRequest) => request.messages.filter(message => ["system", "developer"].includes(message.role)).map(message => message.content).join("\n");
describe("RPC review flow", () => {
it.each(["Edit", "Discuss"])("automatically proposes a draft, handles %s, then enters the supervisor role on Ready", async (choice) => {
const cwd = mkdtempSync(join(tmpdir(), "pi-goals-rpc-"));
const requests: ModelRequest[] = [];
const plan = "# Plan\n\n## Goals\n\n1. [ ] goal: name the output\n - subtle failure mode: the output has no name\n - discriminator: the plan names the output\n\n## Log\n";
let planPath = "";
const server = createServer(async (request, response) => {
let body = "";
for await (const chunk of request) body += chunk;
requests.push(JSON.parse(body));
if (requests.length === 1) {
streamResponse(response, {
tool_calls: [{
index: 0, id: "write-plan", type: "function",
function: { name: "write", arguments: JSON.stringify({ path: planPath, content: plan }) },
}],
}, "tool_calls");
return;
}
streamResponse(response, { content: "Plan inspected." }, "stop");
});
await new Promise<void>((done) => server.listen(0, "127.0.0.1", done));
const address = server.address();
if (!address || typeof address === "string") throw new Error("Offline model did not bind a TCP port.");
const pi = spawn(resolve("node_modules/.bin/pi"), [
"--mode", "rpc", "--no-session", "--no-extensions", "--model", "offline/test",
"-e", resolve("test/fixtures/offline-model.ts"),
"-e", resolve("test/fixtures/subagent-schema.ts"),
"-e", resolve("src/index.ts"),
], {
cwd,
env: {
// Pi/gpt-6-astra: test the parent role even when vitest itself runs in a worker.
...Object.fromEntries(Object.entries(process.env).filter(([name]) => !name.startsWith("PI_SUBAGENT_") && !name.startsWith("PI_GOALS_"))),
PI_CODING_AGENT_DIR: join(cwd, ".agent"),
PI_GOALS_OFFLINE_MODEL_URL: `http://127.0.0.1:${address.port}`,
},
});
const client = new RpcClient(pi);
const exited = once(pi, "exit");
try {
client.send({ type: "get_state", id: "state" });
const state = await client.waitFor((message) => message.type === "response" && message.id === "state");
const sessionId = (state.data as { sessionId: string }).sessionId;
planPath = join(cwd, ".pi", "plan", `${sessionId}-main.md`);
client.send({ type: "prompt", id: "goals", message: "/goals new work out the thing" });
const review = await client.waitFor(isSelect);
expect(review.options).toEqual(["Ready", "Discuss", "Edit", "Cancel"]);
expect(review.title).toContain(planPath);
const proposal = client.messages.find(message => message.type === "message_end" && (message.message as { customType?: string })?.customType === "goal-plan-proposal");
expect(proposal?.message).toMatchObject({ content: plan, display: true });
expect(readFileSync(planPath, "utf8")).toBe(plan);
expect(requests).toHaveLength(2);
expect(systemText(requests[0])).toContain("Plan only in");
const choiceStart = client.messages.length;
client.send({ type: "extension_ui_response", id: review.id, value: choice });
let approvedPlan = plan;
if (choice === "Edit") {
const editor = await client.waitFor(isEditor, choiceStart);
expect(editor.prefill).toBe(plan);
expect(requests).toHaveLength(2);
approvedPlan = plan.replace("the plan names the output", "the plan names output.txt and its exact bytes");
const editStart = client.messages.length;
client.send({ type: "extension_ui_response", id: editor.id, value: approvedPlan });
await client.waitFor(message => message.type === "extension_ui_request" && message.method === "setWidget", editStart);
expect(readFileSync(planPath, "utf8")).toBe(approvedPlan);
expect(requests).toHaveLength(2);
} else {
await client.waitFor(message => message.type === "agent_end", choiceStart);
expect(requests).toHaveLength(3);
expect(systemText(requests[2])).toContain("Plan only in");
expect(JSON.stringify(requests[2].messages.at(-1))).toContain("Discuss the current draft");
expect(client.messages.slice(choiceStart).filter(isEditor)).toEqual([]);
}
const beforeReady = requests.length;
const reopenStart = client.messages.length;
client.send({ type: "prompt", id: "review", message: "/goals review" });
const ready = await client.waitFor(isSelect, reopenStart);
expect(requests).toHaveLength(beforeReady);
const readyStart = client.messages.length;
client.send({ type: "extension_ui_response", id: ready.id, value: "Ready" });
await client.waitFor(message => message.type === "agent_end", readyStart);
expect(requests).toHaveLength(beforeReady + 1);
const supervisor = requests.at(-1)!;
expect(systemText(supervisor)).toContain("You are the goal supervisor in the main chat");
expect(systemText(supervisor)).not.toContain("Plan only in");
expect(JSON.stringify(supervisor.messages)).toContain(JSON.stringify(approvedPlan).slice(1, -1));
expect(client.messages.filter(message => message.type === "tool_execution_start").map(message => message.toolName)).toEqual(["write"]);
expect(client.messages.filter(message => message.type === "extension_error")).toEqual([]);
console.log(`RPC ${choice}: visible automatic proposal; ${choice === "Edit" ? "editor saved exact plan without model call" : "discussion retained planning role without editor"}; Ready request used supervisor role; only write executed.`);
} finally {
pi.kill();
await exited;
await new Promise<void>((done) => server.close(() => done()));
rmSync(cwd, { recursive: true, force: true });
}
}, 25_000);
});
+34
View File
@@ -0,0 +1,34 @@
import { expect, it } from "vitest";
import { report, summarize } from "../scripts/session-usage.mjs";
const start = "2026-09-10T06:00:00.000Z";
const end = "2026-09-10T07:00:00.000Z";
const request = (timestamp: string, model = "a") => ({
type: "message", timestamp, message: { role: "assistant", provider: "test", model,
usage: { input: 10, cacheRead: 100, cacheWrite: 5, output: 20, reasoning: 8, totalTokens: 135 } },
});
it("excludes inherited history and counts repeated cached input without adding reasoning twice", () => {
const result = summarize([request("2026-09-09T06:00:00.000Z"), request(start), request(end, "b")], start, end);
expect(result).toMatchObject({ calls: 2, input: 20, cacheRead: 200, cacheWrite: 10, output: 40, totalTokens: 270 });
expect(result.models.map((m: any) => m.model)).toEqual(["test/a", "test/b"]);
});
it("uses the latest planning start and the same interval for both sessions", () => {
const supervisor = { entries: [
{ type: "custom", customType: "pi-goals-main-supervisor-v1", timestamp: start, id: "boundary", data: { mode: "planning", plan: "plan.md" } },
request(start),
] };
const result = report(supervisor, { entries: [request(end)] }, end);
expect(result.since).toBe(start);
expect(result.elapsedHours).toBe(1);
expect(result.sessions.map((s: any) => s.output)).toEqual([20, 20]);
expect(() => report({ entries: [] }, { entries: [] }, end)).toThrow("No recorded planning start");
});
it("reports missing usage and rejects invalid recorded token counts", () => {
const missing = { type: "message", timestamp: start, message: { role: "assistant" } };
expect(summarize([missing], start, end).missingUsage).toBe(1);
const invalid = request(start); invalid.message.usage.input = Number.NaN;
expect(() => summarize([invalid], start, end)).toThrow("Invalid usage.input");
});
+4 -2
View File
@@ -6,10 +6,12 @@
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"jsx": "react-jsx",
"outDir": "dist",
"rootDir": "src",
"declaration": true
},
"include": ["src/**/*.ts", "src/**/*.tsx"]
"include": [
"src/**/*.ts",
"src/**/*.tsx"
]
}