diff --git a/AGENTS.md b/AGENTS.md index 50c2f16..c57972d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -28,18 +28,22 @@ The following preferences are the user's words, recorded on ## Confirmed product preferences - One installable pi-goals package, not separately configured supervisor packages. -- Run the same Pi profile and package set in two real interactive sessions, visible beside each other in Herdr. Supervisor mode is a role, not a separately assembled profile or installation. Fork the main planning session, activate supervisor mode and compact the fork. +- Run the full normal Pi profile and package set in two real interactive sessions, visible beside each other in Herdr. Keep extensions, skills, prompt templates, themes, configuration and authentication. Do not silently launch a reduced profile with `--no-extensions`; honor deliberate worker resource choices. Supervisor mode changes the role/model and enforces its inspection-only policy, not a separately assembled installation. Fork the main planning session, activate supervisor mode and compact the fork. - The supervisor retains the compacted planning session. Its repeated review loop reminds it that it is the supervisor and supplies the current canonical plan. Keep those directly available rather than relying on the compaction summary alone. - Load Intercom once per Pi process. Do not add a second copy or extra standalone supervisor package when activating supervisor mode. - The worker carries implementation detail. The supervisor gets incremental high-level views and retains its judgments, user intent and decisions. - Remember the last model selected separately for planning, working and supervising. -- Default to at least three task-specific alignment questions before the final plan. An explicit current-plan request to skip questions waives that round, not future plans. +- Ask material unresolved alignment questions, not a fixed quota. Inspect technical facts yourself; do not ask for confirmation of ordinary implementation details or repeat answered questions. Batch high-impact questions with context and a recommendation. An explicit current-plan request to skip optional questions applies only to that plan; it does not grant missing permission. - Discuss returns the review menu to normal chat, preserving the draft. Do not immediately reopen the menu while the conversation is unfinished. Ready is the human's approval to start work. - Live supervision must not use headless Pi RPC mode or a bespoke plan-lifecycle RPC layer. Use the real sessions and Intercom messages for views, steering and plan-bound approval checkpoints. - Keep the design simple and reliable. The user reports that it is constantly breaking; adding more orchestration or approval forms is not progress. Preserve working components and remove unnecessary layers. - Show useful supervisor assessments, advice and perspective, not only hidden tool arguments or delivery receipts. Keep the assessment brief. The supervisor's job is judgment and helping the worker stay on course, not filling forms; transport and approval bookkeeping are supporting details. - After the initial fork compaction, compact the supervisor again above 100k current-context tokens (not cumulative usage), respecting the model's context limit. This is the latest user clarification of the earlier approximate 150k preference. Token/cost savings and the usefulness of advice need a real task trial; passing protocol tests alone does not establish either. +- Supervise autonomously until the agreed result is achieved and inspected. Investigate claims of being blocked, waiting, unable to proceed, or done; change ineffective steering, and keep authorized independent work moving. Respect genuine dependencies, explicit human pauses, scope and permission limits. Do not make the human drive routine progress. +- Keep supervision instructions generic and outcome-focused. Approval bookkeeping supports delivery; it is not the deliverable. Inspect actual artifacts and execution evidence, not just summaries, checked boxes, or test counts. Manual completion checkboxes are claims until CompleteGoal records sign-off. Accepted inconclusive remains explicitly uncertain under fail-forward policy. +- Git status is a review guideline, not a hard acceptance gate. Unrelated dirty files and ignored output directories can be legitimate. Do not force cleanup, commits, or a dirty-state fingerprint framework. + These are user preferences, not a claim that the current implementation satisfies every point. Validate them in real panes as well as automated tests; record remaining gaps and actual per-role usage. @@ -47,6 +51,8 @@ Validate them in real panes as well as automated tests; record remaining gaps an Run `npm test` before a commit. It includes unit and flow tests plus the RPC review test. +For an explicitly approved trusted-package update, a command-scoped npm release-age exception is allowed. Keep the default policy intact; do not turn a targeted update into a general package upgrade. The user approved pi-subagents 0.66.0 on 2026-09-09. + - `test/*.test.ts` unit and flow tests use a small Pi API mock. They check plan state, tool gates, and plan-file updates. - `npm run test:rpc` runs `test/rpc-review.test.ts`. It starts the installed Pi executable in RPC mode, uses Pi's real `select` protocol and conversational Discuss flow, and uses a local deterministic HTTP model. It does not need a credential or spend API credits. This is the closest automated session test. - Use tmux for visual TUI debugging when the RPC test fails or a terminal-only problem is reported: @@ -59,3 +65,17 @@ Run `npm test` before a commit. It includes unit and flow tests plus the RPC rev - `pi -p` has no UI, so it cannot test `Ready`, `Discuss`, `Edit`, or `Cancel`. - `npm run test:supervisor` runs the inherited `node:test` supervisor regressions. `npm test` also includes the always-enabled packed-artifact Intercom flow; Linux requires Unix-socket support. + +## Functional acceptance: real isolated Herdr workflow + +Pi/OpenAI procedure, requested by wassname; adapted from `8953dce`. Automated tests do not replace this check. + +1. Read `herdr --skill` and confirm `HERDR_ENV=1`. Create a separate test pane with `--no-focus` and an isolated temporary Git repo. Never operate the user's existing worker or supervisor panes. Record the code revision and uncommitted changes being tested. Use a packed test package without replacing the active installation. +2. Start real interactive Pi with that package and available, different worker and supervisor models. Test the full normal profile on both sides, not two equally stripped profiles. Isolate only the candidate pi-goals package selection; do not load both old and candidate copies. Keep global settings untouched; if temporary non-secret role preferences must change, save and restore all three with guarded cleanup. +3. Use `/goals` for a trivial, bounded two-goal task: two exact-content files plus saved byte-verification output, with an ignored output directory and an unrelated dirty file to preserve. No GPU, dependencies or unrelated work. Read the planning conversation, verify that questions are material, inspect the draft, exercise ordinary-chat Discuss, and select Ready through the actual UI. +4. Confirm Ready opens a visible supervisor pane and the worker starts. Read both panes. Verify exact supervisor advice is visible, reaches the worker, and helps progress. Delivery receipts alone are not proof. Let the same pair stay active between both goals. +5. Let the worker produce artifacts and real verification output, then complete both CompleteGoal calls (retained supervisor `review_goal`, followed by the fresh evidence judge). Do not perform the worker's task. Record each manual nudge or repair as an intervention, not autonomous success. Conclusive and accepted-inconclusive results are not equivalent. +6. Inspect the actual artifacts and saved execution evidence, final plan, and both sessions. Handwritten output or a manual tick does not prove execution. Success means the requested results and both observed sign-offs, not tests passing or messages exchanged. Verify no commit/cleanup was forced for ignored outputs or unrelated dirt. +7. Exercise worker-only, supervisor-only and both-side reloads; fresh-shell resume without special launcher environment; drafting/Discuss; Ready/startup compaction; pending completion; and a stopped pairing. Preserve the plan, role, restrictions and peer identity. Interrupted decisions must fail visibly and allow retry, not approve stale work or create duplicate panes. Planning reload must not become approval or get permanently stuck. Record any recovery action needed and commands actually available. +8. If a stage fails, read both panes and the exact error before diagnosing it. Fix the cause, reload only the test instance, and retry that stage. After a prompt change use a fresh task. Repeated status checks are not a repair; wait-output timeouts/matches only signal that the pane needs inspection. +9. Save pane captures, session/artifact/log paths, revision, interventions, remaining failures, and separate role usage under `docs/slop/reviews/`. Assess usefulness and actual token/cost use, not a test-count substitute. Only close panes you created. Do not release, merge, or replace the user's installation as part of acceptance. diff --git a/README.md b/README.md index f83239b..d5713f8 100644 --- a/README.md +++ b/README.md @@ -67,10 +67,11 @@ pi -e . `/goals` enters plan mode and starts a conversation; the objective is an optional seed. From there: -1. Align. The agent inspects technical facts read-only, then asks at least three task-specific - questions in one chat round about the expected result, scope/constraints, and success/failure - criteria. It waits for your answers before proposing the final plan. An explicit “no questions” - or “skip questions” clause in the current objective waives this round for that plan only. +1. Align. The agent inspects technical facts read-only, then asks only material unresolved + questions about outcome, scope, constraints, or success criteria. There is no fixed quota or + confirmation ritual for ordinary implementation details. It waits for required answers before + proposing the final plan. An explicit “no questions” or “skip questions” clause waives optional + questions for that plan only, not missing permissions. “No q's” and “skip q's” are also supported. Negated instructions (“do not skip questions”) and quoted feature references (“add a 'skip questions' button”) do not waive alignment. 2. Review. When alignment is complete, the agent requests review and the full draft is printed. @@ -138,8 +139,24 @@ calls, with no plan-lifecycle RPC dispatcher or headless live Pi process. Discon pending approval and is shown explicitly; a send does not prove receipt or execution. One `CompleteGoal` call asks this supervisor about direction and scope, then runs the normal fresh -read-only evidence judge. A prematurely checked submitted goal is reopened before review; only accepted -sign-off checks it again. The judge's checks section accepts ordinary numbered and indented Markdown lists, +read-only evidence judge. Use one unique exact goal subject (case and surrounding whitespace do not +matter); ambiguous or drifted wording gets an actionable retry, not a manual-tick fallback. +Manual `[x]` marks are visible completion claims, not sign-off, even before this tool is called or +after reload. A prematurely checked submitted goal is reopened before review. Only accepted sign-off +checks it again and persists a per-goal record; observed reopening invalidates that record. The widget +and supervisor distinguish conclusive acceptance from **accepted inconclusive** (judge failure or no +verdict). Inconclusive still permits fail-forward, but is not verified completion. Git status is context, +not a gate: the judge can inspect cited uncommitted and ignored files directly. No commit or clean +worktree is required unless the goal itself requires it. + +Older sessions have no trusted per-goal records. Their existing checkboxes/evidence/logs are preserved +as “legacy completion — sign-off not recorded,” not rejected or automatically reimplemented. Use normal +CompleteGoal re-review if needed; editable historical log text is not imported as trusted sign-off. +Stopped pairings remain stopped. New worker views include current completion claims and whether the +canonical plan changed; a manual tick cannot end supervision. Ordinary supervisor prose and genuine +questions no longer suppress later worker direction. Explicit human pauses remain instructions to +respect, not a reason to discard new views; idle responses do not immediately retry themselves. + The judge's checks section accepts ordinary numbered and indented Markdown lists, but an empty section cannot borrow a list from a later heading. Approving one goal does not finish supervision. Cancelled, stale or mismatched replies do not sign off goals. Goal/revision identity is bound in code to the checkpoint actually presented to the supervisor, not copied into a form by the model. Supervisor model checkpoints diff --git a/docs/reviews/2026-09-09_autonomy-validation.md b/docs/reviews/2026-09-09_autonomy-validation.md new file mode 100644 index 0000000..fa85b95 --- /dev/null +++ b/docs/reviews/2026-09-09_autonomy-validation.md @@ -0,0 +1,30 @@ +# Autonomous supervision: implementation checked, live acceptance pending + +Base: `cecb1e9`, branch `feature/simple-visible-supervision`. Follow-up changes are uncommitted. +Scope: [approved plan](../slop/plans/20260909_autonomous-supervision-acceptance.md). + +## Implemented + +- Outcome-focused supervisor instructions: investigate blockers, change ineffective steering, inspect actual results, and keep authorized work moving. VCC and existing lifecycle protections remain. +- Ordinary prose, empty responses and genuine questions do not discard later worker views or direction. Monitoring does not authorize restarting human-paused work. No immediate idle retry loop. +- Manual checkmarks are claims; CompleteGoal records conclusive or inconclusive sign-off separately. Exact goal identity, cancellation and fresh-judge behavior remain. Legacy completion is labelled without inventing approval or restarting old work. +- Dirty Git state is context, not an acceptance gate. Cited ignored output files are valid inspection targets; no forced cleanup or commit. +- Material planning questions replace the quota; Discuss remains ordinary chat. AGENTS.md includes the other branch's relevant user preferences and real-Herdr testing procedure. + +## Parent validation + +- [Permitted Vitest subset](evidence/2026-09-09-autonomy/parent-permitted-vitest.log): `Tests 94 passed (94)`, across 13 files. Explicitly excludes `test/rpc-supervisor.test.ts`; this is not a passing full suite. +- [Supervisor regressions](evidence/2026-09-09-autonomy/parent-supervisor.log): `ℹ tests 176`, `ℹ pass 176`, `ℹ fail 0`. +- [Typecheck](evidence/2026-09-09-autonomy/parent-typecheck.log), [lint](evidence/2026-09-09-autonomy/parent-lint.log), and [build](evidence/2026-09-09-autonomy/parent-build.log) exited successfully. `git diff --check` passed on source changes. +- Review found conflicting advice to prune completed goal lines and stale fuzzy-match descriptions. Parent corrected the instructions and added regressions. [Red](evidence/2026-09-09-autonomy/housekeeping-red.log) shows the two prompt failures; [green](evidence/2026-09-09-autonomy/housekeeping-green.log) records 43 passing tests, including preserved conclusive/inconclusive records after moving detail into the appendix. +- [Read-only recheck](evidence/2026-09-09-autonomy/housekeeping-recheck.md): “Both previous findings are resolved; the narrow fixes are approved.” + +## Still required + +Full `npm test` did not pass: [broker diagnostics](evidence/2026-09-09-autonomy/broker-diagnostic.log) show Unix-socket `listen EPERM` in the tsx launcher. Parent Herdr control independently returned `PermissionDenied: Operation not permitted`. No TMPDIR/IPC workaround was authorized or used to bypass the restriction. + +The fresh two-goal Herdr trial has not run. It must show both actual artifacts and verification output, both CompleteGoal results, useful visible supervision, same-pair continuity, explicit-pause/reload behavior, ignored output files and preserved unrelated dirty work. Record every operator intervention and separate worker/supervisor usage. Do not treat deterministic tests as evidence of live judgment or savings. + +No role preferences, active installation, existing panes, or unrelated root-worktree files were changed in this follow-up. No commit, push, merge or release yet. + +Recorded by Pi (OpenAI) from observed command output and the independent source review. diff --git a/docs/reviews/2026-09-09_recovery-commands.md b/docs/reviews/2026-09-09_recovery-commands.md new file mode 100644 index 0000000..c896543 --- /dev/null +++ b/docs/reviews/2026-09-09_recovery-commands.md @@ -0,0 +1,44 @@ +# Recovery commands + +Implemented directly by Pi/OpenAI at the user's request, on `feature/simple-visible-supervision`, base `cecb1e9` plus the existing uncommitted autonomy changes. The active global installation was not replaced. + +## Commands + +- `/goals help`: available commands and limits. +- `/goals status`: phase, peer connectivity, recorded panes/session files and last runtime failure. +- `/goals stop`: stop goal continuation and supervision; retain the plan and pair. +- `/goals exit`: stop and return to ordinary chat, retaining files. Supervisor sessions remain inspection-only. +- `/goals reconnect`: send the existing pair's identity handshake. Does not fork or authorize work. +- `/goals resume`: in the worker, resume previously authorized work after peer acknowledgement. A stopped draft or changed plan returns to planning and still needs Ready. In the supervisor, directs the human to the worker for authorization. +- `/goals supervisor`, `/goals worker`, `/goals zoom`: existing pane navigation. + +Stop/exit cancel startup, an outstanding Ready selection and sign-off. They persist across reload. They do not kill independently running processes. Peer notification is best-effort and explicitly unconfirmed; use the other pane's stop command if it is disconnected. A pause identity prevents an old resume request from undoing a newer stop. Permanently ended pairings are not revived by reconnect. + +## Observed interactive behaviour + +Used real Pi 0.85.1 in a dedicated Herdr pane, with the normal global extensions, skills, prompts and themes. Project settings replaced only the old pi-goals package selection with the candidate. No global settings or package installation changed. The pane was closed after the check. + +This was an **operator-seeded unapproved draft**, not an autonomous task or a model-produced plan. There were no model responses beyond the explicitly labelled fixture marker. The purpose was to exercise the public commands and persisted state in real Pi. + +Saved [verification output](evidence/2026-09-09-recovery/verification.log) reports: + +> PASS: stop persisted paused draft. +> PASS: resume restored planning without Ready authorization. +> PASS: exit persisted ordinary-chat state. +> PASS: no assistant turn beyond operator fixture marker. +> PASS: draft bytes unchanged. +> PASS: actual reload rendered; stopped status retained. +> PASS: reconnect without a pair reports failure instead of launching one. +> PASS: fresh-shell --session retained exit; status reports ordinary chat. + +The [reload capture](evidence/2026-09-09-recovery/reload-pane.txt) shows the actual reload notice and stopped widget. The [resume capture](evidence/2026-09-09-recovery/resume-pane.txt) says “Draft restored; no work started.” The [fresh-shell capture](evidence/2026-09-09-recovery/fresh-status-pane.txt) says “Goals: ordinary chat (goals exited).” These establish command behaviour in the interactive runtime, not supervisor judgment. + +Local fixture and detailed command receipts: `/tmp/pi-goals-recovery-functional/`. Source diff: `/tmp/pi-goals-recovery-implemented.diff`. + +## Automated checks + +The final `npm test`, typecheck, lint, build and `git diff --check` completed successfully. Saved [npm test output](evidence/2026-09-09-recovery/npm-test.log), [typecheck](evidence/2026-09-09-recovery/typecheck.log), [lint](evidence/2026-09-09-recovery/lint.log) and [build](evidence/2026-09-09-recovery/build.log) are supporting checks, not substitutes for paired functional acceptance. + +## Remaining acceptance + +Real paired worker/supervisor recovery during bootstrap compaction and sign-off, and the autonomous two-goal trial, remain pending. Paired transport regressions use an in-process broker harness; they do not replace those checks. The encrypted-compaction replay mismatch is a separate unresolved issue; this change does not disable its guard or claim to fix it. diff --git a/docs/reviews/evidence/2026-09-09-autonomy/broker-diagnostic.log b/docs/reviews/evidence/2026-09-09-autonomy/broker-diagnostic.log new file mode 100644 index 0000000..1f10c20 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/broker-diagnostic.log @@ -0,0 +1,20 @@ +node:net:1986 + const error = new UVExceptionWithHostPort(rval, 'listen', address, port); + ^ + +Error: listen EPERM: operation not permitted /tmp/claude/tsx-1000/28.pipe + at Server.setupListenHandle [as _listen2] (node:net:1986:21) + at listenInCluster (node:net:2065:12) + at Server.listen (node:net:2187:5) + at file:///tmp/pi-goals-broker-diagnostic-1T7ENA/package/node_modules/tsx/dist/cli.mjs:53:31472 + at new Promise () + at createIpcServer (file:///tmp/pi-goals-broker-diagnostic-1T7ENA/package/node_modules/tsx/dist/cli.mjs:53:31450) + at async file:///tmp/pi-goals-broker-diagnostic-1T7ENA/package/node_modules/tsx/dist/cli.mjs:55:542 { + code: 'EPERM', + errno: -1, + syscall: 'listen', + address: '/tmp/claude/tsx-1000/28.pipe', + port: -1 +} + +Node.js v25.8.1 diff --git a/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-green.log b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-green.log new file mode 100644 index 0000000..c236780 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-green.log @@ -0,0 +1,9 @@ + + RUN v4.1.9 /home/ubuntu/.pi/agent/worktrees/pi-goals-simple-visible-supervision + + + Test Files 3 passed (3) + Tests 43 passed (43) + Start at 09:06:46 + Duration 2.31s (transform 729ms, setup 86ms, import 2.98s, tests 844ms, environment 0ms) + diff --git a/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-recheck.md b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-recheck.md new file mode 100644 index 0000000..8c7252c --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-recheck.md @@ -0,0 +1,28 @@ +## Review + +Read-only recheck limited to the parent’s fixes for the previously reported P1/P2. + +- **Fixed — P1 resolved:** `src/prompts.ts:199–201` now requires retaining every goal line, completion status, and evidence references above `## Log`; only verbose settled detail moves into the Appendix. This matches the unchanged identity/status filtering in `src/index.ts:220–228`. The pruning instruction and Git-history assumption are gone. + - `test/prompts.test.ts:32–38` guards the corrected instructions. + - `test/goals-flow.test.ts:183–201` moves supporting detail below the fold, exercises the reload hook, and confirms both accept/inconclusive records, the `2/2` count, and inconclusive disclosure remain intact without emitting new work. + +- **Fixed — P2 resolved:** `src/prompts.ts:291–294` now directs the judge to review the unique exact subject and reject missing or ambiguous identity without substituting another goal. The overview in `src/index.ts:21–30` and descriptions in `test/tick-goal.test.ts:16–25` agree with that contract. `test/prompts.test.ts:40–44` guards against restoring fuzzy-match wording. + +- **Correct:** These are bounded prompt/documentation and regression changes. They resolve the conflicts without changing sign-off invalidation semantics, inferring historical approval, or introducing a Git gate. + +**No issues found.** + +### Validation + +Inspected the parent’s saved logs: + +- `housekeeping-red.log`: the two new prompt assertions failed against the former wording. +- `housekeeping-green.log`: **3 test files / 43 tests passed**. + +No commands were run or files edited by this reviewer. + +### Merge verdict: OK with notes + +Both previous findings are resolved; the narrow fixes are approved. This supersedes the previous source-review block. + +Full `npm test` and real isolated Herdr acceptance remain environment-blocked/pending. The targeted tests establish prompt and lifecycle behavior, not real-model judgment or completed two-goal acceptance. Broader parent validation was not attested by this recheck. \ No newline at end of file diff --git a/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-red.log b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-red.log new file mode 100644 index 0000000..cdd3b7a --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/housekeeping-red.log @@ -0,0 +1,81 @@ + + RUN v4.1.9 /home/ubuntu/.pi/agent/worktrees/pi-goals-simple-visible-supervision + + ❯ test/prompts.test.ts (6 tests | 2 failed | 4 skipped) 23ms + × keeps signed-off goal identities during plan housekeeping 18ms + × gives the judge an exact subject rather than a fuzzy-match fallback 3ms + +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 2 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL test/prompts.test.ts > planning prompt > keeps signed-off goal identities during plan housekeeping +AssertionError: expected '\nYour plan (.pi/pla…' to contain 'keep every goal line and its completi…' + +- Expected ++ Received + +- keep every goal line and its completion status above ## Log ++ ++ Your plan (.pi/plan/test.md, above the fold; the log, learnings and appendix are in the file): ++ ++ plan ++ ++ Keep it current as you work, with your normal edit tool: ++ - tick finished subtasks ([/] in progress), add discovered ones ++ - append ONE short line to ## Log, and a line to ## Learnings for a gotcha worth keeping ++ - when the active goal's discriminator is satisfied, fill its evidence: list (each item = a durable ++ artifact + a verbatim quote you actually observed + a short read of it), then call CompleteGoal. ++ Don't tick a goal [x] before CompleteGoal accepts; the sign-off log line is the audit trail. ++ - if the working set has grown long, prune finished goals (their evidence lives in git history and ++ ## Log) and move settled detail down to ## Appendix, which is unlimited ++ - the human's latest message outranks this plan. If it corrects the deliverable or scope, amend the ++ user-visible result, user voice, and affected goals before continuing; don't defend the old plan ++ - otherwise keep working toward the active goal; don't stop to ask unless genuinely blocked ++ + + ❯ test/prompts.test.ts:34:16 + 32| it("keeps signed-off goal identities during plan housekeeping", () =>… + 33| const text = reminder("plan", ".pi/plan/test.md"); + 34| expect(text).toContain("keep every goal line and its completion stat… + | ^ + 35| expect(text).toContain("evidence references beside each goal"); + 36| expect(text).not.toContain("prune finished goals"); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/2]⎯ + + FAIL test/prompts.test.ts > planning prompt > gives the judge an exact subject rather than a fuzzy-match fallback +AssertionError: expected 'The working agent claims this goal is…' to contain 'unique exact goal subject' + +- Expected ++ Received + +- unique exact goal subject ++ The working agent claims this goal is complete: ++ ++ goal: first ++ ++ Below is the full plan file (plan.md). Find that goal in it (tolerate small wording drift; if ++ you cannot find a matching goal at all, reject and say so). Read User-visible result and User voice ++ first, then its discriminator, subtle failure modes, verify command, and evidence list. ++ ++ --- plan file --- ++ 1. [ ] goal: first ++ --- end plan file --- ++ ++ Read the cited artifacts (you cannot execute anything), then give your VERDICT. + + ❯ test/prompts.test.ts:42:16 + 40| it("gives the judge an exact subject rather than a fuzzy-match fallba… + 41| const text = judgeUser({ goal: "first", plan: "1. [ ] goal: first", … + 42| expect(text).toContain("unique exact goal subject"); + | ^ + 43| expect(text).not.toContain("tolerate small wording drift"); + 44| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/2]⎯ + + + Test Files 1 failed | 1 passed (2) + Tests 2 failed | 1 passed | 37 skipped (40) + Start at 09:06:07 + Duration 1.42s (transform 314ms, setup 63ms, import 1.30s, tests 38ms, environment 0ms) + diff --git a/docs/reviews/evidence/2026-09-09-autonomy/parent-build.log b/docs/reviews/evidence/2026-09-09-autonomy/parent-build.log new file mode 100644 index 0000000..7b91ae8 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/parent-build.log @@ -0,0 +1,4 @@ + +> @wassname2/pi-goals@0.2.2 build +> tsc -p tsconfig.build.json + diff --git a/docs/reviews/evidence/2026-09-09-autonomy/parent-lint.log b/docs/reviews/evidence/2026-09-09-autonomy/parent-lint.log new file mode 100644 index 0000000..6cc7679 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/parent-lint.log @@ -0,0 +1,5 @@ + +> @wassname2/pi-goals@0.2.2 lint +> biome check src/ test/ + +Checked 36 files in 91ms. No fixes applied. diff --git a/docs/reviews/evidence/2026-09-09-autonomy/parent-permitted-vitest.log b/docs/reviews/evidence/2026-09-09-autonomy/parent-permitted-vitest.log new file mode 100644 index 0000000..2ff302c --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/parent-permitted-vitest.log @@ -0,0 +1,9 @@ + + RUN v4.1.9 /home/ubuntu/.pi/agent/worktrees/pi-goals-simple-visible-supervision + + + Test Files 13 passed (13) + Tests 94 passed (94) + Start at 09:07:25 + Duration 21.60s (transform 1.68s, setup 395ms, import 20.36s, tests 34.15s, environment 2ms) + diff --git a/docs/reviews/evidence/2026-09-09-autonomy/parent-supervisor.log b/docs/reviews/evidence/2026-09-09-autonomy/parent-supervisor.log new file mode 100644 index 0000000..e6657d6 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/parent-supervisor.log @@ -0,0 +1,188 @@ + +> @wassname2/pi-goals@0.2.2 test:supervisor +> node --import tsx --test test/internal-supervisor/*.test.ts + +✔ retries intercom registration when pi-intercom loads after pi-supervise (3.974028ms) +✔ a directive with no text is rejected, so the worker never sees undefined (0.293578ms) +✔ a directive from the paired supervisor becomes a real user message (29.040859ms) +✔ a directive to a busy worker interrupts, instead of waiting for the whole task (11.791265ms) +✔ a directive from an unpaired session is dropped (19.421131ms) +✔ a second pair takes over, and the first supervisor is told it lost the worker (20.715922ms) +✔ only the paired worker can end a run (6.353483ms) +✔ the programmatic pairing API waits for the worker acknowledgement (1.387398ms) +✔ the worker acknowledges a pair, so the supervisor knows it was heard (5.41818ms) +✔ a goal the supervisor inferred reaches the worker, which owns the view header (132.824229ms) +✔ the second view carries only what happened after the first (435.544291ms) +✔ a message addressed to a different session is ignored (10.263168ms) +✔ on settle the worker publishes a view built from the live branch (55.816252ms) +✔ the view is built from the live branch, not from every entry in the session (55.656998ms) +✔ an unpaired session publishes nothing on settle (0.544096ms) +✔ supervision never stops itself: no round limit at all (5.166058ms) +✔ goal, pairing and the steer count all survive a reload together (0.627375ms) +✔ a view arriving during unrelated supervisor thinking waits for a fresh complete overview (12.193906ms) +✔ the nudge repeats neither the instructions already sent nor the verdict rules (6.136686ms) +✔ a multi-line goal returns to supervisor context every fifth review and after compaction (32.391613ms) +✔ a one-line goal is not redundantly reinserted (27.14758ms) +✔ a check in and a worker that stopped ask for different things (11.315795ms) +✔ a loop still gets named after the supervisor compacts, from restored state (0.916959ms) +✔ a session that does not answer the roll call is not offered as a worker (501.50719ms) +✔ a child run stays out of the roll call, so it can never be picked (4.989808ms) +✔ a session already paired stays out of the roll call, and a free one answers (15.007393ms) +✔ /supervise look asks the worker for a fresh view, rather than the supervisor guessing (306.09392ms) +✔ let_it_run says the turn is over, so it is not called four times running (0.665336ms) +✔ a sign-off verdict is answered, not aborted, and a runaway is still cut (0.661296ms) +✔ every verdict result names the way to end the turn, steer included (0.41022ms) +✔ an old view is dropped from context once its verdict is in, and the verdict is kept (1.049552ms) +✔ a worker session never has its context rewritten (0.312445ms) +✔ a newly presented view starts a fresh look (5.462915ms) +✔ a tool a worker cannot use never aborts its turn (0.423344ms) +✔ a resume onto a session that is gone drops the pairing and says so (5.570522ms) +✔ a resume onto a live worker keeps supervising, and takes the writers back off (6.022216ms) +✔ state written before recentSteers existed still loads (0.180461ms) +✔ done unpairs the worker, so it stops publishing views (428.081704ms) +✔ with no goal the supervisor cannot steer, it must ask the human (0.830044ms) +✔ set_goal binds an inferred goal, and steering then works (501.633016ms) +✔ a goal given at pair time still allows steering (0.554833ms) +✔ done is refused while the worker has an unanswered tool call (12.761991ms) +✔ done is allowed once nothing is outstanding (6.532494ms) +✔ steer refuses when the session is not supervising (0.399461ms) +✔ a reworded repeat of an earlier instruction is sent, and named back to the supervisor (0.593568ms) +✔ overlap scores rewording high and a different instruction low (0.128624ms) +✔ the view of the old worker cannot be used to judge the new one (6.184772ms) +✔ with one other session here, /supervise needs no target and the whole line is the goal (501.541112ms) +✔ naming the worker still works, and the rest of the line is the goal (0.679012ms) +✔ with two free sessions here, /supervise asks which one, and pairs with the choice (500.988416ms) +✔ a goal that is a path is read from the file, so it is not pasted every run (2.052271ms) +✔ a long goal is one short line above the picker, and reaches the worker whole (501.890317ms) +✔ a session that stayed quiet is still on the list, because 0 free is a dead end (501.763389ms) +✔ a cancelled picker pairs with nothing (501.758475ms) +✔ supervising takes the writing tools away, and stopping gives them back (500.746419ms) +✔ stopping gives back the writers without undoing another extension's tools (501.771845ms) +✔ a first word that names no session is refused, rather than folded into the goal (0.654949ms) +✔ a goal with spaces needs no target, and @name takes the rest of the line as the goal (502.193535ms) +✔ the brief starts no turn, so there is no answer before the first view (6.184913ms) +✔ /supervise goal changes the goal without breaking the pairing (0.754771ms) +✔ the footer says which side of a pairing this session is, and clears when it ends (508.144262ms) +✔ a session that is not supervising never sees the supervisor tools (5.351772ms) +✔ worker_view refuses when there is no worker, rather than implying a pairing (1.56818ms) +✔ the view names the worker's model and how full its context is (55.815989ms) +✔ supervising a second session is refused while the first is still paired (0.644348ms) +✔ the supervisor gets a look at a working worker every half hour, without being asked (973.448527ms) +✔ a human message in the worker session is not a reason to stand back (6.268048ms) +✔ letting a stopped worker run says plainly that the worker stays stopped (10.365848ms) +✔ a stopped worker is looked at again, so let_it_run cannot silence the pairing (925.887224ms) +✔ a worker that pairs at the prompt and never takes a turn is still watched (605.51902ms) +✔ a worker that reloads at the prompt starts watching itself again (604.743394ms) +✔ a timer look at a worker that has not moved is not sent, until it has been skipped three times (2429.468038ms) +✔ the worker counts reviews in a row where nothing changed (577.985914ms) +✔ an unacknowledged pair gives up, and a takeover cancels that timer (5.178547ms) +✔ duplicate standalone Intercom registries are diagnosed and cannot bootstrap a plan (0.832525ms) +✔ retained non-plan supervision reconnects after the worker reloads, not on unrelated peer traffic (344.975935ms) +✔ retained non-plan supervision reconnects after the supervisor reloads, not on unrelated peer traffic (627.895644ms) +✔ busy supervisor retains its active view and requests one complete overview after settling (31.331632ms) +✔ ordinary prose leaves future reviews live: On course; the saved check is the next useful evidence. (31.299092ms) +✔ ordinary prose leaves future reviews live: Which output format do you want? (7.054922ms) +✔ manual last-goal ticks cannot end plan supervision before CompleteGoal (3.367731ms) +✔ completion counts are plan-bound and distinguish inconclusive sign-off when ending supervision (4.579827ms) +✔ each checkpoint freezes fresh worker evidence and direction without replacing a busy assessment (9.345858ms) +✔ checkpoint snapshot building cannot publish after abort (2.874272ms) +✔ checkpoint snapshot building cannot publish after stop (21.699085ms) +✔ checkpoint snapshot building cannot publish after reload (2.94349ms) +✔ checkpoint snapshot building cannot publish after plan change (3.193298ms) +✔ checkpoint capture includes user direction arriving while tracked work is queried (2.828249ms) +✔ checkpoint payload fits the serialized channel limit without truncating its identity (77.119836ms) +✔ settled empty final response fails the checkpoint explicitly ([]) (5.826315ms) +✔ settled empty final response fails the checkpoint explicitly (["text"]) (2.399616ms) +✔ settled empty final response fails the checkpoint explicitly (["text"]) (1.983875ms) +✔ settled empty final response fails the checkpoint explicitly (["thinking"]) (1.55404ms) +✔ failed assessment resumes on later worker progress without a human poke (stop) (1.767669ms) +✔ failed assessment resumes on later worker progress without a human poke (error) (1.397948ms) +✔ empty low-level response may continue through compaction, ask a human, or finish with a tool verdict (2.553043ms) +✔ a successful goal tool verdict is not undone by an empty final response (1.959099ms) +✔ a duplicate checkpoint rejection never echoes its snapshot or changes the active view (2.140537ms) +✔ an empty routine assessment cannot fail a separately queued checkpoint (2.391144ms) +✔ plan bootstrap compacts only the supervisor and pairing alone never starts a worker or a review (1.738156ms) +✔ goal decisions are correlated, preserve the pair across two goals, and cannot call overall done (2.61962ms) +✔ abort and stop cancel pending requests; late decisions cannot approve a replacement (1.943861ms) +✔ 50 actual model turns trigger one view, independent of the number of messages (2.077155ms) +✔ unknown background providers are not proof of quiescence (0.224206ms) +✔ stale plan content invalidates a pending goal review (2.190617ms) +✔ small forks skip compaction, but real compaction failure prevents pairing (1.770236ms) +✔ native small-history result permits startup after an attempted compaction (50000) (1.149776ms) +✔ native small-history result permits startup after an attempted compaction (null) (0.989274ms) +✔ the hour timer and a coincident turn checkpoint produce a single view (4.810231ms) +✔ settled checks wait for tracked processes and subagents to finish (2.321824ms) +✔ bootstrap stop cannot resurrect a supervisor after compaction completes (1.887956ms) +✔ a restarted worker reconnects by exact saved session identity without a new supervisor (2.651845ms) +✔ unknown initial context must compact instead of taking the known-small shortcut (0.884878ms) +✔ null post-compaction usage cannot raise the next configured 100k checkpoint (1.897635ms) +✔ stopping a routine view during compaction invalidates its suspended continuation (1.841943ms) +✔ restart of a provisional bootstrap resumes compaction and pairing in the same saved session (2.414359ms) +✔ command preserves a stopped supervisor across reload (1.739858ms) +✔ done preserves a stopped supervisor across reload (2.194354ms) +✔ same-binding replay retains activation when the supervisor lost its acknowledgement (2.993195ms) +✔ model-unavailable stop validates binding, cancels pending reviews and ignores old directives (2.274799ms) +✔ supervisor mode is a native inspection allowlist at visibility and execution, including reload and stopped forks (2.183069ms) +✔ a cancelled checkpoint's delayed verdict cannot approve the replacement checkpoint (2.297678ms) +✔ each model call reanchors the canonical plan and role without losing compacted planning context or judgments (1.91378ms) +✔ routine assessments and steering display the actual advice rather than only a receipt (1.950418ms) +✔ current context above 100k compacts; cumulative usage and exactly 100k do not (1.806256ms) +✔ absent optional trackers count as zero tracked work; installed failed or busy trackers remain non-quiet (0.346121ms) +✔ a settled worker with no optional trackers sends one review, not repeated idle wakes (2.802864ms) +✔ disconnect cancels a checkpoint and blocks steering; local stop still clears ownership (2.023057ms) +✔ registered malformed subagent tracker stays unknown even when the process tracker is absent (0.352469ms) +✔ a busy supervisor defers the 100k compaction until its own run settles (1.919556ms) +✔ a healthy goal review can take longer than ten minutes without cancellation or another pairing (2.175237ms) +✔ busy plan supervisor refreshes cumulative VCC evidence without an idle feedback loop (4.336383ms) +✔ progress during a refresh remains pending with only one look in flight (8.493816ms) +✔ complete refresh uses the current worker branch after compaction (4.751519ms) +✔ complete refresh uses the current worker branch after rewind (3.59423ms) +✔ complete refresh uses the current worker branch after reload (3.022736ms) +✔ explicit checkpoints wait separately from routine coalescing and do not replace an active assessment (3.084625ms) +✔ actual supervisor provider failure is returned explicitly without another pair or an elapsed-time cancellation (1.926238ms) +✔ a provider error followed by native retry success does not cancel the supervisor checkpoint (1.872509ms) +✔ a failed full overview waits without spinning and recovers on routine progress (4.173687ms) +✔ a failed full overview waits without spinning and recovers on explicit look (2.924851ms) +✔ overview display metadata uses native context without requiring a remote roster lookup (1.976883ms) +✔ a superseded advance re-drives the current checkpoint after deferred compaction success (3.037736ms) +✔ a superseded advance re-drives the current checkpoint after deferred compaction rejection (2.425301ms) +✔ stopping current work during an obsolete advance prevents re-drive after compaction success (2.481801ms) +✔ stopping current work during an obsolete advance prevents re-drive after compaction rejection (2.130017ms) +✔ goal_review accepts an optional bounded snapshot but rejects malformed snapshots (1.89367ms) +✔ worker completion metadata is optional but cannot claim malformed counts (0.451557ms) +✔ checkpoint snapshot bounding counts JSON escapes and does not split Unicode characters (5.01807ms) +✔ a child process named pi is found by ps, and stops being found when it exits (413.232236ms) +✔ the check is a snapshot, so it cannot hold up the worker's settle (324.687926ms) +✔ latest user direction survives bounded summaries, compaction and later supervisor echoes (114.635148ms) +✔ oversized user direction is visibly bounded with a source reference (2.69704ms) +✔ a one-line goal stays whole while a multi-line goal has a locator (1.094539ms) +✔ a view carries only the turns the supervisor has not been sent (1.084425ms) +✔ the last two reasoning blocks stay in the narrative, and older ones drop out (1.002457ms) +✔ a compaction restarts the view, so no turn falls into the gap (0.457013ms) +✔ pi-vcc reports the files the worker wrote, and separates them from the ones it read (1.503099ms) +✔ progressKey is unchanged when a review produced no new file or commit (0.732168ms) +✔ progressKey still sees a new file past pi-vcc's ten path display cap (1.020185ms) +✔ a commit counts as progress, even when no file was written since (8.584645ms) +✔ outstandingWork finds tool calls that never got a result (3.568082ms) +✔ buildView reports a tool call with no result, so done can be refused (2.667216ms) +✔ the view says how many reviews in a row changed nothing, and says nothing at zero (1.134394ms) +✔ the view merges the worker's compaction summary with the turns after it (0.422717ms) +✔ a turn the compaction summary already covers is not sent twice (1.946043ms) +✔ pi-vcc's sections and its transcript land on the right sides of the split (2.32661ms) +✔ the view does not tell the supervisor to use vcc_recall, a tool it does not have (0.296905ms) +✔ supervisor directives are not sent back as worker evidence (0.792699ms) +✔ bookkeeping tool calls are kept out of the transcript (0.44677ms) +✔ buildView reports the goal, status, and files without historical failures (0.427038ms) +✔ how long the worker has been quiet, measured from its own last entry (0.436767ms) +✔ buildView keeps the newest turns when it has to cut for the channel limit (45.163847ms) +✔ pi's own branch logic drops the abandoned fork, on a session file (2196.625029ms) +✔ a long goal cannot push the view past the broker limit (0.457776ms) +✔ a bounded complete overview explicitly labels a truncated worker compaction summary (0.265551ms) +ℹ tests 176 +ℹ suites 0 +ℹ pass 176 +ℹ fail 0 +ℹ cancelled 0 +ℹ skipped 0 +ℹ todo 0 +ℹ duration_ms 21219.975303 diff --git a/docs/reviews/evidence/2026-09-09-autonomy/parent-typecheck.log b/docs/reviews/evidence/2026-09-09-autonomy/parent-typecheck.log new file mode 100644 index 0000000..4d9a05e --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-autonomy/parent-typecheck.log @@ -0,0 +1,4 @@ + +> @wassname2/pi-goals@0.2.2 typecheck +> tsc -p tsconfig.build.json --noEmit + diff --git a/docs/reviews/evidence/2026-09-09-recovery/build.log b/docs/reviews/evidence/2026-09-09-recovery/build.log new file mode 100644 index 0000000..7b91ae8 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/build.log @@ -0,0 +1,4 @@ + +> @wassname2/pi-goals@0.2.2 build +> tsc -p tsconfig.build.json + diff --git a/docs/reviews/evidence/2026-09-09-recovery/exit-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/exit-pane.txt new file mode 100644 index 0000000..78bae4f --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/exit-pane.txt @@ -0,0 +1,34 @@ +wdl, woodside, workflow-diagram-tracker, ws-roadmap-slide, wsl-proxy + +[Prompts] + /council, /gather-context-and-clarify, /parallel-cleanup, /parallel-research, /parallel-review, /review-loop + +[Extensions] + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web + +[Themes] + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night, +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + Goals mode exited. Files retained; no work approved. Supervisor sessions remain inspection-only. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) \ No newline at end of file diff --git a/docs/reviews/evidence/2026-09-09-recovery/fresh-status-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/fresh-status-pane.txt new file mode 100644 index 0000000..47029ce --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/fresh-status-pane.txt @@ -0,0 +1,70 @@ + │ +[Extensions] │ + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, │ +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web │ + │ +[Themes] │ + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, │ +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, │ +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night│ +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc │ + │ + │ + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. │ + │ + pi-better-compaction loaded • debug artifacts → /home/ubuntu/.pi/agent/artifacts/pi-better-compaction/sessions/01a084a5-a0f4-74c9-99fc-f9e5ae61adec/lifecycle/2026-09-09T05-35-21-169Z-lifecycle.json │ + │ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ + │ + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. │ + │ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ + ┃ + Operator fixture marker, not a model response. ┃ + ┃ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┃ + Package Updates Available ┃ + Package updates are available. Run pi update --extensions ┃ + Packages: ┃ + - pi-sandbox ┃ + - pi-zentui ┃ + - github.com/wassname/pi-annotated-journal ┃ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┃ + ┃ + Goals: ordinary chat (goals exited). ┃ + Pair: disconnected/unconfirmed; inactive. ↓ Jump to latest message · End │ + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + pi-better-compaction loaded • debug artifacts → /home/ubuntu/.pi/agent/artifacts/pi-better-compaction/sessions/01a084a5-a0f4-74c9-99fc-f9e5ae61adec/lifecycle/2026-09-09T05-35-21-169Z-lifecycle.json + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + Package Updates Available + Package updates are available. Run pi update --extensions + Packages: + - pi-sandbox + - pi-zentui + - github.com/wassname/pi-annotated-journal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Goals: ordinary chat (goals exited). + Pair: disconnected/unconfirmed; inactive. + Worker: not recorded · no paired session + Supervisor: not recorded · no paired session + Last failure: none recorded in this runtime + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) diff --git a/docs/reviews/evidence/2026-09-09-recovery/lint.log b/docs/reviews/evidence/2026-09-09-recovery/lint.log new file mode 100644 index 0000000..93e1e2f --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/lint.log @@ -0,0 +1,5 @@ + +> @wassname2/pi-goals@0.2.2 lint +> biome check src/ test/ + +Checked 36 files in 97ms. No fixes applied. diff --git a/docs/reviews/evidence/2026-09-09-recovery/npm-test.log b/docs/reviews/evidence/2026-09-09-recovery/npm-test.log new file mode 100644 index 0000000..f9a7291 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/npm-test.log @@ -0,0 +1,204 @@ + +> @wassname2/pi-goals@0.2.2 test +> vitest run && npm run test:supervisor + + + RUN v4.1.9 /home/ubuntu/.pi/agent/worktrees/pi-goals-simple-visible-supervision + + + Test Files 14 passed (14) + Tests 101 passed (101) + Start at 13:38:22 + Duration 23.27s (transform 2.37s, setup 395ms, import 21.82s, tests 50.33s, environment 5ms) + + +> @wassname2/pi-goals@0.2.2 test:supervisor +> node --import tsx --test test/internal-supervisor/*.test.ts + +✔ retries intercom registration when pi-intercom loads after pi-supervise (3.513793ms) +✔ a directive with no text is rejected, so the worker never sees undefined (0.472691ms) +✔ a directive from the paired supervisor becomes a real user message (27.644033ms) +✔ a directive to a busy worker interrupts, instead of waiting for the whole task (13.94769ms) +✔ a directive from an unpaired session is dropped (15.908473ms) +✔ a second pair takes over, and the first supervisor is told it lost the worker (21.082703ms) +✔ only the paired worker can end a run (7.780322ms) +✔ the programmatic pairing API waits for the worker acknowledgement (1.376736ms) +✔ the worker acknowledges a pair, so the supervisor knows it was heard (6.283045ms) +✔ a goal the supervisor inferred reaches the worker, which owns the view header (131.134854ms) +✔ the second view carries only what happened after the first (436.748028ms) +✔ a message addressed to a different session is ignored (10.19648ms) +✔ on settle the worker publishes a view built from the live branch (54.951438ms) +✔ the view is built from the live branch, not from every entry in the session (54.932155ms) +✔ an unpaired session publishes nothing on settle (0.5274ms) +✔ supervision never stops itself: no round limit at all (5.101803ms) +✔ goal, pairing and the steer count all survive a reload together (0.596129ms) +✔ a view arriving during unrelated supervisor thinking waits for a fresh complete overview (11.19975ms) +✔ the nudge repeats neither the instructions already sent nor the verdict rules (6.045907ms) +✔ a multi-line goal returns to supervisor context every fifth review and after compaction (32.277779ms) +✔ a one-line goal is not redundantly reinserted (26.746878ms) +✔ a check in and a worker that stopped ask for different things (10.249474ms) +✔ a loop still gets named after the supervisor compacts, from restored state (1.006545ms) +✔ a session that does not answer the roll call is not offered as a worker (502.069337ms) +✔ a child run stays out of the roll call, so it can never be picked (5.882165ms) +✔ a session already paired stays out of the roll call, and a free one answers (15.515028ms) +✔ /supervise look asks the worker for a fresh view, rather than the supervisor guessing (306.997507ms) +✔ let_it_run says the turn is over, so it is not called four times running (0.614118ms) +✔ a sign-off verdict is answered, not aborted, and a runaway is still cut (0.766328ms) +✔ every verdict result names the way to end the turn, steer included (0.601185ms) +✔ an old view is dropped from context once its verdict is in, and the verdict is kept (1.039013ms) +✔ a worker session never has its context rewritten (0.303001ms) +✔ a newly presented view starts a fresh look (5.626017ms) +✔ a tool a worker cannot use never aborts its turn (0.604727ms) +✔ a resume onto a session that is gone drops the pairing and says so (5.917942ms) +✔ a resume onto a live worker keeps supervising, and takes the writers back off (5.276493ms) +✔ state written before recentSteers existed still loads (0.221018ms) +✔ done unpairs the worker, so it stops publishing views (423.568859ms) +✔ with no goal the supervisor cannot steer, it must ask the human (0.742368ms) +✔ set_goal binds an inferred goal, and steering then works (501.288202ms) +✔ a goal given at pair time still allows steering (0.55747ms) +✔ done is refused while the worker has an unanswered tool call (12.979101ms) +✔ done is allowed once nothing is outstanding (5.279803ms) +✔ steer refuses when the session is not supervising (0.367787ms) +✔ a reworded repeat of an earlier instruction is sent, and named back to the supervisor (0.586311ms) +✔ overlap scores rewording high and a different instruction low (0.138077ms) +✔ the view of the old worker cannot be used to judge the new one (6.59538ms) +✔ with one other session here, /supervise needs no target and the whole line is the goal (501.480943ms) +✔ naming the worker still works, and the rest of the line is the goal (0.566102ms) +✔ with two free sessions here, /supervise asks which one, and pairs with the choice (500.820463ms) +✔ a goal that is a path is read from the file, so it is not pasted every run (2.17768ms) +✔ a long goal is one short line above the picker, and reaches the worker whole (501.878571ms) +✔ a session that stayed quiet is still on the list, because 0 free is a dead end (501.668999ms) +✔ a cancelled picker pairs with nothing (500.656172ms) +✔ supervising takes the writing tools away, and stopping gives them back (501.514864ms) +✔ stopping gives back the writers without undoing another extension's tools (501.710321ms) +✔ a first word that names no session is refused, rather than folded into the goal (0.552837ms) +✔ a goal with spaces needs no target, and @name takes the rest of the line as the goal (501.716252ms) +✔ the brief starts no turn, so there is no answer before the first view (6.049466ms) +✔ /supervise goal changes the goal without breaking the pairing (0.644898ms) +✔ the footer says which side of a pairing this session is, and clears when it ends (506.479387ms) +✔ a session that is not supervising never sees the supervisor tools (6.510238ms) +✔ worker_view refuses when there is no worker, rather than implying a pairing (0.773467ms) +✔ the view names the worker's model and how full its context is (55.333572ms) +✔ supervising a second session is refused while the first is still paired (0.619021ms) +✔ the supervisor gets a look at a working worker every half hour, without being asked (974.700755ms) +✔ a human message in the worker session is not a reason to stand back (6.119204ms) +✔ letting a stopped worker run says plainly that the worker stays stopped (11.325288ms) +✔ a stopped worker is looked at again, so let_it_run cannot silence the pairing (928.080952ms) +✔ a worker that pairs at the prompt and never takes a turn is still watched (605.615767ms) +✔ a worker that reloads at the prompt starts watching itself again (605.010128ms) +✔ a timer look at a worker that has not moved is not sent, until it has been skipped three times (2434.182048ms) +✔ the worker counts reviews in a row where nothing changed (580.138243ms) +✔ an unacknowledged pair gives up, and a takeover cancels that timer (5.648124ms) +✔ duplicate standalone Intercom registries are diagnosed and cannot bootstrap a plan (0.839488ms) +✔ retained non-plan supervision reconnects after the worker reloads, not on unrelated peer traffic (345.059118ms) +✔ retained non-plan supervision reconnects after the supervisor reloads, not on unrelated peer traffic (628.825545ms) +✔ busy supervisor retains its active view and requests one complete overview after settling (31.386612ms) +✔ human pause survives reload/reconnect; only explicit resume reactivates the same pair (33.112188ms) +✔ stop cancels an in-flight checkpoint even when the cancel notification cannot be sent (6.210007ms) +✔ supervisor pause cancels checkpoint and a disconnected local stop still persists (4.374153ms) +✔ ordinary prose leaves future reviews live: On course; the saved check is the next useful evidence. (8.733354ms) +✔ ordinary prose leaves future reviews live: Which output format do you want? (4.512364ms) +✔ manual last-goal ticks cannot end plan supervision before CompleteGoal (2.39126ms) +✔ completion counts are plan-bound and distinguish inconclusive sign-off when ending supervision (4.338931ms) +✔ each checkpoint freezes fresh worker evidence and direction without replacing a busy assessment (6.420082ms) +✔ checkpoint snapshot building cannot publish after abort (2.288189ms) +✔ checkpoint snapshot building cannot publish after stop (8.410958ms) +✔ checkpoint snapshot building cannot publish after reload (1.97679ms) +✔ checkpoint snapshot building cannot publish after plan change (1.998333ms) +✔ checkpoint capture includes user direction arriving while tracked work is queried (2.027935ms) +✔ checkpoint payload fits the serialized channel limit without truncating its identity (76.173014ms) +✔ settled empty final response fails the checkpoint explicitly ([]) (4.469246ms) +✔ settled empty final response fails the checkpoint explicitly (["text"]) (1.484115ms) +✔ settled empty final response fails the checkpoint explicitly (["text"]) (1.620401ms) +✔ settled empty final response fails the checkpoint explicitly (["thinking"]) (1.455427ms) +✔ failed assessment resumes on later worker progress without a human poke (stop) (2.139792ms) +✔ failed assessment resumes on later worker progress without a human poke (error) (1.490922ms) +✔ empty low-level response may continue through compaction, ask a human, or finish with a tool verdict (2.368309ms) +✔ a successful goal tool verdict is not undone by an empty final response (2.000919ms) +✔ a duplicate checkpoint rejection never echoes its snapshot or changes the active view (1.967399ms) +✔ an empty routine assessment cannot fail a separately queued checkpoint (2.00475ms) +✔ plan bootstrap compacts only the supervisor and pairing alone never starts a worker or a review (1.394665ms) +✔ goal decisions are correlated, preserve the pair across two goals, and cannot call overall done (2.282643ms) +✔ abort and stop cancel pending requests; late decisions cannot approve a replacement (1.660897ms) +✔ 50 actual model turns trigger one view, independent of the number of messages (1.928283ms) +✔ unknown background providers are not proof of quiescence (0.198076ms) +✔ stale plan content invalidates a pending goal review (2.144203ms) +✔ small forks skip compaction, but real compaction failure prevents pairing (1.584189ms) +✔ native small-history result permits startup after an attempted compaction (50000) (0.988832ms) +✔ native small-history result permits startup after an attempted compaction (null) (0.75011ms) +✔ the hour timer and a coincident turn checkpoint produce a single view (4.330684ms) +✔ settled checks wait for tracked processes and subagents to finish (2.062665ms) +✔ bootstrap stop cannot resurrect a supervisor after compaction completes (1.529383ms) +✔ a restarted worker reconnects by exact saved session identity without a new supervisor (2.505604ms) +✔ unknown initial context must compact instead of taking the known-small shortcut (1.267376ms) +✔ null post-compaction usage cannot raise the next configured 100k checkpoint (2.338111ms) +✔ stopping a routine view during compaction invalidates its suspended continuation (1.707556ms) +✔ restart of a provisional bootstrap resumes compaction and pairing in the same saved session (2.48264ms) +✔ command preserves a stopped supervisor across reload (1.840487ms) +✔ done preserves a stopped supervisor across reload (2.952022ms) +✔ same-binding replay retains activation when the supervisor lost its acknowledgement (3.738066ms) +✔ model-unavailable stop validates binding, cancels pending reviews and ignores old directives (2.963335ms) +✔ supervisor mode is a native inspection allowlist at visibility and execution, including reload and stopped forks (3.2371ms) +✔ a cancelled checkpoint's delayed verdict cannot approve the replacement checkpoint (3.477147ms) +✔ each model call reanchors the canonical plan and role without losing compacted planning context or judgments (2.473904ms) +✔ routine assessments and steering display the actual advice rather than only a receipt (2.783747ms) +✔ current context above 100k compacts; cumulative usage and exactly 100k do not (2.379765ms) +✔ absent optional trackers count as zero tracked work; installed failed or busy trackers remain non-quiet (0.421638ms) +✔ a settled worker with no optional trackers sends one review, not repeated idle wakes (3.307215ms) +✔ disconnect cancels a checkpoint and blocks steering; local stop still clears ownership (2.01323ms) +✔ registered malformed subagent tracker stays unknown even when the process tracker is absent (0.311552ms) +✔ a busy supervisor defers the 100k compaction until its own run settles (1.785982ms) +✔ a healthy goal review can take longer than ten minutes without cancellation or another pairing (2.14447ms) +✔ busy plan supervisor refreshes cumulative VCC evidence without an idle feedback loop (3.204869ms) +✔ progress during a refresh remains pending with only one look in flight (3.93267ms) +✔ complete refresh uses the current worker branch after compaction (2.418211ms) +✔ complete refresh uses the current worker branch after rewind (2.364448ms) +✔ complete refresh uses the current worker branch after reload (2.477204ms) +✔ explicit checkpoints wait separately from routine coalescing and do not replace an active assessment (2.757015ms) +✔ actual supervisor provider failure is returned explicitly without another pair or an elapsed-time cancellation (1.60961ms) +✔ a provider error followed by native retry success does not cancel the supervisor checkpoint (1.625357ms) +✔ a failed full overview waits without spinning and recovers on routine progress (3.834412ms) +✔ a failed full overview waits without spinning and recovers on explicit look (2.610627ms) +✔ overview display metadata uses native context without requiring a remote roster lookup (1.794804ms) +✔ a superseded advance re-drives the current checkpoint after deferred compaction success (2.809267ms) +✔ a superseded advance re-drives the current checkpoint after deferred compaction rejection (2.450308ms) +✔ stopping current work during an obsolete advance prevents re-drive after compaction success (2.096725ms) +✔ stopping current work during an obsolete advance prevents re-drive after compaction rejection (1.957782ms) +✔ goal_review accepts an optional bounded snapshot but rejects malformed snapshots (1.551675ms) +✔ worker completion metadata is optional but cannot claim malformed counts (0.459292ms) +✔ checkpoint snapshot bounding counts JSON escapes and does not split Unicode characters (5.01943ms) +✔ a child process named pi is found by ps, and stops being found when it exits (361.27963ms) +✔ the check is a snapshot, so it cannot hold up the worker's settle (327.919292ms) +✔ latest user direction survives bounded summaries, compaction and later supervisor echoes (82.002344ms) +✔ oversized user direction is visibly bounded with a source reference (2.652949ms) +✔ a one-line goal stays whole while a multi-line goal has a locator (1.298412ms) +✔ a view carries only the turns the supervisor has not been sent (1.16182ms) +✔ the last two reasoning blocks stay in the narrative, and older ones drop out (1.009051ms) +✔ a compaction restarts the view, so no turn falls into the gap (0.491271ms) +✔ pi-vcc reports the files the worker wrote, and separates them from the ones it read (1.551769ms) +✔ progressKey is unchanged when a review produced no new file or commit (0.723151ms) +✔ progressKey still sees a new file past pi-vcc's ten path display cap (0.853796ms) +✔ a commit counts as progress, even when no file was written since (1.124031ms) +✔ outstandingWork finds tool calls that never got a result (2.699571ms) +✔ buildView reports a tool call with no result, so done can be refused (1.30884ms) +✔ the view says how many reviews in a row changed nothing, and says nothing at zero (0.553271ms) +✔ the view merges the worker's compaction summary with the turns after it (0.439182ms) +✔ a turn the compaction summary already covers is not sent twice (0.975019ms) +✔ pi-vcc's sections and its transcript land on the right sides of the split (1.84901ms) +✔ the view does not tell the supervisor to use vcc_recall, a tool it does not have (0.311653ms) +✔ supervisor directives are not sent back as worker evidence (0.566437ms) +✔ bookkeeping tool calls are kept out of the transcript (0.395583ms) +✔ buildView reports the goal, status, and files without historical failures (0.43785ms) +✔ how long the worker has been quiet, measured from its own last entry (0.458098ms) +✔ buildView keeps the newest turns when it has to cut for the channel limit (39.471288ms) +✔ pi's own branch logic drops the abandoned fork, on a session file (1939.636064ms) +✔ a long goal cannot push the view past the broker limit (0.551003ms) +✔ a bounded complete overview explicitly labels a truncated worker compaction summary (0.286204ms) +ℹ tests 179 +ℹ suites 0 +ℹ pass 179 +ℹ fail 0 +ℹ cancelled 0 +ℹ skipped 0 +ℹ todo 0 +ℹ duration_ms 20775.476362 diff --git a/docs/reviews/evidence/2026-09-09-recovery/reconnect-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/reconnect-pane.txt new file mode 100644 index 0000000..5a8d1e7 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/reconnect-pane.txt @@ -0,0 +1,34 @@ +[Prompts] + /council, /gather-context-and-clarify, /parallel-cleanup, /parallel-research, /parallel-review, /review-loop + +[Extensions] + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web + +[Themes] + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night, +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + Goals mode exited. Files retained; no work approved. Supervisor sessions remain inspection-only. + + Warning: Recovery failed; no new work authorized: Error: No recoverable pair. Reconnect never creates a supervisor; inspect the recorded pane or select Ready for a new pairing. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) \ No newline at end of file diff --git a/docs/reviews/evidence/2026-09-09-recovery/reload-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/reload-pane.txt new file mode 100644 index 0000000..7aafd06 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/reload-pane.txt @@ -0,0 +1,34 @@ + +[Prompts] + /council, /gather-context-and-clarify, /parallel-cleanup, /parallel-research, /parallel-review, /review-loop + +[Extensions] + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web + +[Themes] + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night, +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + Reloaded keybindings, extensions, skills, prompts, themes, and context files + + Goals stopped by user. Plan and evidence retained. +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF goals stopped · /goals resume | 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) \ No newline at end of file diff --git a/docs/reviews/evidence/2026-09-09-recovery/resume-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/resume-pane.txt new file mode 100644 index 0000000..fc38623 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/resume-pane.txt @@ -0,0 +1,34 @@ + +[Prompts] + /council, /gather-context-and-clarify, /parallel-cleanup, /parallel-research, /parallel-review, /review-loop + +[Extensions] + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web + +[Themes] + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night, +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + Draft restored; no work started. Review the plan and request Ready before working. + + pi-goals: drafting goals +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF planning | 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) \ No newline at end of file diff --git a/docs/reviews/evidence/2026-09-09-recovery/stop-pane.txt b/docs/reviews/evidence/2026-09-09-recovery/stop-pane.txt new file mode 100644 index 0000000..91ca93e --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/stop-pane.txt @@ -0,0 +1,75 @@ +integrated-solutions-architecture, ipynb, jira-search, just, m365-search, marimo, markdown-tables, md-share, mermaid-js, ml-debug, ml-human-in-loop, moa, moa-brainstorm, moa-science, ocr, on-the-record, oracle, pi-intercom, pi-processes, pi-subagents, plan-format, ┃ +platform-alignment, project-brief, pseudopy, react-component-design, react-unit-testing, recommending-pi-extensions, review, roadmap, scout-provision, skill-authoring, skill-bootstrapper, skill-importer, snowflake-query, task-tracker, tufte-viz, uv, vargdown, varglite, ┃ +wdl, woodside, workflow-diagram-tracker, ws-roadmap-slide, wsl-proxy ┃ +[Prompts] ┃ +integrated-solutions-architecture, ipynb, jira-search, just, m365-search, marimo, markdown-tables, md-share, mermaid-js, ml-debug, ml-human-in-loop, moa, moa-brainstorm, moa-science, ocr, on-the-record, oracle, pi-intercom, pi-processes, pi-subagents, plan-format, │ +platform-alignment, project-brief, pseudopy, react-component-design, react-unit-testing, recommending-pi-extensions, review, roadmap, scout-provision, skill-authoring, skill-bootstrapper, skill-importer, snowflake-query, task-tracker, tufte-viz, uv, vargdown, varglite, │ +wdl, woodside, workflow-diagram-tracker, ws-roadmap-slide, wsl-proxy │ +[Prompts] │ + /council, /gather-context-and-clarify, /parallel-cleanup, /parallel-research, /parallel-review, /review-loop │ + │ +[Extensions] │ + @aliou/pi-processes:processes, @aliou/pi-processes:processes-dock, @aliou/pi-processes:processes-logs, @ff-labs/pi-fff:src, herdr-agent-state.ts, pi-annotated-journal, pi-better-compaction, pi-sandbox, pi-schedule-prompt:src, pi-subagents, │ +pi-zed-shift-enter:extension.ts, pi-zentui:zentui, src, wassname/pi-copilot-web │ + │ +[Themes] │ + adventure, adwaita-dark, arcoiris, arthur, atom, aura, black-metal-bathory, black-metal-burzum, black-metal-khold, box, brogrammer, carbonfox, catppuccin-mocha, citruszest, cursor-dark, cutie-pro, dark-modern, dark-pastel, dimmed-monokai, doom-peacock, dracula-plus, │ +earthsong, everforest-dark-hard, fahrenheit, flatland, flexoki-dark, front-end-delight, fun-forrest, galizur, github-dark-colorblind, github-dark-high-contrast, glacier, gruber-darker, gruvbox-dark, gruvbox-dark-hard, gruvbox-material, guezwhoz, hacktober, hardcore, │ +havn-skumring, ic-orange-ppl, iterm2-smoooooth, iterm2-tango-dark, japanesque, jellybeans, kanagawa-wave, kurokula, later-this-evening, lovelace, material-darker, matte-black, mellow, miasma, nvim-dark, popping-and-locking, sea-shells, sleepy-hollow, smyck, tomorrow-night│ +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc │ + │ + │ + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. │ + │ + pi-better-compaction loaded • debug artifacts → /home/ubuntu/.pi/agent/artifacts/pi-better-compaction/sessions/01a084a5-a0f4-74c9-99fc-f9e5ae61adec/lifecycle/2026-09-09T05-32-14-754Z-lifecycle.json │ + │ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────│ + ┃ + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. ┃ + ┃ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┃ + ┃ + Operator fixture marker, not a model response. ┃ + ┃ +────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┃ + Package Updates Available ┃ + Package updates are available. Run pi update --extensions ┃ + Packages: ┃ + - pi-sandbox ┃ + - pi-zentui ┃ + - github.com/wassname/pi-annotated-journal ↓ Jump to latest message · End │ + +tomorrow-night-bright, tomorrow-night-burns, twilight, vague, vesper, xcode-dark-hc + + + Warning: ⚠️ Network sandbox allows all domains because network.allowedDomains contains "*". Only use this intentionally; remove "*" to restore per-domain prompts. + + pi-better-compaction loaded • debug artifacts → /home/ubuntu/.pi/agent/artifacts/pi-better-compaction/sessions/01a084a5-a0f4-74c9-99fc-f9e5ae61adec/lifecycle/2026-09-09T05-32-14-754Z-lifecycle.json + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator-created recovery command fixture. The plan is a draft, not approved. Do not implement it. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Operator fixture marker, not a model response. + +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + Package Updates Available + Package updates are available. Run pi update --extensions + Packages: + - pi-sandbox + - pi-zentui + - github.com/wassname/pi-annotated-journal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Goals stopped. Use /goals resume in the worker when ready. Detached processes are not killed. + + Goals stopped by user. Plan and evidence retained. +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── +│ +│ +│ +│ gpt-6-astra Github Copilot minimal +───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── + trial-nAjUyF goals stopped · /goals resume | 🔒 Sandbox: all domains, 4 write paths | [░░░░░░░░░░] 0.0%/400k (auto) | $0.000 (sub) diff --git a/docs/reviews/evidence/2026-09-09-recovery/typecheck.log b/docs/reviews/evidence/2026-09-09-recovery/typecheck.log new file mode 100644 index 0000000..4d9a05e --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/typecheck.log @@ -0,0 +1,4 @@ + +> @wassname2/pi-goals@0.2.2 typecheck +> tsc -p tsconfig.build.json --noEmit + diff --git a/docs/reviews/evidence/2026-09-09-recovery/verification.log b/docs/reviews/evidence/2026-09-09-recovery/verification.log new file mode 100644 index 0000000..2f2a764 --- /dev/null +++ b/docs/reviews/evidence/2026-09-09-recovery/verification.log @@ -0,0 +1,8 @@ +PASS: stop persisted paused draft. +PASS: resume restored planning without Ready authorization. +PASS: exit persisted ordinary-chat state. +PASS: no assistant turn beyond operator fixture marker. +PASS: draft bytes unchanged. +PASS: actual reload rendered; stopped status retained. +PASS: reconnect without a pair reports failure instead of launching one. +PASS: fresh-shell --session retained exit; status reports ordinary chat. diff --git a/docs/slop/plans/20260909_autonomous-supervision-acceptance.md b/docs/slop/plans/20260909_autonomous-supervision-acceptance.md new file mode 100644 index 0000000..6d5362f --- /dev/null +++ b/docs/slop/plans/20260909_autonomous-supervision-acceptance.md @@ -0,0 +1,33 @@ +# Autonomous supervision and two-goal acceptance + +User requested the other branch's better instructions, manual-checkmark handling, and a completed real trial. Keep the current reliability base and VCC. Do not import Git cleanliness, commit, or ignored-output approval gates. + +- [ ] goal: Keep useful supervision running without invented human waits + - Adopt the latest `565b272` prompt structure: stage-specific check-ins, outcome-first tool guidance, a clear distinction between recap and sent instruction, applicable AGENTS/skills, and short review context versus full orientation. Keep current-plan freshness and checkpoint evidence. + - Preserve the earlier adopted behavior: investigate blockers, change ineffective steering, keep authorized independent work moving, inspect the result, and give brief visible judgments. + - Take the new VCC adapter safeguards: separate budgets for extracted context/recent actions, partial-call handling, and explicit omitted-result/reference notices. Keep the current recovery and coalescing implementation. + - Remove the assumption that ordinary prose without a verdict tool means a human decision is needed. New worker direction must reach the supervisor even after a real question. + - failure modes: normal prose silently stops reviews; alternatively, failure causes an immediate retry loop or the supervisor ignores an explicit user pause. + - deliverable: focused regressions for ordinary prose, empty/error responses, subsequent worker updates, genuine user pauses and no idle loop. +- [ ] goal: Treat manual completion marks as claims while retaining artifact-based review + - Keep unsigned manual `[x]` claims visible and supervised, including reload. CompleteGoal remains the sign-off path; preserve independent judge policy and current plan identity/cancellation protections. + - Git status is review context, not an acceptance gate. Ignored output files may be evidence. Do not force commits or cleanup. + - failure modes: a manual tick ends supervision; dirty/ignored artifacts are rejected merely due to Git status; legitimate signed-off goals reopen on reload. + - deliverable: widget/lifecycle/sign-off regressions, including dirty worktree and ignored-output cases. +- [ ] goal: Finish a real two-goal workflow with the chosen implementation + - Carry relevant user preferences and the other branch's actual-Herdr testing procedure into AGENTS.md. Ask material questions, not a quota; retain ordinary-chat Discuss. + - Run the functional trial with full normal Pi profiles, isolating only candidate pi-goals selection, and different worker/supervisor models. Inspect both panes and actual artifacts; do not substitute bundled-only loading or test counts for acceptance. + - Exercise worker/supervisor/both reloads, fresh-shell resume, drafting/Discuss, Ready/startup, pending completion and stopped sessions. Preserve plan/role/peer; interrupted decisions need a visible retry path, not stale acceptance, duplicate panes or a permanent wait. + - Diagnose failures from exact logs, fix them and retry; after prompt changes use a fresh task. Do not do the worker's artifact work for it. + - failure modes: only unit tests pass; one goal is left unsigned; fabricated logs pass as execution; operator repairs are described as autonomous success. + - deliverable: saved logs and pane evidence of both CompleteGoal results, actual output verification, interventions, remaining limitations, and separate role usage. + +## UAT / Verification + +- Success: both artifacts match the task, real verification output exists, both sign-offs are observed, and the same pair remains active between goals. +- Likely failure: startup/reload/approval stalls. Read both panes and the error, repair the cause, then repeat that stage. +- Sneaky failure: manual checkmarks or handwritten output appear complete. Inspect the underlying artifact and actual execution record, not just the widget or worker summary. +- Include an ignored output directory and preserved unrelated dirty file; neither should force a commit or block valid completion. +- Keep original sessions/settings untouched. Temporary changes to the three non-secret role preferences require guarded restoration. No release, merge, or active-installation replacement. + +Recorded by Pi (OpenAI). [Earlier implementation validation](../../reviews/2026-09-09_autonomy-validation.md) remains a source check, not functional acceptance. The user has since disabled the sandbox and authorized the latest prompt/VCC adoption. pi-subagents was upgraded from 0.62.0 to 0.66.0 with a command-only release-age exception; other direct package versions and the default policy were unchanged. After reload, prioritize the real two-goal trial. Review suggestions for the other branch are posted in [issue #6](https://github.com/wassname/pi-goals/issues/6). diff --git a/src/index.ts b/src/index.ts index 7a6b07a..0dd26ef 100644 --- a/src/index.ts +++ b/src/index.ts @@ -19,15 +19,16 @@ * and the always-present CompleteGoal description carries the contract instead. * 2. format — a skeleton convention taught in planDrafting (prompts.ts), not validated * 3. eyes — CompleteGoal spawns a strictly read-only pi subprocess (--no-session, no bash) - * that gets the whole plan file plus the claimed goal, finds the goal itself - * (tolerates wording drift), checks the evidence (including the agent's saved - * verify output) against the repo, and returns VERDICT: accept|reject + * that gets the whole plan file plus a unique exact goal subject, checks the + * evidence (including the agent's saved verify output) against the repo, and + * returns VERDICT: accept|reject. CompleteGoal rejects ambiguous or drifted + * subjects before review and asks the worker to retry with the exact subject. * - * The judge subsumes what v1 did in code: goal matching (no findGoal), evidence validation (a - * placeholder gets rejected in words), and format reading. The extension's only + * The judge reads the goal's format and validates evidence (a placeholder gets rejected in + * words); code binds sign-off to one exact goal identity. The extension's only * writes are the sign-off: append a log line to ## Log (the audit trail) and tick the goal [x] when - * an exact goal line matches (on drift the agent ticks, and the result says so). A hand-tick - * without a matching tool-written log line is visible in the diff either way. + * one unique exact goal subject matches. CompleteGoal persists conclusive/inconclusive sign-offs + * separately from checkboxes, so manual ticks stay visible claims even across reload. * * Judge ran but failed/errored/timed out, or returned no VERDICT line => accepted_inconclusive: the * working agent is never blocked on judge infra; the log line says the judge ran but failed. There @@ -137,8 +138,16 @@ interface PlanState { /** Distinguishes explicit preferences from the old opt-in defaults. */ defaultsVersion: 1; phase: Phase; + /** Recovery state is separate from ordinary auto-continue backoff. */ + pausedFrom?: "planning" | "working"; + resumeHash?: string; + exited?: boolean; reviewRequested: boolean; questionsWaived: boolean; + /** Only CompleteGoal records these; a plan checkbox alone is a claim. */ + signedOffGoals: Array<{ subject: string; outcome: "accept" | "inconclusive" }>; + /** Pre-tracking checkboxes stay historical claims, never automatic reimplementation work. */ + legacyCompletionClaims: string[]; /** Ready captured its fork, but worker model recovery is still pending (also across reload). */ modelRecovery: "worker" | null; /** Optional model ref for the sign-off judge; unset => current session model, else pi's default. */ @@ -159,6 +168,8 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { phase: null, reviewRequested: false, questionsWaived: false, + signedOffGoals: [], + legacyCompletionClaims: [], modelRecovery: null, judgeModel: null, planVersion: null, @@ -170,6 +181,11 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { let planningContextPending = false; let supervisorOnly = false; let operation: AbortController | null = null; + let reviewAbort = new AbortController(); + let workGeneration = 0; + let goalAbort = new AbortController(); + let latestContext: ExtensionContext; + let lastRecoveryError: string | undefined; const lifetime = new AbortController(); // The reminder sees only the working set. A repeated Log line must not look like progress. let turnsStale = 0; @@ -210,8 +226,29 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { autoTimer = null; } + function refreshSignoffs(plan: string): void { + const goals = scanGoals(plan); + const signedOffGoals = state.signedOffGoals.filter(signoff => { + const matches = goals.filter(goal => goal.subject.toLowerCase() === signoff.subject); + return matches.length === 1 && matches[0].status === "done"; + }); + const legacyCompletionClaims = state.legacyCompletionClaims.filter(subject => goals.some(goal => goal.subject.toLowerCase() === subject && goal.status === "done")); + if (signedOffGoals.length !== state.signedOffGoals.length || legacyCompletionClaims.length !== state.legacyCompletionClaims.length) { + state = { ...state, signedOffGoals, legacyCompletionClaims }; persist(); + } + } + + function signedOff(subject: string): boolean { + return state.signedOffGoals.some(signoff => signoff.subject === subject.toLowerCase()); + } + + function pendingGoals(plan: string) { + refreshSignoffs(plan); + return scanGoals(plan).filter(goal => goal.status !== "cancelled" && (goal.status !== "done" || !signedOff(goal.subject))); + } + function activeGoals(ctx: ExtensionContext): boolean { - return scanGoals(readPlan(ctx)).some((goal) => goal.status === "active" || goal.status === "open"); + return pendingGoals(readPlan(ctx)).some(goal => !state.legacyCompletionClaims.includes(goal.subject.toLowerCase())); } function scheduleAutoContinue(ctx: ExtensionContext, delayMs = state.autoIntervalMs): void { @@ -263,6 +300,11 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { function updateWidget(ctx: ExtensionContext): void { const tools = pi.getActiveTools().filter(tool => tool !== "RequestPlanReview"); pi.setActiveTools(state.phase === "planning" && !supervisorOnly ? [...tools, "RequestPlanReview"] : tools); + if (state.pausedFrom || state.exited) { + ctx.ui.setStatus(STATUS_KEY, state.exited ? undefined : "goals stopped · /goals resume"); + ctx.ui.setWidget(WIDGET_KEY, state.exited ? undefined : ["Goals stopped by user. Plan and evidence retained."]); + return; + } if (state.phase === "planning") { ctx.ui.setStatus(STATUS_KEY, ctx.ui.theme.fg("warning", state.modelRecovery ? models.ready ? "retry Ready" : "worker model paused" : "planning")); ctx.ui.setWidget(WIDGET_KEY, ["pi-goals: drafting goals"]); @@ -273,23 +315,29 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { ctx.ui.setWidget(WIDGET_KEY, ["pi-goals: starting the supervisor session"]); return; } - const goals = scanGoals(readPlan(ctx)); + const plan = readPlan(ctx); + refreshSignoffs(plan); + const goals = scanGoals(plan); if (goals.length === 0) { ctx.ui.setStatus(STATUS_KEY, undefined); ctx.ui.setWidget(WIDGET_KEY, undefined); return; } - const done = goals.filter((g) => g.status === "done").length; + const done = goals.filter((g) => g.status === "done" && signedOff(g.subject)).length; + const claimed = goals.filter(g => g.status === "done" && !signedOff(g.subject)); const auto = state.autoPaused ? " · waiting for user" : state.autoIntervalMs === null ? "" : ` · auto ${state.autoIntervalMs / 60_000}m`; - const steward = state.stewardEnabled ? " · supervisor" : ""; + const steward = (state.stewardEnabled ? " · supervisor" : "") + (claimed.length ? ` · ${claimed.length} claimed, awaiting review` : ""); ctx.ui.setStatus(STATUS_KEY, ctx.ui.theme.fg("accent", `◷ ${done}/${goals.length} goals${auto}${steward}`)); const mark: Record = { done: "✔", active: "▸", open: "◻", cancelled: "✗" }; // Only live goals get lines so finished work never pushes current work off screen. The active // goal also shows its open subtasks: this file is the task list, so the widget is the task list. // No path line: the session id makes it 47 chars, too long to be worth a widget row. The // human opens the file from the Ready menu, and every injected reminder still names it. - const plan = readPlan(ctx); const lines: string[] = state.autoPaused ? [ctx.ui.theme.fg("warning", "⏸ waiting for user")] : []; + lines.push(...claimed.map(g => state.legacyCompletionClaims.includes(g.subject.toLowerCase()) + ? `? legacy completion — sign-off not recorded: ${g.subject}` + : `? claimed complete; awaiting CompleteGoal: ${g.subject}`)); + lines.push(...state.signedOffGoals.filter(s => s.outcome === "inconclusive").map(s => `? accepted inconclusive (not verified): ${s.subject}`)); for (const g of goals.filter((g) => g.status === "active" || g.status === "open")) { lines.push(`${mark[g.status]} ${g.subject}`); if (g.status === "active") lines.push(...openSubtasks(plan, g.line).slice(0, 3).map((s) => ctx.ui.theme.fg("muted", ` ◦ ${s}`))); @@ -319,6 +367,15 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { state = { ...state, phase: "starting" }; persist(); updateWidget(ctx); try { + // A stopped, never-attached bootstrap cannot be rejoined. A new explicit Ready + // may replace it; reconnect/resume never create another fork. + if (state.supervisor) { + const previous = await supervisor.status(signal); + if (previous.binding?.paused && previous.role === "none") { + await supervisor.stop(state.supervisor.id); + state = { ...state, supervisor: null }; persist(); + } + } if (recoveringWorker) { const workerReady = await models.enter("worker", ctx); if (signal.aborted || state.planVersion !== version || !state.stewardEnabled) return; @@ -343,7 +400,9 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { if (signal.aborted || state.planVersion !== version || state.supervisor?.id !== binding.id || !state.stewardEnabled) return; if (planHash(readPlan(ctx)) !== approvedDraft) { await supervisor.stop(binding.id); state = { ...state, supervisor: null }; throw new Error("The plan changed during model restoration; select Ready again"); } state = { ...state, modelRecovery: null }; persist(); - await supervisor.activate(binding.id, signal); + const peerStatus = await supervisor.status(signal); + if (peerStatus.binding?.paused) await supervisor.resume(binding.id, approvedDraft, signal); + else await supervisor.activate(binding.id, signal); if (signal.aborted || state.planVersion !== version || state.supervisor?.id !== binding.id || !state.stewardEnabled) return; if (planHash(readPlan(ctx)) !== approvedDraft) { await supervisor.stop(binding.id); state = { ...state, supervisor: null }; throw new Error("The plan changed during activation; select Ready again"); } state = { ...state, phase: "working" }; @@ -354,7 +413,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { state = { ...state, phase: "planning", modelRecovery: null }; persist(); updateWidget(ctx); ctx.ui.notify(`Could not initialize the supervisor: ${String(error)}. Use /goals supervisor to inspect startup, or /goals steward off and retry Ready.`, "error"); } finally { - if (!lifetime.signal.aborted && state.phase !== "working" && !state.modelRecovery) { + if (!lifetime.signal.aborted && state.phase !== "working" && !state.modelRecovery && !state.pausedFrom && !state.exited) { await models.enter("planning", ctx); if (!state.phase) models.leave(); } @@ -362,11 +421,99 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { } } - // --- /goals: enter plan mode (or clear / set judge / set steward) ------------------------------- + function pauseGoals(ctx: ExtensionContext, exit: boolean): void { + workGeneration++; + goalAbort.abort(); goalAbort = new AbortController(); + operation?.abort(); operation = null; + reviewAbort.abort(); reviewAbort = new AbortController(); + clearAutoTimer(); + planningContextPending = false; resyncReason = null; + if (!supervisorOnly) { + state = { ...state, pausedFrom: state.pausedFrom ?? (state.phase === "working" ? "working" : state.planVersion !== null ? "planning" : undefined), + resumeHash: state.resumeHash ?? (state.phase === "working" ? planHash(readPlan(ctx)) : undefined), + phase: null, reviewRequested: false, modelRecovery: null, autoPaused: true, exited: exit }; + persist(); updateWidget(ctx); models.leave(); + } + } + + const recoveryHelp = [ + "/goals status — phase, connection, recorded sessions and last failure", + "/goals stop — stop goal work and supervision; keep plan, evidence and pair", + "/goals exit — leave goals/planning mode for chat; keep files; no approval", + "/goals reconnect — reconnect the existing pair; never launch or authorize work", + "/goals resume — worker only: resume authorized work, or return a draft to planning (Ready still required)", + "/goals supervisor | worker | zoom — focus the recorded pane", + "/goals clear — disconnect the plan; keep its file", + "/goals plan — start a new draft; free-text objectives still work", + "Stop/exit do not kill detached processes. A disconnected peer is unconfirmed; stop it in its own pane. Supervisor chat stays inspection-only.", + ].join("\n"); + + async function recoveryCommand(arg: string, ctx: ExtensionContext): Promise { + const verb = arg.split(/\s+/)[0]; + if (!["help", "status", "stop", "exit", "reconnect", "resume"].includes(verb)) return false; + if (arg !== verb) { ctx.ui.notify(`Use /goals ${verb} without arguments. To draft an objective, use /goals plan .`, "warning"); return true; } + if (verb === "help") { ctx.ui.notify(recoveryHelp, "info"); return true; } + if (verb === "stop" || verb === "exit") { + pauseGoals(ctx, verb === "exit"); + try { supervisor.pause(verb === "exit"); } catch (error) { lastRecoveryError = String(error); ctx.ui.notify(`Stopped locally; peer unconfirmed: ${error}`, "warning"); } + ctx.abort(); + ctx.ui.notify(verb === "exit" ? "Goals mode exited. Files retained; no work approved. Supervisor sessions remain inspection-only." : "Goals stopped. Use /goals resume in the worker when ready. Detached processes are not killed.", "info"); + return true; + } + try { + if (verb === "status") { + const status = await supervisor.status(); + const binding = status.binding ?? state.supervisor ?? supervisorBootstrap(ctx)?.binding; + ctx.ui.notify([`Goals: ${supervisorOnly ? "supervisor" : state.exited ? "ordinary chat (goals exited)" : state.pausedFrom ? "stopped by user" : state.phase ?? "ordinary chat"}.`, + `Pair: ${status.connected ? "connected" : "disconnected/unconfirmed"}; ${status.activity ?? "inactive"}.`, + `Worker: ${binding?.workerPane ?? "not recorded"} · ${binding?.workerSession ?? "no paired session"}`, + `Supervisor: ${binding?.supervisorPane ?? "not recorded"} · ${binding?.supervisorSession ?? "no paired session"}`, + `Last failure: ${lastRecoveryError ?? status.lastFailure ?? "none recorded in this runtime"}`].join("\n"), "info"); + } else if (verb === "reconnect") { + await supervisor.reconnect(lifetime.signal); + ctx.ui.notify("Reconnect handshake sent to the existing pair. No pane created or work authorized. Use /goals status to inspect connectivity.", "info"); + } else if (supervisorOnly) { + ctx.ui.notify("Resume must be authorized in the worker: /goals worker, then /goals resume. Reconnect alone never starts work.", "info"); + } else if (!state.pausedFrom) { + ctx.ui.notify("No user-stopped plan to resume. Use /goals status; a draft needs Ready.", "info"); + } else if (state.pausedFrom === "planning" || state.resumeHash !== planHash(readPlan(ctx))) { + const generation = workGeneration; + if (!await models.enter("planning", ctx) || generation !== workGeneration || lifetime.signal.aborted) return true; + state = { ...state, phase: "planning", pausedFrom: undefined, resumeHash: undefined, exited: false, reviewRequested: false }; + planningContextPending = true; persist(); updateWidget(ctx); + ctx.ui.notify("Draft restored; no work started. Review the plan and request Ready before working.", "info"); + } else { + if (operation) throw new Error("Recovery already in progress"); + const controller = new AbortController(); operation = controller; + const signal = AbortSignal.any([controller.signal, lifetime.signal]); + const generation = workGeneration; + const hash = state.resumeHash; + try { + if (!await models.enter("worker", ctx)) return true; + signal.throwIfAborted(); + if (state.stewardEnabled) { + if (!state.supervisor) throw new Error("No recorded pair; return to planning and choose Ready, or turn the steward off explicitly."); + await supervisor.resume(state.supervisor.id, hash, signal); + } + signal.throwIfAborted(); + if (generation !== workGeneration || hash !== planHash(readPlan(ctx))) throw new Error("Plan changed during resume; review before working."); + state = { ...state, phase: "working", pausedFrom: undefined, resumeHash: undefined, exited: false, autoPaused: false }; + persist(); updateWidget(ctx); + pi.sendUserMessage(workMessage(ctx), { deliverAs: "followUp" }); + } finally { if (operation === controller) operation = null; } + } + } catch (error) { lastRecoveryError = String(error); ctx.ui.notify(`Recovery failed; no new work authorized: ${error}`, "warning"); } + return true; + } + + // --- /goals: plan entry and explicit human recovery ----------------------------------------- pi.registerCommand("goals", { - description: `Plan mode: draft goals into ${PLAN_SHAPE}, review, then work them. /goals plan | /goals model current | /goals supervisor | /goals worker | /goals zoom | /goals | /goals clear | /goals auto [minutes|off] | /goals judge | /goals steward [on|off|status]`, + description: `Plan mode: draft goals into ${PLAN_SHAPE}, review, then work them. /goals help | status | stop | exit | reconnect | resume | /goals plan | /goals model current | /goals supervisor | /goals worker | /goals zoom | /goals | /goals clear | /goals auto [minutes|off] | /goals judge | /goals steward [on|off|status]`, + getArgumentCompletions: (prefix) => ["help", "status", "stop", "exit", "reconnect", "resume", "supervisor", "worker", "zoom", "clear", "plan", "steward", "auto"].filter(value => value.startsWith(prefix)).map(value => ({ value, label: value })), handler: async (args, ctx) => { + latestContext = ctx; + if (await recoveryCommand(args.trim(), ctx)) return; if (args.trim() === "model current") { if (await models.useCurrent(ctx)) modelRecovered(ctx); return; @@ -397,6 +544,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { state = { ...state, phase: null, + pausedFrom: undefined, resumeHash: undefined, exited: false, modelRecovery: null, planVersion: null, autoPaused: false, @@ -477,14 +625,20 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { ctx.ui.notify(ref ? `Sign-off judge model set to ${ref}` : "Sign-off judge reset to the session model", "info"); return; } + if (!explicitPlan && (/^(?:--|\/)/.test(arg) || /^(?:connect|disconnect|pause|quit|restart|reset|resum|reconect|stpo|stats)(?:\s|$)/i.test(arg))) { + ctx.ui.notify("Unknown recovery command. Use /goals help; for an objective use /goals plan .", "warning"); return; + } await stopSupervisor(ctx); if (!await models.enter("planning", ctx)) return; state = { ...state, phase: "planning", + pausedFrom: undefined, resumeHash: undefined, exited: false, modelRecovery: null, reviewRequested: false, questionsWaived: waivesAlignment(arg), + signedOffGoals: [], + legacyCompletionClaims: [], planVersion: nextVersion(ctx), supervisor: null, }; @@ -512,7 +666,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { resyncReason = null; return why; }; - if (state.phase === "planning" || state.phase === "starting") return null; + if (state.phase !== "working" || state.pausedFrom || state.exited) return null; if (!plan.trim()) return null; const why = drainResync(); if (why) return resync(plan, planRel(ctx), why); @@ -523,7 +677,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { // widget, no injection, no reminders). Say so instead -- cooperative but confused. return `\n${planRel(ctx)} exists but has no goal line pi-goals recognizes. A goal is a checkbox list line starting "goal:", e.g. "1. [ ] goal: " ([ ] open, [/] active, [x] done, [-] cancelled). Reformat it if it's meant to be the plan.\n`; } - if (!goals.some((g) => g.status === "active" || g.status === "open")) return null; + if (!pendingGoals(plan).length) return null; return reminder(foldPlan(plan), planRel(ctx)); } @@ -553,11 +707,13 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { // PI: Human plan-mode replies are durable evidence of the interview, not model summaries. pi.on("input", async (event, ctx) => { + if ((state.pausedFrom || state.exited) && event.source === "extension" && + (event.text.startsWith("Work the goals in ") || event.text.startsWith("Auto-continue") || event.text.startsWith("[supervisor]") || event.text === discussPlan || event.text.startsWith("We're in plan mode."))) return { action: "handled" as const }; if (!models.ready) { ctx.ui.notify("Role model unavailable. Select a different model with /model, explicitly use the current one with /goals model current, or configure the saved model and reload.", "error"); return { action: "handled" as const }; } if (event.source !== "extension") { clearAutoTimer(); autoImmediateUsed = false; - if (state.autoPaused) { + if (state.autoPaused && !state.pausedFrom && !state.exited) { state = { ...state, autoPaused: false }; persist(); updateWidget(ctx); @@ -628,7 +784,10 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { printed = plan; pi.sendMessage({ customType: "plan", content: plan, display: true }); } - const choice = await ctx.ui.select(`Plan drafted in ${planRel(ctx)}.`, ["Ready", "Discuss", "Edit", "Cancel"]); + const version = state.planVersion; + const generation = workGeneration; + const choice = await ctx.ui.select(`Plan drafted in ${planRel(ctx)}.`, ["Ready", "Discuss", "Edit", "Cancel"], { signal: reviewAbort.signal }); + if (state.phase !== "planning" || generation !== workGeneration || version !== state.planVersion || lifetime.signal.aborted) return; if (choice === "Discuss" || choice === undefined) { if (state.modelRecovery && !await models.enter("planning", ctx)) return; state = { ...state, modelRecovery: null }; @@ -639,6 +798,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { } if (choice === "Edit") { const edited = await ctx.ui.editor("Edit the plan", plan); + if (generation !== workGeneration || version !== state.planVersion || state.phase !== "planning") return; if (edited !== undefined && edited !== plan) writePlan(ctx, edited); continue; } @@ -663,7 +823,6 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { await startPlanSupervisor(ctx); return; } - const version = state.planVersion; const approvedDraft = planHash(readPlan(ctx)); state = { ...state, modelRecovery: "worker" }; persist(); const workerReady = await models.enter("worker", ctx); @@ -685,6 +844,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { pi.on("agent_settled", async (_event, ctx) => reviewPlan(ctx)); pi.on("session_start", async (_event, ctx) => { + latestContext = ctx; const bootstrap = supervisorBootstrap(ctx); if (bootstrap) { supervisorOnly = true; @@ -693,7 +853,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { return; } const last = ctx.sessionManager - .getEntries() + .getBranch() .filter((e: { type?: string; customType?: string }) => e.type === "custom" && e.customType === STATE) .pop() as { data?: PlanState } | undefined; // Upgrade cleared/unused legacy sessions, but never attach supervision mid-plan. @@ -701,9 +861,12 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { const useNewDefaults = saved?.defaultsVersion !== 1 && saved?.planVersion == null; state = { defaultsVersion: 1, - phase: last?.data?.phase === "working" ? "working" : last?.data?.phase ? "planning" : null, + phase: saved?.pausedFrom || saved?.exited ? null : saved?.phase === "working" ? "working" : saved?.phase ? "planning" : null, + pausedFrom: saved?.pausedFrom, resumeHash: saved?.resumeHash, exited: saved?.exited, reviewRequested: last?.data?.reviewRequested ?? true, questionsWaived: last?.data?.questionsWaived ?? false, + signedOffGoals: last?.data?.signedOffGoals ?? [], + legacyCompletionClaims: last?.data?.legacyCompletionClaims ?? [], modelRecovery: last?.data?.modelRecovery ?? null, judgeModel: last?.data?.judgeModel ?? null, planVersion: last?.data?.planVersion ?? null, @@ -712,6 +875,10 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { stewardEnabled: useNewDefaults ? true : saved?.stewardEnabled ?? true, supervisor: last?.data?.supervisor ?? null, }; + if (saved && saved.signedOffGoals === undefined) { + state.legacyCompletionClaims = scanGoals(readPlan(ctx)).filter(goal => goal.status === "done").map(goal => goal.subject.toLowerCase()); + persist(); + } lastSeenWorkingSet = foldPlan(readPlan(ctx)); autoLastWorkingSet = lastSeenWorkingSet; planningContextPending = state.phase === "planning" || state.phase === "starting"; @@ -723,6 +890,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { pi.on("session_shutdown", async () => { lifetime.abort(); + goalAbort.abort(); reviewAbort.abort(); operation?.abort(); clearAutoTimer(); }); @@ -730,7 +898,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { pi.registerTool({ name: "RequestPlanReview", label: "Review plan", - description: "Planning only: after task-specific alignment questions have been answered (or explicitly waived for this objective), and the final plan is ready, show the human Ready / Discuss / Edit / Cancel. Do not call while waiting for answers. Call again after discussion is finished, even for an unchanged draft. This does not approve or start work.", + description: "Planning only: after material unresolved questions have been answered (no fixed quota), and the final plan is ready, show the human Ready / Discuss / Edit / Cancel. Do not call while waiting for answers. Call again after discussion is finished, even for an unchanged draft. This does not approve or start work.", parameters: Type.Object({}), async execute(_id, _params, _signal, _update, ctx) { if (supervisorOnly || state.phase !== "planning") return result("Only a planning session can request plan review.", true); @@ -750,6 +918,9 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { goal: Type.String({ description: completeGoalParamDescription }), }), async execute(_id, params, signal, onUpdate, ctx) { + signal = AbortSignal.any([lifetime.signal, goalAbort.signal, ...(signal ? [signal] : [])]); + if (state.pausedFrom || state.exited) return result("Goals stopped. Resume explicitly before signing off.", true); + const generation = workGeneration; if (state.phase === "planning" || state.phase === "starting") return result("Planning is not approved. Wait for the steward or choose Ready before signing off a goal.", true); let plan = readPlan(ctx); if (!plan.trim()) return result(`No plan file at ${planRel(ctx)}. Run /goals to draft one.`, true); @@ -759,7 +930,9 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { // A model may tick before calling this tool. The submitted goal is still under review; // a rejection or cancellation must not leave that premature success visible. const submitted = scanGoals(plan).filter(goal => goal.subject.toLowerCase() === params.goal.trim().toLowerCase()); - if (submitted.length === 1 && submitted[0].status === "done") { + if (submitted.length !== 1) return result("CompleteGoal requires one unique exact goal subject (ignoring case and surrounding whitespace). Copy the text after 'goal:' from the current plan; give duplicate goals distinct subjects before retrying. No sign-off recorded.", true); + refreshSignoffs(plan); + if (submitted[0].status === "done") { const lines = plan.split("\n"); lines[submitted[0].line] = lines[submitted[0].line].replace(/\[[xX]\]/, "[/]"); plan = lines.join("\n"); @@ -778,6 +951,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { } catch (error) { return result(`Supervisor review failed: ${String(error)}`, true); } } + if (generation !== workGeneration) return result("Sign-off stopped; no goal signed off.", true); const reviewedPlanHash = planHash(plan); const reviewedVersion = state.planVersion; const reviewedPairing = state.supervisor?.id; @@ -801,21 +975,24 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { writeFileSync(join(ctx.cwd, rel), `goal: ${params.goal}\nmodel: ${judgeModel ?? "pi default"}\nerror: ${raw.error ?? "none"}\n\n${raw.output}\n`); transcriptNote = ` (${rel})`; } + if (signal?.aborted || lifetime.signal.aborted || generation !== workGeneration) return result("Sign-off aborted; no goal was signed off.", true); if (state.planVersion !== reviewedVersion || state.supervisor?.id !== reviewedPairing || planHash(readPlan(ctx)) !== reviewedPlanHash) return result("The plan or supervisor changed during evidence review; no goal was signed off. Retry.", true); if (outcome.logEntry) { - // Sign-off write: tick the goal [x] (exact-subject match; dogfood showed agent bookkeeping - // is the drift point) and append the audit log line, one write. On wording drift the tick - // falls to the agent and the result says so -- both paths are explicit, never silent. let updated = readPlan(ctx); let tickNote = ""; - if (outcome.logEntry.startsWith("signed off")) { + const signedOutcome = outcome.status === "accept" || outcome.status === "inconclusive" ? outcome.status : undefined; + if (signedOutcome) { const ticked = tickGoal(updated, params.goal); - updated = ticked ?? updated; - tickNote = ticked - ? `\n\nGoal ticked [x] in ${planRel(ctx)}.` - : `\n\nNo exact goal line matched your wording -- tick it [x] in ${planRel(ctx)} yourself.`; + if (!ticked) return result("Goal identity changed; retry with one unique exact goal subject. No sign-off recorded.", true); + updated = ticked; + tickNote = `\n\nGoal ticked [x] in ${planRel(ctx)}.`; } writePlan(ctx, appendLog(updated, `${stamp()} ${outcome.logEntry}${transcriptNote}`)); + if (signedOutcome) { + const subject = submitted[0].subject.toLowerCase(); + state = { ...state, signedOffGoals: [...state.signedOffGoals.filter(s => s.subject !== subject), { subject, outcome: signedOutcome }] }; + persist(); + } updateWidget(ctx); return result(outcome.resultText + tickNote, outcome.isError); } @@ -823,7 +1000,15 @@ export default function piGoalsExtension(pi: ExtensionAPI): void { }, }); // Registered after role restoration, so rejoin cannot start a supervisor turn on the worker model. - const supervisor = supervise(pi, () => models.ready); + const supervisor = supervise(pi, () => models.ready, (plan) => { + const pending = pendingGoals(plan); + const claims = pending.filter(goal => goal.status === "done"); + const inconclusive = state.signedOffGoals.filter(signoff => signoff.outcome === "inconclusive"); + return { + completion: { planHash: planHash(plan), total: scanGoals(plan).length, pending: pending.length, inconclusive: inconclusive.length }, + summary: `CompleteGoal records: ${state.signedOffGoals.length - inconclusive.length} conclusive, ${inconclusive.length} accepted inconclusive (not verified); ${pending.length} goals await sign-off.\nUnsigned completion claims: ${claims.map(goal => goal.subject).join(", ") || "none"}. Legacy completions without tracking: ${state.legacyCompletionClaims.join(", ") || "none"}; preserve their history and use CompleteGoal re-review if needed, not automatic reimplementation. A checkbox is not proof; inspect the current plan and artifacts.`, + }; + }, exit => { if (latestContext) pauseGoals(latestContext, exit); }); function modelRecovered(ctx: ExtensionContext): void { if (supervisorOnly) { const bootstrap = supervisorBootstrap(ctx); @@ -879,6 +1064,7 @@ export interface SignOffInput { /** The outcome of a sign-off: the reply text, whether it's a hard error, and the one ## Log line to * append (null when nothing should be written, e.g. aborted before any verdict). */ export interface SignOffOutcome { + status: "accept" | "reject" | "inconclusive" | "aborted"; resultText: string; isError: boolean; logEntry: string | null; @@ -898,13 +1084,14 @@ export async function decideSignOff( const task = judgeUser({ goal: input.goal, plan: input.plan, planPath: input.planRel }); const judge = await runJudgeFn(task); - if (signal?.aborted) return { resultText: "Sign-off aborted.", isError: true, logEntry: null }; + if (signal?.aborted) return { status: "aborted", resultText: "Sign-off aborted.", isError: true, logEntry: null }; // Judge ran but failed/errored/timed out: fail forward, say so in the log. if (judge.error) { const partial = judge.output ? `\n\npartial judge output:\n${judge.output}` : ""; return { - resultText: `Judge ran but failed (${judge.error}). Accepted inconclusive — logged.${partial}`, + status: "inconclusive", + resultText: `Judge ran but failed (${judge.error}). Accepted inconclusive — logged; this is not verified completion.${partial}`, isError: false, logEntry: `signed off "${input.goal}" (judge inconclusive: ran but failed: ${oneLine(judge.error)})`, }; @@ -923,12 +1110,14 @@ export async function decideSignOff( const checks = /^[ \t]*(?:[-*]|\d+[.)])[ \t]+\S.*$/m.test(checksBody); if (!checks) { return { + status: "reject", resultText: `Sign-off REJECTED. Missing:\nchecked-artifact list before VERDICT: accept\n\n--- judge ---\n${reasoning}`, isError: true, logEntry: `reject "${input.goal}": judge accept had no checked-artifact list`, }; } return { + status: "accept", resultText: `Sign-off ACCEPTED (log line appended).\n\n--- judge ---\n${reasoning}`, isError: false, logEntry: `signed off "${input.goal}" (judge accept)`, @@ -937,6 +1126,7 @@ export async function decideSignOff( if (verdict === "reject") { const missing = judge.output.match(/missing\s*:\s*([\s\S]*)$/i)?.[1].trim() || judge.output.slice(-500); return { + status: "reject", resultText: `Sign-off REJECTED. Missing:\n${missing}\n\n--- judge ---\n${reasoning}`, isError: true, logEntry: `reject "${input.goal}": ${oneLine(missing)}`, @@ -944,15 +1134,16 @@ export async function decideSignOff( } // No VERDICT line: same fail-forward as a judge error -- the judge ran but didn't answer. return { - resultText: `Judge returned no VERDICT line. Accepted inconclusive — logged.\n\n--- judge ---\n${reasoning || "(no output)"}`, + status: "inconclusive", + resultText: `Judge returned no VERDICT line. Accepted inconclusive — logged; this is not verified completion.\n\n--- judge ---\n${reasoning || "(no output)"}`, isError: false, logEntry: `signed off "${input.goal}" (judge inconclusive: no VERDICT line)`, }; } /** Tick the goal line whose subject exactly matches `goal` (trimmed, case-insensitive) to [x]. - * Null when there is no unique exact match (wording drift / duplicates) -- the caller then asks the - * agent to tick it itself. Reuses GOAL_LINE; deliberately NOT fuzzy, that's the judge's job. */ + * Null when there is no unique exact match (wording drift / duplicates). CompleteGoal refuses + * ambiguous identity before review; this helper never chooses a different goal. */ export function tickGoal(plan: string, goal: string): string | null { const lines = plan.split("\n"); const want = goal.trim().toLowerCase(); diff --git a/src/internal/supervisor/index.ts b/src/internal/supervisor/index.ts index 2203e89..e6795d5 100644 --- a/src/internal/supervisor/index.ts +++ b/src/internal/supervisor/index.ts @@ -7,6 +7,7 @@ * triggers its own turn locally with pi.sendUserMessage. The broker stamps fromSessionId from its * own registry, so pairing on that ID cannot be forged by a payload. */ +import { randomUUID } from "node:crypto"; import { readFileSync } from "node:fs"; import { resolve } from "node:path"; import type { ExtensionContext } from "@earendil-works/pi-coding-agent"; @@ -18,7 +19,7 @@ import { Type } from "typebox"; import { loadBundledIntercom } from "../../intercom.js"; import { type Bootstrap, pendingReply, planHash, planText, type SupervisorBinding, type SupervisorController, validBinding } from "../../supervisor.js"; import { backgroundState } from "./background.js"; -import type { GoalDecision, GoalReview, GoalReviewRequest } from "./protocol.js"; +import type { GoalDecision, GoalReview, GoalReviewRequest, PlanCompletion } from "./protocol.js"; /** * Public event names from the bundled pi-intercom/extension-api.ts protocol. The channel types @@ -53,6 +54,7 @@ import { STEER_ACK, TOOL_DONE, TOOL_LET_IT_RUN, + TOOL_REVIEW_GOAL, TOOL_STEER, VIEW_PRUNED, } from "./prompts.js"; @@ -166,7 +168,7 @@ const INSPECTION_TOOLS = new Set(["read", "grep", "find", "ls"]); */ const SUPERVISOR_TOOLS = ["worker_view", "set_goal", "steer", "let_it_run", "done", "review_goal"]; -export default function (pi: any, modelReady: () => boolean = () => true): SupervisorController { +export default function (pi: any, modelReady: () => boolean = () => true, planProgress?: (plan: string) => { completion: PlanCompletion; summary: string }, onPause?: (exit: boolean) => void): SupervisorController { let duplicateIntercom = false; let channel: IntercomExtensionChannel | undefined; let sessionInitialized = false; @@ -187,12 +189,15 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super let peerDisconnected = false; let activeAssessment: "view" | "goal" | undefined; let assessmentFailure: string | undefined; + let lastFailure: string | undefined; + let resuming: ReturnType> | undefined; let assessmentEmpty = false; let routineDirty = false; let refreshInFlight = false; let advancing = false; let checkpointQueued = false; - let awaitingUser = false; + let latestCompletion: PlanCompletion | undefined; + let lastPublishedPlanHash: string | undefined; let lookPending = false; let publishAgain = false; let publishingLeaf: string | number | undefined; @@ -258,11 +263,13 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super } else reject(error); }, })); + showStatus(); try { await compacting; } - finally { compacting = undefined; } + finally { compacting = undefined; showStatus(); } } function cancelGoalRequests(reason: string) { + resuming?.finish(undefined, new Error(reason)); resuming = undefined; attached?.finish(undefined, new Error(reason)); attached = undefined; outgoingReview?.finish(undefined, new Error(reason)); outgoingReview = undefined; if (incomingReview && state.pairedId && channel?.snapshot().connected) send({ t: "goal_cancel", to: state.pairedId, requestId: incomingReview.requestId, bindingId: incomingReview.bindingId }); @@ -294,10 +301,51 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super const controller: SupervisorController = { async status(signal) { - await ready(signal, true); - const live = await channel!.listSessions(); - return { workerId: await resolveOwnId(), role: state.role, binding: state.plan, - connected: state.role === "worker" && !!state.plan && live.some(s => s.id === state.pairedId) }; + signal?.throwIfAborted(); + if (duplicateIntercom) throw new Error("Multiple Intercom runtimes are loaded. Keep one Intercom installation and reload."); + const available = !!channel?.snapshot().connected && !!channel.snapshot().supported; + const live = available ? await channel!.listSessions() : []; + return { workerId: available ? await resolveOwnId() : ownId, role: state.role, binding: state.plan, + connected: available && !!state.pairedId && !peerDisconnected && live.some(s => s.id === state.pairedId), + activity: state.plan?.stopped ? "ended" : state.plan?.paused ? "stopped by user" : compacting ? "compacting" : bootstrapPending ? "starting" : activeAssessment ? "reviewing" : refreshInFlight ? "requesting overview" : state.plan?.active ? "monitoring" : "inactive", + lastFailure }; + }, + pause(exit = false) { + pauseLocally(exit); + if (state.plan && state.pairedId) { + try { send({ t: "plan_pause", to: state.pairedId, bindingId: state.plan.id, exit, pauseId: state.plan.pauseId! }); } + catch (error) { lastFailure = String(error); } + ctx.ui.notify("Stopped locally. Peer stop requested but not confirmed; inspect the other pane. Independently running processes are not killed.", "warning"); + } + }, + async reconnect(signal) { + await ready(signal); + if (!state.plan || state.plan.stopped || !state.pairedId) throw new Error("No recoverable pair. Reconnect never creates a supervisor; inspect the recorded pane or select Ready for a new pairing."); + await rejoinOrDrop(); + signal?.throwIfAborted(); + }, + async resume(bindingId, hash, signal) { + await ready(signal); + const plan = workerPlan(bindingId); + if (plan.stopped) throw new Error("This pairing ended; it cannot be resumed. Review the plan with Ready."); + if (hash !== planHash(planText(plan))) throw new Error("Plan changed before resume; review it with Ready."); + if (resuming) throw new Error("Resume already pending"); + const generation = pairingGeneration; + const pending = pendingReply(signal, () => {}, 10_000); + resuming = pending; + try { + send({ t: "plan_resume", to: state.pairedId, bindingId, requestId: pending.requestId, planHash: hash, pauseId: plan.pauseId }); + if (!await pending.promise) throw new Error("Peer rejected resume; inspect its state and the plan."); + signal?.throwIfAborted(); + if (generation !== pairingGeneration || state.plan?.id !== bindingId || hash !== planHash(planText(state.plan))) throw new Error("Plan or pairing changed during resume"); + state = { ...state, plan: { ...state.plan, paused: false, active: true } }; save(); + startWatch(); + } catch (error) { + pending.finish(false); + // A lost acknowledgement must not leave the worker running or the peer silently active. + controller.pause(); + throw error; + } finally { if (resuming === pending) resuming = undefined; } }, async prepare(binding, signal) { await ready(signal, true); @@ -314,6 +362,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super if (state.role === "supervisor" && state.plan?.id === binding.id && state.planInitialized && state.pairedId) { await rejoinOrDrop(); return state.plan; } + if (state.plan?.id === binding.id && state.plan.paused) return binding; bootstrapPending = true; const generation = ++pairingGeneration; const cancelled = () => stopping || generation !== pairingGeneration || state.plan?.id !== binding.id || state.plan.stopped || signal?.aborted; @@ -354,6 +403,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super }, async activate(bindingId, signal) { await ready(signal); workerPlan(bindingId); + if (state.plan?.paused || state.plan?.stopped) throw new Error("Pairing stopped; use explicit /goals resume from the worker."); state = { ...state, plan: { ...state.plan!, active: true } }; save(); send({ t: "plan_activate", to: state.pairedId, bindingId }); startWatch(); @@ -362,6 +412,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super async review(bindingId, goal, hash, signal) { await ready(signal); const plan = workerPlan(bindingId); + if (plan.paused || plan.stopped) throw new Error("Goal work stopped; resume explicitly before sign-off."); if (outgoingReview) throw new Error("A goal review is already pending"); if (!goal || hash !== planHash(planText(plan))) throw new Error("The plan changed before goal review"); const peer = state.pairedId; @@ -408,13 +459,14 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super */ function showStatus() { if (!ctx?.hasUI) return; + if (state.plan?.paused) { ctx.ui.setStatus(STATUS_ID, "supervision stopped by user"); return; } if (state.role === "none") { ctx.ui.setStatus(STATUS_ID, undefined); return; } if (peerDisconnected || !channel?.snapshot().connected) { ctx.ui.setStatus(STATUS_ID, "supervision disconnected"); return; } const text = state.role === "supervisor" - ? awaitingUser ? "waiting for user" : refreshInFlight ? "waiting for worker overview" : `watching ${state.steerRounds}${routineDirty ? " · update pending" : ""}` + ? compacting ? "compacting supervisor context" : lastFailure ? "assessment failed · /goals status" : refreshInFlight ? "waiting for worker overview" : `watching ${state.steerRounds}${routineDirty ? " · update pending" : ""}` : "watched"; ctx.ui.setStatus(STATUS_ID, ctx.ui.theme.fg("accent", `${EYE} `) + ctx.ui.theme.fg("dim", text)); } @@ -515,6 +567,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super pi.on("tool_call", (event: any, context: any) => { initializeSession(context); + if (state.plan?.paused && SUPERVISOR_TOOLS.includes(event.toolName)) return { block: true, reason: "Supervision stopped by user. Resume explicitly from the worker." }; if (supervisorMode && !inspectionAllowed(event.toolName)) return { block: true, reason: "Supervisor mode is inspection-only. Steer the worker; do not execute, delegate, schedule, or mutate files." }; }); pi.on("user_bash", () => { @@ -545,12 +598,26 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super * The broker's registry is the truth about who exists now, so ask it. Either way this prints, * because "supervising, waiting" and "the other session is gone" look identical otherwise. */ + function pauseLocally(exit: boolean, pauseId: string = randomUUID()) { + pairingGeneration++; + if (state.plan) state = { ...state, plan: { ...state.plan, paused: true, active: false, pauseId } }; + // Local state wins even when cancellation cannot be delivered to the peer. + try { cancelGoalRequests("Stopped by user"); } catch (error) { lastFailure = String(error); } + clearAssessments(); + clearInterval(watchTimer); clearTimeout(pairTimer); + watchTimer = undefined; pairTimer = undefined; + finishPair(new Error("Stopped by user")); + save(); + onPause?.(exit); + ctx?.abort(); + } + async function rejoinOrDrop() { if (!modelReady() || duplicateIntercom || bootstrapPending || !state.role || !state.pairedId || (state.plan && state.role === "supervisor" && !state.planInitialized)) return; const live = await channel!.listSessions(); if (state.plan) { // Pi session files identify the pair across process restarts; broker IDs identify live peers. - send({ t: "plan_hello", to: "*", bindingId: state.plan.id, role: state.role as "worker" | "supervisor", sessionFile: ctx.sessionManager.getSessionFile() }); + send({ t: "plan_hello", to: "*", bindingId: state.plan.id, role: state.role as "worker" | "supervisor", sessionFile: ctx.sessionManager.getSessionFile(), paused: state.plan.paused, pauseId: state.plan.pauseId }); if (!live.some((s: any) => s.id === state.pairedId)) { if (state.role === "supervisor") stripWriters(); ctx.ui.notify("Plan supervisor peer is disconnected; retaining this session and waiting for reconnect.", "warning"); @@ -561,6 +628,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super reset(`intercom-supervisor: ${state.pairedId.slice(0, 8)} is gone, so the pairing is dropped. Run /supervise to start again.`); return; } + if (state.plan?.paused) { showStatus(); return; } if (state.role !== "supervisor") { // A worker reloaded at the prompt takes no turn, so this is its only chance to start watching. startWatch(); @@ -626,15 +694,16 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super routineDirty = false; refreshInFlight = false; checkpointQueued = false; - awaitingUser = false; + latestCompletion = undefined; + lastPublishedPlanHash = undefined; lookPending = false; publishAgain = false; } function failAssessment(reason: string, includeQueued = true) { + lastFailure = reason; // A failed assessment is not a human-input dependency. Resume on later worker // progress/cadence, without retrying the same dirty view from agent_settled. - awaitingUser = false; routineDirty = false; refreshInFlight = false; ctx.ui.notify(reason, "error"); @@ -647,7 +716,7 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super } async function advanceSupervisor() { - if (stopping || advancing || activeAssessment || !ctx?.isIdle() || !modelReady() || peerDisconnected || !channel?.snapshot().connected || state.role !== "supervisor" || (state.plan && ((!state.plan.active && !incomingReview) || !state.planInitialized))) return; + if (stopping || state.plan?.paused || advancing || activeAssessment || !ctx?.isIdle() || !modelReady() || peerDisconnected || !channel?.snapshot().connected || state.role !== "supervisor" || (state.plan && ((!state.plan.active && !incomingReview) || !state.planInitialized))) return; advancing = true; const generation = pairingGeneration; try { @@ -663,14 +732,14 @@ export default function (pi: any, modelReady: () => boolean = () => true): Super workerStopped = false; pi.sendMessage({ customType: "supervisor_checkpoint", display: true, details: { requestId: review.requestId }, content: `Goal sign-off: ${review.goal} -Assess direction and scope against the current plan and evidence. Give a brief useful assessment, then call review_goal. The fresh evidence judge still checks artifacts independently. This is one goal, not the end of supervision. +Inspect the actual result against the agreed outcome, discriminator and evidence. A completed task or existing file is not enough. Give a brief useful judgment, then call review_goal; if the goal is unmet, name the next useful work or check. The fresh evidence judge still checks artifacts independently. This is one goal, not the end of supervision. Worker context frozen for this checkpoint (not a live view): ${latestView}` }, { triggerTurn: true, deliverAs: "followUp" }); return; } - if (incomingReview || awaitingUser || refreshInFlight || !routineDirty) return; + if (incomingReview || refreshInFlight || !routineDirty) return; routineDirty = false; refreshInFlight = true; showStatus(); @@ -681,7 +750,7 @@ ${latestView}` }, advancing = false; // Reconnect/checkpoint handlers may have hit the guard while this obsolete compaction // was pending. Re-drive only actual current work, never a same-generation failure loop. - if (generation !== pairingGeneration && !stopping && ((incomingReview && checkpointQueued) || (!incomingReview && !awaitingUser && !refreshInFlight && routineDirty))) void advanceSupervisor(); + if (generation !== pairingGeneration && !stopping && ((incomingReview && checkpointQueued) || (!incomingReview && !refreshInFlight && routineDirty))) void advanceSupervisor(); } } @@ -711,7 +780,8 @@ ${latestView}` }, const returning = peerDisconnected; peerDisconnected = false; state = { ...state, pairedId: from }; save(); - if (wire.t === "plan_hello") send({ t: "plan_hello_ack", to: from, bindingId: wire.bindingId, role: state.role as "worker" | "supervisor", sessionFile: ctx.sessionManager.getSessionFile() }); + if (wire.paused && (!state.plan?.paused || (wire.pauseId ?? "") > (state.plan.pauseId ?? ""))) pauseLocally(false, wire.pauseId); + if (wire.t === "plan_hello") send({ t: "plan_hello_ack", to: from, bindingId: wire.bindingId, role: state.role as "worker" | "supervisor", sessionFile: ctx.sessionManager.getSessionFile(), paused: state.plan?.paused, pauseId: state.plan?.pauseId }); if (state.role === "worker") { attached?.finish(state.plan); attached = undefined; startWatch(); } else if (returning) requestFreshView(); return; @@ -736,7 +806,7 @@ ${latestView}` }, clearTimeout(pairTimer); pairTimer = undefined; // A same-binding replay must not undo Ready already completed by this worker. - const plan = wire.plan ? { ...wire.plan, active: state.plan?.active ?? wire.plan?.active } : undefined; + const plan = wire.plan ? { ...wire.plan, active: state.plan?.active ?? wire.plan?.active, paused: state.plan?.paused ?? wire.plan.paused } : undefined; peerDisconnected = false; state = { ...EMPTY_STATE, role: "worker", pairedId: from, goal: plan ? planText(plan) : wire.goal, ...(plan ? { plan } : {}) }; latestView = ""; @@ -769,6 +839,15 @@ ${latestView}` }, peerDisconnected = false; showStatus(); } + if (wire.t === "plan_pause" && state.plan) { pauseLocally(wire.exit, wire.pauseId); return; } + if (wire.t === "plan_resume" && state.role === "supervisor" && state.plan) { + const accepted = wire.pauseId === state.plan.pauseId && !state.plan.stopped && !!state.planInitialized && !bootstrapPending && modelReady() && wire.planHash === planHash(planText(state.plan)); + if (accepted) { state = { ...state, plan: { ...state.plan, active: true, paused: false } }; lastFailure = undefined; save(); } + send({ t: "plan_resumed", to: from, bindingId: wire.bindingId, requestId: wire.requestId, accepted }); + return; + } + if (wire.t === "plan_resumed" && resuming?.requestId === wire.requestId) { resuming.finish(wire.accepted); return; } + if (state.plan?.paused && !["goal_cancel", "unpair", "done"].includes(wire.t)) return; if (wire.t === "plan_activate") { state = { ...state, plan: { ...state.plan!, active: true } }; save(); return; } if (wire.t === "goal_cancel") { if (outgoingReview?.requestId === wire.requestId) outgoingReview.finish(undefined, new Error("Supervisor cancelled the review; retry")); @@ -794,7 +873,6 @@ ${latestView}` }, } incomingReview = { ...reviewIdentity(wire), view: wire.view }; checkpointQueued = true; - awaitingUser = false; if (refreshInFlight) { refreshInFlight = false; routineDirty = true; } await advanceSupervisor(); return; @@ -845,11 +923,12 @@ ${latestView}` }, if (!refreshInFlight) { routineDirty = true; return; } refreshInFlight = false; } else if (refreshInFlight) { routineDirty = true; return; } - if (activeAssessment || !ctx.isIdle() || advancing || compacting || incomingReview || awaitingUser) { + if (activeAssessment || !ctx.isIdle() || advancing || compacting || incomingReview) { routineDirty = true; return; } activeAssessment = "view"; + lastFailure = undefined; assessmentFailure = undefined; assessmentEmpty = false; const generation = pairingGeneration; @@ -861,6 +940,7 @@ ${latestView}` }, } if (stopping || generation !== pairingGeneration || state.role !== "supervisor" || from !== state.pairedId || state.plan?.id !== bindingId || peerDisconnected || !channel?.snapshot().connected || (state.plan && !state.plan.active)) return; latestView = wire.view; + latestCompletion = wire.completion; showStatus(); workerStopped = wire.stopped; if (!state.plan && reviewsSinceGoal >= GOAL_REVIEW_INTERVAL - 1) tellGoal(); @@ -893,9 +973,6 @@ ${latestView}` }, verdictsThisLook = 0; }); - pi.on("input", (event: any) => { - if (event.source !== "extension") awaitingUser = false; - }); pi.on("agent_end", (event: any) => { if (state.role !== "supervisor" || !activeAssessment) return; const last = event.messages?.findLast((message: any) => message.role === "assistant"); @@ -926,7 +1003,7 @@ ${latestView}` }, */ const LOOKS_KEPT = 3; pi.on("context", (event: any) => { - if (state.role !== "supervisor") return; + if (state.role !== "supervisor" || state.plan?.paused) return; const checkpoint = event.messages.findLast((m: any) => m.customType === "supervisor_checkpoint"); presentedReview = checkpoint?.details?.requestId === incomingReview?.requestId ? incomingReview : undefined; const messages = event.messages.filter((m: any) => m.customType !== "supervisor_plan"); @@ -963,7 +1040,7 @@ ${planText(state.plan)}` }); ctx = context; // Pi retries overflow recovery immediately after this event. Queue the rubric for its next // ordinary turn, so the failed assistant remains final and can be removed. - if (state.role === "supervisor") { + if (state.role === "supervisor" && !state.plan?.paused) { if (event.willRetry) reviewsSinceGoal = GOAL_REVIEW_INTERVAL - 1; else tellGoal(); } @@ -1073,6 +1150,9 @@ ${planText(state.plan)}` }); const subagents = state.plan ? [] : await childPiProcesses(); const entries = context.sessionManager.getBranch(); const idle = !starting && context.isIdle() && (!background || background.quiet); + const currentPlan = state.plan ? planText(state.plan) : undefined; + const progress = currentPlan === undefined ? undefined : planProgress?.(currentPlan); + const currentPlanHash = currentPlan === undefined ? undefined : planHash(currentPlan); const view = buildView({ goal: state.goal, sourceSession: context.sessionManager.getSessionFile?.(), @@ -1082,8 +1162,9 @@ ${planText(state.plan)}` }); subagents, model: workerModel(context), background: background?.description, + planReview: currentPlanHash ? `Canonical plan ${currentPlanHash === lastPublishedPlanHash ? "unchanged" : "changed since the previous published view (or first view)"}; read ${state.plan!.planPath} for scope and goal-state changes.\n${progress?.summary ?? "CompleteGoal sign-off tracking unavailable; checked boxes alone do not establish completion."}` : undefined, }); - return { view, idle, entries }; + return { view, idle, entries, completion: progress?.completion, currentPlanHash }; } /** @@ -1110,7 +1191,7 @@ ${planText(state.plan)}` }); // Claimed before the await, not after. ps takes long enough that a second timer tick would // otherwise start its own look while this one is still waiting. lastLook = Date.now(); - const { view, idle, entries } = await captureWorkerView(context, why, since, starting); + const { view, idle, entries, completion, currentPlanHash } = await captureWorkerView(context, why, since, starting); if (generation !== pairingGeneration || state.pairedId !== peer || state.plan?.id !== bindingId || stopping || peerDisconnected || outgoingReview) return; // A timer look at a worker that has done nothing since the last one wakes the supervisor to read // a view it has already read. Session 019ffa73: 13 of 92 verdicts were "check-in with no new @@ -1125,8 +1206,9 @@ ${planText(state.plan)}` }); } looksSkipped = 0; } - send({ t: "view", to: state.pairedId, view, stopped: idle, ...(refreshed ? { refreshed: true } : {}) }); + send({ t: "view", to: state.pairedId, view, stopped: idle, ...(refreshed ? { refreshed: true } : {}), ...(completion ? { completion } : {}) }); published = true; + lastPublishedPlanHash = currentPlanHash; if (refreshed) lookPending = false; lastViewBody = bodyOf(view); modelTurns = 0; @@ -1202,15 +1284,14 @@ ${planText(state.plan)}` }); if (state.role === "supervisor") { const unresolved = (activeAssessment === "goal" && incomingReview && !checkpointQueued) || (activeAssessment === "view" && verdictsThisLook === 0); if (activeAssessment && assessmentFailure) failAssessment(assessmentFailure, false); - else if (unresolved && assessmentEmpty) failAssessment("Supervisor assessment incomplete: empty final response without a verdict or user question. Retry the checkpoint.", false); + else if (unresolved && assessmentEmpty) failAssessment("Supervisor assessment incomplete: empty final response without a verdict. Retry the checkpoint.", false); + else if (activeAssessment === "goal" && incomingReview && !checkpointQueued) failAssessment("Supervisor checkpoint incomplete: no review_goal decision. Read the visible assessment and retry when ready; prose alone is not an approval or a human-input dependency.", false); assessmentFailure = undefined; assessmentEmpty = false; - const completed = activeAssessment; + // Ordinary prose (including questions) is not a transport latch. Respect human pause + // instructions in context, but keep later views flowing. Only genuinely queued new work + // can trigger another look here; an idle response never retries itself. activeAssessment = undefined; - if ((completed === "goal" && incomingReview && !checkpointQueued) || (completed === "view" && verdictsThisLook === 0)) { - awaitingUser = true; - context.ui.notify("Supervisor is waiting for user input or an explicit checkpoint decision; routine refresh is paused.", "info"); - } await advanceSupervisor(); return; } @@ -1308,7 +1389,6 @@ ${planText(state.plan)}` }); context.ui?.notify?.("intercom-supervisor: not supervising, so there is nothing to look at", "error"); return; } - awaitingUser = false; requestFreshView(true); context.ui?.notify?.(`requested a complete view from ${state.pairedId.slice(0, 8)}; waiting for the worker overview`, "info"); return; @@ -1333,7 +1413,6 @@ ${planText(state.plan)}` }); // waiting up to half an hour for the next look. tellSupervisor(GOAL_CHANGED(goal)); reviewsSinceGoal = 0; - awaitingUser = false; requestFreshView(); context.ui?.notify?.(`goal changed: ${goal}`, "info"); return; @@ -1430,7 +1509,7 @@ ${planText(state.plan)}` }); pi.registerTool({ name: "review_goal", label: "Review goal", - description: "Assess the current goal checkpoint: approve trajectory/scope, request further work, or name a human decision. The extension binds the decision to the exact goal and revision you were shown. This does not complete the goal or stop supervision.", + description: TOOL_REVIEW_GOAL, parameters: Type.Object({ decision: Type.String({ enum: ["approve", "needs_work", "needs_user"] }), reason: Type.String() }), execute: async (_id: string, params: { decision: GoalDecision["decision"]; reason: string }) => { const review = incomingReview; @@ -1443,7 +1522,7 @@ ${planText(state.plan)}` }); send({ ...reviewIdentity(review), decision: params.decision, reason: params.reason, t: "goal_decision", to: state.pairedId }); tellSupervisor(`Goal assessment — ${review.goal}: ${params.decision}. ${params.reason}`); incomingReview = undefined; - awaitingUser = params.decision === "needs_user"; + // A human decision blocks its dependent work, not delivery of later worker direction. return { content: [{ type: "text", text: "Goal decision sent. Supervision remains active. End this response." }], terminate: true }; }, }); @@ -1572,7 +1651,7 @@ Tell them in your reply, quoting it, so they can correct it.`, return { content: [{ type: "text", text: "Not supervising." }], isError: true }; } if (routineDirty || refreshInFlight) throw new Error("New worker progress awaits a complete overview; do not finish from the older assessment."); - if (state.plan && (incomingReview || /^\s*(?:\d+\.|[-*])\s*\[[ /]\]\s*goal:/im.test(planText(state.plan)))) throw new Error("Open plan goals remain. Use review_goal for an individual goal request."); + if (state.plan && (incomingReview || !latestCompletion || latestCompletion.planHash !== planHash(planText(state.plan)) || !latestCompletion.total || latestCompletion.pending > 0)) throw new Error("Open plan goals or unsigned claims remain, or fresh CompleteGoal tracking is unavailable. Use review_goal for an individual goal request; manual checkboxes are not sign-off."); if (state.plan && /tracked background work:.*unknown|tracked background work:.*(?:processes|subagents): [1-9]/.test(latestView)) throw new Error("Tracked background work is active or unknown"); // "done" while a delegated tool call has no result is a false completion: the worker settled // but its subagent or background job is still spending. This proves only that no tracked @@ -1585,6 +1664,7 @@ Tell them in your reply, quoting it, so they can correct it.`, isError: true, }; } + if (latestCompletion?.inconclusive) tellSupervisor(`${latestCompletion.inconclusive} goal(s) accepted inconclusive under fail-forward policy, not independently verified. Ending supervision does not turn those records into conclusive success.`); send({ t: "done", to: state.pairedId, reason: params.reason }); const rounds = state.steerRounds; reset(`supervision finished: ${params.reason}`, true); diff --git a/src/internal/supervisor/prompts.ts b/src/internal/supervisor/prompts.ts index a6d86f3..dda498d 100644 --- a/src/internal/supervisor/prompts.ts +++ b/src/internal/supervisor/prompts.ts @@ -61,13 +61,13 @@ The view names the worker's model and how full its context is. A small or fast m small step per instruction. A worker near the top of its context is about to compact, so tell it to write down what matters before it loses the detail. -There is no round limit and no budget. Supervision runs until the human stops it. Ending early is -the failure this exists to prevent, so never stop because it feels like enough. +Supervise until the agreed result is delivered and inspected, or the human stops supervision. +Do not abandon unfinished work or prolong completed work for optional polish. Respect permission +and spending limits; autonomous supervision is not an unlimited budget. -The human typing in the worker session is not a handover, and it is not a reason to stand back. -They say a word and go to bed; the worker is then stopped with nobody driving it, which is the -state you exist for. They will stop you themselves when they want you stopped. Judge the worker -against the goal and nothing else. +A human message is new direction, not an automatic handover. Respect an explicit human pause or +required approval: do not steer the paused work until authorized. Keep independent authorized work +moving. Later views still deserve assessment; a question or pause is not a reason to discard them. Every view says how long the worker has gone with no new turn. A worker that has produced nothing for a long time is stuck, or waiting for you, or in one command that will not return. Say which @@ -116,7 +116,7 @@ export const TOOL_LET_IT_RUN = + " The call sends no message to the worker. Call it once, then end the current supervisor response." // Repeated here because a tool description survives a compaction and the brief does not. The // live failure was a let_it_run reasoned "human is actively directing", two hours before dawn. - + " A human message does not end supervision; only an explicit stop command ends supervision."; + + " A human message does not end supervision. Respect explicit human pauses and required decisions; assess new views without restarting paused work."; /** * How a look ends, and it must appear in every verdict's result. @@ -146,7 +146,8 @@ ${workerStopped ? STOPPED_WARNING : "A later worker view starts the next supervi export const STOPPED_WARNING = `The current worker view reports that worker execution stopped. A stopped worker does not resume without a new user or supervisor message. If the goal remains unmet, send a concrete continuation -instruction. A human message does not end supervision. A later worker view will report the worker state.`; +instruction unless a verified dependency or explicit human pause prevents it. Respect the pause; +keep independent authorized work moving. A later worker view will report the worker state.`; /** The answer to a second let_it_run in one look. Costs a round trip and no error line. */ export const LET_IT_RUN_AGAIN = @@ -158,9 +159,11 @@ export const STEER_ACK = (round: number, workerId: string) => `Supervisor instruction ${round} was sent to worker session ${workerId}. Worker receipt and execution are not confirmed. The supervisor has completed its verdict for the current worker view. ${END_TURN}`; export const TOOL_STEER = - "Send one concrete next action to the worker. The extension sends the message to the paired worker session; worker receipt and execution require a later worker view."; + "Send one concrete next action and its purpose toward the agreed goal. Use it to resume authorized work, request a needed check, or correct drift. A recap alone does not send an instruction. Do not interrupt productive work or repeat ineffective steering without changing the approach. Worker receipt and execution require a later worker view."; +export const TOOL_REVIEW_GOAL = + "Judge the presented goal against the user's intended outcome and discriminator, not merely task completion or file existence. Approve only when the inspected evidence warrants it; otherwise give needs_work with the next useful work/check, or needs_user for a specific unresolved human decision. This records your judgment, not proof from mechanical checks. The fresh evidence judge still runs independently; this does not end supervision of remaining goals."; export const TOOL_DONE = - "Declare the goal met and stop supervising. Only call this with quoted evidence from the view."; + "Finish supervision after inspecting the agreed result. For a plan, every non-cancelled goal needs a CompleteGoal record, not a manual checkbox. Accepted inconclusive is fail-forward, not proof of success; disclose that distinction. Quote evidence for any completion claim."; /** * Sent with every view, so it is deliberately short. @@ -194,15 +197,16 @@ export const REVIEW_NUDGE = (view: string, rounds: number, stopped: boolean) => ${view} -${rounds} instructions so far. The status line says how long it has had no new turn. It will not start -again by itself. You are the supervisor, not the worker. Give a brief visible progress assessment -and helpful perspective. Steer with a concrete continuation if work remains, or say what human -decision is needed. Use let_it_run when no intervention is useful, or done when complete.` +${rounds} instructions so far. Judge the actual result against the outcome and discriminator; say your assessment briefly. If work remains, investigate the stop and steer a useful authorized continuation with its purpose. A recap alone does not restart work. Respect explicit human pauses; name real dependencies and how to observe their resolution. Manual ticks are claims. Use let_it_run only if no instruction helps; done needs the agreed result and recorded sign-offs.` : `${VIEW_CHECKIN} ${view} -You are the supervisor, not the worker. Briefly assess progress and the most useful next consideration in visible text. Use let_it_run when on course; steer only when the evidence calls for a concrete correction.`; +You are the supervisor, not the worker. Is the work on track toward the user's intended outcome? +Give a brief visible judgment. Use let_it_run when on course; steer only when the evidence +shows drift, mistaken assumptions or wasted effort. If this is +the Ready handoff, send a concrete starting instruction unless work has already begun. Review plan +changes against user intent, preserving authorized changes rather than treating every edit as failure.`; /** Refusal shown when done is called while the worker still has work running. */ export const DONE_BLOCKED = (what: string) => @@ -218,68 +222,53 @@ their phone.`; * Default supervisor prompt. A project SUPERVISOR.md overrides it, same as @monotykamary/pi-supervisor. * Unlike that extension there is no JSON verdict to parse, because the verdict is a tool call. */ -export const DEFAULT_SUPERVISOR_PROMPT = `You supervise a coding agent from outside its session. -Your job is to help it reach the agreed goal without unnecessary human intervention. -At each check visibly assess how the work is tracking and offer useful perspective in a few sentences. -Inspect, judge and steer; never execute work, delegate it, schedule it, or mutate the worker's files. +export const DEFAULT_SUPERVISOR_PROMPT = `You are the visible, read-only supervisor of another Pi session. +The worker carries implementation detail; you retain user intent, decisions and high-level judgment. +At startup and after compaction, read applicable AGENTS.md instructions and relevant skills. Do not +assume a particular project or workflow. Infer ordinary implementation details without replacing the +agreed outcome or inventing restrictions. Make consequential uncertainty and disagreement visible; +respect reasonable user preferences without making the user repeatedly justify them. +Supervise autonomously until the agreed goal is achieved and you have inspected the actual result. +Identify the missing user-visible outcome and steer the next useful action through to delivery. +Approval records support the work; they are not the outcome. Seek justified confidence, not +certainty at any cost. Investigate uncertainty with the cheapest useful check, then decide. +Do not prolong completed work for optional polish. -Judge from the view only. You cannot see the worker's files unless you read them yourself. +Treat "blocked", "waiting", "impossible", and "already done" as claims to investigate, not +conclusions to repeat. Check the evidence and whether the claimed dependency is real. Consider +mistaken assumptions, bugs, and other authorized ways forward. Never repeat a steer that had no +effect: inspect what happened and change the approach. Keep independent authorized work moving +when it does not depend on the blocker. A verified external dependency can justify waiting; it +does not make an unfinished goal complete. Identify what event resumes progress and how to observe it. -Call steer when the work is incomplete, when the worker asked a question you can answer with a -sensible default, or when it claims success without evidence. One concrete next action per steer. -Never repeat a steer that had no effect; change the approach instead. +Resolve technical choices within agreed scope. Steer one concrete next action when the worker is +idle with unfinished work. If useful work is running, do not invent work or repeat instructions +awaiting execution. Respect explicit human pauses and permission limits, including credentials and +spending. Escalate only a specific unresolved human decision after checking what is already +authorized. New worker views, including answers to earlier questions, still need your judgment; +do not restart paused work without authorization or widen scope to evade a blocker. -The view line "child pi processes still running" means the worker delegated to a subagent that is -still working. It stopped, the subagent did not. Do not call done, it will be refused. Steer the -worker to wait for that subagent and report what it produced. +At each review give a brief visible assessment: what the evidence shows, how work is tracking, and +your judgment about the next step. Add perspective, not unchanged status or delivery receipts. +Distinguish observations from guesses. Inspect, judge and steer; never execute work, delegate it, +schedule it, or mutate files. Let the worker produce both the artifact and its verification output. -The view line "no new file or commit for N reviews in a row" means your last N instructions -moved nothing the worker's session can show. Two or more is your signal to change approach, ask the -human, or check whether the goal is already met. Sometimes it is honest work on one file, so read -the recent turns before you decide. +Ground consequential judgments in verbatim evidence with source paths and enough context to check +the interpretation. Read the actual deliverable against the user's goal. A worker summary, passing +tests, or a checked box alone do not establish success. Repeated summaries are not independent +evidence. Missing evidence stays unknown until inspected. Investigate contradictions and surprising +results with checks that distinguish plausible explanations. Watch for weakened tests, fabricated +measurements, partial runs reported as full ones, and logs that do not demonstrate real execution. +Say what evidence would change your mind. -Call done only when all of these hold: -1. the worker named the artifact file it produced, with a path -2. the worker quoted text from that file, rather than summarising it -3. nothing in the view contradicts the claim +The current canonical plan is the source of truth, subject to newer human direction. Review plan +changes for drift and steer corrections when warranted. Manual completion checkboxes are claims, +not sign-off. CompleteGoal records the retained supervisor's checkpoint and a fresh independent +judge's result. An accepted-inconclusive record preserves fail-forward but is not verified success; +state the uncertainty rather than describing it as conclusive. Git status is review context, not an +acceptance gate: uncommitted changes and ignored output files may be legitimate evidence. Do not +require cleanup or a commit unless the agreed goal requires it. Inspect cited paths directly. -A confident summary is not evidence. When in doubt, steer. - -When the worker does machine learning or data research, a wrong result looks exactly like a right -one. These steers come from wassname's ml-debug skill, roughly in the order they bite. Each one is -something you can see in the view. -- It concluded without reading its data. Steer it to paste the lines it read into the chat, a raw - sample and the metric line, not a summary of them. Quoting is the point, twice over: you and the - human can then check the same text, and an agent that has to quote has to look (Karpathy inspects - the data before touching the model; Nanda: read your data, often it is quite bad). A conclusion - with no quoted output, a ranking with no per-item evidence, or "the method failed" with no sample - of what the output looked like, all mean it has not looked. -- It reports a surprising win. Most true results are boring, so an exciting one is more likely to - be false (Neel Nanda). Steer it to rule out a bug, leakage or a broken evaluation first. -- It reports a failure and moves on, or calls the failure a property of the method. Assume a bug: - bugs are far more common, and far cheaper to find, than a real negative result (Andy Jones). - Steer it to write two or three diagnoses, one of them a bug in its own code, put a rough - probability on each, and run the cheapest test that tells them apart. Broken research code fails - silently and still runs, so "it ran" is not evidence that it worked. -- It is about to start another long run without saying what each outcome would mean. Steer it to - write that prediction first (Rahtz: think more, experiment less). On a shared GPU that is the - cheapest hour you can buy. -- It compares two methods from one run each. Seed variance alone splits identical configurations - into different distributions (Henderson), so steer it to say what varies before it ranks - anything. -- It changed two things in one run and credits one of them. Changing anything changes everything - (Sculley et al., CACE). Steer it to say what it can actually attribute, or to rerun with one - change. -- It saw a number it cannot explain and carried on. An anomaly it did not go looking for is the - cheapest bug it will ever find, so steer it to chase that before anything else. - -Three more ways work gets faked, from @monotykamary/pi-supervisor's cheating list. Steer, and ask -for the output that would settle it. -- the worker edits a test to weaken an assertion, or skips a failing one, and calls that progress -- it reports a number without the command output it came from, or edits the measurement instead of - the thing being measured -- it runs a smaller dataset or part of the suite, then reports as if it ran the whole thing - -Do not answer questions that need real human knowledge: passwords, credentials, spending money, -or a choice between two designs the human cares about. For those, reply in plain text saying what -you need. Your reply reaches the human's phone.`; +Use steer for useful corrections, let_it_run when no instruction is needed, and done when the +agreed work is finished with the required sign-offs. Follow the tools' requirements without letting +bookkeeping replace delivery. Keep visible judgments brief and useful. -- Pi/OpenAI`; diff --git a/src/internal/supervisor/protocol.ts b/src/internal/supervisor/protocol.ts index 4082656..07c8802 100644 --- a/src/internal/supervisor/protocol.ts +++ b/src/internal/supervisor/protocol.ts @@ -8,6 +8,17 @@ import { type SupervisorBinding, validBinding } from "../../supervisor.js"; import { MAX_VIEW_BYTES } from "./view.js"; +/** Compact completion bookkeeping from the worker, bound to the exact plan it observed. */ +export interface PlanCompletion { + planHash: string; + total: number; + pending: number; + inconclusive: number; +} +function validCompletion(value: any): value is PlanCompletion { + return !!value && typeof value.planHash === "string" && [value.total, value.pending, value.inconclusive].every(n => Number.isSafeInteger(n) && n >= 0) && value.pending + value.inconclusive <= value.total; +} + export interface GoalReview { requestId: string; bindingId: string; @@ -44,14 +55,20 @@ export interface GoalDecision extends GoalReview { reason: string; } export type PlanWire = - | { t: "plan_hello" | "plan_hello_ack"; to: string; bindingId: string; role: "worker" | "supervisor"; sessionFile: string } + | { t: "plan_hello" | "plan_hello_ack"; to: string; bindingId: string; role: "worker" | "supervisor"; sessionFile: string; paused?: boolean; pauseId?: string } + | { t: "plan_pause"; to: string; bindingId: string; exit: boolean; pauseId: string } + | { t: "plan_resume"; to: string; bindingId: string; requestId: string; planHash: string; pauseId?: string } + | { t: "plan_resumed"; to: string; bindingId: string; requestId: string; accepted: boolean } | ({ t: "goal_review"; to: string } & GoalReviewRequest) | ({ t: "goal_decision"; to: string } & GoalDecision) | { t: "goal_cancel"; to: string; requestId: string; bindingId: string } | { t: "plan_activate" | "plan_stop"; to: string; bindingId: string }; export function validPlanWire(value: any): value is PlanWire { if (!value || typeof value.to !== "string" || typeof value.bindingId !== "string") return false; - if (value.t === "plan_hello" || value.t === "plan_hello_ack") return ["worker", "supervisor"].includes(value.role) && typeof value.sessionFile === "string"; + if (value.t === "plan_hello" || value.t === "plan_hello_ack") return ["worker", "supervisor"].includes(value.role) && typeof value.sessionFile === "string" && (value.paused === undefined || typeof value.paused === "boolean") && (value.pauseId === undefined || typeof value.pauseId === "string"); + if (value.t === "plan_pause") return typeof value.exit === "boolean" && typeof value.pauseId === "string"; + if (value.t === "plan_resume") return typeof value.requestId === "string" && typeof value.planHash === "string" && (value.pauseId === undefined || typeof value.pauseId === "string"); + if (value.t === "plan_resumed") return typeof value.requestId === "string" && typeof value.accepted === "boolean"; if (value.t === "plan_activate" || value.t === "plan_stop") return true; if (typeof value.requestId !== "string") return false; if (value.t === "goal_cancel") return true; @@ -77,7 +94,7 @@ export type Wire = (PlanWire | { t: "paired"; to: string; plan?: SupervisorBinding } | { t: "goal"; to: string; goal: string } /** stopped: the worker settled, so this is a decision point. false: a check in mid-turn. */ - | { t: "view"; to: string; view: string; stopped: boolean; refreshed?: boolean } + | { t: "view"; to: string; view: string; stopped: boolean; refreshed?: boolean; completion?: PlanCompletion } /** Supervisor asks for a view now. Its own turn cannot make one: the worker publishes them. */ | { t: "look"; to: string } | { t: "directive"; to: string; text: string } @@ -88,11 +105,11 @@ export type Wire = (PlanWire export function isWire(payload: unknown): payload is Wire { if (typeof payload !== "object" || payload === null) return false; if (validPlanWire(payload)) return true; - const { t, to, goal, view, stopped, refreshed, text, reason, plan } = payload as Record; + const { t, to, goal, view, stopped, refreshed, text, reason, plan, completion } = payload as Record; if ((t === "pair" || t === "paired") && plan !== undefined && !validBinding(plan)) return false; if (typeof to !== "string") return false; if (t === "pair" || t === "goal") return typeof goal === "string"; - if (t === "view") return typeof view === "string" && typeof stopped === "boolean" && (refreshed === undefined || typeof refreshed === "boolean"); + if (t === "view") return typeof view === "string" && typeof stopped === "boolean" && (refreshed === undefined || typeof refreshed === "boolean") && (completion === undefined || validCompletion(completion)); if (t === "directive") return typeof text === "string" && text.trim().length > 0; if (t === "done") return typeof reason === "string"; return t === "unpair" || t === "paired" || t === "look" || t === "who" || t === "here"; diff --git a/src/internal/supervisor/view.ts b/src/internal/supervisor/view.ts index 83540f9..5c4bf33 100644 --- a/src/internal/supervisor/view.ts +++ b/src/internal/supervisor/view.ts @@ -248,10 +248,11 @@ export interface ViewInput { */ model?: string; background?: string; + planReview?: string; } /** Render the view, and cut it to MAX_VIEW_BYTES so the broker cannot reject it. */ -export function buildView({ goal, status, entries, since = 0, stale = 0, subagents = [], model = "", sourceSession, background }: ViewInput): string { +export function buildView({ goal, status, entries, since = 0, stale = 0, subagents = [], model = "", sourceSession, background, planReview }: ViewInput): string { const messages = entries.filter((e) => e.type === "message" && e.message); const pending = outstandingWork(messages); const workerMessages = messagesSince(entries); @@ -268,7 +269,7 @@ export function buildView({ goal, status, entries, since = 0, stale = 0, subagen const latestUser = [...entries].reverse().find(e => e.type === "message" && e.message?.role === "user" && textOf(e.message).trim() && !textOf(e.message).startsWith(SUPERVISOR_PREFIX)); const direction = latestUser?.message ? textOf(latestUser.message) : ""; const head = [ - ...(from === 0 && direction ? [ + ...(direction ? [ `# Latest user direction${latestUser?.timestamp ? ` (${latestUser.timestamp})` : ""}`, direction.length > 2000 ? `${direction.slice(0, 2000)}\n[user direction truncated; inspect the worker source session for full text]` : direction, "", @@ -287,6 +288,7 @@ export function buildView({ goal, status, entries, since = 0, stale = 0, subagen `tool calls with no result: ${pending.length ? pending.join(", ") : "none"}`, `child pi processes still running: ${subagents.length ? subagents.join(", ") : "none"}`, ...(background ? [`tracked background work: ${background}`] : []), + ...(planReview ? [`plan review: ${planReview}`] : []), ...(stale > 0 ? [`no new file or commit for ${stale} reviews in a row`] : []), ``, // Sent when this view starts at the compaction boundary, which is the first view and every diff --git a/src/prompts.ts b/src/prompts.ts index bf6aa26..3d38e1b 100644 --- a/src/prompts.ts +++ b/src/prompts.ts @@ -37,15 +37,15 @@ human would need to approve later. If any is uncertain, reduce uncertainty now: search the web when they can answer, then ask the human to confirm your interpretation, pin down the outcome or task, or approve an editorial or other preference choice. Do not present the review menu with a placeholder goal such as "work out the thing", "improve it", or "investigate". -3. By default, before proposing a final plan, ask at least THREE distinct task-specific alignment -questions in ONE chat round: test agreement about the expected result, scope and constraints, and -success/failure criteria. Even if you think you understand, check how far apart your interpretations -are. Wait for the human's answers and use them before declaring the plan final. Do not ask technical -facts that read-only inspection can resolve, or use a generic ritual questionnaire. Additional -questions should materially reduce uncertainty while discovering the right plan. Each question must be short and self-contained: state the relevant -context, use the human's language and ASD-STE100 -Simple Technical English, and give a recommended answer. Record each answer in ## Interview. Do not -make the plan final while material user decisions remain open. +3. Ask material task-specific alignment questions about unresolved outcomes, scope, constraints, or +success/failure criteria. There is no fixed question quota. Inspect technical facts yourself and do +not ask for confirmation of ordinary implementation details or repeat answered questions. Batch +independent high-impact questions in one short round. Each question must be short and self-contained: +state the relevant context, use the human's language and ASD-STE100 Simple Technical English, and +give a recommended answer. Wait for answers to required decisions and use them before declaring +the plan final. Record each answer in ## Interview. Do not make the plan final while material user +decisions remain open; if the requested work is already executable, proceed to review without a +ritual questionnaire. 4. State the user-visible result before the goals: one concrete sentence naming what the human will inspect when this plan is done. Take it from the original request, not from your implementation plan. Every requested artifact and action must survive into this sentence. An agent-inferred constraint may @@ -129,8 +129,9 @@ Conventions: - Make the discriminator a concrete, checkable observation about a real artifact (a file, a test result, a committed diff, a metric), never about the plan file's own checkbox. - evidence stays empty at planning; you fill it at sign-off and a fresh read-only judge checks it. - Cite durable artifacts a future reader can open: committed files, test names, git diffs. .pi/ is - usually gitignored, so files there prove things only at judge time, not in history. + Cite artifacts a reviewer can open: files, test names, saved verification output, git diffs. + Uncommitted and ignored files are valid review evidence when inspected directly. Git history + improves durability; a commit or clean worktree is not a sign-off requirement. - User-visible result: restate the original deliverable, not the proposed implementation. Every goal must contribute to it. Future work may not defer any artifact or action named there. - User voice: quote the human word for word, one line per requirement, as they say it. Never @@ -163,8 +164,8 @@ export function waivesAlignment(objective: string): boolean { export function alignmentPolicy(waived: boolean): string { return waived - ? "Current-plan alignment: the human explicitly waived questions in this objective. Skip the default three-question round for THIS plan only." - : "Current-plan alignment: ask at least THREE task-specific questions in ONE chat round before the final plan; wait for answers and use them. Check the expected result, scope/constraints, and success/failure criteria. A waiver in any previous plan does NOT apply. Do not repeat questions already answered for this plan."; + ? "Current-plan alignment: the human explicitly waived optional questions for THIS plan only. Use the authorized scope; do not invent missing permissions." + : "Current-plan alignment: ask only material unresolved questions about outcome, scope, constraints, or success criteria; there is no fixed quota. Inspect technical facts yourself. Wait for required answers and use them; otherwise present the executable plan for review. A waiver in any previous plan does NOT apply. Do not repeat questions already answered for this plan."; } export const discussPlan = "Continue discussing this draft in normal chat. Ask useful, task-specific alignment questions to check where your understanding differs from the human's: expected result, scope/constraints, and success/failure criteria. Wait for answers; do not open an editor or request review yet. Keep the draft and incorporate answers. When discussion is finished and the plan is ready, call RequestPlanReview, even if the draft is unchanged."; @@ -195,8 +196,9 @@ Keep it current as you work, with your normal edit tool: - when the active goal's discriminator is satisfied, fill its evidence: list (each item = a durable artifact + a verbatim quote you actually observed + a short read of it), then call CompleteGoal. Don't tick a goal [x] before CompleteGoal accepts; the sign-off log line is the audit trail. -- if the working set has grown long, prune finished goals (their evidence lives in git history and - ## Log) and move settled detail down to ## Appendix, which is unlimited +- keep every goal line and its completion status above ## Log; recorded sign-offs depend on those + identities. Keep evidence references beside each goal. If the working set grows long, move only + verbose settled detail down to ## Appendix; do not remove completed goal lines - the human's latest message outranks this plan. If it corrects the deliverable or scope, amend the user-visible result, user voice, and affected goals before continuing; don't defend the old plan - otherwise keep working toward the active goal; don't stop to ask unless genuinely blocked @@ -231,12 +233,13 @@ export const completeGoalDescription = "output. The read must show success POSITIVELY happened, not just that failures were avoided. " + "Check that the claimed result uses the artifact and outcome named in User-visible result and does " + "not substitute an agent-inferred deliverable. Then call this with the goal's text (the line after " + - "'goal:'; small wording drift is fine). A " + + "'goal:'; one unique exact subject, ignoring case and surrounding whitespace). A " + "fresh strictly-read-only judge inspects the LIVE WORKING TREE (uncommitted changes included; " + - "committing first is for durability, not visibility) and returns accept or reject with what's " + + "ignored outputs included; neither a clean worktree nor a commit is required) and returns accept or reject with what's " + "missing. On accept (or if the judge itself failed), a sign-off line is appended to ## Log " + - "and the goal is ticked [x] for you; the result says if you must tick it yourself. On reject the " + - "goal stays open."; + "and the goal is ticked [x] for you with a persisted sign-off record. Judge failure is explicitly " + + "accepted inconclusive, not verified completion. Manual ticks remain claims awaiting this tool. " + + "On reject the goal stays open."; export const completeGoalParamDescription = "The goal's text: the line after 'goal:' in the plan file."; @@ -249,7 +252,9 @@ export const completeGoalParamDescription = "The goal's text: the line after 'go export const judgeSystem = `\ You are a strictly read-only reviewer signing off a coding goal. You cannot execute anything: judge by reading (read/grep/find/ls). Never re-run the work or its verify command -- it may be a 10-hour -job; the agent must bring you its saved output. Your job is evidence discipline, checked in order: +job; the agent must bring you its saved output. Inspect the live working tree, including cited +uncommitted and ignored files via read. Git status is context, not an acceptance gate; do not require +cleanup or a commit unless the agreed goal requires it. Your job is evidence discipline, checked in order: 0. Task fidelity? Read User-visible result and User voice first. Reject if this goal contradicts, replaces, or defers the requested artifact or outcome. Agent-inferred scope is not authority. @@ -283,9 +288,10 @@ The working agent claims this goal is complete: goal: ${p.goal} -Below is the full plan file (${p.planPath}). Find that goal in it (tolerate small wording drift; if -you cannot find a matching goal at all, reject and say so). Read User-visible result and User voice -first, then its discriminator, subtle failure modes, verify command, and evidence list. +Below is the full plan file (${p.planPath}). Review the unique exact goal subject shown above +(ignoring case and surrounding whitespace). If it is missing or ambiguous, reject and request the +exact subject; do not substitute another goal. Read User-visible result and User voice first, then +its discriminator, subtle failure modes, verify command, and evidence list. --- plan file --- ${p.plan} diff --git a/src/supervisor.ts b/src/supervisor.ts index 7c5d508..a73e1c0 100644 --- a/src/supervisor.ts +++ b/src/supervisor.ts @@ -12,12 +12,15 @@ export interface SupervisorBinding { supervisorSession?: string; active?: boolean; stopped?: boolean; + /** User pause: retain the pair, but no autonomous work until explicit resume. */ + paused?: boolean; + pauseId?: string; everyTurns: number; intervalMs: number; compactTokens: number; } export interface Bootstrap { binding: SupervisorBinding; workerId: string } -export interface SupervisorStatus { connected: boolean; binding?: SupervisorBinding; workerId: string; role?: string } +export interface SupervisorStatus { connected: boolean; binding?: SupervisorBinding; workerId: string; role?: string; activity?: string; lastFailure?: string } export interface SupervisorDecision { bindingId: string; goal: string; planHash: string; decision: "approve" | "needs_work" | "needs_user"; reason: string } export function planHash(text: string): string { @@ -32,6 +35,9 @@ export interface SupervisorController { activate(bindingId: string, signal?: AbortSignal): Promise; review(bindingId: string, goal: string, hash: string, signal?: AbortSignal): Promise; stop(bindingId: string): Promise; + pause(exit?: boolean): void; + reconnect(signal?: AbortSignal): Promise; + resume(bindingId: string, hash: string, signal?: AbortSignal): Promise; } export function validBinding(value: unknown): value is SupervisorBinding { @@ -55,7 +61,11 @@ export function pendingReply(signal: AbortSignal | undefined, cancel: () => v settled = true; clearTimeout(timer); signal?.removeEventListener("abort", abort); - if (error) { cancel(); reject(error); } else resolve(value as T); + if (error) { + // Local cancellation must settle even if notifying a disconnected peer throws. + try { cancel(); } catch { /* The caller reports the request failure; remote state is unconfirmed. */ } + reject(error); + } else resolve(value as T); }; signal?.addEventListener("abort", abort, { once: true }); if (signal?.aborted) queueMicrotask(abort); diff --git a/test/goals-flow.test.ts b/test/goals-flow.test.ts index 6c8c34a..2e14ff2 100644 --- a/test/goals-flow.test.ts +++ b/test/goals-flow.test.ts @@ -1,3 +1,5 @@ +import { execFileSync } from "node:child_process"; +import { EventEmitter } from "node:events"; import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; import { tmpdir } from "node:os"; import { join } from "node:path"; @@ -5,7 +7,17 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent"; import { describe, expect, it, vi } from "vitest"; import piGoalsExtension from "../src/index.js"; -vi.mock("../src/internal/supervisor/index.js", () => ({ default: () => {} })); +vi.mock("../src/internal/supervisor/index.js", () => ({ default: () => ({ pause() {}, reconnect: async () => {}, status: async () => ({ connected: false, activity: "inactive" }) }) })); +const judgeRun = vi.hoisted(() => ({ output: "", beforeReply: undefined as (() => void) | undefined, calls: [] as string[][] })); +vi.mock("node:child_process", async (original) => { + const actual = await original(); + return { ...actual, spawn: (_command: string, args: string[]) => { + judgeRun.calls.push(args); + const proc = Object.assign(new EventEmitter(), { stdout: new EventEmitter(), stderr: new EventEmitter(), kill() {} }); + queueMicrotask(() => { judgeRun.beforeReply?.(); proc.stdout.emit("data", judgeRun.output); proc.emit("close", 0); }); + return proc; + } }; +}); function setup( selectChoices: Array, @@ -37,11 +49,12 @@ function setup( modelRegistry: { find: (provider: string, id: string) => ({ provider, id }) }, hasUI: true, isIdle: () => true, + abort: vi.fn(), sessionManager: { getSessionId: () => "session-a", getSessionFile: () => join(cwd, "session-a.jsonl"), getEntries: () => entries, getBranch: () => entries }, ui: { theme: { fg: (_kind: string, text: string) => text }, - setStatus: () => {}, - setWidget: () => {}, + setStatus: vi.fn(), + setWidget: vi.fn(), notify: vi.fn(), select: async () => { events.push("select"); @@ -78,7 +91,203 @@ async function settleDraft(flow: ReturnType) { await flow.hooks.get("agent_settled")({}, flow.ctx); } +describe("/goals recovery", () => { + it("help, status and invalid recovery arguments never start a plan", async () => { + const f = setup([]); + try { + for (const text of ["help", "status", "stop now", "exit now", "reconnect now", "resume now", "--unknown", "connect", "reconect"]) await f.commands.get("goals").handler(text, f.ctx); + expect(f.messages).toHaveLength(0); + expect(f.entries.filter(e => e.customType === "pi-goals-state")).toHaveLength(0); + } finally { rmSync(f.cwd, { recursive: true, force: true }); } + }); + + it.each(["stop", "exit"])("%s keeps the draft, exits the gate and persists through reload and chat", async command => { + const f = setup([]); + try { + await f.commands.get("goals").handler("plan recovery fixture", f.ctx); + const path = join(f.cwd, ".pi/plan/session-a-v1.md"); + writeFileSync(path, "1. [ ] goal: retain this draft\n"); + await f.commands.get("goals").handler(command, f.ctx); + const count = f.messages.length; + await f.hooks.get("session_start")({}, f.ctx); + await f.hooks.get("input")({ text: "ordinary chat", source: "interactive" }, f.ctx); + await f.hooks.get("agent_settled")({}, f.ctx); + expect(f.messages).toHaveLength(count); + expect(await f.hooks.get("input")({ text: "Process finished", source: "extension" }, f.ctx)).toBeUndefined(); + expect(await f.hooks.get("input")({ text: "Work the goals in stale.md", source: "extension" }, f.ctx)).toEqual({ action: "handled" }); + expect(readFileSync(path, "utf8")).toBe("1. [ ] goal: retain this draft\n"); + expect(await f.hooks.get("tool_call")({ toolName: "write", input: { path: "other.txt" } }, f.ctx)).toBeUndefined(); + expect((f.entries.filter(e => e.customType === "pi-goals-state").at(-1)?.data as any).pausedFrom).toBe("planning"); + await f.commands.get("goals").handler("resume", f.ctx); + expect(f.messages).toHaveLength(count); // returning to a draft is NOT Ready + expect((f.entries.filter(e => e.customType === "pi-goals-state").at(-1)?.data as any).phase).toBe("planning"); + } finally { rmSync(f.cwd, { recursive: true, force: true }); } + }); + + it("stop invalidates an outstanding Ready selection", async () => { + const f = setup([]); + try { + await f.commands.get("goals").handler("plan cancel selection", f.ctx); + writeFileSync(join(f.cwd, ".pi/plan/session-a-v1.md"), "1. [ ] goal: not authorized\n"); + let answer!: (s: string) => void; + f.ctx.ui.select = () => new Promise(resolve => { answer = resolve; }); + const selecting = settleDraft(f); + await new Promise(resolve => setImmediate(resolve)); + await f.commands.get("goals").handler("stop", f.ctx); + answer("Ready"); await selecting; + expect(f.messages.some(m => m.content.startsWith("Work the goals"))).toBe(false); + expect((f.entries.filter(e => e.customType === "pi-goals-state").at(-1)?.data as any).phase).toBeNull(); + } finally { rmSync(f.cwd, { recursive: true, force: true }); } + }); + + it("resume starts unchanged authorized work but a changed plan returns to review", async () => { + const f = setup(["Ready"]); + try { + await f.commands.get("goals").handler("plan bounded work", f.ctx); + await f.commands.get("goals").handler("steward off", f.ctx); + const path = join(f.cwd, ".pi/plan/session-a-v1.md"); + writeFileSync(path, "1. [ ] goal: first\n"); + await settleDraft(f); + await f.commands.get("goals").handler("stop", f.ctx); + let count = f.messages.length; + await f.commands.get("goals").handler("resume", f.ctx); + expect(f.messages.length).toBe(count + 1); + await f.commands.get("goals").handler("exit", f.ctx); + writeFileSync(path, "1. [ ] goal: changed scope\n"); + count = f.messages.length; + await f.commands.get("goals").handler("resume", f.ctx); + expect(f.messages).toHaveLength(count); + expect((f.entries.filter(e => e.customType === "pi-goals-state").at(-1)?.data as any).phase).toBe("planning"); + } finally { await f.hooks.get("session_shutdown")({}, f.ctx); rmSync(f.cwd, { recursive: true, force: true }); } + }); +}); + describe("/goals draft flow", () => { + it.each(["accept", "inconclusive", "reject", "abort", "shutdown", "stop", "plan change"])("records only legitimate sign-offs on a dirty repo with ignored artifacts: %s", async (verdict) => { + const flow = setup([]); const fresh = setup([]); const abort = new AbortController(); + try { + execFileSync("git", ["init", "-q", flow.cwd]); + writeFileSync(join(flow.cwd, ".gitignore"), "outputs/\n.pi/\n"); + writeFileSync(join(flow.cwd, "unrelated.txt"), "preserve this unrelated file\n"); + execFileSync("git", ["-C", flow.cwd, "add", "unrelated.txt"]); // isolated fixture only, no commit + writeFileSync(join(flow.cwd, "unrelated.txt"), "preserve this unrelated dirty edit\n"); + mkdirSync(join(flow.cwd, "outputs")); + writeFileSync(join(flow.cwd, "outputs/artifact.txt"), "hello\n"); + const observed = execFileSync(process.execPath, ["-e", "const fs=require('fs'); if(fs.readFileSync('outputs/artifact.txt','utf8')!=='hello\\n') process.exit(1); console.log('PASS exact bytes');"], { cwd: flow.cwd, encoding: "utf8" }); + writeFileSync(join(flow.cwd, "outputs/verify.log"), observed); + expect(execFileSync("git", ["-C", flow.cwd, "check-ignore", "outputs/verify.log"], { encoding: "utf8" })).toContain("outputs/verify.log"); + const dirty = execFileSync("git", ["-C", flow.cwd, "status", "--porcelain"], { encoding: "utf8" }); + expect(dirty).toContain("AM unrelated.txt"); + mkdirSync(join(flow.cwd, ".pi/plan"), { recursive: true }); + const file = join(flow.cwd, ".pi/plan/session-a-v1.md"); + writeFileSync(file, "# Plan\n1. [x] goal: exact output\n - evidence: outputs/artifact.txt `hello`; outputs/verify.log `PASS exact bytes` from node byte check\n"); + flow.entries.push({ type: "custom", customType: "pi-goals-state", data: { phase: "working", planVersion: 1, signedOffGoals: [], stewardEnabled: false, autoIntervalMs: null } }); + await flow.hooks.get("session_start")({}, flow.ctx); + judgeRun.calls = []; + judgeRun.output = verdict === "inconclusive" ? "No verdict from judge" : verdict === "reject" ? "VERDICT: reject\nmissing: test failed" : `checks:\n- outputs/verify.log: \`${readFileSync(join(flow.cwd, "outputs/verify.log"), "utf8").trim()}\`; execution passed\nVERDICT: accept\nmissing:`; + judgeRun.beforeReply = () => { + if (verdict === "abort") abort.abort(); + if (verdict === "shutdown") void flow.hooks.get("session_shutdown")({}, flow.ctx); + if (verdict === "stop") void flow.commands.get("goals").handler("stop", flow.ctx); + if (verdict === "plan change") writeFileSync(file, readFileSync(file, "utf8") + "\nNew scope requiring review\n"); + }; + const outcome = await flow.tools.get("CompleteGoal").execute("", { goal: " EXACT OUTPUT " }, abort.signal, undefined, flow.ctx); + const signed = verdict === "accept" || verdict === "inconclusive"; + expect(outcome.isError).toBe(!signed); + expect(judgeRun.calls).toHaveLength(1); + expect(judgeRun.calls[0]).toContain("--no-extensions"); + expect(judgeRun.calls[0]).toContain("read,grep,find,ls"); + expect(readFileSync(file, "utf8").includes("[x] goal: exact output")).toBe(signed); + if (verdict === "inconclusive") expect(outcome.content[0].text).toContain("not verified completion"); + expect(readFileSync(join(flow.cwd, "unrelated.txt"), "utf8")).toBe("preserve this unrelated dirty edit\n"); + expect(execFileSync("git", ["-C", flow.cwd, "status", "--porcelain"], { encoding: "utf8" })).toBe(dirty); + fresh.entries.push(...structuredClone(flow.entries)); + fresh.ctx.cwd = flow.cwd; // genuinely new extension instance, same plan and persisted entries + await fresh.hooks.get("session_start")({}, fresh.ctx); + expect(fresh.ctx.ui.setStatus.mock.lastCall?.[1]).toContain(verdict === "stop" ? "goals stopped" : signed ? "1/1" : "0/1"); + expect(fresh.messages).toHaveLength(0); + } finally { + judgeRun.beforeReply = undefined; judgeRun.calls = []; + await flow.hooks.get("session_shutdown")({}, flow.ctx); await fresh.hooks.get("session_shutdown")({}, fresh.ctx); + rmSync(flow.cwd, { recursive: true, force: true }); rmSync(fresh.cwd, { recursive: true, force: true }); + } + }); + + it("preserves legacy completion on reload without inventing sign-off or restarting work", async () => { + vi.useFakeTimers(); const flow = setup([]); + try { + mkdirSync(join(flow.cwd, ".pi/plan"), { recursive: true }); + const file = join(flow.cwd, ".pi/plan/session-a-v1.md"); + const plan = '1. [x] goal: historical result\n\n## Log\n- signed off "historical result" (judge accept)\n'; + writeFileSync(file, plan); + flow.entries.push({ type: "custom", customType: "pi-goals-state", data: { phase: "working", planVersion: 1, stewardEnabled: false, autoIntervalMs: 1000 } }); + for (let n = 0; n < 2; n++) { + await flow.hooks.get("session_start")({}, flow.ctx); + await flow.hooks.get("agent_settled")({}, flow.ctx); + await vi.advanceTimersByTimeAsync(5000); + expect(flow.ctx.ui.setWidget.mock.lastCall?.[1]?.join("\n")).toContain("legacy completion — sign-off not recorded"); + expect(flow.messages).toHaveLength(0); + expect(readFileSync(file, "utf8")).toBe(plan); + } + } finally { await flow.hooks.get("session_shutdown")({}, flow.ctx); vi.useRealTimers(); rmSync(flow.cwd, { recursive: true, force: true }); } + }); + it("keeps unsigned manual ticks visible across reload without reopening legitimate sign-offs", async () => { + const flow = setup([]); + try { + flow.entries.push({ type: "custom", customType: "pi-goals-state", data: { phase: "working", planVersion: 1, stewardEnabled: false, autoIntervalMs: null, signedOffGoals: [{ subject: "first", outcome: "accept" }, { subject: "second", outcome: "inconclusive" }] } }); + mkdirSync(join(flow.cwd, ".pi/plan"), { recursive: true }); + const file = join(flow.cwd, ".pi/plan/session-a-v1.md"); + writeFileSync(file, "# Plan\n1. [x] goal: first\n2. [x] goal: second\n3. [x] goal: manual claim\n"); + for (let i = 0; i < 2; i++) { + await flow.hooks.get("session_start")({}, flow.ctx); + expect(flow.ctx.ui.setStatus.mock.lastCall?.[1]).toContain("2/3"); + expect(flow.ctx.ui.setWidget.mock.lastCall?.[1]?.join("\n")).toMatch(/claimed.*manual claim/); + expect(flow.ctx.ui.setWidget.mock.lastCall?.[1]?.join("\n")).toMatch(/inconclusive.*second/); + } + writeFileSync(file, readFileSync(file, "utf8").replace("[x] goal: first", "[/] goal: first")); + await flow.hooks.get("turn_end")({}, flow.ctx); + writeFileSync(file, readFileSync(file, "utf8").replace("[/] goal: first", "[x] goal: first")); + await flow.hooks.get("session_start")({}, flow.ctx); + expect(flow.ctx.ui.setStatus.mock.lastCall?.[1]).toContain("1/3"); + expect(flow.ctx.ui.setWidget.mock.lastCall?.[1]?.join("\n")).toMatch(/claimed.*first/); + } finally { await flow.hooks.get("session_shutdown")({}, flow.ctx); rmSync(flow.cwd, { recursive: true, force: true }); } + }); + + it("retains conclusive and inconclusive sign-offs when settled detail moves below the fold", async () => { + const flow = setup([]); + try { + const signoffs = [{ subject: "first", outcome: "accept" }, { subject: "second", outcome: "inconclusive" }]; + flow.entries.push({ type: "custom", customType: "pi-goals-state", data: { phase: "working", planVersion: 1, stewardEnabled: false, autoIntervalMs: null, signedOffGoals: signoffs } }); + mkdirSync(join(flow.cwd, ".pi/plan"), { recursive: true }); + const file = join(flow.cwd, ".pi/plan/session-a-v1.md"); + const goals = "# Plan\n1. [x] goal: first\n - evidence: outputs/first.log\n2. [x] goal: second\n - evidence: outputs/second.log\n"; + writeFileSync(file, `${goals} - settled detail: implementation notes\n\n## Log\n`); + await flow.hooks.get("session_start")({}, flow.ctx); + writeFileSync(file, `${goals}\n## Log\n\n## Appendix\nSettled detail: implementation notes\n`); + await flow.hooks.get("turn_end")({}, flow.ctx); + await flow.hooks.get("session_start")({}, flow.ctx); + expect(flow.ctx.ui.setStatus.mock.lastCall?.[1]).toContain("2/2"); + expect(flow.ctx.ui.setWidget.mock.lastCall?.[1]?.join("\n")).toMatch(/inconclusive.*second/); + expect((flow.entries.at(-1)?.data as any).signedOffGoals).toEqual(signoffs); + expect(flow.messages).toHaveLength(0); + } finally { await flow.hooks.get("session_shutdown")({}, flow.ctx); rmSync(flow.cwd, { recursive: true, force: true }); } + }); + + it.each(["missing", "first with wording drift", "duplicate"])("requires unique goal identity before review: %s", async (goal) => { + const flow = setup([]); + try { + flow.entries.push({ type: "custom", customType: "pi-goals-state", data: { phase: "working", planVersion: 1, stewardEnabled: true, autoIntervalMs: null } }); + mkdirSync(join(flow.cwd, ".pi/plan"), { recursive: true }); + const file = join(flow.cwd, ".pi/plan/session-a-v1.md"); + const plan = "1. [ ] goal: first\n2. [x] goal: duplicate\n3. [ ] goal: duplicate\n"; + writeFileSync(file, plan); + await flow.hooks.get("session_start")({}, flow.ctx); + const outcome = await flow.tools.get("CompleteGoal").execute("", { goal }, undefined, undefined, flow.ctx); + expect(outcome.isError).toBe(true); + expect(outcome.content[0].text).toMatch(/unique exact goal/i); + expect(readFileSync(file, "utf8")).toBe(plan); + } finally { await flow.hooks.get("session_shutdown")({}, flow.ctx); rmSync(flow.cwd, { recursive: true, force: true }); } + }); it("reopens a prematurely ticked submitted goal before a failed sign-off", async () => { const flow = setup([]); try { diff --git a/test/internal-supervisor/index.test.ts b/test/internal-supervisor/index.test.ts index 5dff7d6..bb6448e 100644 --- a/test/internal-supervisor/index.test.ts +++ b/test/internal-supervisor/index.test.ts @@ -366,7 +366,8 @@ test("the second view carries only what happened after the first", async () => { await new Promise(resolve => setTimeout(resolve, 50)); const second = worker.published.filter((p) => p.t === "view").at(-1).view; assert.match(second, /THE SECOND THING/); - assert.doesNotMatch(second, /THE FIRST INSTRUCTION/, "the supervisor already read this one"); + assert.match(second, /# Latest user direction\nTHE FIRST INSTRUCTION/, "direction and human pauses remain available outside incremental history"); + assert.doesNotMatch(second.split("# New turns since your last look")[1], /THE FIRST INSTRUCTION/, "turn history stays incremental"); }); test("a message addressed to a different session is ignored", async () => { @@ -1345,13 +1346,16 @@ test("a human message in the worker session is not a reason to stand back", asyn await sup.start(); await sup.run("supervise", "@worker make the results table"); - assert.match(sup.contextMessages.at(-1)!.content, /not a handover, and it is not a reason to stand back/); + assert.match(sup.contextMessages.at(-1)!.content, /new direction, not an automatic handover/); + assert.match(sup.contextMessages.at(-1)!.content, /do not steer the paused work until authorized/); + assert.match(sup.contextMessages.at(-1)!.content, /question or pause is not a reason to discard them/); assert.match(sup.tools.get("let_it_run")!.description, /A human message does not end supervision/); // And on the view that carries a stopped worker, where the excuse actually got used. sup.deliver(WORKER_ID, { t: "view", to: SUPER_ID, view: "worker view", stopped: true }); await new Promise((r) => setTimeout(r, 5)); - assert.match(sup.userMessages[0].content, /concrete continuation if work remains/); + assert.match(sup.userMessages[0].content, /steer a useful authorized continuation/); + assert.match(sup.userMessages[0].content, /Respect explicit human pauses/); }); test("letting a stopped worker run says plainly that the worker stays stopped", async () => { diff --git a/test/internal-supervisor/plan.test.ts b/test/internal-supervisor/plan.test.ts index e7382bd..369437e 100644 --- a/test/internal-supervisor/plan.test.ts +++ b/test/internal-supervisor/plan.test.ts @@ -8,6 +8,62 @@ import extension from "../../src/internal/supervisor/index.js"; import { type SupervisorBinding as PlanBinding, planHash } from "../../src/supervisor.js"; const tick = () => new Promise(resolve => setImmediate(resolve)); + +test("human pause survives reload/reconnect; only explicit resume reactivates the same pair", async () => { + const h = pairHarness(); + try { + await h.start(); + await h.worker.controller.activate(h.binding.id); + await tick(); + h.worker.controller.pause(); await tick(); + assert.equal((await h.worker.controller.status()).binding.paused, true); + assert.equal((await h.supervisor.controller.status()).binding.active, false); + const reloaded = await h.restart(h.supervisor, "supervisor-reloaded"); + const views = h.wire.filter(w => w.t === "view").length; + await h.worker.controller.reconnect(); await tick(); + assert.equal((await reloaded.controller.status()).binding.paused, true); + assert.equal(h.wire.filter(w => w.t === "view").length, views); + const pairs = h.wire.filter(w => w.t === "pair").length; + await h.worker.controller.resume(h.binding.id, planHash(readFileSync(h.planPath, "utf8"))); + assert.equal((await reloaded.controller.status()).binding.paused, false); + assert.equal((await h.worker.controller.status()).binding.active, true); + assert.equal(h.wire.filter(w => w.t === "pair").length, pairs); + const oldResume = h.wire.findLast(w => w.t === "plan_resume"); + h.worker.controller.pause(); await tick(); + reloaded.receive({ type: "message", fromSessionId: "worker", payload: oldResume }); await tick(); + assert.equal((await reloaded.controller.status()).binding.paused, true, "a delayed old resume cannot undo a newer stop"); + } finally { await h.close(); } +}); + +test("stop cancels an in-flight checkpoint even when the cancel notification cannot be sent", async () => { + const h = pairHarness(); + try { + await h.start(); + const review = h.worker.controller.review(h.binding.id, "first", planHash(readFileSync(h.planPath, "utf8"))); + const cancelled = assert.rejects(review, /Stopped/); + await tick(); + h.worker.connected = false; + h.worker.controller.pause(); + await cancelled; + assert.equal((await h.worker.controller.status()).binding.paused, true); + } finally { await h.close(); } +}); + +test("supervisor pause cancels checkpoint and a disconnected local stop still persists", async () => { + const h = pairHarness(); + try { + await h.start(); + const pending = h.worker.controller.review(h.binding.id, "first", planHash(readFileSync(h.planPath, "utf8"))); + const cancelled = assert.rejects(pending, /cancelled|Stopped/); + await tick(); + h.supervisor.controller.pause(true); await tick(); await cancelled; + assert.equal((await h.worker.controller.status()).binding.paused, true); + h.worker.connected = false; + assert.doesNotThrow(() => h.worker.controller.pause()); + assert.equal((await h.worker.controller.status()).activity, "stopped by user"); + await assert.rejects(h.worker.controller.resume(h.binding.id, planHash(readFileSync(h.planPath, "utf8"))), /Intercom/); + } finally { await h.close(); } +}); function pairHarness() { const cwd = mkdtempSync(join(tmpdir(), "supervise-plan-")); const peers: any[] = []; @@ -40,7 +96,10 @@ function pairHarness() { const ctx: any = { cwd, hasUI: true, model: { provider: "native", id: "test", contextWindow: 200_000 }, isIdle: () => peer.idle, abort: () => { peer.aborts++; }, getContextUsage: () => ({ tokens: peer.tokens }), compact({ onComplete }: any) { peer.compactions++; peer.tokens = 10_000; onComplete({}); }, ui: { notify() {}, setStatus() {}, theme: { fg: (_: any, text: string) => text } }, sessionManager: { getEntries: () => entries, getBranch: () => entries, getSessionFile: () => sessionFile } }; peer.pi = pi; peer.ctx = ctx; peer.hook = async (name: string, event = {}) => { let result: any; for (const fn of hooks.get(name) ?? []) result = await fn(event, ctx) ?? result; return result; }; - peers.push(peer); const controller = extension(pi, () => peer.modelReady); + peers.push(peer); const controller = extension(pi, () => peer.modelReady, (plan) => { + const goals = plan.match(/^\d+\. \[[ x/]\] goal:/gm) ?? []; + return { completion: { planHash: planHash(plan), total: goals.length, pending: goals.length - (peer.signedOffCount ?? 0), inconclusive: peer.inconclusiveCount ?? 0 }, summary: `Manual completion claims await CompleteGoal; ${peer.signedOffCount ?? 0} recorded sign-offs.` }; + }); peer.controller = controller; peer.request = (method: string, p: any = {}) => { switch (method) { @@ -70,6 +129,51 @@ function pairHarness() { return { worker, supervisor, binding, wire, planPath, async restart(peer: any, id: string) { await peer.hook("session_shutdown"); peers.splice(peers.indexOf(peer), 1); const replacement = make(id, structuredClone(peer.entries), peer.ctx.sessionManager.getSessionFile()); await replacement.hook("session_start"); await tick(); return replacement; }, async start() { await worker.hook("session_start"); await supervisor.hook("session_start"); await worker.request("prepare", { binding }); const attached = worker.request("attached", { bindingId: binding.id }); void attached.catch(() => {}); await supervisor.request("bootstrap", { binding, workerId: "worker" }); await attached; }, async close() { for (const peer of peers) await peer.hook("session_shutdown"); rmSync(cwd, { recursive: true, force: true }); } }; } +for (const assessment of ["On course; the saved check is the next useful evidence.", "Which output format do you want?"]) test(`ordinary prose leaves future reviews live: ${assessment}`, async () => { + const h = pairHarness(); try { + await h.start(); await h.worker.controller.activate(h.binding.id); await tick(); + const before = h.supervisor.messages.length; + const looks = h.wire.filter((w: any) => w.t === "look").length; + await h.supervisor.hook("agent_end", { messages: [{ role: "assistant", stopReason: "stop", content: [{ type: "text", text: assessment }] }] }); + for (let i = 0; i < 3; i++) { await h.supervisor.hook("agent_settled"); await tick(); } + assert.equal(h.supervisor.messages.length, before, "no immediate idle retry"); + assert.equal(h.wire.filter((w: any) => w.t === "look").length, looks); + h.worker.entries.push({ type: "message", message: { role: "user", content: "Use the agreed text format. Pause deployment until I authorize it; continue independent checks." } }); + await h.supervisor.receive({ type: "message", fromSessionId: "worker", payload: { t: "view", to: "supervisor", bindingId: h.binding.id, view: "New worker direction and progress", stopped: true } }); + await tick(); + assert.equal(h.supervisor.messages.length, before + 1, "new worker progress is not discarded after prose or a real question"); + await h.supervisor.finishAssessment(); + await h.supervisor.commands.get("supervise").handler("look", h.supervisor.ctx); await tick(); + assert.match(h.supervisor.messages.at(-1).text, /Pause deployment until I authorize it/); + } finally { await h.close(); } +}); + +test("manual last-goal ticks cannot end plan supervision before CompleteGoal", async () => { + const h = pairHarness(); try { + await h.start(); + writeFileSync(h.planPath, "# Plan\n1. [x] goal: first\n2. [x] goal: second\n"); + await h.worker.controller.activate(h.binding.id); await tick(); + await assert.rejects(h.supervisor.tools.get("done").execute("", { reason: "All boxes checked" }, undefined, undefined, h.supervisor.ctx), /sign.off|claim|Open plan goals/i); + assert.equal((await h.worker.controller.status()).connected, true); + assert.equal(h.wire.filter((w: any) => w.t === "done").length, 0); + } finally { await h.close(); } +}); + +test("completion counts are plan-bound and distinguish inconclusive sign-off when ending supervision", async () => { + const h = pairHarness(); try { + await h.start(); + writeFileSync(h.planPath, "# Plan\n1. [x] goal: first\n2. [x] goal: second\n"); + h.worker.signedOffCount = 2; h.worker.inconclusiveCount = 1; + await h.worker.controller.activate(h.binding.id); await tick(); + const plan = readFileSync(h.planPath, "utf8"); + writeFileSync(h.planPath, plan + "3. [x] goal: unreviewed addition\n"); + await assert.rejects(h.supervisor.tools.get("done").execute("", { reason: "Old counts said complete" }, undefined, undefined, h.supervisor.ctx), /fresh CompleteGoal tracking/); + writeFileSync(h.planPath, plan); + await h.supervisor.tools.get("done").execute("", { reason: "Conclusive first goal; second accepted inconclusive" }, undefined, undefined, h.supervisor.ctx); + assert.ok(h.supervisor.contexts.some((m: any) => /accepted inconclusive.*not independently verified/.test(m.content))); + } finally { await h.close(); } +}); + test("each checkpoint freezes fresh worker evidence and direction without replacing a busy assessment", async () => { const h = pairHarness(); try { await h.start(); @@ -212,13 +316,14 @@ test("empty low-level response may continue through compaction, ask a human, or assert.equal(h.wire.filter((w: any) => w.t === "goal_decision").length, 0); await h.supervisor.hook("agent_end", { messages: [{ role: "assistant", stopReason: "stop", content: [{ type: "text", text: "Should we keep the output format?" }] }] }); await h.supervisor.hook("agent_settled"); await tick(); - assert.equal(h.wire.filter((w: any) => w.t === "goal_decision").length, 0, "nonempty human question still waits"); + assert.equal((await pending).decision, "needs_work", "prose alone cannot leave a checkpoint pending indefinitely"); + const retry = h.worker.controller.review(h.binding.id, "first", planHash(readFileSync(h.planPath, "utf8"))); await tick(); await h.supervisor.hook("context", { messages: h.supervisor.contexts }); await h.supervisor.tools.get("review_goal").execute("approved", { decision: "approve", reason: "The user confirmed the format" }); await h.supervisor.hook("agent_end", empty); await h.supervisor.hook("agent_settled"); await tick(); - assert.equal((await pending).decision, "approve"); - assert.equal(h.wire.filter((w: any) => w.t === "goal_decision").length, 1); + assert.equal((await retry).decision, "approve"); + assert.equal(h.wire.filter((w: any) => w.t === "goal_decision").length, 2); } finally { await h.close(); } }); @@ -502,6 +607,7 @@ for (const stop of ["command", "done"]) test(`${stop} preserves a stopped superv if (stop === "command") await h.supervisor.commands.get("supervise").handler("stop", h.supervisor.ctx); else { writeFileSync(h.planPath, "# Plan\n\n1. [x] goal: first\n2. [x] goal: second\n"); + h.worker.signedOffCount = 2; await h.worker.request("activate", { bindingId: h.binding.id }); await tick(); await h.supervisor.tools.get("done").execute("", { reason: "All goals accepted" }, undefined, undefined, h.supervisor.ctx); } @@ -803,10 +909,15 @@ test("explicit checkpoints wait separately from routine coalescing and do not re await h.supervisor.review("needs_user", "The user must choose the output format"); assert.equal((await pending).decision, "needs_user"); const looks = h.wire.filter((wire: any) => wire.t === "look").length; - h.supervisor.receive({ type: "message", fromSessionId: "worker", payload: { t: "view", to: "supervisor", bindingId: h.binding.id, view: "More routine status", stopped: false } }); - await h.supervisor.hook("agent_settled"); await tick(); - assert.equal(h.wire.filter((wire: any) => wire.t === "look").length, looks, "awaiting a user decision must not wake routine reviews"); - assert.equal(h.supervisor.messages.length, 1); + await h.supervisor.receive({ type: "message", fromSessionId: "worker", payload: { t: "view", to: "supervisor", bindingId: h.binding.id, view: "More routine status", stopped: false } }); + await tick(); + for (let n = 0; n < 3; n++) { await h.supervisor.hook("agent_settled"); await tick(); } + assert.equal(h.wire.filter((wire: any) => wire.t === "look").length, looks, "a settled response must not poll idle work"); + assert.equal(h.supervisor.messages.length, 2, "a real needs_user decision must not discard subsequent views"); + await h.supervisor.finishAssessment(); + h.worker.entries.push({ type: "message", message: { role: "user", content: "Use text output; continue goal two now." } }); + await h.supervisor.commands.get("supervise").handler("look", h.supervisor.ctx); await tick(); + assert.match(h.supervisor.messages.at(-1).text, /Use text output; continue goal two now/); } finally { await h.close(); } }); diff --git a/test/internal-supervisor/protocol.test.ts b/test/internal-supervisor/protocol.test.ts index 41f3cc3..5f154e9 100644 --- a/test/internal-supervisor/protocol.test.ts +++ b/test/internal-supervisor/protocol.test.ts @@ -12,6 +12,14 @@ test("goal_review accepts an optional bounded snapshot but rejects malformed sna } }); +test("worker completion metadata is optional but cannot claim malformed counts", () => { + const view = { t: "view", to: "supervisor", view: "Worker evidence", stopped: true }; + const completion = { planHash: "hash", total: 2, pending: 1, inconclusive: 1 }; + assert.equal(isWire(view), true); + assert.equal(isWire({ ...view, completion }), true); + for (const bad of [null, {}, { ...completion, total: -1 }, { ...completion, pending: 0.5 }, { ...completion, inconclusive: 2 }, { ...completion, total: Infinity }, { ...completion, total: "2" }, { ...completion, planHash: false }]) assert.equal(isWire({ ...view, completion: bad }), false); +}); + test("checkpoint snapshot bounding counts JSON escapes and does not split Unicode characters", () => { const review = { requestId: "request", bindingId: "binding", goal: "goal ".repeat(1800), planHash: "hash" }; const wire = goalReviewWire("supervisor", review, '😀\\"\n'.repeat(2000)); diff --git a/test/prompts.test.ts b/test/prompts.test.ts index eb8efc5..9705031 100644 --- a/test/prompts.test.ts +++ b/test/prompts.test.ts @@ -1,15 +1,15 @@ import { describe, expect, it } from "vitest"; -import { alignmentPolicy, judgeSystem, planDrafting, planningState, reminder, resync, waivesAlignment } from "../src/prompts.js"; +import { alignmentPolicy, judgeSystem, judgeUser, planDrafting, planningState, reminder, resync, waivesAlignment } from "../src/prompts.js"; describe("planning prompt", () => { it("requires fact finding or a focused question before a goal", () => { expect(planDrafting).toContain("Use read-only repository tools or web search when either can\nresolve a fact."); expect(planDrafting).toContain("ask the human to confirm your interpretation"); expect(planDrafting).toContain("approve an editorial or other preference choice"); - expect(planDrafting).toContain("ask at least THREE distinct task-specific alignment"); - expect(planDrafting).toContain("questions in ONE chat round"); - expect(planDrafting).toContain("Wait for the human's answers and use them"); - expect(planDrafting).toContain("self-contained: state the relevant\ncontext, use the human's language and ASD-STE100"); + expect(planDrafting).toContain("Ask material task-specific alignment questions"); + expect(planDrafting).toContain("There is no fixed question quota"); + expect(planDrafting).toContain("Wait for answers to required decisions"); + expect(planDrafting).toContain("self-contained:\nstate the relevant context, use the human's language and ASD-STE100"); expect(planDrafting).toContain("placeholder goal such as \"work out the thing\""); expect(planDrafting).toContain("object, observable result, settled scope, and required approval"); }); @@ -29,6 +29,20 @@ describe("planning prompt", () => { expect(planningState(".pi/plan/test.md")).toContain("self-contained round with relevant context and a recommendation"); }); + it("keeps signed-off goal identities during plan housekeeping", () => { + const text = reminder("plan", ".pi/plan/test.md"); + expect(text).toContain("keep every goal line and its completion status above ## Log"); + expect(text).toContain("evidence references beside each goal"); + expect(text).not.toContain("prune finished goals"); + expect(text).not.toContain("evidence lives in git history"); + }); + + it("gives the judge an exact subject rather than a fuzzy-match fallback", () => { + const text = judgeUser({ goal: "first", plan: "1. [ ] goal: first", planPath: "plan.md" }); + expect(text).toContain("unique exact goal subject"); + expect(text).not.toContain("tolerate small wording drift"); + }); + it("anchors work and sign-off to the user-visible result", () => { expect(planDrafting).toContain("## User-visible result"); expect(planDrafting).toContain("Take it from the original request, not from your implementation plan"); @@ -37,5 +51,7 @@ describe("planning prompt", () => { expect(resync("plan", ".pi/plan/test.md", "Compacted.")).toContain("amend the plan rather than preserving an obsolete decision"); expect(judgeSystem).toContain("Task fidelity?"); expect(judgeSystem).toContain("Agent-inferred scope is not authority"); + expect(judgeSystem).toContain("uncommitted and ignored files via read"); + expect(judgeSystem).toContain("Git status is context, not an acceptance gate"); }); }); diff --git a/test/supervisor-integration.test.ts b/test/supervisor-integration.test.ts index cae1306..365b97c 100644 --- a/test/supervisor-integration.test.ts +++ b/test/supervisor-integration.test.ts @@ -210,7 +210,13 @@ describe("actual goals and supervisor package hooks (Herdr and judge mocked)", ( const completed = await completion; expect(completed.isError, JSON.stringify(completed)).toBe(false); expect(readFileSync(path, "utf8")).toContain(`[x] goal: ${goal}`); + const records = manager.getBranch().findLast((entry: any) => entry.customType === "pi-goals-state") as any; + expect(records.data.signedOffGoals).toContainEqual({ subject: goal, outcome: "accept" }); } + await supervisor.commands.get("supervise").handler("look", supervisor.ctx); await tick(); + expect(wires.findLast(w => w.t === "view").completion).toMatchObject({ total: 2, pending: 0, inconclusive: 0 }); + await supervisor.tools.get("let_it_run").execute("reviewed", { reason: "Both goals signed off" }, undefined, undefined, supervisor.ctx); + await supervisor.hook("agent_settled"); expect(judge.calls).toHaveLength(2); expect(judge.calls.every(args => args.includes("--no-extensions") && args.includes("offline/judge"))).toBe(true); expect(worker.ctx.model.id).toBe("worker"); expect(supervisor.ctx.model.id).toBe("supervisor"); @@ -223,6 +229,8 @@ describe("actual goals and supervisor package hooks (Herdr and judge mocked)", ( await tick(); await worker.commands.get("goals").handler("steward off", worker.ctx); expect((await pending).isError).toBe(true); await tick(); expect(judge.calls).toHaveLength(2); expect(supervisor.aborts).toBeGreaterThan(0); + const cancelledRecords = manager.getBranch().findLast((entry: any) => entry.customType === "pi-goals-state") as any; + expect(cancelledRecords.data.signedOffGoals).toEqual([{ subject: "second", outcome: "accept" }]); } finally { for (const peer of peers) await peer.hook("session_shutdown"); rmSync(cwd, { recursive: true, force: true }); } }, 15_000); }); diff --git a/test/tick-goal.test.ts b/test/tick-goal.test.ts index f5655c8..4faeba0 100644 --- a/test/tick-goal.test.ts +++ b/test/tick-goal.test.ts @@ -13,7 +13,7 @@ const plan = `# Plan ## Log `; -describe("tickGoal (sign-off ticks the goal; agent only ticks on wording drift)", () => { +describe("tickGoal (recorded sign-off ticks one exact goal)", () => { it("ticks the exact-matching goal line, case-insensitive, leaving subtasks alone", () => { const out = tickGoal(plan, "implement the CACHE layer"); expect(out).toContain("1. [x] goal: Implement the cache layer"); @@ -21,7 +21,7 @@ describe("tickGoal (sign-off ticks the goal; agent only ticks on wording drift)" expect(out).toContain("2. [ ] goal: Ship the docs"); // other goal untouched }); - it("returns null on wording drift (fuzzy matching is the judge's job, not TypeScript's)", () => { + it("returns null on wording drift so the caller can request the exact subject", () => { expect(tickGoal(plan, "Implement caching")).toBeNull(); });