Make the main session supervise a retained worker

Co-Authored-By: Pi Codex <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
wassname
2026-09-05 17:12:24 +08:00
co-authored by Pi Codex
parent b32f4af11f
commit f87b8aac2f
18 changed files with 818 additions and 1088 deletions
+20 -18
View File
@@ -1,6 +1,6 @@
# pi-goals
Make a short list of goals in one Markdown plan file. A persistent read-only subagent keeps the high-level context, reviews progress, and checks whether each goal is complete.
Make a short list of goals in one Markdown plan file. The main Pi agent supervises a cheaper retained worker through pi-subagents.
The plan file looks like this:
@@ -43,15 +43,16 @@ The plan file looks like this:
Like [pi-milestones](https://github.com/Neuron-Mr-White/UniPi/tree/main/packages/milestone) and
[burneikis/pi-plan](https://github.com/burneikis/pi-plan), it guides rather than guards. The
reminder cadence is copied from [tintinweb/pi-tasks](https://github.com/tintinweb/pi-tasks) and the
resync-after-compaction from [tmonk/pi-goal-x](https://github.com/tmonk/pi-goal-x).
plan resync after compaction follows [tmonk/pi-goal-x](https://github.com/tmonk/pi-goal-x).
## Install
Requires `pi-subagents` 0.65.1 or newer.
Requires `pi-subagents` 0.65.1 or newer. Install `pi-processes` so the supervisor can check managed processes. Install `pi-vcc` as a Pi extension for main-session compaction; pi-goals also loads its package in the worker.
```bash
pi install npm:pi-subagents
pi install npm:@aliou/pi-processes
pi install npm:@sting8k/pi-vcc
pi install npm:@wassname2/pi-goals
```
@@ -59,7 +60,7 @@ Or for development:
```bash
git clone https://github.com/wassname/pi-goals && cd pi-goals && npm install
pi -e npm:pi-subagents -e ./src/index.ts
pi -e npm:pi-subagents -e npm:@sting8k/pi-vcc -e ./src/index.ts
```
## Use
@@ -74,28 +75,29 @@ pi -e npm:pi-subagents -e ./src/index.ts
2. Review. After Pi settles, the full plan is printed in the transcript. Check that User-visible
result names the final artifact or behavior you expect. The menu offers Ready, Refine, Edit, or
Cancel. Refine collects short notes. Edit opens the full plan in Pi's editor.
3. Work. Ready is the only review action that starts work. It also starts the goal steward. The
agent ticks subtasks, appends to `## Log` and `## Learnings`, fills `evidence:`, and calls
`CompleteGoal` when a discriminator is satisfied. The goal steward rereads the full plan on each
review. `CompleteGoal` resumes the same steward lineage instead of starting a fresh reviewer.
Every human reply and Refine note in plan mode is saved verbatim under `## Interview`. After eight
turns without a change above `## Log`, the worker gets a reminder and the steward gets a progress
checkpoint.
3. Work. Ready forks the approved-plan conversation into a cheaper `goal-worker`. pi-vcc compacts
inherited context before the first worker turn when there is enough context to compact; a small
exact fork is recorded as already below the compaction minimum. The main agent becomes the research
supervisor. Worker completion wakes it through pi-subagents. It calls `CheckGoalWork` before deciding
that all subagents and managed processes stopped, steers or resumes the retained worker with
`GuideGoalWorker`, reads the evidence, and calls `CompleteGoal` to sign off. FleetView and
`/subagents-fleet` show the worker. Every human reply and Refine note in plan mode is saved verbatim
under `## Interview`.
Other commands: `/goals clear` disconnects this session from its active plan, preserving the
versioned file on disk; `/goals auto [minutes|off]` continues active goals after the agent settles
and then on that interval. It pauses after two automatic wakes with no working-plan change; `/goals
model <model-ref>` picks the steward model (default: the pi-subagents agent model). These are TUI
subcommands, not CLI flags.
versioned file on disk. `/goals auto [minutes|off]` changes the supervisor check interval; Ready
enables a 60-minute interval. `/goals model <model-ref>` picks the cheaper worker model. Select the
stronger supervisor with Pi's normal `/model` command. Checks continue until all goals close, the
human uses `auto off`, or the plan is cleared.
## Prompts
Worker prompts live in [`src/prompts.ts`](src/prompts.ts). The read-only supervisor prompt and review requests live in [`src/steward.ts`](src/steward.ts).
Planning and sign-off prompts live in [`src/prompts.ts`](src/prompts.ts). Worker registration and pi-subagents RPC calls live in [`src/worker.ts`](src/worker.ts). [`src/worker-runtime.ts`](src/worker-runtime.ts) compacts the initial fork with pi-vcc.
## Develop
```bash
pi -e npm:pi-subagents -e ./src/index.ts # load locally
pi -e npm:pi-subagents -e npm:@sting8k/pi-vcc -e ./src/index.ts # load locally
npm test # all unit, flow, and Pi RPC tests
npm run test:rpc # Pi RPC review flow with a local offline model
npm run typecheck
+14 -189
View File
File diff suppressed because it is too large Load Diff
+5 -2
View File
@@ -1,7 +1,7 @@
{
"name": "@wassname2/pi-goals",
"version": "0.2.2",
"description": "One plan file per session with a persistent read-only goal steward powered by pi-subagents.",
"description": "One plan file per session with a main research supervisor and retained pi-subagents worker.",
"author": "wassname",
"license": "MIT",
"type": "module",
@@ -18,9 +18,12 @@
"proof",
"uat",
"evidence",
"steward",
"supervisor",
"subagent"
],
"dependencies": {
"@sting8k/pi-vcc": "^0.6.0"
},
"peerDependencies": {
"@earendil-works/pi-coding-agent": "*",
"typebox": "*"
@@ -0,0 +1,9 @@
{
"childSession": "/tmp/pi-goals-worker-runtime-wLPtUJ/.agent/sessions/--tmp-pi-goals-worker-runtime-wLPtUJ--/2026-09-05T09-09-45-566Z_01a070d4-bdda-7529-be64-eb1ecc38c053/forks/2026-09-05T09-10-07-859Z_01a070d5-14f3-7529-be64-eb21d70ab321.jsonl",
"marker": {
"version": 1,
"compacted": false,
"reason": "below-compactable-size"
},
"cwd": "/tmp/pi-goals-worker-runtime-wLPtUJ"
}
@@ -0,0 +1,45 @@
$ npm test
> @wassname2/pi-goals@0.2.2 test
> vitest run
RUN v4.1.9 /home/code/.pi/agent/git/github.com/wassname/pi-goals
Test Files 8 passed (8)
Tests 37 passed (37)
Start at 17:10:33
Duration 1.55s (transform 942ms, setup 0ms, import 1.95s, tests 1.35s, environment 1ms)
$ npm run typecheck
> @wassname2/pi-goals@0.2.2 typecheck
> tsc --noEmit
$ npm run lint
> @wassname2/pi-goals@0.2.2 lint
> biome check src/ test/
Checked 13 files in 20ms. No fixes applied.
$ git diff --check
$ npm pack --dry-run
npm notice 📦 @wassname2/pi-goals@0.2.2
npm notice Tarball Contents
npm notice 4.1kB README.md
npm notice 1.5kB package.json
npm notice 27.3kB src/index.ts
npm notice 13.3kB src/prompts.ts
npm notice 1.2kB src/worker-runtime.ts
npm notice 6.2kB src/worker.ts
npm notice Tarball Details
npm notice name: @wassname2/pi-goals
npm notice version: 0.2.2
npm notice filename: wassname2-pi-goals-0.2.2.tgz
npm notice package size: 18.0 kB
npm notice unpacked size: 53.5 kB
npm notice shasum: e90e0986903238e68218ea42dd0e00fec6c56ddd
npm notice integrity: sha512-CpJr+pjGMxZ0v[...]vSTliF8yvr+Rg==
npm notice total files: 6
wassname2-pi-goals-0.2.2.tgz
+51 -55
View File
@@ -1,71 +1,67 @@
# Supervisor through pi-subagents
# Main-agent supervisor with a retained pi-subagents worker
User, this session: "Another take on my pi-intercom-supervisor but using pi-subagents not a seperate user started terminal. I want to keep it simple by using pi-vcc pi-subagents where possibe"
Branch: `experiment/subagent-supervisor`. Supersedes the custom session-switching plan. Authored by Pi/Codex.
The main Pi session keeps the high-level research context and uses the stronger model. A cheaper `goal-worker` child implements the plan. pi-subagents owns the child fork, continuation, background completion, and Fleet visibility. pi-vcc compacts long context.
## User-visible result
The main agent is the cheaper worker. A stronger supervisor keeps the high-level intent and gives research direction from compressed context. It checks every 60 minutes and when the worker settles without background work, and judges goal sign-off against the plan and work history. Inspect and guide it through normal pi-subagents controls. No new terminal or custom session-switch command.
After Ready, the main session supervises a retained worker: it reviews every 60 minutes, responds when the worker finishes and no other work remains, and signs off goals from inspected evidence.
- [ ] goal: Keep supervisor context through ordinary pi-subagents continuation
- [ ] At Ready, fork the approved-plan conversation and compact the supervisor's copy before its first review. Leave worker context unchanged. Keep runtime registration and a separately selected supervisor model; retain the latest run ID and reconcile on reload.
- [ ] Resume that supervisor with VCC summaries of worker updates, current work status, and the full plan path. Include the worker compaction summary when an update crosses a compaction boundary; do not replay raw worker turns at check-ins.
- [ ] Compact supervisor context at about 100k tokens, earlier if its model requires it. Preserve human intent, research decisions, rejected ideas and reasons, unresolved questions, and evidence references. Prefer package compaction support; verify the child-specific API before implementation.
- [ ] Preserve supervisor decisions through continuation; avoid duplicate worker snapshots. Reuse relevant pi-supervise view logic, not its intercom transport or process scan.
- [ ] Enable normal progress visibility; remove the bespoke visit/role-restoration design and arbitrary read-tool cap.
- failure modes: fresh reviews forget earlier corrections; stale snapshots hide a human correction; review instructions start another supervisor.
- deliverable: a real resumed child refers to an earlier correction and a later worker update in its next decision, visible through Fleet.
## User voice
- [ ] goal: Use one supervisor check path for hourly and idle checks
- [ ] Replace direct auto-continue and stale-turn review dispatch with an hourly timer plus settled/background-completion events.
- [ ] Query pi-subagents status and pi-processes' request/reply protocol; do not infer activity from tool names, elapsed time, or missing results.
- [ ] Exclude the supervisor itself; include nested work and queued completion delivery. Unknown status is an error, not proof of idle.
- [ ] Serialize checks; recheck idle before dispatch and before waking the worker. Apply redirects once to the owning worker session.
- [ ] Remove the two-wake pause. Human messages do not disable supervision. Keep the hourly check armed while the worker is stopped or busy; explicit off/clear or completed goals end checks.
- failure modes: duplicate supervisors; normal async waiting mistaken for abandonment; supervisor completion triggers an endless self-review loop.
- deliverable: hourly checks while busy, idle checks only without remaining work, and a supervisor redirect that actually restarts a stopped worker.
- > “Another take on my pi-intercom-supervisor but using pi-subagents not a seperate user started terminal. I want to keep it simple by using pi-vcc pi-subagents where possibe”
- > “fork and compact because I hope to use a smarter model for the supervisor with better research taste.”
- > “only check in sees a summary, and ideally compacts at 100k”
- [ ] goal: Sign off from the same informed supervisor
- [ ] Route CompleteGoal through the same serialized check path with explicit sign-off mode; retain existing plan evidence and ticking behavior.
- [ ] Only accept/reject decides sign-off; checkpoint advice cannot tick a goal. A failed review cannot approve anything.
- failure modes: the supervisor accepts the worker's claim without reading the cited artifacts; a sign-off request received during a checkpoint is never processed.
- deliverable: reject a plausible claim without evidence, then accept its proven result using the retained supervisor context.
## Goals
1. [x] goal: Keep the stronger main session as supervisor and retain a cheaper child worker
- subtle failure mode: each instruction starts a fresh child that forgets earlier work.
- discriminator: a real Ready flow forks `goal-worker`; later guidance resumes its latest run ID and Fleet shows the child.
- tasks:
1. [x] register `goal-worker` through the pi-subagents runtime API with a separate `/goals model` setting
2. [x] fork on Ready and store each replacement run ID returned by resume
3. [x] keep implementation and plan edits in the child; keep evidence review and `CompleteGoal` in the main session
- evidence:
- `slop/audits/20260905_goal-worker-runtime-proof.json`: a real Ready flow created a child fork and wrote `"reason": "below-compactable-size"`.
- `slop/audits/20260905_subagent-supervisor-validation.txt`: `Tests 37 passed (37)`, including spawn, continuation, steering, and a completion-before-RPC-reply race.
2. [x] goal: Compact context without adding a custom transport
- subtle failure mode: the worker inherits the full expensive transcript, or a different compactor silently handles it.
- discriminator: the child runtime records either pi-vcc compaction or that the exact fork is below Pi's compaction minimum; the main session requests pi-vcc near 100k tokens.
- tasks:
1. [x] load pi-vcc in the child and compact the initial fork before its first turn
2. [x] fail if a different child compactor reports success
3. [x] compact the main supervisor near 100k and warn if pi-vcc did not handle it
- evidence:
- `slop/audits/20260905_goal-worker-runtime-proof.json`: the real child runtime recorded its explicit below-minimum outcome rather than silently claiming compaction.
- `test/worker-runtime.test.ts` covers pi-vcc success, below-minimum context, retained resume, and wrong-compactor failure; `test/goals-flow.test.ts` covers the 100k main-session request.
3. [x] goal: Check hourly and after worker completion without a self-review loop
- subtle failure mode: each supervisor settle schedules a new immediate review, or a child completion is mistaken for all work being idle.
- discriminator: one timer keeps its original hourly cadence; pi-subagents completion wakes the main session once; `CheckGoalWork` reports exact subagent and process state before a restart decision.
- tasks:
1. [x] keep one 60-minute timer active until goals close, auto is disabled, or the plan is cleared
2. [x] use native pi-subagents completion delivery instead of a second idle wake
3. [x] query public pi-subagents and pi-processes status; treat omitted or missing status as unknown
- evidence:
- `test/goals-flow.test.ts` keeps the first timer deadline across an intervening settle and checks idle status after native completion.
- `test/worker.test.ts` distinguishes active, nested-active, idle, omitted/unknown, and missing pi-processes status.
- `slop/audits/20260905_subagent-supervisor-validation.txt`: typecheck, lint, whitespace check, and package dry-run passed.
## UAT / verification
- Context: verify the first review receives a compacted fork, later checks receive summaries, and supervisor compaction occurs near 100k tokens. Success retains earlier corrections through both agents' compactions; likely failure forgets them; subtle failure retains the plan but omits a newer user correction. Inspect actual model context, not just the saved transcript.
- Cost and usefulness: record worker/supervisor input, cached input, output, compaction usage, and cost separately. Test the hypothesis of roughly 10x lower supervisor token use; do not infer it from context size alone. Record concrete supervisor corrections and worker outcomes. A cheaper run alone does not demonstrate better research decisions.
- Triggers: fake-clock tests exercise 60 minutes during busy work and after repeated stops. Success also waits for an already-running child/process; likely failure never wakes; subtle failure counts its own supervisor or treats a completed launch call as finished work. Test nested work, completion delivery, overlap, reload, off, and all-goals-done.
- Sign-off: exercise missing evidence, positive artifact evidence, and concurrent checkpoint/sign-off. Check the saved decision and actual plan checkbox, not only mocked helper output.
- Run repo tests, typecheck, lint, and a real spawn/resume/redirect/sign-off test with normal Fleet visibility. Save local run artifacts outside tracked research/source files. If a check fails, inspect its source events and fix the cause; do not substitute a separate terminal.
- Run `npm test`, `npm run typecheck`, `npm run lint`, `git diff --check`, and `npm pack --dry-run`; save exact output.
- Real runtime: load pi-goals with pi-subagents and pi-vcc, approve a plan, observe a real `goal-worker` fork and the fork-preparation record, then stop the test run.
- Inspect the final diff for duplicate wake paths, silent status fallbacks, and instructions that tell the main supervisor to implement worker tasks.
## Appendix: ownership and flow
## Appendix (context, not approved)
```text
Ready: fork approved-plan conversation -> compact child copy -> stronger supervisor
Every 60 minutes: request checkpoint, even if worker has been busy
Worker settled / background work finished: reconcile; request idle check if no work remains
CompleteGoal: request sign-off with claim + plan + VCC worker view
A normal Pi TUI cannot switch into the child's full interactive session. Fleet can inspect and steer it. This design keeps the visitable persistent context in the main session and uses retained child continuation for implementation.
One check at a time -> resume with VCC update -> store latest run ID
Supervisor context near 100k tokens -> compact its context, retaining decisions
checkpoint/idle -> let_run or redirect worker with one concrete next action
sign-off -> reject with missing evidence, or accept and tick goal
```
pi-subagents always sends an async completion to the parent. Therefore worker completion is the idle-review wake. Adding another `agent_settled` wake would create a completion → supervisor → settle loop.
User clarification: "fork and compact because I hope to use a smarter model for the supervisor with better research taste." The initial supervisor input must be compacted, even if the worker has not compacted. Do not compact the worker as a side effect. Verify that inherited extension state does not start nested supervision.
`CheckGoalWork` sees parent-process pi-processes state. A process started inside the worker remains covered indirectly because the top-level worker run stays active while its child work runs.
Cost hypothesis, not an observed result: roughly one tenth of the token use could fund a model with roughly ten times the per-token price at comparable total cost. Measure cache pricing, output, and compaction separately. The intended benefit is better worker research decisions, not merely fewer supervisor tokens.
Sources read: pi-subagents `README.md`, `docs/extension-api.md`, `docs/observability.md`, `docs/workflows.md`, execution controls; pi-processes request/list client; pi-vcc package behavior; `pi-supervise/RESEARCH_JOURNAL.md`.
These are conceptual operations, not invented package API names. The parent extension supplies triggers and context; pi-subagents owns child execution, persistence, resume, control, and UI. The supervisor judges direction and evidence. It may finish each review turn; the next check continues its stored context. No claim that the same child process stays running.
Use a small parent-session timer: installed RPC `manage` does not expose schedule creation, and package schedules launch fresh workflows without automatically capturing current parent context. Do not use the package's budget-bound mission goal loop for this unlimited-duration supervision policy.
Sources read for this plan:
- [Supervisor research journal](../../../pi-supervise/RESEARCH_JOURNAL.md): user preferences, premature-stop bug, VCC context, stale-snapshot accumulation.
- [pi-subagents extension API](/home/code/.pi/agent/npm/node_modules/pi-subagents/docs/extension-api.md): runtime registration, RPC spawn/resume/status, exact-session completions.
- [Execution controls](/home/code/.pi/agent/npm/node_modules/pi-subagents/skills/pi-subagents/references/execution-controls.md): retained continuation, latest run identity, schedules, native parent contact.
- [Observability](/home/code/.pi/agent/npm/node_modules/pi-subagents/docs/observability.md): Fleet transcripts, `s` guidance to live async children. Enter opens an optional Herdr inspector, not a TUI session switch.
- [Process client](/home/code/.pi/agent/npm/node_modules/@aliou/pi-processes/extensions/processes/client.ts): existing request/reply list protocol. Verify ownership/scope in the handler before integrating; never silently interpret a missing reply as an empty list.
- [Existing VCC view](/home/code/.pi/agent/git/github.com/wassname/pi-supervise/src/view.ts): reuse context construction selectively; its process scan and intercom byte limits do not belong in this branch.
-- Pi/Codex
+198 -276
View File
@@ -1,34 +1,32 @@
/**
* PI: pi-goals owns one versioned plan per session. A persistent, read-only pi-subagents child
* keeps the high-level context, reviews progress, and decides CompleteGoal sign-off.
* PI: pi-goals owns one versioned plan per session. The main agent supervises a cheaper retained
* pi-subagents worker, reviews progress, and decides CompleteGoal sign-off.
*
* Each /goals call makes `.pi/plan/<session_id>-vN.md`. The selected version survives resume and
* compaction. Old plans stay on disk but inactive. A session with no selected plan has no widget,
* injections, steward reviews, or CompleteGoal sign-off.
* supervision, worker, or CompleteGoal sign-off.
*
* TypeScript reads only goal checkbox lines for the widget. Models read the plan as prose. The
* worker alone edits it. The steward receives the plan path on every review, rereads the complete
* file, and can inspect cited artifacts with read-only tools. pi-subagents owns child sessions,
* persistence, resume, completion events, structured output, and contact with the parent.
* worker edits the project and records evidence. The main agent keeps the high-level context and
* directs the worker. pi-subagents owns the worker session, fork, persistence, resume, events,
* and Fleet controls.
*
* — Pi/Codex
* -- Pi/Codex
*/
import { existsSync, mkdirSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { join, resolve } from "node:path";
import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
import { completeGoalDescription, completeGoalParamDescription, planDrafting, planningState, reminder, resync } from "./prompts.js";
import { completeGoalDescription, completeGoalParamDescription, planDrafting, planningState, resync } from "./prompts.js";
import {
checkpointReview,
readyReview,
registerStewardAgent,
resumeSteward,
runStewardReview,
type StewardDecision,
signoffReview,
startSteward,
} from "./steward.js";
processWorkState,
registerGoalWorker,
resumeGoalWorker,
startGoalWorker,
steerGoalWorker,
subagentWorkState,
} from "./worker.js";
const STATE = "pi-goals-state";
const STATUS_KEY = "pi-goals";
@@ -37,23 +35,18 @@ const PLANNING_CONTEXT = "pi-goals-planning-context";
const PLAN_DIR = ".pi/plan";
// For static text (the /goals description) where there is no ctx to resolve the session id.
const PLAN_SHAPE = `${PLAN_DIR}/<session_id>-vN.md`;
// Plan mode is read-only by convention AND a light gate: edit/write are blocked (except the plan
// file, the deliverable). bash stays open — the prompt says don't mutate; guide, not gate (spec D3).
// Plan mode blocks edit/write except for its plan file. bash remains available for read-only inspection. -- Pi/Codex
const PLAN_MODE_BLOCKED_TOOLS = ["edit", "write"];
// A plan reminder is only useful after a substantial run of work that has not changed the working
// set. Log and learning entries do not count as progress. Unlike pi-tasks, goals have no dedicated
// progress tool, so this cadence repeats until the working set changes.
const STALE_TURNS = 8;
const AUTO_DEFAULT_INTERVAL_MS = 60 * 60 * 1_000;
const AUTO_MAX_WAKES_WITHOUT_PROGRESS = 2;
const SUPERVISOR_COMPACT_TOKENS = 100_000;
// A checkbox line beginning "goal:", for the widget and the "any goals open?" reminder condition.
// A checkbox line beginning "goal:", used by the widget and supervisor scheduling.
// Everything else reads the file as prose.
const GOAL_LINE = /^\s*(?:\d+\.|[-*])\s*\[([ xX/-])\]\s*goal:\s*(.*)$/i;
// An indented checkbox line that isn't a goal: a subtask. Only the widget reads these, so the human
// sees the next action and not just the goal -- this file IS the task list.
const SUBTASK_LINE = /^\s+(?:\d+\.|[-*])\s*\[([ xX/-])\]\s*(.*)$/;
// The fold. Above it: the working set that gets re-sent. Below it: durable memory.
// The fold separates current goals from the longer research record.
const FOLD_LINE = /^##\s+Log\s*$/im;
type GoalStatus = "open" | "active" | "done" | "cancelled";
const CHAR_TO_STATUS: Record<string, GoalStatus> = { " ": "open", "/": "active", x: "done", "-": "cancelled" };
@@ -67,8 +60,7 @@ function scanGoals(plan: string): Array<{ status: GoalStatus; subject: string; l
return goals;
}
/** The working set: everything above "## Log". Log, Learnings and Appendix below it are durable
* memory -- unlimited, read on demand, pushed back only by a resync. Exported for the unit test. */
/** Return the short current-goal section above "## Log". Exported for the unit test. */
export function foldPlan(plan: string): string {
const m = FOLD_LINE.exec(plan);
return (m ? plan.slice(0, m.index) : plan).trimEnd();
@@ -100,39 +92,33 @@ type Phase = "planning" | "working" | null;
interface PlanState {
phase: Phase;
stewardModel: string | null;
stewardRunId: string | null;
stewardPending: boolean;
workerModel: string | null;
workerRunId: string | null;
workerPending: boolean;
planVersion: number | null;
autoIntervalMs: number | null;
autoPaused: boolean;
}
export default function piGoalsExtension(pi: ExtensionAPI): void {
if (process.env.PI_SUBAGENT_CHILD === "1") return;
let state: PlanState = {
phase: null,
stewardModel: null,
stewardRunId: null,
stewardPending: false,
workerModel: null,
workerRunId: null,
workerPending: false,
planVersion: null,
autoIntervalMs: null,
autoPaused: false,
};
let planningContextPending = false;
// The reminder sees only the working set. A repeated Log line must not look like progress.
let turnsStale = 0;
let lastSeenWorkingSet = "";
let autoTimer: ReturnType<typeof setTimeout> | null = null;
let autoWakeInFlight = false;
let autoWakesWithoutProgress = 0;
let autoLastWorkingSet = "";
let autoImmediateUsed = false;
let runStartedBackgroundWork = false;
let stewardRegistration: { dispose(): void } | null = null;
let stewardRegistrationError: string | null = null;
let unsubscribeStewardCompletion: (() => void) | null = null;
// Set on session start and after a compaction; drained by the next LLM call, which then carries
// the WHOLE file (appendix included) instead of just the working set.
let supervisorWakePending = false;
let supervisorCompactionPending = false;
let workerRegistration: { dispose(): void } | null = null;
let workerRegistrationError: string | null = null;
let unsubscribeWorkerCompletion: (() => void) | null = null;
let workerLaunchPending = false;
const workerCompletionsDuringLaunch = new Set<string>();
// Set on session start and after compaction; the next supervisor call receives the whole plan.
let resyncReason: string | null = "New session.";
const planRel = (ctx: ExtensionContext) => (state.planVersion === null ? PLAN_SHAPE : `${PLAN_DIR}/${ctx.sessionManager.getSessionId()}-v${state.planVersion}.md`);
@@ -152,44 +138,52 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
pi.appendEntry<PlanState>(STATE, state);
}
function setupSteward(ctx: ExtensionContext): void {
stewardRegistration?.dispose();
stewardRegistration = null;
stewardRegistrationError = null;
function setupWorker(ctx: ExtensionContext): void {
workerRegistration?.dispose();
workerRegistration = null;
workerRegistrationError = null;
try {
stewardRegistration = registerStewardAgent(pi.events, state.stewardModel);
workerRegistration = registerGoalWorker(pi.events, state.workerModel);
} catch (error) {
stewardRegistrationError = error instanceof Error ? error.message : String(error);
if (state.phase === "working") ctx.ui.notify(`Goal steward unavailable: ${stewardRegistrationError}`, "warning");
workerRegistrationError = error instanceof Error ? error.message : String(error);
if (state.phase === "working") ctx.ui.notify(`Goal worker unavailable: ${workerRegistrationError}`, "warning");
}
}
function rememberStewardRun(runId: string): void {
state = { ...state, stewardRunId: runId, stewardPending: true };
function rememberWorkerRun(runId: string): void {
state = { ...state, workerRunId: runId, workerPending: !workerCompletionsDuringLaunch.delete(runId) };
persist();
}
async function reviewInBackground(ctx: ExtensionContext, task: string): Promise<void> {
if (!stewardRegistration) {
ctx.ui.notify(`Goal steward unavailable: ${stewardRegistrationError ?? "pi-subagents is not ready"}. Install pi-subagents and reload Pi.`, "warning");
return;
}
if (state.stewardPending) return;
async function startOrResumeWorker(ctx: ExtensionContext, task: string, signal?: AbortSignal): Promise<string> {
if (workerLaunchPending) throw new Error("A goal-worker launch is already in progress.");
if (!workerRegistration) setupWorker(ctx);
if (!workerRegistration) throw new Error(`Goal worker unavailable: ${workerRegistrationError ?? "pi-subagents is not ready"}.`);
workerLaunchPending = true;
workerCompletionsDuringLaunch.clear();
try {
const runId = state.stewardRunId
? await resumeSteward(pi.events, state.stewardRunId, task)
: await startSteward(pi.events, ctx.cwd, task);
rememberStewardRun(runId);
} catch (error) {
ctx.ui.notify(`Goal steward could not start: ${error instanceof Error ? error.message : String(error)}`, "warning");
const runId = state.workerRunId
? await resumeGoalWorker(pi.events, state.workerRunId, task, signal)
: await startGoalWorker(pi.events, ctx.cwd, task, signal);
rememberWorkerRun(runId);
return runId;
} finally {
workerLaunchPending = false;
workerCompletionsDuringLaunch.clear();
}
}
function watchStewardCompletion(): void {
unsubscribeStewardCompletion?.();
unsubscribeStewardCompletion = pi.events.on("subagent:async-complete", (raw) => {
if (!raw || typeof raw !== "object" || (raw as { runId?: string }).runId !== state.stewardRunId) return;
state = { ...state, stewardPending: false };
function watchWorkerCompletion(): void {
unsubscribeWorkerCompletion?.();
unsubscribeWorkerCompletion = pi.events.on("subagent:async-complete", (raw) => {
if (!raw || typeof raw !== "object") return;
const runId = (raw as { runId?: string }).runId;
if (!runId) return;
if (runId !== state.workerRunId) {
if (workerLaunchPending) workerCompletionsDuringLaunch.add(runId);
return;
}
state = { ...state, workerPending: false };
persist();
});
}
@@ -203,49 +197,42 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
return scanGoals(readPlan(ctx)).some((goal) => goal.status === "active" || goal.status === "open");
}
function scheduleAutoContinue(ctx: ExtensionContext, delayMs = state.autoIntervalMs): void {
clearAutoTimer();
if (delayMs === null || state.phase !== "working" || state.autoIntervalMs === null || state.autoPaused || !activeGoals(ctx)) return;
function wakeSupervisor(ctx: ExtensionContext, reason: string): void {
if (supervisorWakePending || state.phase !== "working" || !activeGoals(ctx)) return;
supervisorWakePending = true;
pi.sendUserMessage(
`<system-reminder>${reason}\nYou are the goal supervisor. Read ${planRel(ctx)} and inspect the cited evidence. Use CheckGoalWork before concluding that work stopped, and GuideGoalWorker to steer or resume the cheaper worker. Sign off a goal only after its discriminator is positively proved. Do not perform implementation work yourself.</system-reminder>`,
{ deliverAs: "followUp" },
);
}
function scheduleSupervisorCheck(ctx: ExtensionContext): void {
if (autoTimer !== null || state.phase !== "working" || state.autoIntervalMs === null || !activeGoals(ctx)) return;
autoTimer = setTimeout(() => {
autoTimer = null;
if (state.phase !== "working" || state.autoPaused || !ctx.isIdle() || !activeGoals(ctx)) return;
autoWakeInFlight = true;
pi.sendUserMessage(
`<system-reminder>Auto-continue is enabled by the human. Continue the active goal in ${planRel(ctx)}. Work from the open subtasks and observed artifacts. Keep the plan current, including useful Log entries. If you need a human decision, ask one direct question and leave the goal active.</system-reminder>`,
{ deliverAs: "followUp" },
);
}, delayMs);
scheduleSupervisorCheck(ctx);
wakeSupervisor(ctx, `The ${state.autoIntervalMs! / 60_000}-minute supervisor check is due.`);
}, state.autoIntervalMs);
autoTimer.unref();
}
function settleAuto(ctx: ExtensionContext): void {
if (state.phase !== "working" || state.autoIntervalMs === null || state.autoPaused || !activeGoals(ctx)) return;
const workingSet = foldPlan(readPlan(ctx));
const changed = workingSet !== autoLastWorkingSet;
if (changed) {
autoLastWorkingSet = workingSet;
autoWakesWithoutProgress = 0;
autoImmediateUsed = false;
}
if (autoWakeInFlight) {
autoWakeInFlight = false;
if (!changed) autoWakesWithoutProgress++;
if (autoWakesWithoutProgress >= AUTO_MAX_WAKES_WITHOUT_PROGRESS) {
state = { ...state, autoPaused: true };
persist();
updateWidget(ctx);
ctx.ui.notify("Goal auto-continue paused; waiting for user after two wakes without working-plan progress.", "warning");
return;
}
scheduleAutoContinue(ctx);
return;
}
if (!runStartedBackgroundWork && !autoImmediateUsed) {
autoImmediateUsed = true;
scheduleAutoContinue(ctx, 0);
return;
}
scheduleAutoContinue(ctx);
function compactSupervisor(ctx: ExtensionContext): void {
const usage = ctx.getContextUsage();
if (supervisorCompactionPending || usage?.tokens === null || usage?.tokens === undefined || usage.tokens < SUPERVISOR_COMPACT_TOKENS) return;
supervisorCompactionPending = true;
ctx.compact({
customInstructions: "__pi_vcc__ keep:1",
onComplete: (result) => {
supervisorCompactionPending = false;
if ((result.details as { compactor?: string } | undefined)?.compactor !== "pi-vcc") {
ctx.ui.notify("Goal supervisor compaction did not use pi-vcc. Install and load @sting8k/pi-vcc.", "warning");
}
},
onError: (error) => {
supervisorCompactionPending = false;
ctx.ui.notify(`Goal supervisor compaction failed: ${error.message}`, "warning");
},
});
}
function updateWidget(ctx: ExtensionContext): void {
@@ -261,15 +248,14 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
return;
}
const done = goals.filter((g) => g.status === "done").length;
const auto = state.autoPaused ? " · waiting for user" : state.autoIntervalMs === null ? "" : ` · auto ${state.autoIntervalMs / 60_000}m`;
const auto = state.autoIntervalMs === null ? "" : ` · supervise ${state.autoIntervalMs / 60_000}m`;
ctx.ui.setStatus(STATUS_KEY, ctx.ui.theme.fg("accent", `◷ ${done}/${goals.length} goals${auto}`));
const mark: Record<GoalStatus, string> = { done: "✔", active: "▸", open: "◻", cancelled: "✗" };
// Only live goals get lines so finished work never pushes current work off screen. The active
// goal also shows its open subtasks: this file is the task list, so the widget is the task list.
// No path line: the session id makes it 47 chars, too long to be worth a widget row. The
// human opens the file from the Ready menu, and every injected reminder still names it.
// No path line: the session id makes it too long to be useful in the widget.
const plan = readPlan(ctx);
const lines: string[] = state.autoPaused ? [ctx.ui.theme.fg("warning", "⏸ waiting for user")] : [];
const lines: string[] = [];
for (const g of goals.filter((g) => g.status === "active" || g.status === "open")) {
lines.push(`${mark[g.status]} ${g.subject}`);
if (g.status === "active") lines.push(...openSubtasks(plan, g.line).slice(0, 3).map((s) => ctx.ui.theme.fg("muted", ` ◦ ${s}`)));
@@ -277,7 +263,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
ctx.ui.setWidget(WIDGET_KEY, lines);
}
// --- /goals: enter plan mode (or clear / configure the steward) — Pi/Codex ---------------------
// --- /goals: enter plan mode or configure supervision -- Pi/Codex -----------------------------
pi.registerCommand("goals", {
description: `Plan mode: draft goals into ${PLAN_SHAPE}, review, then work them. /goals <objective> | /goals clear | /goals auto [minutes|off] | /goals model <model>`,
@@ -290,7 +276,7 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
}
const currentPlan = planRel(ctx);
clearAutoTimer();
state = { ...state, phase: null, stewardRunId: null, stewardPending: false, planVersion: null, autoIntervalMs: null, autoPaused: false };
state = { ...state, phase: null, workerRunId: null, workerPending: false, planVersion: null, autoIntervalMs: null };
persist();
updateWidget(ctx);
ctx.ui.notify(`Disconnected from ${currentPlan}; the file remains on disk.`, "info");
@@ -300,14 +286,14 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
const value = arg.slice("auto".length).trim();
if (value === "off") {
clearAutoTimer();
state = { ...state, autoIntervalMs: null, autoPaused: false };
state = { ...state, autoIntervalMs: null };
persist();
updateWidget(ctx);
ctx.ui.notify("Goal auto-continue disabled.", "info");
ctx.ui.notify("Hourly goal supervision disabled.", "info");
return;
}
if (state.phase !== "working") {
ctx.ui.notify("Approve a plan with Ready before enabling auto-continue.", "warning");
ctx.ui.notify("Approve a plan with Ready before enabling supervision.", "warning");
return;
}
const minutes = value ? Number(value) : AUTO_DEFAULT_INTERVAL_MS / 60_000;
@@ -315,25 +301,23 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
ctx.ui.notify("Use /goals auto [whole minutes], or /goals auto off.", "warning");
return;
}
autoWakeInFlight = false;
autoWakesWithoutProgress = 0;
autoLastWorkingSet = foldPlan(readPlan(ctx));
state = { ...state, autoIntervalMs: minutes * 60_000, autoPaused: false };
clearAutoTimer();
state = { ...state, autoIntervalMs: minutes * 60_000 };
persist();
updateWidget(ctx);
scheduleAutoContinue(ctx);
ctx.ui.notify(`Goal auto-continue enabled every ${minutes}m.`, "info");
scheduleSupervisorCheck(ctx);
ctx.ui.notify(`Goal supervision will check every ${minutes}m.`, "info");
return;
}
if (arg === "model" || arg.startsWith("model ")) {
const ref = arg.slice("model".length).trim();
state = { ...state, stewardModel: ref || null, stewardRunId: null, stewardPending: false };
state = { ...state, workerModel: ref || null, workerRunId: null, workerPending: false };
persist();
setupSteward(ctx);
ctx.ui.notify(ref ? `Goal-steward model set to ${ref}` : "Goal-steward model reset to pi-subagents default", "info");
setupWorker(ctx);
ctx.ui.notify(ref ? `Goal-worker model set to ${ref}` : "Goal-worker model reset to pi-subagents default", "info");
return;
}
state = { ...state, phase: "planning", stewardRunId: null, stewardPending: false, planVersion: nextVersion(ctx) };
state = { ...state, phase: "planning", workerRunId: null, workerPending: false, planVersion: nextVersion(ctx), autoIntervalMs: null };
planningContextPending = true;
resyncReason = null;
writePlan(ctx, "");
@@ -351,31 +335,23 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
// --- hooks --------------------------------------------------------------------------------------
/** What this LLM call should carry, if anything: a one-shot resync, or a staleness reminder. */
/** Restore the complete plan once after session start or compaction. */
function dueInjection(ctx: ExtensionContext, plan: string): string | null {
const drainResync = (): string | null => {
const why = resyncReason;
resyncReason = null;
return why;
};
if (state.phase === "planning") return null;
if (!plan.trim()) return null;
const why = drainResync();
if (why) return resync(plan, planRel(ctx), why);
if (turnsStale < STALE_TURNS) return null;
const goals = scanGoals(plan);
if (goals.length === 0) {
// Non-empty plan but no recognizable goal line: the harness would go silently inert (no
// widget, no injection, no reminders). Say so instead -- cooperative but confused.
return `<system-reminder>\n${planRel(ctx)} exists but has no goal line pi-goals recognizes. A goal is a checkbox list line starting "goal:", e.g. "1. [ ] goal: <imperative>" ([ ] open, [/] active, [x] done, [-] cancelled). Reformat it if it's meant to be the plan.\n</system-reminder>`;
}
if (!goals.some((g) => g.status === "active" || g.status === "open")) return null;
return reminder(foldPlan(plan), planRel(ctx));
if (state.phase === "planning" || !plan.trim() || !resyncReason) return null;
const why = resyncReason;
resyncReason = null;
return resync(plan, planRel(ctx), why);
}
// The phase snapshot enters context only when planning starts or context was lost.
pi.on("before_agent_start", async (_event, ctx) => {
if (state.phase !== "planning" || !planningContextPending) return;
supervisorWakePending = false;
if (state.phase === "working") {
return {
systemPrompt: `${ctx.getSystemPrompt()}\n\nYou are the research supervisor for ${planRel(ctx)}. Keep the high-level goal and the human's intent stable. The retained goal-worker owns implementation and plan updates; direct it with GuideGoalWorker instead of implementing work yourself. Read cited artifacts before calling CompleteGoal. A pi-subagents completion can wake you while nested subagents or managed processes are still running, so call CheckGoalWork before concluding that work stopped. At hourly checks, inspect progress and steer the worker only when a concrete correction is useful. Keep work going until all goals are proved or the human stops it. -- Pi/Codex`,
};
}
if (!planningContextPending) return;
planningContextPending = false;
return { message: { customType: PLANNING_CONTEXT, content: planningState(planPath(ctx)), display: false } };
});
@@ -384,50 +360,26 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
// before_agent_start, so context restores the planning snapshot exactly once in that path.
pi.on("context", async (event, ctx) => {
const messages = state.phase === "planning" ? event.messages : event.messages.filter((message) => (message as { customType?: string }).customType !== PLANNING_CONTEXT);
const removedPlanningContext = messages.length !== event.messages.length;
if (state.phase === "planning" && planningContextPending) {
planningContextPending = false;
return { messages: [...messages, { role: "user" as const, content: [{ type: "text" as const, text: planningState(planPath(ctx)) }], timestamp: Date.now() }] };
}
const text = dueInjection(ctx, readPlan(ctx));
if (!text) return messages === event.messages ? undefined : { messages };
turnsStale = 0;
if (!text) return removedPlanningContext ? { messages } : undefined;
return { messages: [...messages, { role: "user" as const, content: [{ type: "text" as const, text }], timestamp: Date.now() }] };
});
// PI: Human plan-mode replies are durable evidence of the interview, not model summaries.
pi.on("input", async (event, ctx) => {
if (event.source !== "extension") {
clearAutoTimer();
autoImmediateUsed = false;
if (state.autoPaused) {
state = { ...state, autoPaused: false };
persist();
updateWidget(ctx);
}
}
if (state.phase === "planning" && event.source !== "extension") writePlan(ctx, appendInterview(readPlan(ctx), event.text));
});
// The staleness clock sees only the working set. Log updates are durable evidence, not progress.
pi.on("turn_end", async (_event, ctx) => {
const workingSet = foldPlan(readPlan(ctx));
if (workingSet === lastSeenWorkingSet) {
turnsStale++;
return;
}
lastSeenWorkingSet = workingSet;
turnsStale = 0;
updateWidget(ctx);
});
pi.on("agent_start", async () => {
runStartedBackgroundWork = false;
});
pi.on("tool_call", async (event, ctx) => {
if (state.phase === "working" && (event.toolName === "subagent" || (event.toolName === "process" && (event.input as { action?: string }).action === "start"))) {
runStartedBackgroundWork = true;
}
if (state.phase !== "planning") return;
if (PLAN_MODE_BLOCKED_TOOLS.includes(event.toolName)) {
const target = (event.input as { path?: string }).path;
@@ -448,8 +400,8 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
// PI: Print after Pi settles. agent_end is still streaming, so its message queues behind the menu.
pi.on("agent_settled", async (_event, ctx) => {
if (state.phase === "working") {
if (turnsStale >= STALE_TURNS && activeGoals(ctx)) await reviewInBackground(ctx, checkpointReview(planRel(ctx), turnsStale));
settleAuto(ctx);
compactSupervisor(ctx);
scheduleSupervisorCheck(ctx);
return;
}
if (state.phase !== "planning" || !ctx.hasUI) return;
@@ -480,18 +432,24 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
}
if (choice === "Cancel") {
rmSync(planPath(ctx), { force: true });
state = { ...state, phase: null, stewardRunId: null, stewardPending: false, planVersion: null };
state = { ...state, phase: null, workerRunId: null, workerPending: false, planVersion: null, autoIntervalMs: null };
persist();
updateWidget(ctx);
ctx.ui.notify("Plan discarded.", "info");
return;
}
if (choice !== "Ready") return;
state = { ...state, phase: "working" };
state = { ...state, phase: "working", autoIntervalMs: AUTO_DEFAULT_INTERVAL_MS };
resyncReason = "The plan was approved.";
persist();
updateWidget(ctx);
await reviewInBackground(ctx, readyReview(planRel(ctx)));
pi.sendUserMessage(`Work the goals in ${planPath(ctx)}. Pick an open goal, mark it active ([/]), work its subtasks, and when its discriminator is satisfied fill its evidence: list, then call CompleteGoal with the goal's text. Keep the plan file current as you go.`, { deliverAs: "followUp" });
try {
await startOrResumeWorker(ctx, `Work the goals in ${planRel(ctx)}. Mark one open goal active, execute its subtasks, and record exact evidence in the plan Log. Report progress and evidence to the main research supervisor. Do not sign off goals.`);
scheduleSupervisorCheck(ctx);
} catch (error) {
ctx.ui.notify(`Goal worker could not start: ${error instanceof Error ? error.message : String(error)}`, "warning");
wakeSupervisor(ctx, "The approved goal worker failed to start.");
}
return;
}
});
@@ -503,32 +461,66 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
.pop() as { data?: PlanState } | undefined;
state = {
phase: last?.data?.phase ?? null,
stewardModel: last?.data?.stewardModel ?? null,
stewardRunId: last?.data?.stewardRunId ?? null,
stewardPending: last?.data?.stewardPending ?? false,
workerModel: last?.data?.workerModel ?? null,
workerRunId: last?.data?.workerRunId ?? null,
workerPending: last?.data?.workerPending ?? false,
planVersion: last?.data?.planVersion ?? null,
autoIntervalMs: last?.data?.autoIntervalMs ?? null,
autoPaused: last?.data?.autoPaused ?? false,
};
watchStewardCompletion();
setupSteward(ctx);
lastSeenWorkingSet = foldPlan(readPlan(ctx));
autoLastWorkingSet = lastSeenWorkingSet;
watchWorkerCompletion();
setupWorker(ctx);
planningContextPending = state.phase === "planning";
resyncReason = state.phase === "working" ? "New session." : null;
updateWidget(ctx);
scheduleAutoContinue(ctx);
scheduleSupervisorCheck(ctx);
});
pi.on("session_shutdown", async () => {
clearAutoTimer();
stewardRegistration?.dispose();
stewardRegistration = null;
unsubscribeStewardCompletion?.();
unsubscribeStewardCompletion = null;
workerRegistration?.dispose();
workerRegistration = null;
unsubscribeWorkerCompletion?.();
unsubscribeWorkerCompletion = null;
});
// --- the one blessed tool: CompleteGoal ---------------------------------------------------------
pi.registerTool({
name: "CheckGoalWork",
label: "Check goal work",
description: "Check whether pi-subagents or pi-processes still has active work before deciding that the goal worker stopped.",
parameters: Type.Object({}),
async execute(_id, _params, _signal, _onUpdate, _ctx) {
try {
const [subagents, processes] = await Promise.all([subagentWorkState(pi.events), Promise.resolve(processWorkState(pi.events))]);
const unknown = subagents === "unknown" || processes === "unknown";
return result(`subagents=${subagents}; processes=${processes}`, unknown);
} catch (error) {
return result(`Goal work status failed: ${error instanceof Error ? error.message : String(error)}`, true);
}
},
});
pi.registerTool({
name: "GuideGoalWorker",
label: "Guide goal worker",
description: "Send one concrete instruction to the retained goal-worker. A live worker is steered; a completed worker is resumed with its saved context.",
parameters: Type.Object({
instruction: Type.String({ description: "The next research or implementation action, with the evidence that should distinguish success from failure." }),
}),
async execute(_id, params, signal, _onUpdate, ctx) {
if (state.phase !== "working") return result("Approve a plan with Ready before directing the goal worker.", true);
try {
if (state.workerPending) {
if (!state.workerRunId) throw new Error("Goal-worker state says running but has no run ID.");
await steerGoalWorker(pi.events, state.workerRunId, params.instruction, signal);
return result(`Instruction delivered to live goal worker ${state.workerRunId}.`);
}
const runId = await startOrResumeWorker(ctx, params.instruction, signal);
return result(`Goal worker resumed as ${runId}.`);
} catch (error) {
return result(`Goal-worker guidance failed: ${error instanceof Error ? error.message : String(error)}`, true);
}
},
});
pi.registerTool({
name: "CompleteGoal",
@@ -537,54 +529,15 @@ export default function piGoalsExtension(pi: ExtensionAPI): void {
parameters: Type.Object({
goal: Type.String({ description: completeGoalParamDescription }),
}),
async execute(_id, params, signal, onUpdate, ctx) {
if (state.phase === "planning") return result("Planning is not approved. Choose Ready before signing off a goal.", true);
async execute(_id, params, _signal, _onUpdate, ctx) {
if (state.phase !== "working") return result("Planning is not approved. Choose Ready before signing off a goal.", true);
const plan = readPlan(ctx);
if (!plan.trim()) return result(`No plan file at ${planRel(ctx)}. Run /goals to draft one.`, true);
if (!stewardRegistration) {
return result(`Goal steward unavailable: ${stewardRegistrationError ?? "pi-subagents is not ready"}. Install pi-subagents and reload Pi.`, true);
}
if (state.stewardPending) return result("The goal steward is still reviewing the previous checkpoint. Retry CompleteGoal after its result arrives.", true);
onUpdate?.({ content: [{ type: "text", text: `Persistent goal steward inspecting: ${params.goal}` }], details: {} });
let reviewRunId = "";
let decision: StewardDecision;
try {
const review = await runStewardReview(
pi.events,
ctx.cwd,
state.stewardRunId,
signoffReview(planRel(ctx), params.goal),
signal,
600_000,
rememberStewardRun,
);
reviewRunId = review.runId;
decision = review.decision;
} catch (error) {
return result(`Goal-steward review failed: ${error instanceof Error ? error.message : String(error)}`, true);
}
state = { ...state, stewardPending: false };
persist();
const outcome = decideStewardSignOff(params.goal, decision, reviewRunId);
if (outcome.logEntry) {
// Sign-off write: tick the goal [x] (exact-subject match; dogfood showed agent bookkeeping
// is the drift point) and append the audit log line, one write. On wording drift the tick
// falls to the agent and the result says so -- both paths are explicit, never silent.
let updated = readPlan(ctx);
let tickNote = "";
if (outcome.logEntry.startsWith("signed off")) {
const ticked = tickGoal(updated, params.goal);
updated = ticked ?? updated;
tickNote = ticked
? `\n\nGoal ticked [x] in ${planRel(ctx)}.`
: `\n\nNo exact goal line matched your wording -- tick it [x] in ${planRel(ctx)} yourself.`;
}
writePlan(ctx, appendLog(updated, `${stamp()} ${outcome.logEntry}`));
updateWidget(ctx);
return result(outcome.resultText + tickNote, outcome.isError);
}
return result(outcome.resultText, outcome.isError);
const ticked = tickGoal(plan, params.goal);
if (!ticked) return result(`No unique exact goal line matched "${params.goal}" in ${planRel(ctx)}.`, true);
writePlan(ctx, appendLog(ticked, `${stamp()} signed off "${params.goal}" by the main research supervisor`));
updateWidget(ctx);
return result(`Sign-off accepted. Goal ticked [x] in ${planRel(ctx)}.`);
},
});
}
@@ -608,39 +561,8 @@ function stamp(): string {
return `${d.getFullYear()}-${p(d.getMonth() + 1)}-${p(d.getDate())} ${p(d.getHours())}:${p(d.getMinutes())}`;
}
function oneLine(s: string): string {
return s.replace(/\s+/g, " ").trim().slice(0, 200);
}
export interface SignOffOutcome {
resultText: string;
isError: boolean;
logEntry: string;
}
export function decideStewardSignOff(goal: string, decision: StewardDecision, runId: string): SignOffOutcome {
if (decision.verdict === "accept") {
return {
resultText: `Sign-off ACCEPTED.\n\nGoal steward: ${decision.summary}\nRun: ${runId}`,
isError: false,
logEntry: `signed off "${goal}" (steward accept; run ${runId})`,
};
}
const missing = decision.verdict === "reject"
? decision.missingEvidence?.join("; ") || decision.summary
: decision.verdict === "redirect"
? decision.nextAction ?? decision.summary
: `The steward returned let_run instead of a sign-off verdict: ${decision.summary}`;
return {
resultText: `Sign-off REJECTED. Missing:\n${missing}\n\nGoal steward: ${decision.summary}\nRun: ${runId}`,
isError: true,
logEntry: `reject "${goal}": ${oneLine(missing)} (steward run ${runId})`,
};
}
/** Tick the goal line whose subject exactly matches `goal` (trimmed, case-insensitive) to [x].
* Null when there is no unique exact match (wording drift / duplicates) -- the caller then asks the
* agent to tick it itself. Reuses GOAL_LINE; deliberately not fuzzy; the steward reads prose. */
* Null when there is no unique exact match. Reuses GOAL_LINE and is deliberately not fuzzy. */
export function tickGoal(plan: string, goal: string): string | null {
const lines = plan.split("\n");
const want = goal.trim().toLowerCase();
+26 -56
View File
@@ -2,22 +2,19 @@
* pi-goals v2 — all model-facing text, in flow order.
*
* Design: the plan file is for LLMs and the human, not for TypeScript. No parser and no schema;
* the skeleton below is a convention the drafting prompt teaches, the working agent maintains with
* its normal Edit tool, and the goal steward reads natively. The harness does three things for a
* cooperative-but-confused model: memory (a transient re-send of the plan when it goes stale),
* format guidance (the skeleton), and supervision (a persistent read-only pi-subagents child).
* the skeleton below is a convention the drafting prompt teaches, the worker maintains with its
* normal Edit tool, and the main research supervisor reads natively. The harness provides format
* guidance, one full-plan resync after context loss, and a retained pi-subagents worker.
*
* THE FOLD: everything above "## Log" is the working set (title, user voice, goals,
* discriminators) and is what gets re-sent on the reminder cadence. Everything below it (Log,
* Learnings, Appendix) is durable memory: unlimited, read on demand, and re-sent in full only at
* session start and after a compaction, which is where the settled context is actually needed.
* THE FOLD: everything above "## Log" is the short current-goal section. Everything below it
* (Log, Learnings, Appendix) is durable memory: unlimited, read on demand, and sent in full at
* session start and after compaction.
*
* Flow:
* SETUP (plan mode) 1. planDrafting — draft goals into the plan file (read-only), sent once
* EXEC, on cadence 2. reminder — the folded plan + upkeep nudge when it went stale
* EXEC, after compact 3. resync — the WHOLE file back, once
* SIGN-OFF, agent-side 4. completeGoal* — the one blessed tool's description
* SUPERVISION steward.ts — persistent review and sign-off prompts
* EXEC, after compact 2. resync — the WHOLE file back, once
* SIGN-OFF, agent-side 3. completeGoal* — the one blessed tool's description
* SUPERVISION worker.ts - retained implementation worker
*
* The goal's test is the DISCRIMINATOR: the concrete observation that tells real success from the
* named subtle failure mode. Evidence is empty at planning and filled at sign-off.
@@ -58,13 +55,13 @@ Detail that doesn't change a goal or a discriminator belongs in the appendix, no
Right-size it:
- One goal per distinct judgeable outcome. Group related goals when it helps judge them together
and readability. The count flows from the outcomes.
- Describe outcomes in qualitative terms the steward and user can discriminate.
- Describe outcomes in qualitative terms the supervisor and user can discriminate.
- Use the users language or more precise don't transform "MV" into "knob" as it looses precision and is overloaded
- Don't invent metrics or thresholds for problems you haven't explored yet — the steward should know it when it sees the outcome.
- Don't invent metrics or thresholds for problems you haven't explored yet - the supervisor should know it when it sees the outcome.
- Quantitative gates are fine only when you are certain they survive contact with reality.
- Subtasks are the steps inside a goal; add them when a goal has 3+ distinct steps, skip otherwise.
- Two goals that share one discriminator are one goal. Merge them.
- Keep the goal subject short. Put its important scope, failure modes, discriminator, tasks, and evidence in the indented block beneath it. The steward reads the whole block and the whole plan.
- Keep the goal subject short. Put its important scope, failure modes, discriminator, tasks, and evidence in the indented block beneath it. The supervisor reads the whole block and the whole plan.
- Keep the working set under 50 lines, excluding ## User voice. ## User voice has no line limit: quote
the human fully rather than shorten or paraphrase them. Everything below "## Log" is unlimited.
@@ -72,7 +69,7 @@ Style: Make it easy for a busy and forgetfull user to review. Use ASD-STE100 Sim
the same word for the same thing, and define a new terms at first use. Use redundant context for skim readers e.g. "our output - the cells, CV tag" is easy to read and reminds context. This covers the context
paragraph and the appendix too, not just the checklist. No all-caps headers and no bold spam. Just write less, add your voice less, persuade less, and burden the reader less.
Write the plan file in roughly this shape -- the file is read directly by the human and the goal steward, so clarity beats conformance; small deviations are fine):
Write the plan file in roughly this shape -- the file is read directly by the human and the main supervisor, so clarity beats conformance; small deviations are fine):
# <short plan title>
@@ -92,7 +89,7 @@ Write the plan file in roughly this shape -- the file is read directly by the hu
- subtle failure mode: <a way this could look done but isn't>
- discriminator: <the concrete observation that tells real success from that failure>
- verify: <optional shell command that exits 0 only when the discriminator passes; omit if not
testable. YOU run it at sign-off time and save its output as evidence; the steward only reads>
testable. The worker runs it and saves its output; the main supervisor reads the evidence>
- tasks:
1. [ ] <subtask>
- evidence: (empty until sign-off)
@@ -122,9 +119,9 @@ Conventions:
none of the failure modes could fake. Ruling out failures is necessary, not sufficient.
- Make the discriminator a concrete, checkable observation about a real artifact (a file, a test
result, a committed diff, a metric), never about the plan file's own checkbox.
- evidence stays empty at planning; you fill it at sign-off and the read-only goal steward checks it.
- evidence stays empty at planning; the worker fills it and the main research supervisor checks it.
Cite durable artifacts a future reader can open: committed files, test names, git diffs. .pi/ is
usually gitignored, so files there prove things only at steward review time, not in history.
usually gitignored, so files there prove things only at supervisor review time, not in history.
- User-visible result: restate the original deliverable, not the proposed implementation. Every goal
must contribute to it. Future work may not defer any artifact or action named there.
- User voice: quote the human word for word, one line per requirement, as they say it. Never
@@ -142,12 +139,6 @@ Conventions:
When the goals are drafted, present them and say the plan is final. Do not begin execution.`;
/* ─────────────────────────────────────────────────────────────────────────
* 3. reminder — EXEC. Transient, never persisted, and only when the plan went stale for a couple of
* turns. pi-tasks tried a per-turn injection and deleted it: "wallpaper noise that trains the
* model to ignore the task block" (tintinweb/pi-tasks CHANGELOG.md:149). Carries the folded plan
* (above ## Log), because a nudge with no plan in it makes the model go read the file anyway.
* ──────────────────────────────────────────────────────────────────────── */
export function planningState(planPath: string): string {
return `\
[PLANNING MODE]
@@ -160,45 +151,25 @@ work, mark a goal [/] or [x], or sign off a goal. The plan is not approved until
Ready.`;
}
export function reminder(foldedPlan: string, planRel: string): string {
return `\
<system-reminder>
Your plan (${planRel}, above the fold; the log, learnings and appendix are in the file):
${foldedPlan}
Keep it current as you work, with your normal edit tool:
- tick finished subtasks ([/] in progress), add discovered ones
- append ONE short line to ## Log, and a line to ## Learnings for a gotcha worth keeping
- when the active goal's discriminator is satisfied, fill its evidence: list (each item = a durable
artifact + a verbatim quote you actually observed + a short read of it), then call CompleteGoal.
Don't tick a goal [x] before CompleteGoal accepts; the sign-off log line is the audit trail.
- if the working set has grown long, prune finished goals (their evidence lives in git history and
## Log) and move settled detail down to ## Appendix, which is unlimited
- the human's latest message outranks this plan. If it corrects the deliverable or scope, amend the
user-visible result, user voice, and affected goals before continuing; don't defend the old plan
- otherwise keep working toward the active goal; don't stop to ask unless genuinely blocked
</system-reminder>`;
}
/* ─────────────────────────────────────────────────────────────────────────
* 3b. resync — EXEC, one-shot at session start and after a compaction: the WHOLE file back,
* 2. resync — EXEC, one-shot at session start and after a compaction: the WHOLE file back,
* appendix included. Modelled on pi-goal-x's [POST-COMPACTION RESYNC] one-shot. This is the
* only place the below-the-fold sections are pushed; otherwise the agent reads them on demand.
* ──────────────────────────────────────────────────────────────────────── */
export function resync(plan: string, planRel: string, why: string): string {
return `\
<system-reminder>
${why} This is the whole plan file (${planRel}), appendix included. Keep working the active goal;
edit the file directly as you go. The human's latest message outranks the plan: if it corrects the
deliverable or scope, amend the plan rather than preserving an obsolete decision.
${why} This is the whole plan file (${planRel}), appendix included. You are the main research
supervisor. Keep the high-level goal and human intent stable; direct the retained goal-worker rather
than doing its implementation. The human's latest message outranks the plan: if it corrects the
deliverable or scope, direct the worker to amend the plan rather than preserving an obsolete decision.
${plan}
</system-reminder>`;
}
/* ─────────────────────────────────────────────────────────────────────────
* 4. completeGoal — SIGN-OFF, agent-side: the one blessed tool
* 3. completeGoal — SIGN-OFF, agent-side: the one blessed tool
* ──────────────────────────────────────────────────────────────────────── */
export const completeGoalDescription =
"Sign off a goal once its discriminator is satisfied. First fill the goal's evidence: list in the " +
@@ -207,13 +178,12 @@ export const completeGoalDescription =
"output you actually observed; never reconstruct numbers from memory. If you couldn't see an " +
"output, rerun it or write that you couldn't -- an honest gap beats a plausible fabrication. If " +
"the goal names a verify: command, run it yourself first and save its output to a file cited in " +
"the evidence: the goal steward cannot execute anything and will reject a claimed pass with no saved " +
"the evidence. The main research supervisor must reject a claimed pass with no saved " +
"output. The read must show success POSITIVELY happened, not just that failures were avoided. " +
"Check that the claimed result uses the artifact and outcome named in User-visible result and does " +
"not substitute an agent-inferred deliverable. Then call this with the goal's text (the line after " +
"'goal:'; small wording drift is fine). The persistent read-only goal steward rereads the complete " +
"plan, inspects the LIVE WORKING TREE (uncommitted changes included), and returns accept or reject " +
"with what is missing. On accept, a sign-off line is appended to ## Log and the goal is ticked [x] " +
"for you; the result says if you must tick it yourself. On reject or steward failure, the goal stays open.";
"'goal:'). You are the main research supervisor: reread the complete plan and inspect the live " +
"working tree, including uncommitted changes, before calling. The tool appends the sign-off to ## Log " +
"and ticks the goal [x]. If the evidence is missing, direct the goal-worker to obtain it instead.";
export const completeGoalParamDescription = "The goal's text: the line after 'goal:' in the plan file.";
-246
View File
@@ -1,246 +0,0 @@
import { randomUUID } from "node:crypto";
const REGISTER_EVENT = "pi-subagents:runtime-agent-register:v1";
const RPC_REQUEST_EVENT = "subagents:rpc:v1:request";
const RPC_REPLY_PREFIX = "subagents:rpc:v1:reply:";
const ASYNC_COMPLETE_EVENT = "subagent:async-complete";
const RPC_VERSION = 1;
const RPC_TIMEOUT_MS = 15_000;
export const STEWARD_AGENT = "goal-steward";
interface EventBus {
on(event: string, handler: (data: unknown) => void): () => void;
emit(event: string, data: unknown): void;
}
interface Registration {
dispose(): void;
}
interface RpcData {
text: string;
details?: Record<string, unknown>;
}
export interface StewardDecision {
verdict: "let_run" | "redirect" | "accept" | "reject";
summary: string;
nextAction?: string;
missingEvidence?: string[];
}
export const stewardOutputSchema = {
type: "object",
properties: {
verdict: { type: "string", enum: ["let_run", "redirect", "accept", "reject"] },
summary: { type: "string", maxLength: 800 },
nextAction: { type: "string", maxLength: 400 },
missingEvidence: { type: "array", maxItems: 8, items: { type: "string", maxLength: 400 } },
},
required: ["verdict", "summary"],
additionalProperties: false,
} as const;
export const stewardSystemPrompt = `You are the read-only goal steward for one Pi work session.
Act as its supervisor, mentor, project manager, and skeptical board member. Keep the high-level goal
and the human's stated result stable while the worker handles implementation detail.
The plan path arrives in every review. Read the complete plan from disk every time, including after
compaction. Treat User-visible result and User voice as the authority. The plan is maintained by the
worker; never edit it. Use read-only tools to inspect cited files when this changes your decision.
Spend few tokens. Call structured_output as soon as the evidence is sufficient. Do not send prose
before that call, write a review essay, restate the plan, or narrate routine progress. Return one verdict:
- let_run: progress follows the plan and no instruction is useful
- redirect: drift, a missed failure mode, or a specific better next action needs worker attention
- accept: only for a sign-off review whose evidence positively proves the discriminator
- reject: only for a sign-off review that names the missing evidence
For redirect, include one concrete nextAction. For reject, include missingEvidence. Contact the
parent only when an immediate decision is needed. Do not accept a confident summary as evidence.
You are advisory and read-only; pi-goals alone writes and signs off the plan.
— Pi/Codex`;
export function registerStewardAgent(events: EventBus, model: string | null): Registration {
const request: Record<string, unknown> = {
version: 1,
name: STEWARD_AGENT,
definition: {
description: "Persistent read-only supervisor for one pi-goals plan.",
systemPrompt: stewardSystemPrompt,
tools: ["read", "grep", "find", "ls", "contact_supervisor"],
excludeTools: ["bash", "edit", "write", "subagent"],
allowNestedSubagents: false,
...(model ? { model } : {}),
thinking: "low",
systemPromptMode: "replace",
inheritProjectContext: false,
inheritGlobalContext: false,
inheritSkills: false,
defaultContext: "fresh",
defaultAsync: true,
defaultTimeoutMs: 300_000,
acceptanceRole: "read-only",
defaultProgress: false,
toolBudget: { soft: 8, hard: 12, block: ["read", "grep", "find", "ls"] },
},
};
events.emit(REGISTER_EVENT, request);
const result = request.result as { ok?: boolean; registration?: Registration; error?: Error } | undefined;
if (!result) throw new Error("pi-subagents is not installed or not ready.");
if (!result.ok || !result.registration) throw result.error ?? new Error("pi-subagents rejected the goal-steward agent.");
return result.registration;
}
export function readyReview(planPath: string): string {
return `Review reason: plan approved\nPlan path: ${planPath}\n\nRead the complete plan now. Check that its goals still match the user-visible result and that the first work step is sensible. This is not sign-off: return only let_run or redirect.`;
}
export function checkpointReview(planPath: string, staleTurns: number): string {
return `Review reason: progress checkpoint\nPlan path: ${planPath}\n\nThe worker completed ${staleTurns} turns without changing the plan's working set. Read the complete plan now and inspect only files needed to decide whether one concrete redirect would help. This is not sign-off: return only let_run or redirect.`;
}
export function signoffReview(planPath: string, goal: string): string {
return `Review reason: goal sign-off\nPlan path: ${planPath}\nClaimed goal: ${goal}\n\nRead the complete plan now. Check User-visible result, User voice, this goal's discriminator, failure mode, and evidence. Inspect the cited files. Return accept only when the evidence positively proves the requested result; otherwise return reject and name the missing evidence.`;
}
async function rpc(events: EventBus, method: "spawn" | "resume", params: Record<string, unknown>, signal?: AbortSignal): Promise<RpcData> {
if (signal?.aborted) throw new Error("Goal-steward request aborted.");
const requestId = randomUUID();
return new Promise((resolve, reject) => {
let timer: ReturnType<typeof setTimeout>;
const replyEvent = `${RPC_REPLY_PREFIX}${requestId}`;
const cleanup = () => {
clearTimeout(timer);
unsubscribe();
signal?.removeEventListener("abort", onAbort);
};
const onAbort = () => {
cleanup();
reject(new Error("Goal-steward request aborted."));
};
const unsubscribe = events.on(replyEvent, (raw) => {
const reply = raw as { success?: boolean; data?: RpcData; error?: { message?: string } };
cleanup();
if (!reply.success || !reply.data) reject(new Error(reply.error?.message ?? `pi-subagents ${method} failed.`));
else resolve(reply.data);
});
timer = setTimeout(() => {
cleanup();
reject(new Error(`pi-subagents ${method} did not reply within ${RPC_TIMEOUT_MS / 1000}s.`));
}, RPC_TIMEOUT_MS);
timer.unref();
signal?.addEventListener("abort", onAbort, { once: true });
events.emit(RPC_REQUEST_EVENT, { version: RPC_VERSION, requestId, method, params, source: { extension: "pi-goals" } });
});
}
function asyncRunId(data: RpcData): string {
const runId = data.details?.asyncId ?? data.details?.runId;
if (typeof runId !== "string" || !runId) throw new Error("pi-subagents returned no async run ID.");
return runId;
}
export async function startSteward(events: EventBus, cwd: string, task: string, signal?: AbortSignal): Promise<string> {
const data = await rpc(events, "spawn", {
agent: STEWARD_AGENT,
task,
cwd,
context: "fresh",
async: true,
mission: false,
outputSchema: stewardOutputSchema,
}, signal);
return asyncRunId(data);
}
export async function resumeSteward(events: EventBus, runId: string, task: string, signal?: AbortSignal): Promise<string> {
const data = await rpc(events, "resume", { id: runId, message: task }, signal);
return asyncRunId(data);
}
function completionRunId(raw: unknown): string | null {
if (!raw || typeof raw !== "object") return null;
const runId = (raw as Record<string, unknown>).runId;
return typeof runId === "string" ? runId : null;
}
export async function runStewardReview(
events: EventBus,
cwd: string,
previousRunId: string | null,
task: string,
signal?: AbortSignal,
timeoutMs = 600_000,
onRunId?: (runId: string) => void,
): Promise<{ runId: string; decision: StewardDecision }> {
if (signal?.aborted) throw new Error("Goal-steward review aborted.");
let expectedRunId: string | null = null;
const earlyCompletions: unknown[] = [];
let settle: (raw: unknown) => void = () => {};
let fail: (error: Error) => void = () => {};
let timer: ReturnType<typeof setTimeout> | undefined;
const completion = new Promise<StewardDecision>((resolve, reject) => {
settle = (raw) => {
try {
resolve(parseStewardDecision(raw));
} catch (error) {
reject(error);
}
};
fail = reject;
});
const unsubscribe = events.on(ASYNC_COMPLETE_EVENT, (raw) => {
const completedRunId = completionRunId(raw);
if (expectedRunId === null) {
earlyCompletions.push(raw);
return;
}
if (completedRunId === expectedRunId) settle(raw);
});
const onAbort = () => fail(new Error("Goal-steward review aborted."));
try {
expectedRunId = previousRunId
? await resumeSteward(events, previousRunId, task, signal)
: await startSteward(events, cwd, task, signal);
onRunId?.(expectedRunId);
if (signal?.aborted) throw new Error("Goal-steward review aborted.");
signal?.addEventListener("abort", onAbort, { once: true });
timer = setTimeout(() => fail(new Error(`Goal-steward review timed out after ${timeoutMs / 1000}s.`)), timeoutMs);
timer.unref();
const early = earlyCompletions.find((raw) => completionRunId(raw) === expectedRunId);
if (early) settle(early);
return { runId: expectedRunId, decision: await completion };
} finally {
if (timer) clearTimeout(timer);
unsubscribe();
signal?.removeEventListener("abort", onAbort);
}
}
export function parseStewardDecision(raw: unknown): StewardDecision {
if (!raw || typeof raw !== "object") throw new Error("Goal-steward completion was not an object.");
const results = (raw as Record<string, unknown>).results;
if (!Array.isArray(results) || results.length !== 1 || !results[0] || typeof results[0] !== "object") {
throw new Error("Goal-steward completion did not contain exactly one result.");
}
const child = results[0] as Record<string, unknown>;
if (child.success === false) throw new Error(typeof child.error === "string" ? child.error : "Goal-steward run failed.");
const value = child.structuredOutput;
if (!value || typeof value !== "object" || Array.isArray(value)) throw new Error("Goal-steward returned no structured verdict.");
const decision = value as Record<string, unknown>;
if (!(["let_run", "redirect", "accept", "reject"] as unknown[]).includes(decision.verdict) || typeof decision.summary !== "string" || !decision.summary.trim()) {
throw new Error("Goal-steward returned an invalid structured verdict.");
}
if (decision.verdict === "redirect" && (typeof decision.nextAction !== "string" || !decision.nextAction.trim())) {
throw new Error("Goal-steward redirect omitted nextAction.");
}
if (
decision.verdict === "reject"
&& (!Array.isArray(decision.missingEvidence) || decision.missingEvidence.length === 0 || decision.missingEvidence.some((item) => typeof item !== "string" || !item.trim()))
) {
throw new Error("Goal-steward rejection omitted missingEvidence.");
}
return decision as unknown as StewardDecision;
}
+33
View File
@@ -0,0 +1,33 @@
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
const PREPARED_FORK = "pi-goals-worker-fork-prepared";
export default function workerRuntime(pi: ExtensionAPI): void {
pi.on("session_start", async (_event, ctx) => {
const prepared = ctx.sessionManager
.getEntries()
.some((entry: { type?: string; customType?: string }) => entry.type === "custom" && entry.customType === PREPARED_FORK);
if (prepared) return;
await new Promise<void>((resolve, reject) => {
ctx.compact({
customInstructions: "__pi_vcc__ keep:0",
onComplete: (result) => {
if ((result.details as { compactor?: string } | undefined)?.compactor !== "pi-vcc") {
reject(new Error("Goal-worker fork was not compacted by pi-vcc."));
return;
}
pi.appendEntry(PREPARED_FORK, { version: 1, compacted: true, compactor: "pi-vcc" });
resolve();
},
onError: (error) => {
if (error.message !== "Nothing to compact (session too small)") {
reject(error);
return;
}
pi.appendEntry(PREPARED_FORK, { version: 1, compacted: false, reason: "below-compactable-size" });
resolve();
},
});
});
});
}
+161
View File
@@ -0,0 +1,161 @@
import { randomUUID } from "node:crypto";
import { fileURLToPath } from "node:url";
const REGISTER_EVENT = "pi-subagents:runtime-agent-register:v1";
const RPC_REQUEST_EVENT = "subagents:rpc:v1:request";
const RPC_REPLY_PREFIX = "subagents:rpc:v1:reply:";
const RPC_VERSION = 1;
const RPC_TIMEOUT_MS = 15_000;
export const WORKER_AGENT = "goal-worker";
interface EventBus {
on(event: string, handler: (data: unknown) => void): () => void;
emit(event: string, data: unknown): void;
}
interface Registration {
dispose(): void;
}
interface RpcData {
text: string;
details?: Record<string, unknown>;
asyncSnapshot?: AsyncSnapshot;
}
interface AsyncNode {
id: string;
state: string;
children?: AsyncNode[];
}
interface AsyncSnapshot {
kind: string;
version: number;
omitted: { runs: number; children: number; byteLimitExceeded: boolean };
runs: AsyncNode[];
}
export type WorkState = "active" | "idle" | "unknown";
export const workerSystemPrompt = `You are the implementation worker for one supervised Pi session.
Work autonomously from the approved plan. Keep the plan current, run the real checks, and leave
specific evidence in its Log. The human's latest message outranks the plan; update affected goals
instead of defending an obsolete decision. The main Pi agent is the research supervisor and owns
direction and goal sign-off. Send contact_supervisor progress updates when evidence changes the research direction,
when an hourly check asks for one, or when you need a decision. Do not claim a goal is complete;
report the evidence and let the supervisor decide. Continue until the plan is complete or the human
stops the session. -- Pi/Codex`;
export function registerGoalWorker(events: EventBus, model: string | null): Registration {
const runtimeExtension = fileURLToPath(new URL("./worker-runtime.ts", import.meta.url));
const piVccExtension = fileURLToPath(import.meta.resolve("@sting8k/pi-vcc"));
const request: Record<string, unknown> = {
version: 1,
name: WORKER_AGENT,
definition: {
description: "Implementation worker directed by the main goal supervisor.",
systemPrompt: workerSystemPrompt,
allowNestedSubagents: true,
subagentOnlyExtensions: [piVccExtension, runtimeExtension],
...(model ? { model } : {}),
systemPromptMode: "replace",
inheritProjectContext: true,
inheritGlobalContext: true,
inheritSkills: true,
defaultContext: "fork",
defaultAsync: true,
defaultProgress: true,
},
};
events.emit(REGISTER_EVENT, request);
const result = request.result as { ok?: boolean; registration?: Registration; error?: Error } | undefined;
if (!result) throw new Error("pi-subagents is not installed or not ready.");
if (!result.ok || !result.registration) throw result.error ?? new Error("pi-subagents rejected the goal-worker agent.");
return result.registration;
}
async function rpc(events: EventBus, method: "spawn" | "resume" | "steer" | "status", params: Record<string, unknown>, signal?: AbortSignal): Promise<RpcData> {
if (signal?.aborted) throw new Error("Goal-worker request aborted.");
const requestId = randomUUID();
return new Promise((resolve, reject) => {
let timer: ReturnType<typeof setTimeout>;
const replyEvent = `${RPC_REPLY_PREFIX}${requestId}`;
const cleanup = () => {
clearTimeout(timer);
unsubscribe();
signal?.removeEventListener("abort", onAbort);
};
const onAbort = () => {
cleanup();
reject(new Error("Goal-worker request aborted."));
};
const unsubscribe = events.on(replyEvent, (raw) => {
const reply = raw as { success?: boolean; data?: RpcData; error?: { message?: string } };
cleanup();
if (!reply.success || !reply.data) reject(new Error(reply.error?.message ?? `pi-subagents ${method} failed.`));
else resolve(reply.data);
});
timer = setTimeout(() => {
cleanup();
reject(new Error(`pi-subagents ${method} did not reply within ${RPC_TIMEOUT_MS / 1000}s.`));
}, RPC_TIMEOUT_MS);
timer.unref();
signal?.addEventListener("abort", onAbort, { once: true });
events.emit(RPC_REQUEST_EVENT, { version: RPC_VERSION, requestId, method, params, source: { extension: "pi-goals" } });
});
}
function asyncRunId(data: RpcData): string {
const runId = data.details?.asyncId ?? data.details?.runId;
if (typeof runId !== "string" || !runId) throw new Error("pi-subagents returned no async run ID.");
return runId;
}
export async function startGoalWorker(events: EventBus, cwd: string, task: string, signal?: AbortSignal): Promise<string> {
const data = await rpc(events, "spawn", {
agent: WORKER_AGENT,
task,
cwd,
context: "fork",
async: true,
mission: false,
}, signal);
return asyncRunId(data);
}
export async function resumeGoalWorker(events: EventBus, runId: string, task: string, signal?: AbortSignal): Promise<string> {
return asyncRunId(await rpc(events, "resume", { id: runId, message: task }, signal));
}
export async function steerGoalWorker(events: EventBus, runId: string, task: string, signal?: AbortSignal): Promise<void> {
await rpc(events, "steer", { id: runId, message: task, mode: "steer" }, signal);
}
function activeNode(node: AsyncNode): boolean {
return node.state === "queued" || node.state === "running" || Boolean(node.children?.some(activeNode));
}
export async function subagentWorkState(events: EventBus): Promise<WorkState> {
const snapshot = (await rpc(events, "status", {})).asyncSnapshot;
if (snapshot?.kind !== "pi-subagents.async-status-snapshot" || snapshot.version !== 1) return "unknown";
if (snapshot.omitted.runs > 0 || snapshot.omitted.children > 0 || snapshot.omitted.byteLimitExceeded) return "unknown";
return snapshot.runs.some(activeNode) ? "active" : "idle";
}
export interface ProcessInfo {
status: string;
}
export function processWorkState(events: EventBus): WorkState {
let replied = false;
let processes: ProcessInfo[] = [];
events.emit("processes:request:list", {
reply(value: ProcessInfo[]) {
replied = true;
processes = value;
},
});
if (!replied) return "unknown";
return processes.some((process) => process.status === "running" || process.status === "terminating" || process.status === "terminate_timeout") ? "active" : "idle";
}
-38
View File
@@ -1,38 +0,0 @@
import { describe, expect, it } from "vitest";
import { decideStewardSignOff } from "../src/index.js";
describe("decideStewardSignOff", () => {
it("accepts only the steward's accept verdict", () => {
const out = decideStewardSignOff("produce report", { verdict: "accept", summary: "The cited report contains every required row." }, "run-2");
expect(out.isError).toBe(false);
expect(out.resultText).toContain("Sign-off ACCEPTED");
expect(out.logEntry).toContain("steward accept; run run-2");
});
it("reports the steward's missing evidence", () => {
const out = decideStewardSignOff(
"produce report",
{ verdict: "reject", summary: "The count is not cited.", missingEvidence: ["Saved output with the recursive file count", "A matching table row count"] },
"run-3",
);
expect(out.isError).toBe(true);
expect(out.resultText).toContain("Saved output with the recursive file count; A matching table row count");
expect(out.logEntry).toContain("steward run run-3");
});
it("treats redirect as a rejected sign-off with one next action", () => {
const out = decideStewardSignOff(
"produce report",
{ verdict: "redirect", summary: "The worker inspected only the top level.", nextAction: "Repeat the snapshot recursively." },
"run-4",
);
expect(out.isError).toBe(true);
expect(out.resultText).toContain("Repeat the snapshot recursively.");
});
it("does not turn let_run into acceptance", () => {
const out = decideStewardSignOff("produce report", { verdict: "let_run", summary: "Continue the current work." }, "run-5");
expect(out.isError).toBe(true);
expect(out.resultText).toContain("let_run instead of a sign-off verdict");
});
});
+1 -1
View File
@@ -28,7 +28,7 @@ const plan = `# Plan
## Appendix (context, not approved)
${"filler line\n".repeat(200)}`;
describe("foldPlan (the working set is what gets re-sent; below ## Log is durable memory)", () => {
describe("foldPlan (current goals are above ## Log; durable memory is below it)", () => {
it("keeps the title, user voice and goals", () => {
const folded = foldPlan(plan);
expect(folded).toContain("keep it under 50 lines");
+99 -69
View File
@@ -9,6 +9,8 @@ function setup(
selectChoices: Array<string | undefined>,
editorChoices: Array<string | undefined> = [],
editPlan?: () => Promise<string | undefined>,
contextTokens = 0,
completeWorkerBeforeReply = false,
) {
const cwd = mkdtempSync(join(tmpdir(), "pi-goals-flow-"));
const commands = new Map<string, any>();
@@ -18,6 +20,7 @@ function setup(
const eventLog: string[] = [];
const messages: Array<{ content: string; display?: boolean }> = [];
const rpcRequests: any[] = [];
const compactCalls: any[] = [];
const eventHandlers = new Map<string, Set<(data: unknown) => void>>();
const eventBus = {
on(name: string, handler: (data: unknown) => void) {
@@ -37,13 +40,33 @@ function setup(
eventBus.on("subagents:rpc:v1:request", (raw) => {
const request = raw as any;
rpcRequests.push(request);
if (request.method === "status") {
eventBus.emit(`subagents:rpc:v1:reply:${request.requestId}`, {
success: true,
data: {
text: "status",
asyncSnapshot: { kind: "pi-subagents.async-status-snapshot", version: 1, omitted: { runs: 0, children: 0, byteLimitExceeded: false }, runs: [] },
},
});
return;
}
asyncRun++;
eventBus.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "started", details: { asyncId: `steward-${asyncRun}` } } });
if (completeWorkerBeforeReply) eventBus.emit("subagent:async-complete", { runId: `worker-${asyncRun}`, results: [{ success: true }] });
eventBus.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "started", details: { asyncId: `worker-${asyncRun}` } } });
});
eventBus.on("processes:request:list", (raw) => {
(raw as { reply(value: object[]): void }).reply([]);
});
const ctx = {
cwd,
hasUI: true,
isIdle: () => true,
getContextUsage: () => ({ tokens: contextTokens }),
compact: (options: any) => {
compactCalls.push(options);
options.onComplete?.({ summary: "summary", details: { compactor: "pi-vcc" } });
},
getSystemPrompt: () => "base prompt",
sessionManager: { getSessionId: () => "session-a", getEntries: () => entries },
ui: {
theme: { fg: (_kind: string, text: string) => text },
@@ -73,7 +96,7 @@ function setup(
sendUserMessage: (message: string) => messages.push({ content: message }),
};
piGoalsExtension(pi as unknown as ExtensionAPI);
return { commands, ctx, cwd, entries, events: eventLog, eventBus, hooks, messages, rpcRequests, tools };
return { commands, compactCalls, ctx, cwd, entries, events: eventLog, eventBus, hooks, messages, rpcRequests, tools };
}
describe("/goals draft flow", () => {
@@ -169,10 +192,10 @@ describe("/goals draft flow", () => {
await flow.hooks.get("agent_settled")({}, flow.ctx);
expect(flow.events).toEqual(["display", "select"]);
expect(flow.messages.filter((message) => !message.display)).toHaveLength(2);
expect(flow.messages.at(-1)?.content).toContain("Work the goals");
await flow.hooks.get("session_start")({}, flow.ctx);
expect(await flow.hooks.get("before_agent_start")({}, flow.ctx)).toBeUndefined();
expect(flow.rpcRequests[0]).toMatchObject({ method: "spawn", params: { agent: "goal-worker", context: "fork" } });
expect(flow.messages.filter((message) => !message.display)).toHaveLength(1);
const supervisor = await flow.hooks.get("before_agent_start")({}, flow.ctx);
expect(supervisor.systemPrompt).toContain("research supervisor");
} finally {
rmSync(flow.cwd, { recursive: true, force: true });
}
@@ -197,34 +220,24 @@ describe("/goals draft flow", () => {
}
});
it("reminds every eight unchanged working-set turns, ignoring log-only edits", async () => {
it("resyncs the whole plan once without telling the supervisor to implement it", async () => {
const flow = setup(["Ready"]);
try {
await flow.commands.get("goals").handler("objective", flow.ctx);
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n\n## Log\n");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n\n## Log\n- worker evidence\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
await flow.hooks.get("turn_end")({}, flow.ctx);
for (let turn = 0; turn < 3; turn++) await flow.hooks.get("turn_end")({}, flow.ctx);
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n\n## Log\n- checked input\n");
for (let turn = 0; turn < 5; turn++) await flow.hooks.get("turn_end")({}, flow.ctx);
const reminder = await flow.hooks.get("context")({ messages: [] }, flow.ctx);
expect(reminder.messages.at(-1).content[0].text).toContain(".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n - [x] inspect input\n\n## Log\n- checked input\n");
await flow.hooks.get("turn_end")({}, flow.ctx);
for (let turn = 0; turn < 7; turn++) await flow.hooks.get("turn_end")({}, flow.ctx);
expect((await flow.hooks.get("context")({ messages: [] }, flow.ctx)).messages).toHaveLength(0);
await flow.hooks.get("turn_end")({}, flow.ctx);
expect((await flow.hooks.get("context")({ messages: [] }, flow.ctx)).messages.at(-1).content[0].text).toContain("make the output");
const resync = await flow.hooks.get("context")({ messages: [] }, flow.ctx);
expect(resync.messages.at(-1).content[0].text).toContain("worker evidence");
expect(resync.messages.at(-1).content[0].text).not.toContain("Keep it current as you work");
expect(await flow.hooks.get("context")({ messages: [] }, flow.ctx)).toBeUndefined();
} finally {
rmSync(flow.cwd, { recursive: true, force: true });
}
});
it("auto-continues once on stop, then pauses after two no-progress wakes", async () => {
it("checks every interval without pausing after unchanged work", async () => {
vi.useFakeTimers();
const flow = setup(["Ready"]);
try {
@@ -233,43 +246,56 @@ describe("/goals draft flow", () => {
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
await flow.commands.get("goals").handler("auto 1", flow.ctx);
const checks = () => flow.messages.filter((message) => message.content.includes("supervisor check is due"));
await vi.advanceTimersByTimeAsync(30_000);
await flow.hooks.get("agent_settled")({}, flow.ctx);
await vi.advanceTimersByTimeAsync(0);
const autoMessages = () => flow.messages.filter((message) => message.content.includes("Auto-continue is enabled"));
expect(autoMessages()).toHaveLength(1);
await vi.advanceTimersByTimeAsync(30_000);
expect(checks()).toHaveLength(1);
await flow.hooks.get("before_agent_start")({}, flow.ctx);
await flow.hooks.get("agent_settled")({}, flow.ctx);
await vi.advanceTimersByTimeAsync(60_000);
expect(autoMessages()).toHaveLength(2);
await flow.hooks.get("agent_settled")({}, flow.ctx);
await vi.advanceTimersByTimeAsync(60_000);
expect(autoMessages()).toHaveLength(2);
for (let n = 2; n <= 3; n++) {
await vi.advanceTimersByTimeAsync(60_000);
expect(checks()).toHaveLength(n);
await flow.hooks.get("before_agent_start")({}, flow.ctx);
await flow.hooks.get("agent_settled")({}, flow.ctx);
}
} finally {
vi.useRealTimers();
rmSync(flow.cwd, { recursive: true, force: true });
}
});
it("delays auto-continuation after a known background start", async () => {
vi.useFakeTimers();
const flow = setup(["Ready"]);
it("compacts the main supervisor with pi-vcc near 100k tokens", async () => {
const flow = setup(["Ready"], [], undefined, 100_000);
try {
await flow.commands.get("goals").handler("objective", flow.ctx);
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
await flow.commands.get("goals").handler("auto 1", flow.ctx);
await flow.hooks.get("agent_start")({}, flow.ctx);
await flow.hooks.get("tool_call")({ toolName: "process", input: { action: "start" } }, flow.ctx);
await flow.hooks.get("agent_settled")({}, flow.ctx);
await vi.advanceTimersByTimeAsync(0);
const autoMessages = () => flow.messages.filter((message) => message.content.includes("Auto-continue is enabled"));
expect(autoMessages()).toHaveLength(0);
await vi.advanceTimersByTimeAsync(60_000);
expect(autoMessages()).toHaveLength(1);
expect(flow.compactCalls).toHaveLength(1);
expect(flow.compactCalls[0].customInstructions).toBe("__pi_vcc__ keep:1");
} finally {
rmSync(flow.cwd, { recursive: true, force: true });
}
});
it("checks exact subagent and process status after native worker completion", async () => {
const flow = setup(["Ready"]);
try {
await flow.hooks.get("session_start")({}, flow.ctx);
await flow.commands.get("goals").handler("objective", flow.ctx);
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: make the output\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
flow.eventBus.emit("subagent:async-complete", { runId: "worker-1", results: [{ success: true }] });
const status = await flow.tools.get("CheckGoalWork").execute("", {}, undefined, undefined, flow.ctx);
expect(status.isError).toBe(false);
expect(status.content[0].text).toBe("subagents=idle; processes=idle");
expect(flow.messages.some((message) => message.content.includes("worker stopped"))).toBe(false);
} finally {
vi.useRealTimers();
rmSync(flow.cwd, { recursive: true, force: true });
}
});
@@ -282,7 +308,7 @@ describe("/goals draft flow", () => {
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: produce report\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
flow.eventBus.emit("subagent:async-complete", { runId: "steward-1" });
flow.eventBus.emit("subagent:async-complete", { runId: "worker-1" });
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [x] goal: produce report\n");
await flow.hooks.get("turn_end")({}, flow.ctx);
@@ -295,7 +321,23 @@ describe("/goals draft flow", () => {
}
});
it("resumes the same steward lineage for sign-off and persists the latest run", async () => {
it("records a worker that completes before its launch RPC reply as stopped", async () => {
const flow = setup(["Ready"], [], undefined, 0, true);
try {
await flow.hooks.get("session_start")({}, flow.ctx);
await flow.commands.get("goals").handler("objective", flow.ctx);
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: finish quickly\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
expect(flow.entries.at(-1)?.data).toMatchObject({ workerRunId: "worker-1", workerPending: false });
} finally {
rmSync(flow.cwd, { recursive: true, force: true });
}
});
it("resumes the retained worker and lets the main supervisor sign off", async () => {
const flow = setup(["Ready"]);
try {
await flow.hooks.get("session_start")({}, flow.ctx);
@@ -303,34 +345,22 @@ describe("/goals draft flow", () => {
const planPath = join(flow.cwd, ".pi/plan/session-a-v1.md");
writeFileSync(planPath, "# Plan\n\n## Goals\n\n1. [/] goal: produce report\n - discriminator: report.txt contains PASS\n - evidence:\n - report.txt: `PASS`\n\n## Log\n");
await flow.hooks.get("agent_settled")({}, flow.ctx);
expect(flow.rpcRequests[0]).toMatchObject({ method: "spawn", params: { agent: "goal-steward" } });
flow.eventBus.emit("subagent:async-complete", {
runId: "steward-1",
results: [{ success: true, structuredOutput: { verdict: "let_run", summary: "Start work." } }],
});
expect(flow.rpcRequests[0]).toMatchObject({ method: "spawn", params: { agent: "goal-worker", context: "fork" } });
flow.eventBus.emit("subagent:async-complete", { runId: "worker-1", results: [{ success: true }] });
const signoff = flow.tools.get("CompleteGoal").execute("", { goal: "produce report" }, undefined, undefined, flow.ctx);
await new Promise((resolve) => setImmediate(resolve));
expect(flow.rpcRequests[1]).toMatchObject({ method: "resume", params: { id: "steward-1" } });
flow.eventBus.emit("subagent:async-complete", {
runId: "steward-2",
results: [{ success: true, structuredOutput: { verdict: "accept", summary: "report.txt contains PASS." } }],
});
const outcome = await signoff;
const resumed = await flow.tools.get("GuideGoalWorker").execute("", { instruction: "Verify report.txt." }, undefined, undefined, flow.ctx);
expect(resumed.isError).toBe(false);
expect(flow.rpcRequests[1]).toMatchObject({ method: "resume", params: { id: "worker-1", message: "Verify report.txt." } });
expect(flow.entries.at(-1)?.data).toMatchObject({ workerRunId: "worker-2", workerPending: true });
expect(outcome.isError).toBe(false);
const signoff = await flow.tools.get("CompleteGoal").execute("", { goal: "produce report" }, undefined, undefined, flow.ctx);
expect(signoff.isError).toBe(false);
expect(readFileSync(planPath, "utf-8")).toContain("1. [x] goal: produce report");
expect(flow.entries.at(-1)?.data).toMatchObject({ stewardRunId: "steward-2", stewardPending: false });
expect(readFileSync(planPath, "utf-8")).toContain("signed off \"produce report\" by the main research supervisor");
await flow.hooks.get("session_start")({}, flow.ctx);
const afterReload = flow.tools.get("CompleteGoal").execute("", { goal: "produce report" }, undefined, undefined, flow.ctx);
await new Promise((resolve) => setImmediate(resolve));
expect(flow.rpcRequests[2]).toMatchObject({ method: "resume", params: { id: "steward-2" } });
flow.eventBus.emit("subagent:async-complete", {
runId: "steward-3",
results: [{ success: true, structuredOutput: { verdict: "reject", summary: "Already complete.", missingEvidence: ["No second sign-off needed"] } }],
});
expect((await afterReload).isError).toBe(true);
await flow.tools.get("GuideGoalWorker").execute("", { instruction: "Report current status." }, undefined, undefined, flow.ctx);
expect(flow.rpcRequests.at(-1)).toMatchObject({ method: "steer", params: { id: "worker-2", message: "Report current status." } });
} finally {
rmSync(flow.cwd, { recursive: true, force: true });
}
+5 -5
View File
@@ -1,6 +1,6 @@
import { describe, expect, it } from "vitest";
import { planDrafting, planningState, reminder, resync } from "../src/prompts.js";
import { stewardSystemPrompt } from "../src/steward.js";
import { completeGoalDescription, planDrafting, planningState, resync } from "../src/prompts.js";
import { workerSystemPrompt } from "../src/worker.js";
describe("planning prompt", () => {
it("requires fact finding or a focused question before a goal", () => {
@@ -23,9 +23,9 @@ describe("planning prompt", () => {
expect(planDrafting).toContain("## User-visible result");
expect(planDrafting).toContain("Take it from the original request, not from your implementation plan");
expect(planDrafting).toContain("Future work may not defer any artifact or action named there");
expect(reminder("plan", ".pi/plan/test.md")).toContain("latest message outranks this plan");
expect(resync("plan", ".pi/plan/test.md", "Compacted.")).toContain("amend the plan rather than preserving an obsolete decision");
expect(stewardSystemPrompt).toContain("Treat User-visible result and User voice as the authority");
expect(stewardSystemPrompt).toContain("Do not accept a confident summary as evidence");
expect(workerSystemPrompt).toContain("latest message outranks the plan");
expect(workerSystemPrompt).toContain("main Pi agent is the research supervisor");
expect(completeGoalDescription).toContain("inspect the live working tree");
});
});
-133
View File
@@ -1,133 +0,0 @@
import { describe, expect, it, vi } from "vitest";
import {
checkpointReview,
parseStewardDecision,
readyReview,
registerStewardAgent,
runStewardReview,
signoffReview,
startSteward,
stewardSystemPrompt,
} from "../src/steward.js";
class Events {
private handlers = new Map<string, Set<(data: unknown) => void>>();
on(event: string, handler: (data: unknown) => void): () => void {
const handlers = this.handlers.get(event) ?? new Set();
handlers.add(handler);
this.handlers.set(event, handlers);
return () => handlers.delete(handler);
}
emit(event: string, data: unknown): void {
for (const handler of [...(this.handlers.get(event) ?? [])]) handler(data);
}
}
function completion(runId: string, verdict: "let_run" | "redirect" | "accept" | "reject" = "accept") {
return {
runId,
results: [{ success: true, structuredOutput: { verdict, summary: "Observed the cited artifact." } }],
};
}
describe("goal steward registration", () => {
it("registers one read-only runtime agent through pi-subagents", () => {
const events = new Events();
let definition: Record<string, unknown> | undefined;
events.on("pi-subagents:runtime-agent-register:v1", (raw) => {
const request = raw as { definition: Record<string, unknown>; result?: unknown };
definition = request.definition;
request.result = { ok: true, registration: { dispose() {} } };
});
registerStewardAgent(events, "provider/cheap-model");
expect(definition?.model).toBe("provider/cheap-model");
expect(definition?.tools).toEqual(["read", "grep", "find", "ls", "contact_supervisor"]);
expect(definition?.excludeTools).toEqual(expect.arrayContaining(["bash", "edit", "write", "subagent"]));
expect(definition?.inheritProjectContext).toBe(false);
expect(stewardSystemPrompt).toContain("Read the complete plan from disk every time");
expect(stewardSystemPrompt).toContain("Call structured_output as soon as the evidence is sufficient");
});
it("fails clearly when pi-subagents is absent", () => {
expect(() => registerStewardAgent(new Events(), null)).toThrow("pi-subagents is not installed or not ready");
});
it("reserves accept and reject for sign-off", () => {
expect(readyReview("plan.md")).toContain("return only let_run or redirect");
expect(checkpointReview("plan.md", 8)).toContain("return only let_run or redirect");
expect(signoffReview("plan.md", "goal")).toContain("Return accept only");
});
});
describe("goal steward RPC", () => {
it("starts a detached structured child", async () => {
const events = new Events();
let request: any;
events.on("subagents:rpc:v1:request", (raw) => {
request = raw;
events.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "started", details: { asyncId: "run-1" } } });
});
await expect(startSteward(events, "/repo", "review now")).resolves.toBe("run-1");
expect(request.method).toBe("spawn");
expect(request.params).toMatchObject({ agent: "goal-steward", cwd: "/repo", context: "fresh", async: true, task: "review now" });
expect(request.params.outputSchema.required).toEqual(["verdict", "summary"]);
});
it("resumes the same lineage and accepts only its exact completion", async () => {
const events = new Events();
let request: any;
events.on("subagents:rpc:v1:request", (raw) => {
request = raw;
events.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "resumed", details: { asyncId: "run-2" } } });
queueMicrotask(() => {
events.emit("subagent:async-complete", completion("other-run", "reject"));
events.emit("subagent:async-complete", completion("run-2"));
});
});
const review = await runStewardReview(events, "/repo", "run-1", "sign off");
expect(request.method).toBe("resume");
expect(request.params).toEqual({ id: "run-1", message: "sign off" });
expect(review).toEqual({ runId: "run-2", decision: { verdict: "accept", summary: "Observed the cited artifact." } });
});
it("does not lose a completion emitted before the RPC reply", async () => {
const events = new Events();
events.on("subagents:rpc:v1:request", (raw) => {
const request = raw as { requestId: string };
events.emit("subagent:async-complete", completion("run-fast"));
events.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "started", details: { asyncId: "run-fast" } } });
});
await expect(runStewardReview(events, "/repo", null, "review")).resolves.toMatchObject({ runId: "run-fast", decision: { verdict: "accept" } });
});
it("fails closed when the exact child does not complete", async () => {
vi.useFakeTimers();
try {
const events = new Events();
events.on("subagents:rpc:v1:request", (raw) => {
const request = raw as { requestId: string };
events.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: { text: "started", details: { asyncId: "run-stuck" } } });
});
const review = runStewardReview(events, "/repo", null, "review", undefined, 1_000);
const rejected = expect(review).rejects.toThrow("timed out after 1s");
await vi.advanceTimersByTimeAsync(1_000);
await rejected;
} finally {
vi.useRealTimers();
}
});
});
describe("steward verdict parsing", () => {
it("requires the fields for redirect and reject", () => {
expect(() => parseStewardDecision(completion("run", "redirect"))).toThrow("redirect omitted nextAction");
expect(() => parseStewardDecision(completion("run", "reject"))).toThrow("rejection omitted missingEvidence");
});
});
+45
View File
@@ -0,0 +1,45 @@
import { describe, expect, it } from "vitest";
import workerRuntime from "../src/worker-runtime.js";
function setup(entries: object[], details: object = { compactor: "pi-vcc" }, compactError?: Error) {
const hooks = new Map<string, any>();
const appended: Array<{ type: string; data: unknown }> = [];
const pi = {
on: (name: string, handler: any) => hooks.set(name, handler),
appendEntry: (type: string, data: unknown) => appended.push({ type, data }),
};
const ctx = {
sessionManager: { getEntries: () => [...entries, ...appended.map(({ type, data }) => ({ type: "custom", customType: type, data }))] },
compact: ({ customInstructions, onComplete, onError }: any) => {
if (compactError) onError(compactError);
else onComplete({ summary: "summary", firstKeptEntryId: "", tokensBefore: 1000, details, customInstructions });
},
};
workerRuntime(pi as any);
return { hooks, ctx, appended };
}
describe("goal-worker fork compaction", () => {
it("requires pi-vcc before the first worker turn and records completion", async () => {
const runtime = setup([]);
await runtime.hooks.get("session_start")({}, runtime.ctx);
expect(runtime.appended).toEqual([{ type: "pi-goals-worker-fork-prepared", data: { version: 1, compacted: true, compactor: "pi-vcc" } }]);
});
it("records when the exact fork is already too small to compact", async () => {
const runtime = setup([], undefined, new Error("Nothing to compact (session too small)"));
await runtime.hooks.get("session_start")({}, runtime.ctx);
expect(runtime.appended).toEqual([{ type: "pi-goals-worker-fork-prepared", data: { version: 1, compacted: false, reason: "below-compactable-size" } }]);
});
it("does not prepare the retained worker again after resume", async () => {
const runtime = setup([{ type: "custom", customType: "pi-goals-worker-fork-prepared" }]);
await runtime.hooks.get("session_start")({}, runtime.ctx);
expect(runtime.appended).toEqual([]);
});
it("fails if another compactor handled the fork", async () => {
const runtime = setup([], { compactor: "other" });
await expect(runtime.hooks.get("session_start")({}, runtime.ctx)).rejects.toThrow("not compacted by pi-vcc");
});
});
+106
View File
@@ -0,0 +1,106 @@
import { describe, expect, it } from "vitest";
import {
processWorkState,
registerGoalWorker,
resumeGoalWorker,
startGoalWorker,
steerGoalWorker,
subagentWorkState,
workerSystemPrompt,
} from "../src/worker.js";
class Events {
private handlers = new Map<string, Set<(data: unknown) => void>>();
on(event: string, handler: (data: unknown) => void): () => void {
const handlers = this.handlers.get(event) ?? new Set();
handlers.add(handler);
this.handlers.set(event, handlers);
return () => handlers.delete(handler);
}
emit(event: string, data: unknown): void {
for (const handler of [...(this.handlers.get(event) ?? [])]) handler(data);
}
}
function replyToRpc(events: Events, inspect: (request: any) => object): void {
events.on("subagents:rpc:v1:request", (raw) => {
const request = raw as any;
events.emit(`subagents:rpc:v1:reply:${request.requestId}`, { success: true, data: inspect(request) });
});
}
describe("goal worker registration", () => {
it("registers one retained implementation worker", () => {
const events = new Events();
let definition: Record<string, unknown> | undefined;
events.on("pi-subagents:runtime-agent-register:v1", (raw) => {
const request = raw as { definition: Record<string, unknown>; result?: unknown };
definition = request.definition;
request.result = { ok: true, registration: { dispose() {} } };
});
registerGoalWorker(events, "provider/cheap-model");
expect(definition?.model).toBe("provider/cheap-model");
expect(definition?.defaultContext).toBe("fork");
expect(definition?.defaultProgress).toBe(true);
expect(definition?.allowNestedSubagents).toBe(true);
expect(definition?.subagentOnlyExtensions).toEqual([
expect.stringContaining("pi-vcc"),
expect.stringContaining("worker-runtime.ts"),
]);
expect(workerSystemPrompt).toContain("main Pi agent is the research supervisor");
});
it("fails clearly when pi-subagents is absent", () => {
expect(() => registerGoalWorker(new Events(), null)).toThrow("pi-subagents is not installed or not ready");
});
});
describe("goal worker RPC", () => {
it("starts from a fork, resumes retained context, and steers a live run", async () => {
const events = new Events();
const requests: any[] = [];
replyToRpc(events, (request) => {
requests.push(request);
return { text: "ok", details: { asyncId: `run-${requests.length}` } };
});
await expect(startGoalWorker(events, "/repo", "start")).resolves.toBe("run-1");
await expect(resumeGoalWorker(events, "run-1", "continue")).resolves.toBe("run-2");
await steerGoalWorker(events, "run-2", "report");
expect(requests[0]).toMatchObject({ method: "spawn", params: { agent: "goal-worker", cwd: "/repo", context: "fork", async: true } });
expect(requests[1]).toMatchObject({ method: "resume", params: { id: "run-1", message: "continue" } });
expect(requests[2]).toMatchObject({ method: "steer", params: { id: "run-2", message: "report", mode: "steer" } });
});
it("reports active, idle, and incomplete status snapshots", async () => {
for (const [snapshot, expected] of [
[{ kind: "pi-subagents.async-status-snapshot", version: 1, omitted: { runs: 0, children: 0, byteLimitExceeded: false }, runs: [{ id: "worker", state: "running" }] }, "active"],
[{ kind: "pi-subagents.async-status-snapshot", version: 1, omitted: { runs: 0, children: 0, byteLimitExceeded: false }, runs: [{ id: "worker", state: "complete", children: [{ id: "nested", state: "running" }] }] }, "active"],
[{ kind: "pi-subagents.async-status-snapshot", version: 1, omitted: { runs: 0, children: 0, byteLimitExceeded: false }, runs: [{ id: "worker", state: "complete" }] }, "idle"],
[{ kind: "pi-subagents.async-status-snapshot", version: 1, omitted: { runs: 1, children: 0, byteLimitExceeded: false }, runs: [] }, "unknown"],
] as const) {
const events = new Events();
replyToRpc(events, () => ({ text: "status", asyncSnapshot: snapshot }));
await expect(subagentWorkState(events)).resolves.toBe(expected);
}
});
});
describe("managed process status", () => {
it("does not treat a missing process extension as idle", () => {
expect(processWorkState(new Events())).toBe("unknown");
});
it("uses pi-processes live statuses", () => {
const events = new Events();
events.on("processes:request:list", (raw) => {
(raw as { reply(value: object[]): void }).reply([{ status: "terminate_timeout" }]);
});
expect(processWorkState(events)).toBe("active");
});
});