diff --git a/README.md b/README.md index 490ea73..102266f 100644 --- a/README.md +++ b/README.md @@ -9,6 +9,10 @@ The plan file looks like this: +### User-visible result + + + ### User voice - │ "" @@ -64,8 +68,9 @@ pi -e ./src/index.ts `/goals` enters plan mode and starts a conversation; the objective is an optional seed. From there: 1. Plan. The agent explores read-only and drafts the plan. -2. Review. After Pi settles, the full plan is printed in the transcript. The menu offers Ready, - Refine, Edit, or Cancel. Refine collects short notes. Edit opens the full plan in Pi's editor. +2. Review. After Pi settles, the full plan is printed in the transcript. Check that User-visible + result names the final artifact or behavior you expect. The menu offers Ready, Refine, Edit, or + Cancel. Refine collects short notes. Edit opens the full plan in Pi's editor. 3. Work. Ready is the only review action that starts work. The agent ticks subtasks, appends to `## Log` and `## Learnings`, fills `evidence:`, and calls `CompleteGoal` when a discriminator is satisfied. Every human reply and Refine note in plan mode is saved verbatim under `## Interview`. diff --git a/src/prompts.ts b/src/prompts.ts index 5652f45..28005ff 100644 --- a/src/prompts.ts +++ b/src/prompts.ts @@ -43,7 +43,11 @@ while discovering the right plan. Each question must be short and self-contained context, use the human's language and ASD-STE100 Simple Technical English, and give a recommended answer. Record each answer in ## Interview. Do not make the plan final while material user decisions remain open. -4. When every goal has an object, observable result, settled scope, and required approval, draft the +4. State the user-visible result before the goals: one concrete sentence naming what the human will +inspect when this plan is done. Take it from the original request, not from your implementation plan. +Every requested artifact and action must survive into this sentence. An agent-inferred constraint may +not replace, defer, or contradict it; ask the human if an inference would change the result. +5. When every goal has an object, observable result, settled scope, and required approval, draft the plan file and present it. It should be safe to work overnight and present the requested outcome. How this mode ends: after each settled draft the human gets a menu (Ready / Refine / Edit / Cancel). @@ -74,6 +78,10 @@ Write the plan file in roughly this shape -- the file is read directly by the hu +## User-visible result + + + ## User voice - > "" @@ -117,9 +125,11 @@ Conventions: - evidence stays empty at planning; you fill it at sign-off and a fresh read-only judge checks it. Cite durable artifacts a future reader can open: committed files, test names, git diffs. .pi/ is usually gitignored, so files there prove things only at judge time, not in history. +- User-visible result: restate the original deliverable, not the proposed implementation. Every goal + must contribute to it. Future work may not defer any artifact or action named there. - User voice: quote the human word for word, one line per requirement, as they say it. Never paraphrase there -- a paraphrase drifts, and then the goals churn on the next reply. It is exempt - from the working-set line limit. + from the working-set line limit. Never put an agent inference in User voice. - Interview: every human reply in plan mode is stored here verbatim as a dated blockquote. It is durable memory below the fold, not a substitute for ## User voice. - Rejected options stay visible: ~~struck through~~ with who rejected them and why, so nobody @@ -165,6 +175,8 @@ Keep it current as you work, with your normal edit tool: Don't tick a goal [x] before CompleteGoal accepts; the sign-off log line is the audit trail. - if the working set has grown long, prune finished goals (their evidence lives in git history and ## Log) and move settled detail down to ## Appendix, which is unlimited +- the human's latest message outranks this plan. If it corrects the deliverable or scope, amend the + user-visible result, user voice, and affected goals before continuing; don't defend the old plan - otherwise keep working toward the active goal; don't stop to ask unless genuinely blocked `; } @@ -177,8 +189,9 @@ Keep it current as you work, with your normal edit tool: export function resync(plan: string, planRel: string, why: string): string { return `\ -${why} This is the whole plan file (${planRel}), appendix included, so you don't re-litigate what -was already settled. Keep working the active goal; edit the file directly as you go. +${why} This is the whole plan file (${planRel}), appendix included. Keep working the active goal; +edit the file directly as you go. The human's latest message outranks the plan: if it corrects the +deliverable or scope, amend the plan rather than preserving an obsolete decision. ${plan} `; @@ -196,7 +209,9 @@ export const completeGoalDescription = "the goal names a verify: command, run it yourself first and save its output to a file cited in " + "the evidence: the judge cannot execute anything and will reject a claimed pass with no saved " + "output. The read must show success POSITIVELY happened, not just that failures were avoided. " + - "Then call this with the goal's text (the line after 'goal:'; small wording drift is fine). A " + + "Check that the claimed result uses the artifact and outcome named in User-visible result and does " + + "not substitute an agent-inferred deliverable. Then call this with the goal's text (the line after " + + "'goal:'; small wording drift is fine). A " + "fresh strictly-read-only judge inspects the LIVE WORKING TREE (uncommitted changes included; " + "committing first is for durability, not visibility) and returns accept or reject with what's " + "missing. On accept (or if the judge itself failed), a sign-off line is appended to ## Log " + @@ -216,6 +231,8 @@ You are a strictly read-only reviewer signing off a coding goal. You cannot exec by reading (read/grep/find/ls). Never re-run the work or its verify command -- it may be a 10-hour job; the agent must bring you its saved output. Your job is evidence discipline, checked in order: +0. Task fidelity? Read User-visible result and User voice first. Reject if this goal contradicts, + replaces, or defers the requested artifact or outcome. Agent-inferred scope is not authority. 1. Anything here? An empty or placeholder evidence: list -> reject: "there's nothing here -- fill the evidence and try again." 2. Quoted and attributed? Each item needs a source (file path / command) plus a verbatim quote of @@ -247,8 +264,8 @@ The working agent claims this goal is complete: goal: ${p.goal} Below is the full plan file (${p.planPath}). Find that goal in it (tolerate small wording drift; if -you cannot find a matching goal at all, reject and say so). Read its discriminator, subtle failure -modes, verify command, and evidence list from the file itself. +you cannot find a matching goal at all, reject and say so). Read User-visible result and User voice +first, then its discriminator, subtle failure modes, verify command, and evidence list. --- plan file --- ${p.plan} diff --git a/test/prompts.test.ts b/test/prompts.test.ts index ea91eaf..da62fbd 100644 --- a/test/prompts.test.ts +++ b/test/prompts.test.ts @@ -1,5 +1,5 @@ import { describe, expect, it } from "vitest"; -import { planDrafting, planningState } from "../src/prompts.js"; +import { judgeSystem, planDrafting, planningState, reminder, resync } from "../src/prompts.js"; describe("planning prompt", () => { it("requires fact finding or a focused question before a goal", () => { @@ -17,4 +17,14 @@ describe("planning prompt", () => { expect(planningState(".pi/plan/test.md")).toContain("choice that needs their approval"); expect(planningState(".pi/plan/test.md")).toContain("self-contained round with relevant context and a recommendation"); }); + + it("anchors work and sign-off to the user-visible result", () => { + expect(planDrafting).toContain("## User-visible result"); + expect(planDrafting).toContain("Take it from the original request, not from your implementation plan"); + expect(planDrafting).toContain("Future work may not defer any artifact or action named there"); + expect(reminder("plan", ".pi/plan/test.md")).toContain("latest message outranks this plan"); + expect(resync("plan", ".pi/plan/test.md", "Compacted.")).toContain("amend the plan rather than preserving an obsolete decision"); + expect(judgeSystem).toContain("Task fidelity?"); + expect(judgeSystem).toContain("Agent-inferred scope is not authority"); + }); });