Consolidate supervision and remember models per role

Bundle Intercom/VCC with the internal supervisor. Add default alignment questions and conversational plan review. Cover model recovery, cancellation and packed single-package operation.
This commit is contained in:
wassname2 committed 2026-09-08 08:45:26 +08:00
1 parent 8e44738773
commit d5729ac106
36 files changed
+8119 -305

No files matched your search

@@ -0,0 +1,38 @@
# Single-package and role-model validation
2026-09-07. The parent accepted the implementation after independent review and a targeted recheck. Changes are in the local pi-goals feature worktree; this is not an npm release.
## Delivered
- One pi-goals package: internal supervisor, bundled Intercom/VCC, and one package-root supervisor launch. Herdr remains the terminal host.
- Separate remembered planning, worker and supervisor provider/model choices. Files are under `getAgentDir()/pi-goals/`; the fresh evidence-judge override stays separate.
- At least three task-specific alignment questions by default. An explicit affirmative current-plan waiver skips them; negations and quoted feature names do not.
- Ready / Discuss / Edit / Cancel. Discuss and Escape preserve the draft and return to chat. `RequestPlanReview` reopens review when discussion is finished, including an unchanged draft. Only Ready starts work.
## Observed validation
[Saved full output](evidence/2026-09-07_single-package-final-validation.log) contains:
```text
Test Files 12 passed (12)
Tests 67 passed (67)
...
ℹ tests 118
ℹ pass 118
ℹ fail 0
ℹ skipped 0
...
SINGLE_PACKAGE_FINAL_VALIDATION_PASSED
```
`npm test` includes the packed/extracted production artifact running two real Pi sessions, the actual bundled Intercom broker and an offline fresh judge. It checks distinct actual role models and does not load a companion source checkout. The real-Pi conversational test covers Discuss and same-current-model recovery. Herdr is mocked in automated tests. Typecheck, lint, build and diff checks passed. Comparing common entries in the old/new lockfiles found no changed versions of existing locked packages.
The first review found four defects: stop blocked by model unavailability; activation before worker restoration and wrong recovery role; negated waiver matching; and same-model selection not triggering recovery. All were fixed with regressions. Parent inspection also caught a recovery await before cancellation ownership was captured; two more regressions cover clear/replacement during that await. The targeted review read current source and tests and returned `No issues found.` and `Merge verdict: OK`.
## Migration and remaining limits
The parent removed the two old companion entries from the user's Pi package list after the packed path passed. Only pi-goals remains registered for this workflow. Existing workers must reload before Ready: their old launch arguments still name the standalone extensions. The failed supervisor pane was observed at a shell with no Pi process; no duplicate recovery process was started during implementation.
Live Herdr recovery/navigation, the quality of questions from the user's chosen model, and token savings still need a human trial. The old historical tool-call-without-result issue can still block whole-plan `done`; `/goals clear` explicitly disconnects supervision and preserves history. This change did not weaken outstanding-work checks or alter old session transcripts.
<!-- Parent synthesis from observed commands, source inspection and review output, by Pi. -->
@@ -0,0 +1,157 @@
> @wassname2/pi-goals@0.2.2 test
> vitest run && npm run test:supervisor
RUN v4.1.9 /home/ubuntu/.pi/agent/worktrees/pi-goals-persistent-steward
Test Files 12 passed (12)
Tests 67 passed (67)
Start at 16:14:44
Duration 15.63s (transform 1.67s, setup 307ms, import 22.12s, tests 28.12s, environment 3ms)
> @wassname2/pi-goals@0.2.2 test:supervisor
> node --import tsx --test test/internal-supervisor/*.test.ts
✔ retries intercom registration when pi-intercom loads after pi-supervise (3.540888ms)
✔ a directive with no text is rejected, so the worker never sees undefined (0.288987ms)
✔ a directive from the paired supervisor becomes a real user message (27.396147ms)
✔ a directive to a busy worker interrupts, instead of waiting for the whole task (13.079025ms)
✔ a directive from an unpaired session is dropped (11.977008ms)
✔ a second pair takes over, and the first supervisor is told it lost the worker (21.508059ms)
✔ only the paired worker can end a run (6.314542ms)
✔ the programmatic pairing API waits for the worker acknowledgement (1.485485ms)
✔ the worker acknowledges a pair, so the supervisor knows it was heard (5.359587ms)
✔ a goal the supervisor inferred reaches the worker, which owns the view header (37.477099ms)
✔ the second view carries only what happened after the first (328.068703ms)
✔ a message addressed to a different session is ignored (11.834377ms)
✔ on settle the worker publishes a view built from the live branch (17.888259ms)
✔ the view is built from the live branch, not from every entry in the session (22.049192ms)
✔ an unpaired session publishes nothing on settle (0.498005ms)
✔ supervision never stops itself: no round limit at all (9.442441ms)
✔ goal, pairing and the steer count all survive a reload together (0.791695ms)
✔ a view that arrives while the supervisor is thinking is queued, not dropped (5.643173ms)
✔ the nudge repeats neither the instructions already sent nor the verdict rules (5.269154ms)
✔ a multi-line goal returns to supervisor context every fifth review and after compaction (31.103788ms)
✔ a one-line goal is not redundantly reinserted (26.640178ms)
✔ a check in and a worker that stopped ask for different things (12.287139ms)
✔ a loop still gets named after the supervisor compacts, from restored state (1.450856ms)
✔ a session that does not answer the roll call is not offered as a worker (501.077007ms)
✔ a child run stays out of the roll call, so it can never be picked (5.337167ms)
✔ a session already paired stays out of the roll call, and a free one answers (15.899876ms)
✔ /supervise look asks the worker for a fresh view, rather than the supervisor guessing (306.105484ms)
✔ let_it_run says the turn is over, so it is not called four times running (0.835308ms)
✔ a sign-off verdict is answered, not aborted, and a runaway is still cut (0.609635ms)
✔ every verdict result names the way to end the turn, steer included (0.449781ms)
✔ an old view is dropped from context once its verdict is in, and the verdict is kept (1.046937ms)
✔ a worker session never has its context rewritten (0.289351ms)
✔ a view that arrives mid-answer starts a fresh look (5.824464ms)
✔ a tool a worker cannot use never aborts its turn (0.339873ms)
✔ a resume onto a session that is gone drops the pairing and says so (6.297894ms)
✔ a resume onto a live worker keeps supervising, and takes the writers back off (5.29537ms)
✔ state written before recentSteers existed still loads (0.171283ms)
✔ done unpairs the worker, so it stops publishing views (320.167743ms)
✔ with no goal the supervisor cannot steer, it must ask the human (0.573651ms)
✔ set_goal binds an inferred goal, and steering then works (501.94885ms)
✔ a goal given at pair time still allows steering (0.665903ms)
✔ done is refused while the worker has an unanswered tool call (12.424782ms)
✔ done is allowed once nothing is outstanding (6.182792ms)
✔ steer refuses when the session is not supervising (0.325609ms)
✔ a reworded repeat of an earlier instruction is sent, and named back to the supervisor (0.510602ms)
✔ overlap scores rewording high and a different instruction low (0.136533ms)
✔ the view of the old worker cannot be used to judge the new one (6.045704ms)
✔ with one other session here, /supervise needs no target and the whole line is the goal (501.381162ms)
✔ naming the worker still works, and the rest of the line is the goal (0.554981ms)
✔ with two free sessions here, /supervise asks which one, and pairs with the choice (501.08053ms)
✔ a goal that is a path is read from the file, so it is not pasted every run (2.297652ms)
✔ a long goal is one short line above the picker, and reaches the worker whole (501.854523ms)
✔ a session that stayed quiet is still on the list, because 0 free is a dead end (501.687356ms)
✔ a cancelled picker pairs with nothing (500.579211ms)
✔ supervising takes the writing tools away, and stopping gives them back (501.255208ms)
✔ stopping gives back the writers without undoing another extension's tools (500.601415ms)
✔ a first word that names no session is refused, rather than folded into the goal (0.611801ms)
✔ a goal with spaces needs no target, and @name takes the rest of the line as the goal (501.893405ms)
✔ the brief starts no turn, so there is no answer before the first view (5.995751ms)
✔ /supervise goal changes the goal without breaking the pairing (0.70475ms)
✔ the footer says which side of a pairing this session is, and clears when it ends (506.432358ms)
✔ a session that is not supervising never sees the supervisor tools (5.830284ms)
✔ worker_view refuses when there is no worker, rather than implying a pairing (0.513982ms)
✔ the view names the worker's model and how full its context is (21.061667ms)
✔ supervising a second session is refused while the first is still paired (0.548942ms)
✔ the supervisor gets a look at a working worker every half hour, without being asked (926.124762ms)
✔ a human message in the worker session is not a reason to stand back (6.127037ms)
✔ letting a stopped worker run says plainly that the worker stays stopped (11.377477ms)
✔ a stopped worker is looked at again, so let_it_run cannot silence the pairing (922.993821ms)
✔ a worker that pairs at the prompt and never takes a turn is still watched (604.616558ms)
✔ a worker that reloads at the prompt starts watching itself again (604.845926ms)
✔ a timer look at a worker that has not moved is not sent, until it has been skipped three times (2425.474247ms)
✔ the worker counts reviews in a row where nothing changed (356.432776ms)
✔ an unacknowledged pair gives up, and a takeover cancels that timer (3.893969ms)
✔ duplicate standalone Intercom registries are diagnosed and cannot bootstrap a plan (1.232918ms)
✔ plan bootstrap compacts only the supervisor and pairing alone never starts a worker or a review (18.935035ms)
✔ goal decisions are correlated, preserve the pair across two goals, and cannot call overall done (5.389797ms)
✔ abort and stop cancel pending requests; late decisions cannot approve a replacement (3.265783ms)
✔ 50 actual model turns trigger one view, independent of the number of messages (3.457965ms)
✔ unknown background providers are not proof of quiescence (0.294658ms)
✔ stale plan content invalidates a pending goal review (3.01992ms)
✔ small forks skip compaction, but real compaction failure prevents pairing (2.471879ms)
✔ the hour timer and a coincident turn checkpoint produce a single view (5.559542ms)
✔ settled checks wait for tracked processes and subagents to finish (2.833087ms)
✔ bootstrap stop cannot resurrect a supervisor after compaction completes (2.15529ms)
✔ a restarted worker reconnects by exact saved session identity without a new supervisor (2.812666ms)
✔ unknown initial context must compact instead of taking the known-small shortcut (1.119486ms)
✔ null post-compaction usage cannot raise the next configured 100k checkpoint (2.180167ms)
✔ stopping a routine view during compaction invalidates its suspended continuation (1.977452ms)
✔ restart of a provisional bootstrap resumes compaction and pairing in the same saved session (2.18079ms)
✔ command preserves a stopped supervisor across reload (2.007763ms)
✔ done preserves a stopped supervisor across reload (2.456918ms)
✔ same-binding replay retains activation when the supervisor lost its acknowledgement (2.599438ms)
✔ model-unavailable stop validates binding, cancels pending reviews and ignores old directives (2.091306ms)
✔ a child process named pi is found by ps, and stops being found when it exits (367.036839ms)
✔ the check is a snapshot, so it cannot hold up the worker's settle (321.13146ms)
✔ a one-line goal stays whole while a multi-line goal has a locator (6.226689ms)
✔ a view carries only the turns the supervisor has not been sent (2.16041ms)
✔ the last two reasoning blocks stay in the narrative, and older ones drop out (1.078426ms)
✔ a compaction restarts the view, so no turn falls into the gap (0.455488ms)
✔ pi-vcc reports the files the worker wrote, and separates them from the ones it read (1.361829ms)
✔ progressKey is unchanged when a review produced no new file or commit (0.506944ms)
✔ progressKey still sees a new file past pi-vcc's ten path display cap (0.575315ms)
✔ a commit counts as progress, even when no file was written since (0.960805ms)
✔ outstandingWork finds tool calls that never got a result (1.43676ms)
✔ buildView reports a tool call with no result, so done can be refused (0.87924ms)
✔ the view says how many reviews in a row changed nothing, and says nothing at zero (0.692566ms)
✔ the view merges the worker's compaction summary with the turns after it (0.470565ms)
✔ a turn the compaction summary already covers is not sent twice (0.336722ms)
✔ pi-vcc's sections and its transcript land on the right sides of the split (2.943ms)
✔ the view does not tell the supervisor to use vcc_recall, a tool it does not have (0.291333ms)
✔ supervisor directives are not sent back as worker evidence (0.488463ms)
✔ bookkeeping tool calls are kept out of the transcript (0.390618ms)
✔ buildView reports the goal, status, and files without historical failures (0.42015ms)
✔ how long the worker has been quiet, measured from its own last entry (0.593754ms)
✔ buildView keeps the newest turns when it has to cut for the channel limit (26.316809ms)
✔ pi's own branch logic drops the abandoned fork, on a session file (1793.809954ms)
✔ a long goal cannot push the view past the broker limit (0.55492ms)
ℹ tests 118
ℹ suites 0
ℹ pass 118
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 20178.623102
> @wassname2/pi-goals@0.2.2 typecheck
> tsc -p tsconfig.build.json --noEmit
> @wassname2/pi-goals@0.2.2 lint
> biome check src/ test/
Checked 32 files in 93ms. No fixes applied.
> @wassname2/pi-goals@0.2.2 build
> tsc -p tsconfig.build.json
SINGLE_PACKAGE_FINAL_VALIDATION_PASSED
@@ -0,0 +1,138 @@
# One package, remembered role models, and conversational plan alignment
The user wants supervision inside pi-goals, with a separate remembered model choice for planning,
working and supervising. This replaces the three-extension installation described in the earlier
integration plan. Two follow-up requests require default alignment questions and a clear return to
normal chat from review; the parent supplied these approved additions during implementation.
## User-visible result
Install only pi-goals, discuss a plan and select Ready to open its real supervisor. Each role restores
its last explicitly selected provider/model. Planning asks at least three task-specific alignment
questions by default. Discuss returns the review menu to ordinary conversation without an editor or
an immediately recurring menu; Ready remains the sole handoff to work.
## User voice
- > did you do the part where the last model choice in plan ,vs worker vs supervisor mode is sticky?
- > ok implement. and just put it into pi-goals I guess... seems messy to have a sep package. one clean package
- > note I tried it in another window and it didn't grill me even one quesiton that is bad should be default to at least ask 3 questions to see how far apart in understanding user and agent are
- > and then I wans't sure what to press, edit? no. refine? that's a special prompt. how to go back to chat about plan and ask it to grill. should not have ot
## Goals
1. [x] goal: Install one package for planning and supervision
- scope: Move supervisor code and regression coverage into pi-goals. Bundle the existing Intercom transport through supported Pi packaging; no separate companion registration or source checkout required. Keep Herdr as the terminal host.
- failure modes: A source-checkout test hides missing packed dependencies; duplicate extension registration; lifecycle regressions during the move.
- discriminator: A packed pi-goals artifact alone supplies both real Pi sessions and completes the offline supervised-goal test. Full inherited supervisor regressions run from this repository.
- evidence: `src/internal/supervisor/` imports the implementation from `145c2cb081f85c08b0244c4a2c8a2d9aef8debda`; `THIRD_PARTY_NOTICES.md` and source attribution retain provenance/MIT notices. `package.json` bundles locked Intercom 0.10.0 and VCC 0.5.0; its manifest loads Intercom through `node_modules/`, while VCC remains a compiler rather than a loaded extension. Launch passes only the pi-goals package root. Duplicate registry diagnostics reject ambiguous plan bootstrap. `test/rpc-supervisor.test.ts` packs/extracts outside companion checkouts, runs actual Pi/Intercom and an offline fresh judge, and asserts production resources and no bundled Pi core peers. Only the worker's Herdr exec is adapted; the supervisor loads the untouched extracted manifest. Parent full validation passed 59 Vitest tests and 117 internal tests (116 inherited plus duplicate-registry coverage), zero skipped.
2. [x] goal: Restore the last model selected for each role
- scope: Remember provider/model separately for planning, worker and supervisor in user-scoped pi-goals preferences. First use inherits the current model. Preserve the separate evidence-judge override.
- failure modes: Automatic switching overwrites another role's preference; reload/fork assigns the wrong role; one process overwrites another role's choice; an unavailable model silently changes the saved preference.
- discriminator: Tests select distinct models in each role, enter/re-enter roles, reload and start another plan. Each role restores its own selection; missing models are reported without overwriting preferences.
- evidence: `src/role-models.ts` uses public `model_select` (`set`/`cycle`; ignores `restore` and guarded automatic `setModel` events). Separate atomic `planning-model.json`, `worker-model.json`, and `supervisor-model.json` files under `getAgentDir()/pi-goals` contain only provider/id. Failed authentication or lookup visibly pauses the role, preserving its saved choice. `test/role-models.test.ts` covers deterministic cycling, fresh processes and unavailable/unauthenticated models. `test/goals-flow.test.ts` covers Ready, cancellation during setModel, failed Ready, next plans, fresh-instance resume/reload and inherited supervisor isolation. `test/supervisor-integration.test.ts` proves planning model capture before worker restore, supervisor model before initial compaction, activation cancellation and judge override isolation. The packed RPC flow also verifies actual planning selection and distinct worker/supervisor/judge model IDs.
3. [x] goal: Ask useful alignment questions before the final plan
- scope: Default to at least three distinct task-specific questions in one chat round, about expected result, scope/constraints and success/failure criteria. Resolve technical facts read-only. Wait for answers and use them. Only an explicit current-objective no/skip-questions instruction waives the default.
- failure modes: Generic ritual questions; questions skipped merely because the agent thinks it understands; an old waiver leaking into a new plan.
- discriminator: Prompt and flow regressions require the default, persist the current-plan waiver across resync, reset it on a new plan, and keep the final-review menu closed until review is requested.
- evidence: `src/prompts.ts` defines the default and current-plan policy; `src/index.ts` persists `questionsWaived` per plan. Prompt and flow tests cover default/waiver/next-plan behavior. The deterministic real-Pi `test/rpc-review.test.ts` asks questions before the first review, accepts chat answers, and reaches Ready. This validates protocol/control flow, not semantic question quality from every live model; no semantic question-count framework was added.
4. [x] goal: Return to ordinary plan chat from review
- scope: Ready / Discuss / Edit / Cancel. Discuss preserves the draft and planning role, asks useful questions in chat, and supports multiple answer turns without another modal. Escape also preserves the draft and returns to chat. Edit remains direct editing; explicit Cancel discards the draft.
- failure modes: Refine-notes editor persists under a renamed button; unchanged drafts immediately reopen the menu; discussion is lost on reload; another approval gate starts work.
- discriminator: Discuss -> multiple chat turns without a menu -> completed discussion -> review again (including unchanged draft) -> exactly one Ready handoff. Reload retains discussion state.
- evidence: Planning-only `RequestPlanReview` is the unambiguous signal to offer review; it does not approve work. New plans and Discuss set persisted `reviewRequested: false`. `test/goals-flow.test.ts` proves unchanged-draft re-review and reload. `test/rpc-review.test.ts` uses real Pi select/chat events and asserts no Discuss editor or premature menu. Ready still exclusively starts work.
## Log
Final parent acceptance: [saved validation and review disposition](../../reviews/2026-09-07_single-package-role-models.md). All four review findings and the recovery-cancellation correction passed the targeted recheck (`No issues found.`, `Merge verdict: OK`). Final tests: 67 Vitest and 118 internal tests, no skips. The package-list migration is complete; existing workers need reload before another Ready attempt.
## Verification
Initial (pre-review-fix) parent unsandboxed checkpoint validation, read back from
`/tmp/pi-goals-single-package-parent-validation.log`:
```text
npm test
Test Files 12 passed (12)
Tests 59 passed (59)
internal node:test: tests 117, pass 117, fail 0, skipped 0
npm run typecheck: passed
npm run lint: Checked 32 files. No fixes applied.
npm run build: passed
git diff --check: passed
SINGLE_PACKAGE_PARENT_VALIDATION_PASSED
```
The full test includes production `npm pack` and the real broker/offline supervised flow, without an
optional integration flag or companion checkout. Worker-local validation independently passes 58
non-broker Vitest tests and all 117 internal regressions, plus typecheck/lint/build. Worker full
`npm test` fails only at Intercom startup: a minimal Unix `net.listen()` also returns `EPERM` in this
sandbox. The parent ran the unchanged enabled test outside that restriction and passed it. Do not
confuse the worker environment limitation with a skipped test or a product pass claim.
`tsconfig.build.json` maps only the four used source-only VCC API surfaces to narrow declarations,
because VCC 0.5.0 has upstream Pi-message/Intl type incompatibilities. Pi-goals source remains fully
typechecked/linted; tests execute the actual pinned VCC implementation. The inherited `node:test`
suite has its own package script and is not silently collected or skipped by Vitest.
Final production tarball, file listing, tracked-plus-new-source diff, source copies, status and
validation logs are saved under `/tmp/pi-goals-single-package-review/` for read-only review. No files
are staged. Fresh review and the targeted recheck are complete; parent acceptance is recorded above.
## Accepted review findings and narrow fixes
Parent accepted all four findings in `single-package-review-recovery.md`; the source was frozen
again for unsandboxed validation after this targeted pass:
1. **Stop while a model is unavailable:** the validated `stop` control-plane operation bypasses
model readiness, while binding/channel checks and peer notification remain. AbortSignal review
cancellation also works during the pause. Directives are ignored for paused/inactive or stopped
workers. New actual-module clear/off tests and the internal wrong-binding/abort/stop regression
prove that the old pair cannot revive work.
2. **Restore before activation and recover the right role:** Ready captures/attaches the planning
fork, persists a pending worker recovery target, restores the worker model, then activates and
hands off. Missing lookup/authentication returns to the review UI without activation or a
supervisor review turn. The persisted recovery target remains worker across reload; an explicit
selection writes worker-model.json, not planning-model.json. Both failure paths have supervised
integration regressions; successful recovery reuses the existing fork and hands off once.
3. **Affirmative waivers only:** a small clause recognizer accepts explicit affirmative current-plan
instructions including `no q's` and `skip q's`, but not `do not skip questions`, quoted feature
names, or embedded quoted clauses. No general NLP parser or semantic question-count gate was added.
4. **Already-current model recovery:** `/goals model current` explicitly authenticates and saves the
current model for the paused role. It does not depend on model_select, which Pi suppresses for
an unchanged model. It is never automatic and does nothing if no role is paused. Unit tests
cover failed authentication without preference replacement; real-Pi RPC proves that same-model
selection alone cannot recover, then the explicit command safely recovers and Ready starts once.
The parent identified one cancellation gap in the initial repair: the recovery restore await was
outside Ready's controller/version ownership. It is now inside the captured lifetime/controller and
plan version/hash checks, before any launch call. Deferred recovery -> clear and -> replacement
regressions prove zero additional startup/activation/handoff and preserve the new/null state.
Worker validation on the final fix source: 66 non-broker Vitest tests and 118 internal supervisor tests pass;
`npm run typecheck`, lint (32 files), build and diff checks pass. Worker full `npm test` still fails
only at actual Intercom startup under the same sandbox Unix-socket EPERM limitation. The enabled
packed test has not been skipped or replaced. Parent unsandboxed validation of the four-fix checkpoint passed 65/65 Vitest and 118/118 internal
tests, zero skipped, plus typecheck/lint/build/diff checks; see
`/tmp/pi-goals-single-package-fixed-validation.log`, ending `SINGLE_PACKAGE_FIX_VALIDATION_PASSED`.
That full run predates the final recovery-await cancellation correction and two extra local tests.
Final parent unsandboxed validation of that correction has now passed **67/67 Vitest tests in 12
files and 118/118 internal tests, zero skipped**, plus typecheck, lint (32 files), build and diffcheck.
The read-back log is `/tmp/pi-goals-single-package-final-validation.log`, ending
`SINGLE_PACKAGE_FINAL_VALIDATION_PASSED`. Parent's targeted review subsequently returned `No issues found.` and `Merge verdict: OK`; all named fixes and their immediate regressions were checked.
Fresh complete snapshots are at `/tmp/pi-goals-single-package-fix-review/`, including a fix-only
delta against the previous review snapshot as well as the full tracked-plus-new-source diff. The
normal three-role/Discuss/RequestPlanReview flow is unchanged, and no extra normal-path approval
gate, IPC/package redesign, global setting edit, or legacy orphaned-write fix was made.
## Boundaries and remaining limits
- Sole writer in the goals worktree on `feature/persistent-steward`, based on `8e44738`; the supervisor source checkout was read-only. No commit/push/publish or global settings/auth edits.
- Only pi-goals is loaded after migration. Parent reports that the old standalone supervisor/Intercom global registrations have now been removed; existing workers remain instructed to wait for validated reload/retry guidance before Ready. Old worker code can otherwise launch old `-e` companion arguments. No panes were controlled here.
- Herdr remains the supported host. Automated tests mock its exec adapter; live navigation/rendering/recovery and measured token savings are not established by this change.
- Native fork/compaction, SUPERVISOR.md precedence, incremental VCC memory, correlated review/cancellation/recovery and default-on policy remain. The fresh evidence judge stays separate.
- The previous live trial's apparently stale unresolved-write `done` failure remains pre-existing and undiagnosed, as recorded in `docs/reviews/2026-09-07_supervisor-validation.md`. Outstanding-work safeguards were not weakened and cleanup is not claimed fixed.
- Role files avoid cross-role lost updates; concurrent explicit choices within the same role are intentionally last-write-wins. No thinking-level preferences or credentials are stored.
<!-- Implementation and parent-observed validation by Pi. -->