Plan: Harden CLI Test Runtime Without Weakening Filesystem Contracts
Plan: Harden CLI Test Runtime Without Weakening Filesystem Contracts
Plan State
- Status: Implemented and validated. Mandatory review blockers are resolved;
the scoped operation now prepares non-mutating root authority before inspection
and commits a distinct absent root only after inspection succeeds. One known
Node filesystem limitation is parked and stated without a fail-closed claim.
T2's bounded CI experiment remains reverted after unsafe GitHub evidence. T4
retains GREEN evidence but did not meet its mandated RED-first contract. Both
manual rereviews after the latest fixes are clean. Structured autoreview on
d58f638dreturned one P1 and one P2; the third structured autoreview on head911a73e1was CLEAN with no actionable findings and rated the overall patch correct at0.91. Docs ingest and PR preparation are complete. Delivery has no in-scope blocker; only the explicitly excluded routed Wiki projection remains external. - Scope: Ordinary CLI CI scheduling plus a scoped scaffold-operation seam for
src/scaffold/run.test.tsandsrc/scaffold/stage.test.ts. - Execution mode: One branch and one PR. T1 uses one profiling worker. Parallel implementation is allowed only for T2 and T3 after T1 because their write scopes are disjoint. T5 and T6 each return to one scoped worker.
- Base intent: Create an isolated clean worktree from fresh
origin/main. Do not implement in the current detached worktree or absorb its scaffold drift. - Backlog: User approved completion without backlog sync. The active Linear connector cannot access the Devpunks workspace. IP-319/IP-331 are related completed authorities, not reopened owners.
- Scaffold remediation: Explicitly outside this plan.
Goal
Reduce ordinary CLI test feedback time while preserving every public behavior and every real filesystem security contract. Prove explicitly bounded ordinary CI on GitHub, then stop orchestration and policy tests from repeatedly traversing and materializing the full bundled scaffold through a real host filesystem.
Initial Situation
Commits 129e835a and 113f570b already established the safe update-suite
topology:
- 92 update cases remain four physical wrappers of 23 sequential cases.
- CI runs those wrappers in four exact jobs.
- Local release runs the wrappers with four workers, then ordinary tests with two workers.
- Same-wrapper concurrency and generic two-way Vitest sharding are superseded.
The latest main run used the same ordinary topology:
| Evidence | Result |
|---|---|
| Ordinary job | 79 files, 1,058 tests passed |
| Vitest duration | 586.80s |
| Test execution | 547.21s |
| Collection | 23.31s |
stage.test.ts | 166.749s |
public-output-contract.test.ts | 119.652s |
run.test.ts | 117.779s |
public-context-contract.test.ts | 48.438s |
sync-subagents.test.ts | 43.509s |
The five hot files account for most execution time. This plan attacks
stage.test.ts and run.test.ts, approximately 284.5s in the latest sample.
The ordinary CI command remains serialized because apps/cli/vitest.config.ts
sets fileParallelism: false. The release runner already proves one local
two-worker pass, but GitHub stability is unverified.
The latest overall workflow is red from a pre-existing stale wiki projection. All CLI jobs passed. That docs/scaffold divergence is not a CLI runtime failure and is excluded from this plan.
Accepted Authority
This plan continues the implemented IP-319/IP-331 contract:
- one comprehensive suite; no weaker fast lane;
- real domain policy with deterministic substitutes only at explicit system seams;
- one focused real-adapter contract for every deterministic adapter family;
- mutable filesystem and process state isolated per test scope;
- benchmark-tuned concurrency with comparable before/after evidence;
- no permanent runtime, worker-count, coverage, or release threshold.
Resolved Decision Ledger
| Decision | Resolution | Reason |
|---|---|---|
| Plan boundary | Scheduling experiment, profiling, scoped scaffold-operation seam, bounded run/stage migration, convergence evidence | Targets the dominant filesystem slice without a repository-wide rewrite |
| Filesystem substitute | High-level scoped-operation port with real and recording in-memory adapters | Keeps the fake truthful and the interface deep |
| Low-level filesystem mock | Rejected | A broad fake would encode false POSIX, inode, and containment semantics |
Global node:fs mock | Rejected | Static imports and Vitest mock lifecycle make it brittle and implementation-coupled |
| Real filesystem coverage | Retain bounded OS/security contracts | Links, modes, identity, confinement, atomic replacement, rollback, and races require the host filesystem |
| Performance policy | Three comparable samples per mode; threshold-free evidence | Matches IP-319/IP-331 and avoids arbitrary release policy |
| Backlog sync | Skip with user approval | Devpunks Linear workspace is unavailable to the active connector |
| Scaffold remediation | Excluded | It contains local-edited assets and is a separate authorization boundary |
These rows capture all five annotated approvals: bounded scope; the high-level operation seam; threshold-free measurement; user-approved backlog skip; and no scaffold-remediation expansion.
Assumptions And Constraints
- The 645-file bundled baseline remains production authority; tests may replace repeated orchestration traversals, not baseline correctness or generated-byte proof.
- Existing public CLI results, typed failures, rollback semantics, update case inventory, and release topology cannot change.
- Worker count is benchmark-tuned evidence. The contract requires a positive explicit bound but cannot freeze a numeric value as policy.
- Implementation begins from a fresh clean
origin/mainworktree and must use scoped workers under the repository agent policy. - Scaffold remediation and the stale wiki projection remain separate work.
Research Used
/tmp/harness-cli-test-performance-handoff-20260805.md: consolidated current timing, concurrency, filesystem-cost, and untried-approach research.- Current repository source, workflow, tests, prior IP-319 plan, and completed IP-319/IP-331 acceptance constraints.
- Readonly parallel investigations of current scheduling safety, hot filesystem ownership, alternative seam designs, and task ordering.
- No external web claim is needed for the selected design; repository behavior and measured CI evidence are authoritative. External validation is the planned GitHub-run comparison in T2/T5.
Scope
Included
- Complete per-test and per-file attribution for the two hot scaffold files.
- Serial versus bounded-worker ordinary-suite evidence, beginning at two.
- An explicit
ScopedScaffoldOperation-style interface withinspect,plan, andmaterializeresponsibilities. - A production adapter that composes existing Node/confined filesystem, planning, preflight, project-settings, and materialization behavior.
- A recording in-memory adapter for caller orchestration, exact-plan identity, failure mapping, rollback decisions, and result shaping.
- Vertical migration of
runScaffoldfollowed byrunStageScaffold. - An assertion-ownership ledger for every test moved away from a real filesystem.
- Repeated local and GitHub evidence with exact update exclusions.
- Canonical runbook correction if the bounded scheduling change is retained.
Excluded
public-output-contract.test.tsreduction.public-context-contract.test.tsrestructuring.sync-subagents.test.tscore extraction or fixture reuse.- Dependency-lockfile policy extraction.
- Update-wrapper case migration or same-wrapper concurrency.
- Broad
memfsadoption or a repository-wide filesystem abstraction. - Test deletion, assertion weakening, timeout increases, or a separate fast suite.
- Scaffold baseline remediation, generated skill/hook updates, lint-config reconciliation, or the stale wiki projection repair.
Codebase Findings
runScaffoldalready consumesScaffoldCapabilities, but context/output planning still reads concrete paths before the injected write boundary.runStageScaffolddirectly owns Node reads, root/session binding, planning, preflight, materialization, settings rollback, and cleanup.features/scaffold-state/lifecycle.tsalready expresses in-memorygather -> compile -> applyandobserve -> compare/reconcilesequencing.ObserveScaffoldStateandScaffoldReconciliationApplicatorsdemonstrate typed observation and mutation ports.features/project-settingsalready uses the desired narrow service pattern: one filesystem-facing feature seam, an in-memory adapter, and focused real temporary-filesystem proof.- The generic runtime
Filesystemonly proves basic byte operations and cannot truthfully model confined sessions, hardlinks, parent identity, modes, or atomic publication. It must not be expanded into the primary scaffold seam. - The hot files use unique temporary roots and do not directly mutate parent worker globals. The main two-worker risk is resource contention against bounded subprocess/test timeouts.
Design It Twice
Option A: Deep scoped scaffold operation — selected
Conceptual interface:
interface ScopedScaffoldOperation {
inspect(input: InspectInput): Effect<ScopedInspection, ScopedScaffoldFailure>;
plan(input: PlanInput): Effect<ScopedScaffoldPlan, ScopedScaffoldFailure>;
materialize(plan: ScopedScaffoldPlan): Effect<MaterializedScaffold, ScopedScaffoldFailure>;
}The exact names may follow existing Effect conventions during T3. The contract is semantic:
runInitScaffoldLifecycleand the existing scaffold-state compiler/planners remain the sole semantic lifecycle and plan authorities;- the scoped operation is the root-bound infrastructure port used by that lifecycle, not a second orchestration layer or a second plan model;
inspectsupplies the existing lifecycle's gather/observe input;plandelegates to the existing compiler/planner and includes preflight; it does not introduce another planning algorithm;materializedelegates the existing applicator/output behavior;- the exact plan that passes preflight is the plan materialized;
- failures are typed at the public operation boundary;
- commit state determines whether settings rollback is legal;
- the test adapter records operations and returns caller-supplied semantic results; it does not regenerate scaffold output.
This interface is deep: it hides root/session identity, traversal, preflight,
mutation ordering, and cleanup behind three caller-visible operations.
Once both callers migrate, remove the overlapping root/preflight/write
responsibilities from ScaffoldCapabilities; keep its external tool and
clipboard responsibilities. No old and new orchestration paths may remain in
parallel.
Option B: Low-level scoped filesystem — rejected
A stat/read/list/write/mkdir/remove/link/replace/snapshot interface appears easy to inject but exposes filesystem mechanics across stage, output, preflight, project settings, and cleanup. A correct fake would need to model OS behavior it cannot faithfully reproduce. It would enlarge the test surface while leaving orchestration coupled to filesystem steps.
Adapter boundary
- Real adapter: owns Node and confined-filesystem composition. It remains authoritative for generated bytes and OS/security behavior.
- Recording adapter: owns deterministic semantic inputs/results and an operation log. It proves caller behavior only.
- Shared semantic contract: run the same inspection result, plan identity, typed failure, commit-state, and materialized-result contract against the recording and real adapters.
- POSIX contract split: only the real adapter claims filesystem identity, link, mode, containment, atomicity, and race behavior. The shared semantic contract must not pretend the recording adapter implements POSIX.
Filesystem Contract Matrix
May move to the recording adapter
- pack/provider/wiki selection propagation;
- assume-yes and prompt-result handling;
- required-tool/bootstrap result propagation;
- inspection -> plan -> materialize ordering;
- exact plan object identity across preflight/materialization;
- typed failure-to-result mapping;
- settings rollback before commit and no rollback after commit;
- completion payload, baseline version, and CLI version shaping.
Must remain on a real temporary filesystem
- output-root creation and root replacement;
- symlinked parents, dangling links, literal mirror links, and path containment;
- hardlink-safe atomic replacement;
- file modes and permission preservation;
- stable root/parent identity and TOCTOU defenses;
- bounded regular-file reads and non-regular targets;
- directory backup, publication, cleanup, and interrupted repair;
- stale generated-file deletion and receipt/manifest rollback;
- generated projector subprocesses and generated TypeScript/package checks;
- one representative setup and one init idempotency/materialization smoke.
No test moves until its old assertion, new owning test, and retained real contract are recorded together.
Dependency Readiness
- Status: No Stack Required.
- Base: fresh
origin/mainin an isolated clean worktree. - Reason: This is one bounded follow-on slice with internal task dependencies. No upstream feature branch or stacked PR is required.
- External execution constraint: The current
hi check --jsonreports scaffold divergence. Implementation workers must not rely on the listed drifted skills, hooks, configs, or generated guidance. Remediation remains outside this plan.
Dependency Graph
T1
├── T2 ───────────────┐
└── T3 -> T4 ─────────┤
v
T5 -> T6Parallel Execution Waves
| Wave | Tasks | Start condition | Write-scope rule |
|---|---|---|---|
| 1 | T1 | Immediately | One profiling worker owns disposable diagnostics; all are removed before completion |
| 2 | T2, T3 | T1 complete | T2 owns workflow/root contract; T3 owns CLI feature/platform/run slice |
| 3 | T4 | T3 green | Stage slice and shared scoped-operation seam |
| 4 | T5 | T2 and T4 green | One integration worker owns all convergence edits; parent validates/reviews |
| 5 | T6 | T5 decision recorded | One docs/wiki worker owns plan and canonical docs edits; parent validates/reviews |
T1 uses exactly one profiling worker. T2 and T3 may then run in parallel with two workers. T3-T4 are sequential because they share the scoped-operation interface and hot test files. T5 uses exactly one integration worker; T6 uses exactly one docs/wiki worker. The parent owns cross-task validation and review, not implementation edits.
Tasks
T1: Establish comparable baseline and attribution
- depends_on: []
- location:
apps/cli/scripts/profile-scaffold-tests.mjs(temporary),apps/cli/src/scaffold/{run.test.ts,stage.test.ts}(temporary diagnostic hooks only),.reports/cli-test-runtime/**, thisPLAN.mdevidence fields - description: In an isolated clean worktree, record machine/runner state,
exact SHA, Bun/Node/Vitest versions, and three forced comparable runs of each
hot scaffold file plus the ordinary serial and bounded-worker commands,
beginning with two workers.
Capture complete per-test timing, file timing, subprocess count/duration, and
time spent in fixture construction, baseline traversal, planning, preflight,
and materialization. Add one disposable diagnostic script plus test-local
timed wrappers around those named phases and subprocess calls; remove both
before T1 completes. The script creates every output directory, invokes
Vitest's JSON reporter, and merges reporter and wrapper events. Do not add
production telemetry. Store each sample under
.reports/cli-test-runtime/<run-id>/with a summary containingrunId, exact SHA, runner, tool versions, mode, command, test file/name, status,durationMs, subprocess count/duration when available, phase timings, and temporary root. The script derivesrunIdas<serial|workers-N>-<sample-1..3>-<sha8>and fails rather than overwriting an existing sample. Use the same schema for local and GitHub data. - validation: Evidence distinguishes runner/import time from baseline traversal, host filesystem mutation, and subprocess time. Every run uses the exact update exclusions and identical test inventory. Scoped diff inspection proves the disposable profiler and test hooks are absent at task completion.
- status: Completed
- log: Measured six green samples on
63b9cbf9af6a9468dcee520810a60a9a9c7fbaa9using Bun1.3.5, Nodev22.22.0, and pinned Vitest3.2.4on macOSdarwin 27.0.0, Apple M3 Pro (12 logical CPUs, arm64, 36 GiB), AC power. The disposable profiler used the pinned Vitest executable,CI=true,TZ=UTC, the exact exclusionssrc/update/run.test-cases.test.tsandsrc/update/run.shard-*.test.ts, and either--maxWorkers=1or--fileParallelism --maxWorkers=2. One discarded pre-sample exposed a profiler-only error: aTMPDIRnested below the worktree changed two Git-discovery tests. Moving run-owned roots to the system temp directory restored the ordinary inventory; the invalid sample was excluded and removed. Gate passed: the two candidate assertions each invoke the full scaffold operation twice, while timed full operations account for 94.3%-96.1% of focused-file wall time and direct test subprocesses account for only 0.3%-1.3%. T3/T4 may proceed; the premise was not disproved. - files edited/created:
apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md. Temporary profiler/reporter files, test-local hooks, raw reports, debug log, and run-owned roots were removed before completion. - backlog_item_id: not_synced
- backlog_item_url: not_available
- relation_mode: user-approved-skip
- assigned_skills: [
autoreview,codebase-design,effect-backend-structure,effect,effect-recoverable-actions,improve-codebase-architecture,parallel-research,quality-types,simplify,swarm-planner,tdd,turborepo] - tdd_status: not_applicable
- tdd_target: Produce comparable baseline evidence without changing suite behavior.
- red_command: not_applicable
- expected_red_failure: not_applicable
- green_command:
CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=1 && CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=2 && CI=true TZ=UTC bun run --cwd apps/cli test -- --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' - reason_not_testable: Measurement establishes evidence and does not change observable product behavior.
- red_evidence:
- green_evidence:
CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=1and the same command with--worker-count=2produced six passing summaries. Every focused-run sample executed the same 28-test inventory, every focused-stage sample the same 49-test inventory, and every ordinary sample the same 80-file/1,062-test inventory; SHA-256 hashes were respectively3251e7cfd8b22d7311ab34891efb223b4abc1ca932189a4418fc77df4b73e8c9,854b3efb11c1b61ddbb46a25de2d82f932d9921341b49a20edfa2273e0640ef1, andfe906ce7e79f729d201a7b58f3e1c425bdfac9491a6c803559345140985b81e7across both modes. No timeout changed. - codebase_design_notes: Attribution must measure the owning semantic phases rather than private helper calls. Diagnostic hooks must remain test-only and removable. T3/T4 proceed only if the evidence identifies repeated full baseline traversal/materialization as a dominant cost for the assertions selected there. If attribution disproves that premise, stop and revise this plan instead of adding the operation seam as a performance change.
- review_mode: cli
- runtime_validation: required
- runtime_target: Local ordinary CLI suite and isolated hot-file runs on one fixed SHA.
- runtime_evidence: Serial wall means (min-max, population CV) were focused
run
52.938s(51.856-54.814s,2.52%), focused stage88.885s(79.949-106.033s,13.65%), and ordinary348.641s(326.043-391.707s,8.74%). Two-worker means were52.396s(51.949-52.676s,0.61%),80.129s(79.574-80.653s,0.55%), and170.452s(170.337-170.658s,0.09%): bounded ordinary execution was51.1%faster by the sample means while single-file controls were unchanged. Candidate means were stale-skill regeneration5.263sserial /5.165sbounded and exact planned baseline symlink5.548s/4.884s; source inspection confirms two complete scaffold calls in each. Focused-run samples recorded 29 full operations averaging50.848sserial /50.346sbounded (96.05%/96.09%of wall); focused-stage recorded 45 averaging84.062s/75.682s(94.55%/94.45%). Direct subprocess means were0.638s/0.635sfor focused run,0.299s/0.266sfor focused stage, and0.969s/0.987sordinary. Reporter-to-wall gaps were below2.1s; collection and import were sub-second in the focused modules. All 18 commands passed with identical inventories and no observable behavior difference. - runtime_cleanup: All six external run-owned temp roots were absent after
profiler
finallycleanup. The invalid sample,.reports, disposable profiler/reporter, test-local timing/debug hooks, and session debug log were moved to Trash.rgfound zero diagnostic markers; scopedgit diffforrun.test.ts,stage.test.ts, andapps/cli/scriptswas empty;git diff --checkpassed. The unrelated untrackedIMPLEMENTATION-NOTES.mdwas not modified.
T2: Prove explicit bounded ordinary CI
- depends_on: [T1]
- location:
.github/workflows/behavior-contract.yml,scripts/behavior-contract/root-suite.test.ts - description: Test-drive an ordinary-job-only change that passes
--fileParallelism --maxWorkers=\<N\>with the exact update registrar/wrapper exclusions. Begin measurement atN=2, but keep the behavior contract intentionally numeric-value-agnostic so later evidence can tune the positive bounded value. Preserve the serialized packagetestscript, the release runner, and all four exact update jobs. Do not raise timeouts or introduce a generic shard. - validation: The static behavior contract rejects an ordinary CLI command missing either bounded flag or either update exclusion, and still proves four exact update jobs. Local root contract is green before GitHub validation.
- status: Reverted after unsafe GitHub runtime evidence
- log: Added a value-agnostic root behavior contract before changing the
workflow. The focused RED failed because the ordinary
cli-testsjob lacked--fileParallelism. The workflow then gained only--fileParallelism --maxWorkers=2; package scripts, release runner, timeouts, four exact update jobs, and exclusions remain unchanged. Simplify removed the initial numeric value from the durable contract so it accepts any explicit positive bound while rejectingtest:parallel, generic shards, and missing exclusions. Local validation remained green, but GitHub attempt 1 produced one Vitest timeout and threespawnSync node ETIMEDOUTfailures inpublic-output-contract.test.ts; all four update jobs passed. Per the unsafe sample gate, the workflow and T2-specific contract were reverted byte-for- byte toorigin/main. Package/release/update topology and the scoped operation changes remain intact. - files edited/created: Net production diff: none. The temporary changes to
.github/workflows/behavior-contract.ymlandscripts/behavior-contract/root-suite.test.tswere reverted. - backlog_item_id: not_synced
- backlog_item_url: not_available
- relation_mode: user-approved-skip
- assigned_skills: [
autoreview,codebase-design,parallel-research,simplify,swarm-planner,tdd,turborepo] - tdd_status: required
- tdd_target: The behavior contract rejects the current serial ordinary CLI workflow because it lacks file parallelism with an explicit positive bounded worker value while retaining exact update exclusions.
- red_command:
bun test ./scripts/behavior-contract/root-suite.test.ts -t "delegates ordinary CLI tests to bounded file workers" - expected_red_failure: The workflow command does not contain
--fileParallelismand a positive--maxWorkers=\<N\>, so the new scheduling assertion fails. - green_command:
bun test ./scripts/behavior-contract/root-suite.test.ts && CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' - reason_not_testable:
- red_evidence: Before the workflow edit,
bun test ./scripts/behavior-contract/root-suite.test.ts -t "delegates ordinary CLI tests to bounded file workers"exited 1 atExpected to contain: "--fileParallelism". - green_evidence: The focused contract passed 1/1; the full root contract
passed 11 tests and 286 assertions.
CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts'passed 80 files and 1,062 tests in 170.40s.git diff --checkpassed. - codebase_design_notes: Scheduling stays in the CI adapter. Package and release semantics remain unchanged; no new root or Turbo task is introduced. The retained numeric value is evidence, not a permanent policy threshold.
- review_mode: cli
- runtime_validation: required
- runtime_target: GitHub Actions ordinary
cli-testsjob on one fixed SHA. - runtime_evidence: Local implementation evidence was green, including a
committed-SHA pair at 323.30s serial versus 168.18s bounded with identical
82-file/1,068-test inventory. GitHub run
31054456710, attempt 1, head9fe7cb57f6c1f14b47ba97223fd9d37e38042f11, falsified the safety gate: ordinary job92468806556passed 81 files and 1,066 tests but failed four of 1,070 Linux tests after 507.45s (one 30s Vitest timeout and threespawnSync node ETIMEDOUTfailures). Update jobs92468806561,92468806566,92468806557, and92468806579each passed 23/23. Further bounded attempts were unnecessary because the plan required every sample to be green. Three corrected-SHA serialized workflow attempts subsequently passed, as recorded under T5. - runtime_cleanup: No external fixtures remain. The exact diagnostic run is retained as evidence; no GitHub run was targeted by SHA or deleted.
T3: Introduce the scoped operation through runScaffold
- depends_on: [T1]
- location:
apps/cli/src/features/scaffold-operation/{model.ts,index.ts,service.test.ts},apps/cli/src/platform/scoped-scaffold-operation.ts,apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts},apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts,apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts} - description: Implement the first vertical slice of the deep scoped
operation.
runScaffoldinvokes one injected operation that delegates inspection to the existing lifecycle gather/observe path, delegates planning and preflight to the existing compiler/planners, materializes that exact plan through the existing output path, and returns the existing public result without the caller reading a host repository. Preflight is internal to the operation'splanphase; the caller must not perform a second preflight. Add the real adapter only for behavior required by this slice and a recording in-memory adapter for caller tests. Run one shared semantic contract against both adapters. In this task, move the run-side stale regenerated-skill orchestration assertion to the recording adapter while retaining one bounded real stale-file deletion/containment witness. Record every moved and retained assertion in the task log. Keep planning algorithms and generated bytes authoritative in existing production modules. - validation: A test uses nonexistent virtual roots and proves operation order, exact plan identity, result equality, and zero host fixture creation. The shared contract proves the same inspection result, plan identity, typed failure, and public result semantics for real and recording adapters. Existing representative real stale-file deletion/containment stays green; no test count substitutes for assertion ownership.
- status: Completed
- log: Added the three-operation Effect service, real adapter, recording
adapter, shared semantic contract, and first
runScaffoldvertical slice. Review hardened both adapters to use fresh inspection-bound, one-shot plan capabilities consumed before any mutation; copied, foreign, successful replay, and failed replay are rejected consistently. Stable root/baseline/ shape authority is derived from the owned inspection rather than repeated by callers, and repository/validation failures remain in the typed error channel. The stale-skill candidate now splits orchestration from the OS witness: two virtual recording-adapter reruns prove exact operation order, unique plan identity, and public result shaping without host access; one bounded real projection proves stale deletion, containment via an unchanged outside sentinel, and absence from the managed result. Complete bundled projections in that candidate fell from two to one; observed focused time fell from about4.93sto2.72s. Structured autoreview then raised a P2: the opaque plan did not yet own enough prepared state to guarantee that apply consumed exactly what preflight approved. The corrected rich plan owns the compiled desired input and state, explicit reset directories, the static public result including its deep-owned baseline summary, and an opaque one-shot applier. Preflight and apply consume that same plan. Apply performs no source traversal or replanning. Focused proof removes source bytes after planning, exercises stale reset directories, mutates the original summary, and attempts replay; materialization still uses the frozen plan and replay is rejected. Both subsequent manual rereviews were clean. The later structured review and its stage/settings remediation are recorded under T4. - files edited/created:
apps/cli/src/features/scaffold-operation/{index.ts,model.ts,service.test.ts},apps/cli/src/platform/scoped-scaffold-operation.ts,apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts},apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts,apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts} - backlog_item_id: not_synced
- backlog_item_url: not_available
- relation_mode: user-approved-skip
- assigned_skills: [
autoreview,codebase-design,effect-backend-structure,effect,effect-recoverable-actions,improve-codebase-architecture,parallel-research,quality-types,simplify,swarm-planner,tdd,turborepo] - tdd_status: required
- tdd_target:
runScaffoldpreflights and materializes one injected plan on virtual roots without reading the host repository. - red_command:
bun run --cwd apps/cli test -- src/scaffold/run.test.ts -t "preflights and materializes the exact injected plan without reading the host repository" - expected_red_failure: Current
runScaffoldperforms concrete context and output planning before the injected capability boundary, so nonexistent virtual roots cannot complete and exact injected plan identity is unavailable. - green_command:
bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/run.test.ts -t "scoped operation|exact injected plan|shared semantic contract|rerun|symlink" && bun run --cwd apps/cli check-types - reason_not_testable:
- red_evidence: Before
runScaffoldconsumed the new operation, the exact virtual-root test failed 1/1 withRepositoryScanErrorfor/virtual/scaffold-input, proving the old caller still inspected the host. - green_evidence: The exact T3 suite passed 3 files and 32/32 tests in
50.03s;run.test.tspassed 29/29 in46.56s; the real shared contract passed in2.59s;check-types, scoped oxlint, scoped oxfmt, andgit diff --checkpassed. Assertion ownership is explicit: shared adapters own inspection/plan identity and one-shot lifecycle; virtual caller tests own rerun orchestration/result shaping; the real temporary filesystem owns stale deletion and containment. The historical post-binding adapter contract passed 10/10 in15.95s. After the exact-plan correction, the adapter contract passed 11/11, output passed 12/12, and their combined run passed 23/23. The added proof covers source removal after planning, stale reset-directory application, deep-owned summary immutability, same-plan preflight/apply, and replay rejection without apply-time traversal or replanning. - codebase_design_notes: Keep the public interface at three semantic
operations. Do not expose stat/read/write calls, confined sessions, inode
identities, or the recording adapter through production callers. The
recording adapter is a caller test double, not a scaffold generator. The
operation composes the existing lifecycle/compiler/planner/applicator into a
rich opaque plan. The plan retains compiled desired input/state, explicit
reset directories, a static deep-owned public result, and a one-shot applier;
it does not duplicate planning semantics or coexist with a second root/
preflight/write route in
ScaffoldCapabilitiesafter both callers migrate. - review_mode: cli
- runtime_validation: not_required
- runtime_target: not_applicable
- runtime_evidence: not_applicable
- runtime_cleanup: not_applicable
T4: Route staged initialization through the same scoped operation
- depends_on: [T3]
- location:
apps/cli/src/features/scaffold-operation/**,apps/cli/src/platform/scoped-scaffold-operation.ts,apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts},apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts,apps/cli/src/scaffold/{output.ts,stage.ts,stage.test.ts} - description: Migrate one vertical
runStageScaffoldpath from inspection through materialization. Proveinspect -> plan -> materializeordering, exact plan identity, rollback only when materialization fails before commit, and no rollback when completion rendering fails after committed materialization. Preserve typed failures and existing public stage results. In this task, move the stage-side exact projection rerun and planned baseline symlink rerun orchestration assertions to the recording adapter while retaining bounded real mirror/idempotency and symlink witnesses. Record every moved and retained assertion in the task log. - validation: The RED tests run against a virtual root without
mkdtempSyncand prove exact operation order/identity plus both sides of the commit boundary: rollback before commit and preservation after a downstream failure following commit. The shared real/recording contract covers commit state and typed failure translation. Focused real settings rollback, exact mirror/idempotency, and planned symlink behavior remain green. - status: Completed; mandatory review blockers resolved, GREEN evidence retained, TDD status partial/not-met
- log: Routed
runStageScaffoldthrough the required scoped operation and installed the real service in production/test composition; removed the optionalrunScaffoldfallback and the old caller-owned inspect/root-bind/ preflight/write route. The stage adapter owns settings rollback until its materialization commits; a virtual caller test proves rollback before commit and preservation after downstream presentation failure. The shared semantic contract now runs stage variants against both recording and production adapters and proves owned inspection, exact one-shot plans, committed result, typed pre-commit failure, and replay rejection after success or failure. Repository diagnostics remain unchanged at the scanner authority; the stage boundary preserves its documentedCliDetectionErrorwrapper with concrete typed errors. Assertion ownership was split without losing OS claims: the recording adapter owns both named two-rerun sequences and exact/fresh plan identity; one consolidated real two-operation rerun retains all exact mirror readlinks plus the planned skill symlink, reducing the former four complete projections to two; a separate real mismatched-symlink witness retains containment. Mandatory review then required stable pre-inspection root authority, caller-input snapshots, concrete adapter-owned stage plan state, selected-baseline preservation, separation of publicScaffoldCapabilitiesfrom the internal filesystem adapter, and a stage orchestration test with no host fixture. Final review identified and resolved a deletion/rollback race: the adapter prepares an existing root session or nearest-existing-ancestor authority before inspection, commits a distinct absent root only after inspection succeeds, and performs no mutation when inspection fails. This removes the newly introduced foreign-root deletion path while preserving the prior filesystem security contracts. Structured autoreview ond58f638dthen returned P1 because stage replanned output and read live stage-skill source after planning, plus P2 because scoped settings initialization read a live baseline version. The fix makes stage own and consume one rich output plan. ItsadditionalSkillSelectioncaptures selected-first, bundled-fallback stage-skill bytes, aliases, and reset roots; those exact reset targets are preflighted, and the old post-commit skill copy is removed. Settings initialization and baseline metadata are also planned, and scoped settings consumes thescaffoldOutputbaseline rather than live caller state. Both final manual rereviews were clean. The third structured autoreview on head911a73e1was CLEAN with no actionable findings and an overall patch- correctness score of0.91. It confirmed retained plans and one-shot lifecycle, root authority, stage-skill snapshotting, settings metadata, rollback boundary, and serialized CI remain consistent. - files edited/created:
apps/cli/src/index.ts,apps/cli/src/platform/feature-application-operations.ts,apps/cli/src/platform/scaffold-capabilities.ts,apps/cli/src/platform/scoped-scaffold-operation-adapter.ts,apps/cli/src/features/scaffold-operation/{model.ts,index.ts},apps/cli/src/platform/scoped-scaffold-operation.ts,apps/cli/src/scaffold/capabilities.ts,apps/cli/src/testing/{layers.ts,scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts},apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts,apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts,stage.ts,stage.test.ts,confined-filesystem.ts,confined-filesystem.test.ts} - backlog_item_id: not_synced
- backlog_item_url: not_available
- relation_mode: user-approved-skip
- assigned_skills: [
autoreview,codebase-design,effect-backend-structure,effect,effect-recoverable-actions,improve-codebase-architecture,parallel-research,quality-types,simplify,swarm-planner,tdd,turborepo] - tdd_status: partial — mandated RED not met; GREEN retained
- tdd_target: One injected scoped operation owns staged inspection through commit, and rollback follows commit state rather than downstream rendering.
- red_command:
bun run --cwd apps/cli test -- src/scaffold/stage.test.ts -t "rolls back before commit and preserves committed state after downstream failure" - expected_red_failure: Current
runStageScaffolddirectly creates a confined session and calls concrete planning/materialization code; no injected operation log or commit-aware rollback seam exists. - green_command:
bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/stage.test.ts -t "scoped scaffold operation|rollback|committed state|stale|idempotent|shared semantic contract" && bun run --cwd apps/cli check-types - reason_not_testable: not_applicable; the behavior was testable and the mandated RED was missed
- red_evidence: Not met. The exact mandated target did not exist before T4;
Vitest collected
stage.test.tsand reported 49/49 skipped under the missing name. A missing target and skipped tests are not RED evidence. During convergence, the repository typed-failure contract caught an invalid scanner-cause change; that independent failure was removed and the existing diagnostic contract restored, but it does not satisfy the mandated T4 RED. Later review remediation did produce two independent public REDs without repairing that original TDD miss: after planning, removing the selected stage skill source made the old apply fail withCould not locate baseline asset; the scoped-settings test expected baseline version3.1.3but observed the mutated-after-plan value. The stage metadata regression became GREEN as part of the first fix and is not falsely claimed as RED evidence. - green_evidence: The exact commit-boundary target passed 1/1. Full
stage.test.tspassed 49/49 in72.16s; typed-failure, service, shared adapters, and full run compatibility passed 36/36 in53.06s; the shared adapter contract passed 4/4 with the real stage case around2.8s. Historical T3 exact-plan evidence remains adapter 11/11, output 12/12, and combined 23/23. After the stage/settings correction, the adapter contract passed 13/13 in21.79s, output passed 12/12 in9.73s, and their combined run passed 25/25 in32.29s. Focused settings version and selected/bundled fallback coverage passed 3/3; symlink witnesses passed 2/2; full stage passed 49/49. Typecheck, check, and diff gates passed. - codebase_design_notes: The scoped operation owns phase ordering, plan
provenance, caller-input snapshots, and one-shot plan consumption. The stage
adapter owns concrete
StagePlanStateplus its commit/rollback boundary. The publicScaffoldCapabilitiesservice retains tool, clipboard, settings, and runtime concerns; the internal real adapter owns root-authority preparation, deferred absent-root binding, inspection, preflight, and filesystem materialization. The recording adapter proves orchestration semantics without becoming filesystem authority. Stage owns one rich output plan through preflight and apply.additionalSkillSelectionfreezes selected-first or bundled-fallback bytes, aliases, and reset roots; planned settings freezes initialization and baseline metadata, and scoped settings reads that plannedscaffoldOutputbaseline. Native Nodemkdirfollowed bylstatcannot atomically prove ownership against an exact replacement in between; that pre-observation window is explicitly unsupported rather than claimed fail-closed. Parent/root replacement before commit and later session swaps remain fail-closed. - review_mode: cli
- runtime_validation: not_required
- runtime_target: not_applicable
- runtime_evidence: not_applicable
- runtime_cleanup: not_applicable
T5: Converge scheduling, contracts, and runtime evidence
-
depends_on: [T2, T4]
-
location:
.github/workflows/behavior-contract.yml,scripts/behavior-contract/root-suite.test.ts,apps/cli/src/features/scaffold-operation/**,apps/cli/src/platform/scoped-scaffold-operation.ts,apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts},apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts,apps/cli/src/scaffold/{run.ts,run.test.ts,stage.ts,stage.test.ts,confined-filesystem.test.ts},apps/cli/src/update/run.shards.test.ts,apps/cli/scripts/run-release-tests.mjs, thisPLAN.md -
description: Run the complete compatibility and performance matrix. Repeat serial and bounded-worker ordinary commands on the final SHA, run all four update wrappers and
test:release, then collect three GitHub samples. Retain T2 only when every bounded-worker sample is green with identical inventory and its improvement is repeatable rather than mixed with timeout or variance regressions. Tune the positive bounded worker value when evidence supports it, updating both workflow and value-agnostic contract. If evidence is mixed or unsafe, revert only T2's workflow/root-contract change and keep the scoped-operation improvement. -
validation: Focused feature/adapter tests, real filesystem contracts, serial ordinary suite, bounded-worker ordinary suite, four update wrappers, release suite, CLI build/check/typecheck, root behavior contract, and diff checks are green. Known unrelated wiki projection failure is reported separately.
-
status: Completed; bounded scheduling reverted after unsafe GitHub evidence
-
log: The focused real/adapter matrix, all four independent update wrappers, the two-phase release runner, build/type/check/root contracts, and matched serial/bounded ordinary controls are green. The four wrappers retain exactly 23 unique cases each and 92 unique cases in union.
test:releasevisibly completed the four-worker update phase before its two-worker ordinary phase. The local serial and bounded commands ran the identical 82-file/ 1,068-test inventory; bounded wall time was165.75sversus serial312.63s(46.98%lower), while isolated run/stage controls remained comparable locally. GitHub then exposed subprocess starvation under two workers, so the final decision is to revert T2 scheduling without undoing the scoped-operation seam. The same run also exposed a forbidden runtime barrel;scaffold-operation/index.tsnow owns the Effect service contract directly, and Effect-v4 validation passes without a whitelist. Three corrected-SHA GitHub attempts then passed the serialized ordinary job and every update shard. Their workflow conclusion remained red only because the excluded wiki projection reports four stale paths. -
files edited/created: No new T5 implementation files. Validation covers all T2-T4 files listed above; evidence is recorded in this plan and
IMPLEMENTATION-NOTES.md. -
backlog_item_id: not_synced
-
backlog_item_url: not_available
-
relation_mode: user-approved-skip
-
assigned_skills: [
autoreview,codebase-design,effect-backend-structure,effect,effect-recoverable-actions,improve-codebase-architecture,parallel-research,quality-types,review-phase,simplify,swarm-planner,tdd,turborepo] -
tdd_status: not_applicable
-
tdd_target: Validate the already test-driven slices and make the evidence-based retain/revert decision for CI scheduling.
-
red_command: not_applicable
-
expected_red_failure: not_applicable
-
green_command:
bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/confined-filesystem.test.ts src/adapters/filesystem-adapter-contract.test.ts src/scaffold/stage.test.ts src/scaffold/run.test.ts && bun run --cwd apps/cli test -- --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && bun run --cwd apps/cli test src/update/run.shard-0.test.ts && bun run --cwd apps/cli test src/update/run.shard-1.test.ts && bun run --cwd apps/cli test src/update/run.shard-2.test.ts && bun run --cwd apps/cli test src/update/run.shard-3.test.ts && bun run --cwd apps/cli test:release && bun run --cwd apps/cli build && bun run --cwd apps/cli check-types && bun run --cwd apps/cli check && bun test ./scripts/behavior-contract/root-suite.test.ts && git diff --check -
reason_not_testable: Convergence task validates prior behavior-changing tasks and records runtime evidence; it introduces no new behavior.
-
red_evidence:
-
green_evidence: Focused matrix passed 6 files/113 tests in
129.74s(inventory1bc136663987a19cc031f12445891ae926627fb8cb53acbd1842ae20d605652c). Independent wrappers passed 23/23 each in202.06s,215.89s,181.44s, and211.60s; their union was 92/92 unique titles (inventory3c1f102a7fe17688ecd599a16415bfc5413b27ffe1f86c9b21ae9bc12da221a5).test:releasepassed in429.33s: update 92/92 in260.91s, then bounded ordinary 82 files/1,068 tests in163.62s. Build, check-types, CLI check, root 11/11 with 286 assertions, and diff checks passed. The post-revert root contract passes 10/10 with 280 assertions; Effect-v4 validation passes 94/94. -
codebase_design_notes: Scheduling and filesystem seams remain independent. A failed scheduling experiment must not undo the deeper operation boundary.
-
review_mode: cli
-
runtime_validation: required
-
runtime_target: Final local release suite plus GitHub ordinary/update jobs on one fixed SHA.
-
runtime_evidence: Pre-commit serial and bounded controls passed the same 82-file/1,068-test inventory hash
86fdba61c52dcdc0b380550123b82205cc8db7cd5fb3ed4004bc7090859b0a5d. Serial wall was312.63s; bounded wall was165.75s. Hot controls wererun.test.ts46.955s -> 48.364sandstage.test.ts71.676s -> 74.341s. A committed-SHA pair repeated the local result at323.30s -> 168.18s, but GitHub attempt 1 failed the bounded ordinary job as recorded under T2. The final post-revert local serial control passed 82 files and 1,068 tests in316.19swith the same inventory SHA-25686fdba61c52dcdc0b380550123b82205cc8db7cd5fb3ed4004bc7090859b0a5d.Diagnostic bounded run 31054456710, attempt 1, used head
9fe7cb57f6c1f14b47ba97223fd9d37e38042f11. Its ordinary job completed 81 files and 1,066 tests but failed four of 1,070 after507.45s. Update shards 0, 1, 2, and 3 each passed 23/23 in324.66s,344.27s,243.54s, and320.51s. The root job passed 92 of 94 Effect checks and failed one semantic pass-through-alias test; correcting that public-root ownership produced 94/94 locally and on every corrected attempt.Corrected run 31055566836 used head
4ad8a9b9b66b35bd4eaad84e7184888186b35d79and PR mergede14b84e. Attempts 1-3 passed the serialized ordinary inventory at 82 files/1,070 tests: attempt 1 in420.67s, attempt 2 in359.36s, and attempt 3 in595.58s. Every corrected attempt also passed four 23/23 update jobs: attempt 1 shard 0271.48s, shard 1345.83s, shard 2292.29s, shard 3321.43s; attempt 2 shard 0328.85s, shard 1351.91s, shard 2312.31s, shard 3293.55s; attempt 3 shard 0323.11s, shard 1334.49s, shard 2296.97s, and shard 3250.19s.Corrected root attempts 1, 2, and 3 completed all 94 Effect public-surface checks (93 passed and one focused configuration check intentionally skipped) and passed every root behavior- contract assertion. They failed only
@punks/wiki#check:contentfor four excluded projection paths:apps/wiki/content/docs/project/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md,apps/wiki/content/docs/project/specs/cli/cli-test-runtime-hardening/PLAN.md,apps/wiki/content/docs/project/specs/.source-projection.json, andapps/wiki/content/docs/project/runbooks/hi-cli-scaffolding.md.Final lifecycle run 31067183663 used head
3e81a31c.cli-testsand all four exact update jobs were green; API, backoffice, and wiki Vercel checks were green. Behavior Contract was red only at@punks/wiki#check:content, which requested exactly the excluded projection work: create routedcli-test-runtime-hardening/PLAN.mdandIMPLEMENTATION-NOTES.md, updateapps/wiki/content/docs/project/specs/.source-projection.json, and update the routedhi-cli-scaffoldingrunbook mirror. -
runtime_cleanup: External reports/logs were removed, run-owned roots are absent, and exact run/attempt/job provenance is durable above. No GitHub run was deleted.
T6: Record durable closeout without scaffold remediation
- depends_on: [T5]
- location:
apps/wiki/specs/cli/cli-test-runtime-hardening/{PLAN.md,IMPLEMENTATION-NOTES.md};docs/README.mdanddocs/runbooks/hi-cli-scaffolding.mdfor retained architecture and current test topology - description: Fill every task's RED/GREEN/runtime evidence and assertion
ownership ledger. Document the retained scoped-operation architecture in
docs/README.mdand the canonical runbook. Correct obsolete test-topology wording to the current serialized/four-wrapper arrangement without documenting the reverted bounded scheduling experiment. Do not update the scaffold-managed routed wiki mirror, generated skills/hooks, settings, lint specs, or Oxlint configs. Record the known scaffold/wiki divergence as external and keep CLI acceptance separate from the overall red workflow. - validation: Plan evidence is complete, canonical docs match retained CLI
topology, targeted formatting passes, and scoped
git diff --checkis clean. No excluded scaffold-managed file is modified. - status: Completed
- log: Filled final local and GitHub runtime evidence, recorded the T2
REVERT decision, corrected stale feature locations, and closed the assertion
and acceptance ledgers. Scheduling topology remained serialized, but the
retained scoped-operation architecture changed: canonical docs now record
pre-inspection root-authority preparation, deferred absent-root binding,
one-shot snapshotted plans, adapter-owned stage state,
commit/rollback ownership, and the real/recording contract split. They also
replace the stale two-shard description with the current serialized ordinary
job plus four exact wrapper jobs and mark
test:parallellocal-only. The four scaffold-managed wiki projection paths remain excluded and unmodified. - files edited/created:
apps/wiki/specs/cli/cli-test-runtime-hardening/{PLAN.md,IMPLEMENTATION-NOTES.md},docs/README.md,docs/runbooks/hi-cli-scaffolding.md - backlog_item_id: not_synced
- backlog_item_url: not_available
- relation_mode: user-approved-skip
- assigned_skills: [
autoreview,codebase-design,create-plan,create-spec,docs-onboarding,implement-spec,improve-codebase-architecture,parallel-research,review-phase,simplify,tdd,writing-beats,writing-fragments,writing-great-skills,writing-shape] - tdd_status: not_applicable
- tdd_target: Record non-runtime evidence and canonical maintainer guidance for the retained topology.
- red_command: not_applicable
- expected_red_failure: not_applicable
- green_command:
bunx oxfmt --check docs/README.md docs/runbooks/hi-cli-scaffolding.md apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md apps/wiki/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md && git diff --check -- docs/README.md docs/runbooks/hi-cli-scaffolding.md apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md apps/wiki/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md - reason_not_testable: Evidence/docs closeout does not change runtime behavior.
- red_evidence:
- green_evidence: The canonical root docs and both durable evidence files
pass targeted Oxfmt and scoped
git diff --check. The canonical docs record the retained scoped-operation architecture and current serialized/four-wrapper topology; reverted bounded-CI scheduling is not documented. Scaffold-managed mirror and projection files have no T6 diff. - codebase_design_notes: Closeout documents the selected deep boundary and retained real contract matrix; it does not expand the implementation seam.
- review_mode: cli
- runtime_validation: not_required
- runtime_target: not_applicable
- runtime_evidence: not_applicable
- runtime_cleanup: not_applicable
Testing Strategy
RED/GREEN tracer bullets
- T2 first proves the workflow contract rejects unbounded/serial ordinary CI.
- T3 proves
runScaffolduses one injected semantic plan on virtual roots, moves only its run-side orchestration assertions, and retains their real filesystem witnesses. - T4 proves staged commit/rollback semantics through the same public operation, moves only its stage-side orchestration assertions, and retains their real filesystem witnesses.
T2 and T3 recorded valid public RED before GREEN. T4 retained its later GREEN and assertion-migration evidence, but its mandated target was absent during the purported RED run; 49 skipped tests are not RED evidence. T4 is therefore partial/not-met for TDD even though its implemented behavior and validation are green.
Validation order
- Focused new operation tests.
- Focused
runandstagetests. - Real filesystem and adapter contracts.
- Serial ordinary CLI suite.
- Bounded-worker ordinary CLI suite using the evidence-selected positive value.
- Four exact update wrappers.
- Local
test:release. - CLI build, typecheck, check, behavior contract, and diff checks.
- Three GitHub samples on one fixed SHA.
Evidence rules
- Record exact SHA, runner, tool versions, command, test count, file count, duration components, and per-hot-file timings.
- Compare like with like; never compare a battery/low-power local run directly with a GitHub runner as causal proof.
- One green run proves feasibility, not stability.
- A moved test needs assertion-level replacement evidence, not only preserved aggregate counts.
- Never increase a timeout to make a scheduling experiment appear green.
Validation Gates
Gate A: Measurement readiness
- T1 has three comparable samples for both scheduling modes.
- Bottleneck attribution distinguishes baseline traversal, filesystem work, planning/materialization, and subprocess time.
- T3/T4 proceed only if attribution confirms repeated full-baseline traversal or materialization is a dominant cost for their candidate assertions; otherwise stop and revise the plan.
Gate B: Scoped operation integrity
- Recording adapter proves caller behavior without host paths.
- Real adapter remains authoritative for bytes and OS/security semantics.
- No generic filesystem or
node:fsmock becomes a production/test authority.
Gate C: Coverage preservation
- Assertion ownership ledger is complete.
- Retained real contract matrix is green.
- Four update wrappers still cover 92 unique cases, 23 each.
Gate D: Scheduling decision
- Three local and three GitHub bounded-worker samples are green with identical inventory.
- Median improvement is repeatable and not paired with timeout or variance regressions.
- Mixed evidence reverts T2 only; it does not invalidate completed T3-T4.
Gate E: Closeout
- All task evidence fields are filled.
- Canonical docs reflect retained behavior.
- Excluded scaffold drift remains untouched and explicitly reported.
Risks And Mitigations
| Risk | Mitigation |
|---|---|
| Recording adapter becomes a second scaffold generator | It accepts semantic caller-supplied results and records operations only |
| Deep interface becomes a pass-through wrapper | Keep three semantic operations and hide sessions, filesystem calls, and mutation ordering |
| Real security assertions move to the fake | Enforce the contract matrix and assertion-ownership ledger inside T3/T4 and at T5 review |
| Bounded workers exhaust subprocess timeouts | Repeat on GitHub; retain only with stable evidence; never raise timeouts |
| Interface work and scheduling become coupled | Separate T2 and T3 write scopes; T5 may revert T2 independently |
| Full baseline witnesses disappear | Retain representative setup, rerun/idempotency, generated-byte, and real adapter contracts |
| Current scaffold drift contaminates implementation | Use a clean worktree; exclude listed drifted assets; do not run scaffold remediation |
| Required scoped Effect skills remain drifted | Do not rely on their changed scaffold copies; execution must use verified source/current project patterns or stop and report the prerequisite |
| Overall workflow remains red from wiki projection | Judge CLI jobs separately and record the pre-existing docs/scaffold failure |
Skill Routing
grilling: resolved scope, seam, evidence, backlog, and remediation decisions.parallel-research: supplied current CI state, concurrency audit, hot-file attribution, and untried approaches.codebase-design: selected a deep scoped-operation seam over a broad filesystem-shaped interface.swarm-planner: produced dependency-aware waves with disjoint T2/T3 write scopes and sequential shared-interface work.tdd: attached public RED/GREEN tracer bullets to T2-T4. T2 and T3 met the contract; T4 retained its assertion migration, real witness, and GREEN proof but is partial/not-met because its mandated RED target was absent.turborepo: kept scheduling package-owned and avoided new root task logic or generic graph-bypassing parallelism.- Scoped
apps/cliskills are recorded on implementation tasks. Drifted scaffold copies are an execution prerequisite, not planning authority.
Backlog Sync
- Result: Skipped by explicit user decision.
- Reason: The active Linear connector cannot access the Devpunks workspace; the user requested completion of the plan without backlog mutation.
- Related completed authority: IP-319 and IP-331.
- Created/updated items: None.
- Future optional action: Create a new Harness Intelligence story related to IP-331; do not reopen or overwrite the completed IP-319 plan.
Unresolved Questions
No product questions remain. The known Node directory-publication limitation is
parked: mkdir followed by lstat cannot atomically prove ownership against an
exact intervening replacement without native no-replace publication.
Debug Assessment
The delivery debug assessment is closed from already captured CI runtime logs; no new instrumentation was added because the incident was proven before this phase.
| Hypothesis | Result | Existing evidence |
|---|---|---|
| A. The two-worker ordinary job caused subprocess starvation | CONFIRMED | GitHub attempt 1 failed four ordinary tests, including three spawnSync node ETIMEDOUT failures, while all four update jobs were green. |
| B. Update shards overlapped or repeated work | REJECTED | All four exact wrappers passed 23/23 and their union was unique. |
| C. Serialized ordinary execution contained a behavior regression | REJECTED | The local serialized inventory and all three corrected GitHub serialized inventories were green. |
| D. A generic timeout increase alone would resolve the incident | REJECTED / narrowed | Increasing timeouts does not address subprocess spawn starvation; the reverted serialized topology passed. |
Root cause was resource contention and subprocess starvation introduced by the bounded two-worker ordinary GitHub experiment. The fix reverted only T2 ordinary scheduling to serialized execution and retained the four exact update jobs. The verification evidence is the already recorded local and GitHub matrix.
Stop Condition
T1-T6 delivery work, final mandatory review corrections, focused validation,
clean structured autoreview, and debug assessment are complete. Bounded ordinary
GitHub scheduling remains rejected by the measured safety gate; serialized CI
and the scoped-operation seam remain the selected direction. T4's TDD status is
partial/not-met. The parked Node publication limitation is not represented as
fail-closed. Private/internal docs ingest updated docs/README.md,
docs/runbooks/hi-cli-scaffolding.md, and the existing scaffold-parent-identity
note at .agents/notes/2026-08-04-scaffold-parent-identity.md; routed Wiki
projection was intentionally skipped by scope. Stack status
showed an independent branch with no PR stack. stack sync --dry-run is
unsupported; the default stack sync preview reported the dev stack current,
and no apply was performed. PR #100 is ready for review with its body refreshed.
Run 31067183663 is green for all in-scope CLI/update and deployment checks;
Behavior Contract remains red only for the excluded Wiki projection requests.
Delivery is complete with no in-scope blocker. The projection drift is an
external excluded blocker, and no final branch head is asserted before the
parent commit.