Harness Intelligence Wiki
SpecsCLICli Test Runtime Hardening

Plan: Harden CLI Test Runtime Without Weakening Filesystem Contracts

Plan: Harden CLI Test Runtime Without Weakening Filesystem Contracts

Plan State

  • Status: Implemented and validated. Mandatory review blockers are resolved; the scoped operation now prepares non-mutating root authority before inspection and commits a distinct absent root only after inspection succeeds. One known Node filesystem limitation is parked and stated without a fail-closed claim. T2's bounded CI experiment remains reverted after unsafe GitHub evidence. T4 retains GREEN evidence but did not meet its mandated RED-first contract. Both manual rereviews after the latest fixes are clean. Structured autoreview on d58f638d returned one P1 and one P2; the third structured autoreview on head 911a73e1 was CLEAN with no actionable findings and rated the overall patch correct at 0.91. Docs ingest and PR preparation are complete. Delivery has no in-scope blocker; only the explicitly excluded routed Wiki projection remains external.
  • Scope: Ordinary CLI CI scheduling plus a scoped scaffold-operation seam for src/scaffold/run.test.ts and src/scaffold/stage.test.ts.
  • Execution mode: One branch and one PR. T1 uses one profiling worker. Parallel implementation is allowed only for T2 and T3 after T1 because their write scopes are disjoint. T5 and T6 each return to one scoped worker.
  • Base intent: Create an isolated clean worktree from fresh origin/main. Do not implement in the current detached worktree or absorb its scaffold drift.
  • Backlog: User approved completion without backlog sync. The active Linear connector cannot access the Devpunks workspace. IP-319/IP-331 are related completed authorities, not reopened owners.
  • Scaffold remediation: Explicitly outside this plan.

Goal

Reduce ordinary CLI test feedback time while preserving every public behavior and every real filesystem security contract. Prove explicitly bounded ordinary CI on GitHub, then stop orchestration and policy tests from repeatedly traversing and materializing the full bundled scaffold through a real host filesystem.

Initial Situation

Commits 129e835a and 113f570b already established the safe update-suite topology:

  • 92 update cases remain four physical wrappers of 23 sequential cases.
  • CI runs those wrappers in four exact jobs.
  • Local release runs the wrappers with four workers, then ordinary tests with two workers.
  • Same-wrapper concurrency and generic two-way Vitest sharding are superseded.

The latest main run used the same ordinary topology:

EvidenceResult
Ordinary job79 files, 1,058 tests passed
Vitest duration586.80s
Test execution547.21s
Collection23.31s
stage.test.ts166.749s
public-output-contract.test.ts119.652s
run.test.ts117.779s
public-context-contract.test.ts48.438s
sync-subagents.test.ts43.509s

The five hot files account for most execution time. This plan attacks stage.test.ts and run.test.ts, approximately 284.5s in the latest sample. The ordinary CI command remains serialized because apps/cli/vitest.config.ts sets fileParallelism: false. The release runner already proves one local two-worker pass, but GitHub stability is unverified.

The latest overall workflow is red from a pre-existing stale wiki projection. All CLI jobs passed. That docs/scaffold divergence is not a CLI runtime failure and is excluded from this plan.

Accepted Authority

This plan continues the implemented IP-319/IP-331 contract:

  • one comprehensive suite; no weaker fast lane;
  • real domain policy with deterministic substitutes only at explicit system seams;
  • one focused real-adapter contract for every deterministic adapter family;
  • mutable filesystem and process state isolated per test scope;
  • benchmark-tuned concurrency with comparable before/after evidence;
  • no permanent runtime, worker-count, coverage, or release threshold.

Resolved Decision Ledger

DecisionResolutionReason
Plan boundaryScheduling experiment, profiling, scoped scaffold-operation seam, bounded run/stage migration, convergence evidenceTargets the dominant filesystem slice without a repository-wide rewrite
Filesystem substituteHigh-level scoped-operation port with real and recording in-memory adaptersKeeps the fake truthful and the interface deep
Low-level filesystem mockRejectedA broad fake would encode false POSIX, inode, and containment semantics
Global node:fs mockRejectedStatic imports and Vitest mock lifecycle make it brittle and implementation-coupled
Real filesystem coverageRetain bounded OS/security contractsLinks, modes, identity, confinement, atomic replacement, rollback, and races require the host filesystem
Performance policyThree comparable samples per mode; threshold-free evidenceMatches IP-319/IP-331 and avoids arbitrary release policy
Backlog syncSkip with user approvalDevpunks Linear workspace is unavailable to the active connector
Scaffold remediationExcludedIt contains local-edited assets and is a separate authorization boundary

These rows capture all five annotated approvals: bounded scope; the high-level operation seam; threshold-free measurement; user-approved backlog skip; and no scaffold-remediation expansion.

Assumptions And Constraints

  • The 645-file bundled baseline remains production authority; tests may replace repeated orchestration traversals, not baseline correctness or generated-byte proof.
  • Existing public CLI results, typed failures, rollback semantics, update case inventory, and release topology cannot change.
  • Worker count is benchmark-tuned evidence. The contract requires a positive explicit bound but cannot freeze a numeric value as policy.
  • Implementation begins from a fresh clean origin/main worktree and must use scoped workers under the repository agent policy.
  • Scaffold remediation and the stale wiki projection remain separate work.

Research Used

  • /tmp/harness-cli-test-performance-handoff-20260805.md: consolidated current timing, concurrency, filesystem-cost, and untried-approach research.
  • Current repository source, workflow, tests, prior IP-319 plan, and completed IP-319/IP-331 acceptance constraints.
  • Readonly parallel investigations of current scheduling safety, hot filesystem ownership, alternative seam designs, and task ordering.
  • No external web claim is needed for the selected design; repository behavior and measured CI evidence are authoritative. External validation is the planned GitHub-run comparison in T2/T5.

Scope

Included

  • Complete per-test and per-file attribution for the two hot scaffold files.
  • Serial versus bounded-worker ordinary-suite evidence, beginning at two.
  • An explicit ScopedScaffoldOperation-style interface with inspect, plan, and materialize responsibilities.
  • A production adapter that composes existing Node/confined filesystem, planning, preflight, project-settings, and materialization behavior.
  • A recording in-memory adapter for caller orchestration, exact-plan identity, failure mapping, rollback decisions, and result shaping.
  • Vertical migration of runScaffold followed by runStageScaffold.
  • An assertion-ownership ledger for every test moved away from a real filesystem.
  • Repeated local and GitHub evidence with exact update exclusions.
  • Canonical runbook correction if the bounded scheduling change is retained.

Excluded

  • public-output-contract.test.ts reduction.
  • public-context-contract.test.ts restructuring.
  • sync-subagents.test.ts core extraction or fixture reuse.
  • Dependency-lockfile policy extraction.
  • Update-wrapper case migration or same-wrapper concurrency.
  • Broad memfs adoption or a repository-wide filesystem abstraction.
  • Test deletion, assertion weakening, timeout increases, or a separate fast suite.
  • Scaffold baseline remediation, generated skill/hook updates, lint-config reconciliation, or the stale wiki projection repair.

Codebase Findings

  • runScaffold already consumes ScaffoldCapabilities, but context/output planning still reads concrete paths before the injected write boundary.
  • runStageScaffold directly owns Node reads, root/session binding, planning, preflight, materialization, settings rollback, and cleanup.
  • features/scaffold-state/lifecycle.ts already expresses in-memory gather -> compile -> apply and observe -> compare/reconcile sequencing.
  • ObserveScaffoldState and ScaffoldReconciliationApplicators demonstrate typed observation and mutation ports.
  • features/project-settings already uses the desired narrow service pattern: one filesystem-facing feature seam, an in-memory adapter, and focused real temporary-filesystem proof.
  • The generic runtime Filesystem only proves basic byte operations and cannot truthfully model confined sessions, hardlinks, parent identity, modes, or atomic publication. It must not be expanded into the primary scaffold seam.
  • The hot files use unique temporary roots and do not directly mutate parent worker globals. The main two-worker risk is resource contention against bounded subprocess/test timeouts.

Design It Twice

Option A: Deep scoped scaffold operation — selected

Conceptual interface:

interface ScopedScaffoldOperation {
  inspect(input: InspectInput): Effect<ScopedInspection, ScopedScaffoldFailure>;
  plan(input: PlanInput): Effect<ScopedScaffoldPlan, ScopedScaffoldFailure>;
  materialize(plan: ScopedScaffoldPlan): Effect<MaterializedScaffold, ScopedScaffoldFailure>;
}

The exact names may follow existing Effect conventions during T3. The contract is semantic:

  • runInitScaffoldLifecycle and the existing scaffold-state compiler/planners remain the sole semantic lifecycle and plan authorities;
  • the scoped operation is the root-bound infrastructure port used by that lifecycle, not a second orchestration layer or a second plan model;
  • inspect supplies the existing lifecycle's gather/observe input;
  • plan delegates to the existing compiler/planner and includes preflight; it does not introduce another planning algorithm;
  • materialize delegates the existing applicator/output behavior;
  • the exact plan that passes preflight is the plan materialized;
  • failures are typed at the public operation boundary;
  • commit state determines whether settings rollback is legal;
  • the test adapter records operations and returns caller-supplied semantic results; it does not regenerate scaffold output.

This interface is deep: it hides root/session identity, traversal, preflight, mutation ordering, and cleanup behind three caller-visible operations. Once both callers migrate, remove the overlapping root/preflight/write responsibilities from ScaffoldCapabilities; keep its external tool and clipboard responsibilities. No old and new orchestration paths may remain in parallel.

Option B: Low-level scoped filesystem — rejected

A stat/read/list/write/mkdir/remove/link/replace/snapshot interface appears easy to inject but exposes filesystem mechanics across stage, output, preflight, project settings, and cleanup. A correct fake would need to model OS behavior it cannot faithfully reproduce. It would enlarge the test surface while leaving orchestration coupled to filesystem steps.

Adapter boundary

  • Real adapter: owns Node and confined-filesystem composition. It remains authoritative for generated bytes and OS/security behavior.
  • Recording adapter: owns deterministic semantic inputs/results and an operation log. It proves caller behavior only.
  • Shared semantic contract: run the same inspection result, plan identity, typed failure, commit-state, and materialized-result contract against the recording and real adapters.
  • POSIX contract split: only the real adapter claims filesystem identity, link, mode, containment, atomicity, and race behavior. The shared semantic contract must not pretend the recording adapter implements POSIX.

Filesystem Contract Matrix

May move to the recording adapter

  • pack/provider/wiki selection propagation;
  • assume-yes and prompt-result handling;
  • required-tool/bootstrap result propagation;
  • inspection -> plan -> materialize ordering;
  • exact plan object identity across preflight/materialization;
  • typed failure-to-result mapping;
  • settings rollback before commit and no rollback after commit;
  • completion payload, baseline version, and CLI version shaping.

Must remain on a real temporary filesystem

  • output-root creation and root replacement;
  • symlinked parents, dangling links, literal mirror links, and path containment;
  • hardlink-safe atomic replacement;
  • file modes and permission preservation;
  • stable root/parent identity and TOCTOU defenses;
  • bounded regular-file reads and non-regular targets;
  • directory backup, publication, cleanup, and interrupted repair;
  • stale generated-file deletion and receipt/manifest rollback;
  • generated projector subprocesses and generated TypeScript/package checks;
  • one representative setup and one init idempotency/materialization smoke.

No test moves until its old assertion, new owning test, and retained real contract are recorded together.

Dependency Readiness

  • Status: No Stack Required.
  • Base: fresh origin/main in an isolated clean worktree.
  • Reason: This is one bounded follow-on slice with internal task dependencies. No upstream feature branch or stacked PR is required.
  • External execution constraint: The current hi check --json reports scaffold divergence. Implementation workers must not rely on the listed drifted skills, hooks, configs, or generated guidance. Remediation remains outside this plan.

Dependency Graph

T1
├── T2 ───────────────┐
└── T3 -> T4 ─────────┤
                     v
                    T5 -> T6

Parallel Execution Waves

WaveTasksStart conditionWrite-scope rule
1T1ImmediatelyOne profiling worker owns disposable diagnostics; all are removed before completion
2T2, T3T1 completeT2 owns workflow/root contract; T3 owns CLI feature/platform/run slice
3T4T3 greenStage slice and shared scoped-operation seam
4T5T2 and T4 greenOne integration worker owns all convergence edits; parent validates/reviews
5T6T5 decision recordedOne docs/wiki worker owns plan and canonical docs edits; parent validates/reviews

T1 uses exactly one profiling worker. T2 and T3 may then run in parallel with two workers. T3-T4 are sequential because they share the scoped-operation interface and hot test files. T5 uses exactly one integration worker; T6 uses exactly one docs/wiki worker. The parent owns cross-task validation and review, not implementation edits.

Tasks

T1: Establish comparable baseline and attribution

  • depends_on: []
  • location: apps/cli/scripts/profile-scaffold-tests.mjs (temporary), apps/cli/src/scaffold/{run.test.ts,stage.test.ts} (temporary diagnostic hooks only), .reports/cli-test-runtime/**, this PLAN.md evidence fields
  • description: In an isolated clean worktree, record machine/runner state, exact SHA, Bun/Node/Vitest versions, and three forced comparable runs of each hot scaffold file plus the ordinary serial and bounded-worker commands, beginning with two workers. Capture complete per-test timing, file timing, subprocess count/duration, and time spent in fixture construction, baseline traversal, planning, preflight, and materialization. Add one disposable diagnostic script plus test-local timed wrappers around those named phases and subprocess calls; remove both before T1 completes. The script creates every output directory, invokes Vitest's JSON reporter, and merges reporter and wrapper events. Do not add production telemetry. Store each sample under .reports/cli-test-runtime/<run-id>/ with a summary containing runId, exact SHA, runner, tool versions, mode, command, test file/name, status, durationMs, subprocess count/duration when available, phase timings, and temporary root. The script derives runId as <serial|workers-N>-<sample-1..3>-<sha8> and fails rather than overwriting an existing sample. Use the same schema for local and GitHub data.
  • validation: Evidence distinguishes runner/import time from baseline traversal, host filesystem mutation, and subprocess time. Every run uses the exact update exclusions and identical test inventory. Scoped diff inspection proves the disposable profiler and test hooks are absent at task completion.
  • status: Completed
  • log: Measured six green samples on 63b9cbf9af6a9468dcee520810a60a9a9c7fbaa9 using Bun 1.3.5, Node v22.22.0, and pinned Vitest 3.2.4 on macOS darwin 27.0.0, Apple M3 Pro (12 logical CPUs, arm64, 36 GiB), AC power. The disposable profiler used the pinned Vitest executable, CI=true, TZ=UTC, the exact exclusions src/update/run.test-cases.test.ts and src/update/run.shard-*.test.ts, and either --maxWorkers=1 or --fileParallelism --maxWorkers=2. One discarded pre-sample exposed a profiler-only error: a TMPDIR nested below the worktree changed two Git-discovery tests. Moving run-owned roots to the system temp directory restored the ordinary inventory; the invalid sample was excluded and removed. Gate passed: the two candidate assertions each invoke the full scaffold operation twice, while timed full operations account for 94.3%-96.1% of focused-file wall time and direct test subprocesses account for only 0.3%-1.3%. T3/T4 may proceed; the premise was not disproved.
  • files edited/created: apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md. Temporary profiler/reporter files, test-local hooks, raw reports, debug log, and run-owned roots were removed before completion.
  • backlog_item_id: not_synced
  • backlog_item_url: not_available
  • relation_mode: user-approved-skip
  • assigned_skills: [autoreview, codebase-design, effect-backend-structure, effect, effect-recoverable-actions, improve-codebase-architecture, parallel-research, quality-types, simplify, swarm-planner, tdd, turborepo]
  • tdd_status: not_applicable
  • tdd_target: Produce comparable baseline evidence without changing suite behavior.
  • red_command: not_applicable
  • expected_red_failure: not_applicable
  • green_command: CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=1 && CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=2 && CI=true TZ=UTC bun run --cwd apps/cli test -- --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts'
  • reason_not_testable: Measurement establishes evidence and does not change observable product behavior.
  • red_evidence:
  • green_evidence: CI=true TZ=UTC bun apps/cli/scripts/profile-scaffold-tests.mjs --samples=3 --output=.reports/cli-test-runtime --worker-count=1 and the same command with --worker-count=2 produced six passing summaries. Every focused-run sample executed the same 28-test inventory, every focused-stage sample the same 49-test inventory, and every ordinary sample the same 80-file/1,062-test inventory; SHA-256 hashes were respectively 3251e7cfd8b22d7311ab34891efb223b4abc1ca932189a4418fc77df4b73e8c9, 854b3efb11c1b61ddbb46a25de2d82f932d9921341b49a20edfa2273e0640ef1, and fe906ce7e79f729d201a7b58f3e1c425bdfac9491a6c803559345140985b81e7 across both modes. No timeout changed.
  • codebase_design_notes: Attribution must measure the owning semantic phases rather than private helper calls. Diagnostic hooks must remain test-only and removable. T3/T4 proceed only if the evidence identifies repeated full baseline traversal/materialization as a dominant cost for the assertions selected there. If attribution disproves that premise, stop and revise this plan instead of adding the operation seam as a performance change.
  • review_mode: cli
  • runtime_validation: required
  • runtime_target: Local ordinary CLI suite and isolated hot-file runs on one fixed SHA.
  • runtime_evidence: Serial wall means (min-max, population CV) were focused run 52.938s (51.856-54.814s, 2.52%), focused stage 88.885s (79.949-106.033s, 13.65%), and ordinary 348.641s (326.043-391.707s, 8.74%). Two-worker means were 52.396s (51.949-52.676s, 0.61%), 80.129s (79.574-80.653s, 0.55%), and 170.452s (170.337-170.658s, 0.09%): bounded ordinary execution was 51.1% faster by the sample means while single-file controls were unchanged. Candidate means were stale-skill regeneration 5.263s serial / 5.165s bounded and exact planned baseline symlink 5.548s / 4.884s; source inspection confirms two complete scaffold calls in each. Focused-run samples recorded 29 full operations averaging 50.848s serial / 50.346s bounded (96.05% / 96.09% of wall); focused-stage recorded 45 averaging 84.062s / 75.682s (94.55% / 94.45%). Direct subprocess means were 0.638s / 0.635s for focused run, 0.299s / 0.266s for focused stage, and 0.969s / 0.987s ordinary. Reporter-to-wall gaps were below 2.1s; collection and import were sub-second in the focused modules. All 18 commands passed with identical inventories and no observable behavior difference.
  • runtime_cleanup: All six external run-owned temp roots were absent after profiler finally cleanup. The invalid sample, .reports, disposable profiler/reporter, test-local timing/debug hooks, and session debug log were moved to Trash. rg found zero diagnostic markers; scoped git diff for run.test.ts, stage.test.ts, and apps/cli/scripts was empty; git diff --check passed. The unrelated untracked IMPLEMENTATION-NOTES.md was not modified.

T2: Prove explicit bounded ordinary CI

  • depends_on: [T1]
  • location: .github/workflows/behavior-contract.yml, scripts/behavior-contract/root-suite.test.ts
  • description: Test-drive an ordinary-job-only change that passes --fileParallelism --maxWorkers=\<N\> with the exact update registrar/wrapper exclusions. Begin measurement at N=2, but keep the behavior contract intentionally numeric-value-agnostic so later evidence can tune the positive bounded value. Preserve the serialized package test script, the release runner, and all four exact update jobs. Do not raise timeouts or introduce a generic shard.
  • validation: The static behavior contract rejects an ordinary CLI command missing either bounded flag or either update exclusion, and still proves four exact update jobs. Local root contract is green before GitHub validation.
  • status: Reverted after unsafe GitHub runtime evidence
  • log: Added a value-agnostic root behavior contract before changing the workflow. The focused RED failed because the ordinary cli-tests job lacked --fileParallelism. The workflow then gained only --fileParallelism --maxWorkers=2; package scripts, release runner, timeouts, four exact update jobs, and exclusions remain unchanged. Simplify removed the initial numeric value from the durable contract so it accepts any explicit positive bound while rejecting test:parallel, generic shards, and missing exclusions. Local validation remained green, but GitHub attempt 1 produced one Vitest timeout and three spawnSync node ETIMEDOUT failures in public-output-contract.test.ts; all four update jobs passed. Per the unsafe sample gate, the workflow and T2-specific contract were reverted byte-for- byte to origin/main. Package/release/update topology and the scoped operation changes remain intact.
  • files edited/created: Net production diff: none. The temporary changes to .github/workflows/behavior-contract.yml and scripts/behavior-contract/root-suite.test.ts were reverted.
  • backlog_item_id: not_synced
  • backlog_item_url: not_available
  • relation_mode: user-approved-skip
  • assigned_skills: [autoreview, codebase-design, parallel-research, simplify, swarm-planner, tdd, turborepo]
  • tdd_status: required
  • tdd_target: The behavior contract rejects the current serial ordinary CLI workflow because it lacks file parallelism with an explicit positive bounded worker value while retaining exact update exclusions.
  • red_command: bun test ./scripts/behavior-contract/root-suite.test.ts -t "delegates ordinary CLI tests to bounded file workers"
  • expected_red_failure: The workflow command does not contain --fileParallelism and a positive --maxWorkers=\<N\>, so the new scheduling assertion fails.
  • green_command: bun test ./scripts/behavior-contract/root-suite.test.ts && CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts'
  • reason_not_testable:
  • red_evidence: Before the workflow edit, bun test ./scripts/behavior-contract/root-suite.test.ts -t "delegates ordinary CLI tests to bounded file workers" exited 1 at Expected to contain: "--fileParallelism".
  • green_evidence: The focused contract passed 1/1; the full root contract passed 11 tests and 286 assertions. CI=true TZ=UTC bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' passed 80 files and 1,062 tests in 170.40s. git diff --check passed.
  • codebase_design_notes: Scheduling stays in the CI adapter. Package and release semantics remain unchanged; no new root or Turbo task is introduced. The retained numeric value is evidence, not a permanent policy threshold.
  • review_mode: cli
  • runtime_validation: required
  • runtime_target: GitHub Actions ordinary cli-tests job on one fixed SHA.
  • runtime_evidence: Local implementation evidence was green, including a committed-SHA pair at 323.30s serial versus 168.18s bounded with identical 82-file/1,068-test inventory. GitHub run 31054456710, attempt 1, head 9fe7cb57f6c1f14b47ba97223fd9d37e38042f11, falsified the safety gate: ordinary job 92468806556 passed 81 files and 1,066 tests but failed four of 1,070 Linux tests after 507.45s (one 30s Vitest timeout and three spawnSync node ETIMEDOUT failures). Update jobs 92468806561, 92468806566, 92468806557, and 92468806579 each passed 23/23. Further bounded attempts were unnecessary because the plan required every sample to be green. Three corrected-SHA serialized workflow attempts subsequently passed, as recorded under T5.
  • runtime_cleanup: No external fixtures remain. The exact diagnostic run is retained as evidence; no GitHub run was targeted by SHA or deleted.

T3: Introduce the scoped operation through runScaffold

  • depends_on: [T1]
  • location: apps/cli/src/features/scaffold-operation/{model.ts,index.ts,service.test.ts}, apps/cli/src/platform/scoped-scaffold-operation.ts, apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts}, apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts, apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts}
  • description: Implement the first vertical slice of the deep scoped operation. runScaffold invokes one injected operation that delegates inspection to the existing lifecycle gather/observe path, delegates planning and preflight to the existing compiler/planners, materializes that exact plan through the existing output path, and returns the existing public result without the caller reading a host repository. Preflight is internal to the operation's plan phase; the caller must not perform a second preflight. Add the real adapter only for behavior required by this slice and a recording in-memory adapter for caller tests. Run one shared semantic contract against both adapters. In this task, move the run-side stale regenerated-skill orchestration assertion to the recording adapter while retaining one bounded real stale-file deletion/containment witness. Record every moved and retained assertion in the task log. Keep planning algorithms and generated bytes authoritative in existing production modules.
  • validation: A test uses nonexistent virtual roots and proves operation order, exact plan identity, result equality, and zero host fixture creation. The shared contract proves the same inspection result, plan identity, typed failure, and public result semantics for real and recording adapters. Existing representative real stale-file deletion/containment stays green; no test count substitutes for assertion ownership.
  • status: Completed
  • log: Added the three-operation Effect service, real adapter, recording adapter, shared semantic contract, and first runScaffold vertical slice. Review hardened both adapters to use fresh inspection-bound, one-shot plan capabilities consumed before any mutation; copied, foreign, successful replay, and failed replay are rejected consistently. Stable root/baseline/ shape authority is derived from the owned inspection rather than repeated by callers, and repository/validation failures remain in the typed error channel. The stale-skill candidate now splits orchestration from the OS witness: two virtual recording-adapter reruns prove exact operation order, unique plan identity, and public result shaping without host access; one bounded real projection proves stale deletion, containment via an unchanged outside sentinel, and absence from the managed result. Complete bundled projections in that candidate fell from two to one; observed focused time fell from about 4.93s to 2.72s. Structured autoreview then raised a P2: the opaque plan did not yet own enough prepared state to guarantee that apply consumed exactly what preflight approved. The corrected rich plan owns the compiled desired input and state, explicit reset directories, the static public result including its deep-owned baseline summary, and an opaque one-shot applier. Preflight and apply consume that same plan. Apply performs no source traversal or replanning. Focused proof removes source bytes after planning, exercises stale reset directories, mutates the original summary, and attempts replay; materialization still uses the frozen plan and replay is rejected. Both subsequent manual rereviews were clean. The later structured review and its stage/settings remediation are recorded under T4.
  • files edited/created: apps/cli/src/features/scaffold-operation/{index.ts,model.ts,service.test.ts}, apps/cli/src/platform/scoped-scaffold-operation.ts, apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts}, apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts, apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts}
  • backlog_item_id: not_synced
  • backlog_item_url: not_available
  • relation_mode: user-approved-skip
  • assigned_skills: [autoreview, codebase-design, effect-backend-structure, effect, effect-recoverable-actions, improve-codebase-architecture, parallel-research, quality-types, simplify, swarm-planner, tdd, turborepo]
  • tdd_status: required
  • tdd_target: runScaffold preflights and materializes one injected plan on virtual roots without reading the host repository.
  • red_command: bun run --cwd apps/cli test -- src/scaffold/run.test.ts -t "preflights and materializes the exact injected plan without reading the host repository"
  • expected_red_failure: Current runScaffold performs concrete context and output planning before the injected capability boundary, so nonexistent virtual roots cannot complete and exact injected plan identity is unavailable.
  • green_command: bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/run.test.ts -t "scoped operation|exact injected plan|shared semantic contract|rerun|symlink" && bun run --cwd apps/cli check-types
  • reason_not_testable:
  • red_evidence: Before runScaffold consumed the new operation, the exact virtual-root test failed 1/1 with RepositoryScanError for /virtual/scaffold-input, proving the old caller still inspected the host.
  • green_evidence: The exact T3 suite passed 3 files and 32/32 tests in 50.03s; run.test.ts passed 29/29 in 46.56s; the real shared contract passed in 2.59s; check-types, scoped oxlint, scoped oxfmt, and git diff --check passed. Assertion ownership is explicit: shared adapters own inspection/plan identity and one-shot lifecycle; virtual caller tests own rerun orchestration/result shaping; the real temporary filesystem owns stale deletion and containment. The historical post-binding adapter contract passed 10/10 in 15.95s. After the exact-plan correction, the adapter contract passed 11/11, output passed 12/12, and their combined run passed 23/23. The added proof covers source removal after planning, stale reset-directory application, deep-owned summary immutability, same-plan preflight/apply, and replay rejection without apply-time traversal or replanning.
  • codebase_design_notes: Keep the public interface at three semantic operations. Do not expose stat/read/write calls, confined sessions, inode identities, or the recording adapter through production callers. The recording adapter is a caller test double, not a scaffold generator. The operation composes the existing lifecycle/compiler/planner/applicator into a rich opaque plan. The plan retains compiled desired input/state, explicit reset directories, a static deep-owned public result, and a one-shot applier; it does not duplicate planning semantics or coexist with a second root/ preflight/write route in ScaffoldCapabilities after both callers migrate.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T4: Route staged initialization through the same scoped operation

  • depends_on: [T3]
  • location: apps/cli/src/features/scaffold-operation/**, apps/cli/src/platform/scoped-scaffold-operation.ts, apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts}, apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts, apps/cli/src/scaffold/{output.ts,stage.ts,stage.test.ts}
  • description: Migrate one vertical runStageScaffold path from inspection through materialization. Prove inspect -> plan -> materialize ordering, exact plan identity, rollback only when materialization fails before commit, and no rollback when completion rendering fails after committed materialization. Preserve typed failures and existing public stage results. In this task, move the stage-side exact projection rerun and planned baseline symlink rerun orchestration assertions to the recording adapter while retaining bounded real mirror/idempotency and symlink witnesses. Record every moved and retained assertion in the task log.
  • validation: The RED tests run against a virtual root without mkdtempSync and prove exact operation order/identity plus both sides of the commit boundary: rollback before commit and preservation after a downstream failure following commit. The shared real/recording contract covers commit state and typed failure translation. Focused real settings rollback, exact mirror/idempotency, and planned symlink behavior remain green.
  • status: Completed; mandatory review blockers resolved, GREEN evidence retained, TDD status partial/not-met
  • log: Routed runStageScaffold through the required scoped operation and installed the real service in production/test composition; removed the optional runScaffold fallback and the old caller-owned inspect/root-bind/ preflight/write route. The stage adapter owns settings rollback until its materialization commits; a virtual caller test proves rollback before commit and preservation after downstream presentation failure. The shared semantic contract now runs stage variants against both recording and production adapters and proves owned inspection, exact one-shot plans, committed result, typed pre-commit failure, and replay rejection after success or failure. Repository diagnostics remain unchanged at the scanner authority; the stage boundary preserves its documented CliDetectionError wrapper with concrete typed errors. Assertion ownership was split without losing OS claims: the recording adapter owns both named two-rerun sequences and exact/fresh plan identity; one consolidated real two-operation rerun retains all exact mirror readlinks plus the planned skill symlink, reducing the former four complete projections to two; a separate real mismatched-symlink witness retains containment. Mandatory review then required stable pre-inspection root authority, caller-input snapshots, concrete adapter-owned stage plan state, selected-baseline preservation, separation of public ScaffoldCapabilities from the internal filesystem adapter, and a stage orchestration test with no host fixture. Final review identified and resolved a deletion/rollback race: the adapter prepares an existing root session or nearest-existing-ancestor authority before inspection, commits a distinct absent root only after inspection succeeds, and performs no mutation when inspection fails. This removes the newly introduced foreign-root deletion path while preserving the prior filesystem security contracts. Structured autoreview on d58f638d then returned P1 because stage replanned output and read live stage-skill source after planning, plus P2 because scoped settings initialization read a live baseline version. The fix makes stage own and consume one rich output plan. Its additionalSkillSelection captures selected-first, bundled-fallback stage-skill bytes, aliases, and reset roots; those exact reset targets are preflighted, and the old post-commit skill copy is removed. Settings initialization and baseline metadata are also planned, and scoped settings consumes the scaffoldOutput baseline rather than live caller state. Both final manual rereviews were clean. The third structured autoreview on head 911a73e1 was CLEAN with no actionable findings and an overall patch- correctness score of 0.91. It confirmed retained plans and one-shot lifecycle, root authority, stage-skill snapshotting, settings metadata, rollback boundary, and serialized CI remain consistent.
  • files edited/created: apps/cli/src/index.ts, apps/cli/src/platform/feature-application-operations.ts, apps/cli/src/platform/scaffold-capabilities.ts, apps/cli/src/platform/scoped-scaffold-operation-adapter.ts, apps/cli/src/features/scaffold-operation/{model.ts,index.ts}, apps/cli/src/platform/scoped-scaffold-operation.ts, apps/cli/src/scaffold/capabilities.ts, apps/cli/src/testing/{layers.ts,scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts}, apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts, apps/cli/src/scaffold/{run.ts,run.test.ts,output.ts,stage.ts,stage.test.ts,confined-filesystem.ts,confined-filesystem.test.ts}
  • backlog_item_id: not_synced
  • backlog_item_url: not_available
  • relation_mode: user-approved-skip
  • assigned_skills: [autoreview, codebase-design, effect-backend-structure, effect, effect-recoverable-actions, improve-codebase-architecture, parallel-research, quality-types, simplify, swarm-planner, tdd, turborepo]
  • tdd_status: partial — mandated RED not met; GREEN retained
  • tdd_target: One injected scoped operation owns staged inspection through commit, and rollback follows commit state rather than downstream rendering.
  • red_command: bun run --cwd apps/cli test -- src/scaffold/stage.test.ts -t "rolls back before commit and preserves committed state after downstream failure"
  • expected_red_failure: Current runStageScaffold directly creates a confined session and calls concrete planning/materialization code; no injected operation log or commit-aware rollback seam exists.
  • green_command: bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/stage.test.ts -t "scoped scaffold operation|rollback|committed state|stale|idempotent|shared semantic contract" && bun run --cwd apps/cli check-types
  • reason_not_testable: not_applicable; the behavior was testable and the mandated RED was missed
  • red_evidence: Not met. The exact mandated target did not exist before T4; Vitest collected stage.test.ts and reported 49/49 skipped under the missing name. A missing target and skipped tests are not RED evidence. During convergence, the repository typed-failure contract caught an invalid scanner-cause change; that independent failure was removed and the existing diagnostic contract restored, but it does not satisfy the mandated T4 RED. Later review remediation did produce two independent public REDs without repairing that original TDD miss: after planning, removing the selected stage skill source made the old apply fail with Could not locate baseline asset; the scoped-settings test expected baseline version 3.1.3 but observed the mutated-after-plan value. The stage metadata regression became GREEN as part of the first fix and is not falsely claimed as RED evidence.
  • green_evidence: The exact commit-boundary target passed 1/1. Full stage.test.ts passed 49/49 in 72.16s; typed-failure, service, shared adapters, and full run compatibility passed 36/36 in 53.06s; the shared adapter contract passed 4/4 with the real stage case around 2.8s. Historical T3 exact-plan evidence remains adapter 11/11, output 12/12, and combined 23/23. After the stage/settings correction, the adapter contract passed 13/13 in 21.79s, output passed 12/12 in 9.73s, and their combined run passed 25/25 in 32.29s. Focused settings version and selected/bundled fallback coverage passed 3/3; symlink witnesses passed 2/2; full stage passed 49/49. Typecheck, check, and diff gates passed.
  • codebase_design_notes: The scoped operation owns phase ordering, plan provenance, caller-input snapshots, and one-shot plan consumption. The stage adapter owns concrete StagePlanState plus its commit/rollback boundary. The public ScaffoldCapabilities service retains tool, clipboard, settings, and runtime concerns; the internal real adapter owns root-authority preparation, deferred absent-root binding, inspection, preflight, and filesystem materialization. The recording adapter proves orchestration semantics without becoming filesystem authority. Stage owns one rich output plan through preflight and apply. additionalSkillSelection freezes selected-first or bundled-fallback bytes, aliases, and reset roots; planned settings freezes initialization and baseline metadata, and scoped settings reads that planned scaffoldOutput baseline. Native Node mkdir followed by lstat cannot atomically prove ownership against an exact replacement in between; that pre-observation window is explicitly unsupported rather than claimed fail-closed. Parent/root replacement before commit and later session swaps remain fail-closed.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T5: Converge scheduling, contracts, and runtime evidence

  • depends_on: [T2, T4]

  • location: .github/workflows/behavior-contract.yml, scripts/behavior-contract/root-suite.test.ts, apps/cli/src/features/scaffold-operation/**, apps/cli/src/platform/scoped-scaffold-operation.ts, apps/cli/src/testing/{scoped-scaffold-operation.ts,scoped-scaffold-operation-contract.ts}, apps/cli/src/adapters/scoped-scaffold-operation-adapter-contract.test.ts, apps/cli/src/scaffold/{run.ts,run.test.ts,stage.ts,stage.test.ts,confined-filesystem.test.ts}, apps/cli/src/update/run.shards.test.ts, apps/cli/scripts/run-release-tests.mjs, this PLAN.md

  • description: Run the complete compatibility and performance matrix. Repeat serial and bounded-worker ordinary commands on the final SHA, run all four update wrappers and test:release, then collect three GitHub samples. Retain T2 only when every bounded-worker sample is green with identical inventory and its improvement is repeatable rather than mixed with timeout or variance regressions. Tune the positive bounded worker value when evidence supports it, updating both workflow and value-agnostic contract. If evidence is mixed or unsafe, revert only T2's workflow/root-contract change and keep the scoped-operation improvement.

  • validation: Focused feature/adapter tests, real filesystem contracts, serial ordinary suite, bounded-worker ordinary suite, four update wrappers, release suite, CLI build/check/typecheck, root behavior contract, and diff checks are green. Known unrelated wiki projection failure is reported separately.

  • status: Completed; bounded scheduling reverted after unsafe GitHub evidence

  • log: The focused real/adapter matrix, all four independent update wrappers, the two-phase release runner, build/type/check/root contracts, and matched serial/bounded ordinary controls are green. The four wrappers retain exactly 23 unique cases each and 92 unique cases in union. test:release visibly completed the four-worker update phase before its two-worker ordinary phase. The local serial and bounded commands ran the identical 82-file/ 1,068-test inventory; bounded wall time was 165.75s versus serial 312.63s (46.98% lower), while isolated run/stage controls remained comparable locally. GitHub then exposed subprocess starvation under two workers, so the final decision is to revert T2 scheduling without undoing the scoped-operation seam. The same run also exposed a forbidden runtime barrel; scaffold-operation/index.ts now owns the Effect service contract directly, and Effect-v4 validation passes without a whitelist. Three corrected-SHA GitHub attempts then passed the serialized ordinary job and every update shard. Their workflow conclusion remained red only because the excluded wiki projection reports four stale paths.

  • files edited/created: No new T5 implementation files. Validation covers all T2-T4 files listed above; evidence is recorded in this plan and IMPLEMENTATION-NOTES.md.

  • backlog_item_id: not_synced

  • backlog_item_url: not_available

  • relation_mode: user-approved-skip

  • assigned_skills: [autoreview, codebase-design, effect-backend-structure, effect, effect-recoverable-actions, improve-codebase-architecture, parallel-research, quality-types, review-phase, simplify, swarm-planner, tdd, turborepo]

  • tdd_status: not_applicable

  • tdd_target: Validate the already test-driven slices and make the evidence-based retain/revert decision for CI scheduling.

  • red_command: not_applicable

  • expected_red_failure: not_applicable

  • green_command: bun run --cwd apps/cli test -- src/features/scaffold-operation/service.test.ts src/adapters/scoped-scaffold-operation-adapter-contract.test.ts src/scaffold/confined-filesystem.test.ts src/adapters/filesystem-adapter-contract.test.ts src/scaffold/stage.test.ts src/scaffold/run.test.ts && bun run --cwd apps/cli test -- --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && bun run --cwd apps/cli test -- --fileParallelism --maxWorkers=2 --exclude=src/update/run.test-cases.test.ts '--exclude=src/update/run.shard-*.test.ts' && bun run --cwd apps/cli test src/update/run.shard-0.test.ts && bun run --cwd apps/cli test src/update/run.shard-1.test.ts && bun run --cwd apps/cli test src/update/run.shard-2.test.ts && bun run --cwd apps/cli test src/update/run.shard-3.test.ts && bun run --cwd apps/cli test:release && bun run --cwd apps/cli build && bun run --cwd apps/cli check-types && bun run --cwd apps/cli check && bun test ./scripts/behavior-contract/root-suite.test.ts && git diff --check

  • reason_not_testable: Convergence task validates prior behavior-changing tasks and records runtime evidence; it introduces no new behavior.

  • red_evidence:

  • green_evidence: Focused matrix passed 6 files/113 tests in 129.74s (inventory 1bc136663987a19cc031f12445891ae926627fb8cb53acbd1842ae20d605652c). Independent wrappers passed 23/23 each in 202.06s, 215.89s, 181.44s, and 211.60s; their union was 92/92 unique titles (inventory 3c1f102a7fe17688ecd599a16415bfc5413b27ffe1f86c9b21ae9bc12da221a5). test:release passed in 429.33s: update 92/92 in 260.91s, then bounded ordinary 82 files/1,068 tests in 163.62s. Build, check-types, CLI check, root 11/11 with 286 assertions, and diff checks passed. The post-revert root contract passes 10/10 with 280 assertions; Effect-v4 validation passes 94/94.

  • codebase_design_notes: Scheduling and filesystem seams remain independent. A failed scheduling experiment must not undo the deeper operation boundary.

  • review_mode: cli

  • runtime_validation: required

  • runtime_target: Final local release suite plus GitHub ordinary/update jobs on one fixed SHA.

  • runtime_evidence: Pre-commit serial and bounded controls passed the same 82-file/1,068-test inventory hash 86fdba61c52dcdc0b380550123b82205cc8db7cd5fb3ed4004bc7090859b0a5d. Serial wall was 312.63s; bounded wall was 165.75s. Hot controls were run.test.ts 46.955s -> 48.364s and stage.test.ts 71.676s -> 74.341s. A committed-SHA pair repeated the local result at 323.30s -> 168.18s, but GitHub attempt 1 failed the bounded ordinary job as recorded under T2. The final post-revert local serial control passed 82 files and 1,068 tests in 316.19s with the same inventory SHA-256 86fdba61c52dcdc0b380550123b82205cc8db7cd5fb3ed4004bc7090859b0a5d.

    Diagnostic bounded run 31054456710, attempt 1, used head 9fe7cb57f6c1f14b47ba97223fd9d37e38042f11. Its ordinary job completed 81 files and 1,066 tests but failed four of 1,070 after 507.45s. Update shards 0, 1, 2, and 3 each passed 23/23 in 324.66s, 344.27s, 243.54s, and 320.51s. The root job passed 92 of 94 Effect checks and failed one semantic pass-through-alias test; correcting that public-root ownership produced 94/94 locally and on every corrected attempt.

    Corrected run 31055566836 used head 4ad8a9b9b66b35bd4eaad84e7184888186b35d79 and PR merge de14b84e. Attempts 1-3 passed the serialized ordinary inventory at 82 files/1,070 tests: attempt 1 in 420.67s, attempt 2 in 359.36s, and attempt 3 in 595.58s. Every corrected attempt also passed four 23/23 update jobs: attempt 1 shard 0 271.48s, shard 1 345.83s, shard 2 292.29s, shard 3 321.43s; attempt 2 shard 0 328.85s, shard 1 351.91s, shard 2 312.31s, shard 3 293.55s; attempt 3 shard 0 323.11s, shard 1 334.49s, shard 2 296.97s, and shard 3 250.19s.

    Corrected root attempts 1, 2, and 3 completed all 94 Effect public-surface checks (93 passed and one focused configuration check intentionally skipped) and passed every root behavior- contract assertion. They failed only @punks/wiki#check:content for four excluded projection paths: apps/wiki/content/docs/project/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md, apps/wiki/content/docs/project/specs/cli/cli-test-runtime-hardening/PLAN.md, apps/wiki/content/docs/project/specs/.source-projection.json, and apps/wiki/content/docs/project/runbooks/hi-cli-scaffolding.md.

    Final lifecycle run 31067183663 used head 3e81a31c. cli-tests and all four exact update jobs were green; API, backoffice, and wiki Vercel checks were green. Behavior Contract was red only at @punks/wiki#check:content, which requested exactly the excluded projection work: create routed cli-test-runtime-hardening/PLAN.md and IMPLEMENTATION-NOTES.md, update apps/wiki/content/docs/project/specs/.source-projection.json, and update the routed hi-cli-scaffolding runbook mirror.

  • runtime_cleanup: External reports/logs were removed, run-owned roots are absent, and exact run/attempt/job provenance is durable above. No GitHub run was deleted.

T6: Record durable closeout without scaffold remediation

  • depends_on: [T5]
  • location: apps/wiki/specs/cli/cli-test-runtime-hardening/{PLAN.md,IMPLEMENTATION-NOTES.md}; docs/README.md and docs/runbooks/hi-cli-scaffolding.md for retained architecture and current test topology
  • description: Fill every task's RED/GREEN/runtime evidence and assertion ownership ledger. Document the retained scoped-operation architecture in docs/README.md and the canonical runbook. Correct obsolete test-topology wording to the current serialized/four-wrapper arrangement without documenting the reverted bounded scheduling experiment. Do not update the scaffold-managed routed wiki mirror, generated skills/hooks, settings, lint specs, or Oxlint configs. Record the known scaffold/wiki divergence as external and keep CLI acceptance separate from the overall red workflow.
  • validation: Plan evidence is complete, canonical docs match retained CLI topology, targeted formatting passes, and scoped git diff --check is clean. No excluded scaffold-managed file is modified.
  • status: Completed
  • log: Filled final local and GitHub runtime evidence, recorded the T2 REVERT decision, corrected stale feature locations, and closed the assertion and acceptance ledgers. Scheduling topology remained serialized, but the retained scoped-operation architecture changed: canonical docs now record pre-inspection root-authority preparation, deferred absent-root binding, one-shot snapshotted plans, adapter-owned stage state, commit/rollback ownership, and the real/recording contract split. They also replace the stale two-shard description with the current serialized ordinary job plus four exact wrapper jobs and mark test:parallel local-only. The four scaffold-managed wiki projection paths remain excluded and unmodified.
  • files edited/created: apps/wiki/specs/cli/cli-test-runtime-hardening/{PLAN.md,IMPLEMENTATION-NOTES.md}, docs/README.md, docs/runbooks/hi-cli-scaffolding.md
  • backlog_item_id: not_synced
  • backlog_item_url: not_available
  • relation_mode: user-approved-skip
  • assigned_skills: [autoreview, codebase-design, create-plan, create-spec, docs-onboarding, implement-spec, improve-codebase-architecture, parallel-research, review-phase, simplify, tdd, writing-beats, writing-fragments, writing-great-skills, writing-shape]
  • tdd_status: not_applicable
  • tdd_target: Record non-runtime evidence and canonical maintainer guidance for the retained topology.
  • red_command: not_applicable
  • expected_red_failure: not_applicable
  • green_command: bunx oxfmt --check docs/README.md docs/runbooks/hi-cli-scaffolding.md apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md apps/wiki/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md && git diff --check -- docs/README.md docs/runbooks/hi-cli-scaffolding.md apps/wiki/specs/cli/cli-test-runtime-hardening/PLAN.md apps/wiki/specs/cli/cli-test-runtime-hardening/IMPLEMENTATION-NOTES.md
  • reason_not_testable: Evidence/docs closeout does not change runtime behavior.
  • red_evidence:
  • green_evidence: The canonical root docs and both durable evidence files pass targeted Oxfmt and scoped git diff --check. The canonical docs record the retained scoped-operation architecture and current serialized/four-wrapper topology; reverted bounded-CI scheduling is not documented. Scaffold-managed mirror and projection files have no T6 diff.
  • codebase_design_notes: Closeout documents the selected deep boundary and retained real contract matrix; it does not expand the implementation seam.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

Testing Strategy

RED/GREEN tracer bullets

  1. T2 first proves the workflow contract rejects unbounded/serial ordinary CI.
  2. T3 proves runScaffold uses one injected semantic plan on virtual roots, moves only its run-side orchestration assertions, and retains their real filesystem witnesses.
  3. T4 proves staged commit/rollback semantics through the same public operation, moves only its stage-side orchestration assertions, and retains their real filesystem witnesses.

T2 and T3 recorded valid public RED before GREEN. T4 retained its later GREEN and assertion-migration evidence, but its mandated target was absent during the purported RED run; 49 skipped tests are not RED evidence. T4 is therefore partial/not-met for TDD even though its implemented behavior and validation are green.

Validation order

  1. Focused new operation tests.
  2. Focused run and stage tests.
  3. Real filesystem and adapter contracts.
  4. Serial ordinary CLI suite.
  5. Bounded-worker ordinary CLI suite using the evidence-selected positive value.
  6. Four exact update wrappers.
  7. Local test:release.
  8. CLI build, typecheck, check, behavior contract, and diff checks.
  9. Three GitHub samples on one fixed SHA.

Evidence rules

  • Record exact SHA, runner, tool versions, command, test count, file count, duration components, and per-hot-file timings.
  • Compare like with like; never compare a battery/low-power local run directly with a GitHub runner as causal proof.
  • One green run proves feasibility, not stability.
  • A moved test needs assertion-level replacement evidence, not only preserved aggregate counts.
  • Never increase a timeout to make a scheduling experiment appear green.

Validation Gates

Gate A: Measurement readiness

  • T1 has three comparable samples for both scheduling modes.
  • Bottleneck attribution distinguishes baseline traversal, filesystem work, planning/materialization, and subprocess time.
  • T3/T4 proceed only if attribution confirms repeated full-baseline traversal or materialization is a dominant cost for their candidate assertions; otherwise stop and revise the plan.

Gate B: Scoped operation integrity

  • Recording adapter proves caller behavior without host paths.
  • Real adapter remains authoritative for bytes and OS/security semantics.
  • No generic filesystem or node:fs mock becomes a production/test authority.

Gate C: Coverage preservation

  • Assertion ownership ledger is complete.
  • Retained real contract matrix is green.
  • Four update wrappers still cover 92 unique cases, 23 each.

Gate D: Scheduling decision

  • Three local and three GitHub bounded-worker samples are green with identical inventory.
  • Median improvement is repeatable and not paired with timeout or variance regressions.
  • Mixed evidence reverts T2 only; it does not invalidate completed T3-T4.

Gate E: Closeout

  • All task evidence fields are filled.
  • Canonical docs reflect retained behavior.
  • Excluded scaffold drift remains untouched and explicitly reported.

Risks And Mitigations

RiskMitigation
Recording adapter becomes a second scaffold generatorIt accepts semantic caller-supplied results and records operations only
Deep interface becomes a pass-through wrapperKeep three semantic operations and hide sessions, filesystem calls, and mutation ordering
Real security assertions move to the fakeEnforce the contract matrix and assertion-ownership ledger inside T3/T4 and at T5 review
Bounded workers exhaust subprocess timeoutsRepeat on GitHub; retain only with stable evidence; never raise timeouts
Interface work and scheduling become coupledSeparate T2 and T3 write scopes; T5 may revert T2 independently
Full baseline witnesses disappearRetain representative setup, rerun/idempotency, generated-byte, and real adapter contracts
Current scaffold drift contaminates implementationUse a clean worktree; exclude listed drifted assets; do not run scaffold remediation
Required scoped Effect skills remain driftedDo not rely on their changed scaffold copies; execution must use verified source/current project patterns or stop and report the prerequisite
Overall workflow remains red from wiki projectionJudge CLI jobs separately and record the pre-existing docs/scaffold failure

Skill Routing

  • grilling: resolved scope, seam, evidence, backlog, and remediation decisions.
  • parallel-research: supplied current CI state, concurrency audit, hot-file attribution, and untried approaches.
  • codebase-design: selected a deep scoped-operation seam over a broad filesystem-shaped interface.
  • swarm-planner: produced dependency-aware waves with disjoint T2/T3 write scopes and sequential shared-interface work.
  • tdd: attached public RED/GREEN tracer bullets to T2-T4. T2 and T3 met the contract; T4 retained its assertion migration, real witness, and GREEN proof but is partial/not-met because its mandated RED target was absent.
  • turborepo: kept scheduling package-owned and avoided new root task logic or generic graph-bypassing parallelism.
  • Scoped apps/cli skills are recorded on implementation tasks. Drifted scaffold copies are an execution prerequisite, not planning authority.

Backlog Sync

  • Result: Skipped by explicit user decision.
  • Reason: The active Linear connector cannot access the Devpunks workspace; the user requested completion of the plan without backlog mutation.
  • Related completed authority: IP-319 and IP-331.
  • Created/updated items: None.
  • Future optional action: Create a new Harness Intelligence story related to IP-331; do not reopen or overwrite the completed IP-319 plan.

Unresolved Questions

No product questions remain. The known Node directory-publication limitation is parked: mkdir followed by lstat cannot atomically prove ownership against an exact intervening replacement without native no-replace publication.

Debug Assessment

The delivery debug assessment is closed from already captured CI runtime logs; no new instrumentation was added because the incident was proven before this phase.

HypothesisResultExisting evidence
A. The two-worker ordinary job caused subprocess starvationCONFIRMEDGitHub attempt 1 failed four ordinary tests, including three spawnSync node ETIMEDOUT failures, while all four update jobs were green.
B. Update shards overlapped or repeated workREJECTEDAll four exact wrappers passed 23/23 and their union was unique.
C. Serialized ordinary execution contained a behavior regressionREJECTEDThe local serialized inventory and all three corrected GitHub serialized inventories were green.
D. A generic timeout increase alone would resolve the incidentREJECTED / narrowedIncreasing timeouts does not address subprocess spawn starvation; the reverted serialized topology passed.

Root cause was resource contention and subprocess starvation introduced by the bounded two-worker ordinary GitHub experiment. The fix reverted only T2 ordinary scheduling to serialized execution and retained the four exact update jobs. The verification evidence is the already recorded local and GitHub matrix.

Stop Condition

T1-T6 delivery work, final mandatory review corrections, focused validation, clean structured autoreview, and debug assessment are complete. Bounded ordinary GitHub scheduling remains rejected by the measured safety gate; serialized CI and the scoped-operation seam remain the selected direction. T4's TDD status is partial/not-met. The parked Node publication limitation is not represented as fail-closed. Private/internal docs ingest updated docs/README.md, docs/runbooks/hi-cli-scaffolding.md, and the existing scaffold-parent-identity note at .agents/notes/2026-08-04-scaffold-parent-identity.md; routed Wiki projection was intentionally skipped by scope. Stack status showed an independent branch with no PR stack. stack sync --dry-run is unsupported; the default stack sync preview reported the dev stack current, and no apply was performed. PR #100 is ready for review with its body refreshed. Run 31067183663 is green for all in-scope CLI/update and deployment checks; Behavior Contract remains red only for the excluded Wiki projection requests. Delivery is complete with no in-scope blocker. The projection drift is an external excluded blocker, and no final branch head is asserted before the parent commit.

On this page