Issue 217: CI Cost and Test Signal
Issue 217: CI Cost and Test Signal
Status: evidence-grounded research, extended September 22 with current-code hotspot analysis. Accepted requirements are authoritative in the linked grill status; remaining recommendations are candidates, not measured improvements. No tests, workflows, runner settings, or release policy were changed. The September 22 deep dive below supersedes the original release-recovery build interpretation and reopens technical pruning decisions.
Final accepted scope — R5
The user accepted the consolidated proposals and explicitly removed the 80% target. Optimize actual execution and cache reuse as far as these identified inefficiencies permit, using only 2-vCPU workers and no test sharding. There is no fixed runtime ceiling, monthly minute cap/reserve or test-count deletion quota. Earlier shard experiments and target arithmetic in this historical report are superseded, not acceptance requirements.
The current authority is Issue 217: CI efficiency and test signal, compiled from closed Q1–Q11. The accepted scope covers the full inventory, updater preparation/install waste, minimal real recovery history with full replay on demand, narrower built-CLI proof, precise cache inputs, selected browser provisioning, duplicate root proof, isolated suite infrastructure reuse and wiki runner consolidation. Preserve all identified meaningful safety boundaries, PR/main verification and current release authority. Capability cache tasks remain within existing job topology; no shard fan-out is authorized.
Recommendation
Optimize useful confidence per normalized runner minute. The biggest opportunity is reducing repeated preparation and unnecessary invalidation in the serialized CLI suite. R5 excludes test sharding; bounded execution must respect the two-vCPU resource budget. Deleting registration/prose tests improves test quality but cannot, on present evidence, explain a severe runtime reduction. Sharding alone reduces feedback latency while duplicating setup and usually increasing consumption.
Model the system as a graph from owned invariant → cheapest sufficient proof → actual inputs and capabilities → isolated execution → complete decision. Each proof has one owner and one meaningful public result. Expensive execution should correspond to a real boundary that needs proving: installed executable, filesystem, database, browser, or publication. A test's historical issue number is neither a reason to keep it nor a reason to delete it.
Blacksmith normalizes allowance usage to x64 2-vCPU runners; historical 4-vCPU jobs consume two allowance minutes per wall-clock minute. The user subsequently required only 2-vCPU workers and explicitly rejected a monthly ceiling or reserve. Historical allowance arithmetic below is context, not an accepted budget requirement. Q10 supersedes Q8: no numerical speedup target and no sharding. Compare uncached feedback and total runner work; report warm cache-hit performance separately.
Evidence boundary and method
- Source examined: 7e147b3300c391db979b87e20c4159a5a03d12cb, clean initial working tree. Source paths and line numbers below refer to that revision.
- User request: issue 217, analyzed with Brainstorm, TDD/high-signal-tests, Make TSuite audit, Parallel Research, and Architect Pipeline guidance.
- Four independent readonly lanes: CLI test value; all other test value and selection; execution/caching topology; official Blacksmith/GitHub/Bun/Vitest/Turbo guidance. Coordinator inspected live Actions jobs and consolidated this report.
- Complete source inventory: 151 executable test files, excluding fixtures, snapshots, test support modules, and bundled third-party skill examples: CLI 68; API 29; backoffice 16; wiki 4; web 1; shared packages 21; root/operator tooling 12. This is file inventory, not current executed case count.
- Assertions, fixtures, selectors, and owned seams were inspected. Sensitivity is statically reasoned, not mutation-tested. The full test suite was not rerun; no fresh cold-cache benchmark was commissioned.
- Existing August research was consulted but not copied forward as current fact. That report describes 225 files and a different workflow. Current code already removed much of that topology.
Actual latency and consumption
Completed runs versus misleading samples
GitHub job timestamps were read through the Actions jobs API. Queuing and inter-job scheduling are excluded from individual job durations.
| Run | Affected job | All three jobs, summed | Evidence meaning |
|---|---|---|---|
| 35190271748 | 85 s | 121 s | Successful warm run; issue records 27/31 tasks cached and 2m03s workflow latency. |
| 35185839490 | 2,529 s | 2,566 s | Successful expensive CLI miss; issue records 28/31 tasks cached and 42m57s workflow latency. |
| 35197493495 | 2,800 s | 2,833 s | Successful main run; logs show 30/31 cached, yet Turbo takes 46m0.03s. Task-count hit percentage conceals cost. |
| 35182727045 | 2,657 s | 2,686 s | Failed run providing detailed test timing; not a passing benchmark. |
| 35605561374 | 14 s | 36 s | Fails at frozen-lockfile install; short failure is not a performance improvement. |
The failed detailed job 105078205569 reports 68 files, 723 cases, 2,612.83s total, including 2,600.01s in test bodies. Transform 1.18s, collection 6.48s, preparation 2.07s are negligible compared with test bodies.
| CLI file | Cases in that run | Time | Share of test-body time |
|---|---|---|---|
| update/run.test.ts | 60 | 1,301.638 s | 50.1% |
| scripts/release-recovery-publication.test.ts | 11 | 702.363 s | 27.0% |
| cli/behavioral-portfolio.test.ts | 76 | 415.508 s | 16.0% |
| Other 65 files | 576 | 180.501 s | 6.9% |
147/723 cases account for 93.1% of test-body time. The issue also reports the longest updater case at 83.652s; its 22-minute total is many cases, so capability-based file splitting can help. The two long recovery workers take 394.860s and 306.165s, creating a separate indivisible floor until their preparation is reduced.
Successful run output is suppressed by --output-logs=errors-only; the passing logs establish overall cost but do not expose fresh per-file durations. Do not present the failed run's breakdown as the exact current passing breakdown.
September usage estimate
Query: repository Actions runs created 2026-09-01 through 2026-09-21, retrieved September 21. There are 184 workflow runs: 146 Behavior Contract, 22 protected release, 13 cache witnesses, 3 external drift. The 146 Behavior Contract runs yielded 402 non-skipped Blacksmith jobs with timestamps: 401 on 4 vCPU and one on 2 vCPU.
| Behavior Contract category | Job wall minutes | Estimated normalized 2-vCPU minutes |
|---|---|---|
| Affected Verification | 2,396.58 | 4,793.17 |
| Root Static Verification | 37.40 | 74.80 |
| Stable Aggregate Check | 21.77 | 43.43 |
| Total | 2,455.75 | 4,911.40 |
| Jobs in failed workflow runs | 1,080.72 | 2,161.33 |
| Jobs in cancelled workflow runs | 580.48 | 1,160.97 |
| Jobs in successful workflow runs | 794.55 | 1,589.10 |
The result rows partition the same total; they are not additional consumption. Affected Verification accounts for 97.6%; failed/cancelled workflows account for 67.6%. Failure work is not automatically disposable: it can be valid bug detection. The useful question is how much is avoidable late failure, fixture contention, repeated setup, and superseded work.
This is a GitHub-timestamp estimate, not invoice or current Blacksmith usage. It applies documented CPU weighting, does not assume an undocumented billing-rounding rule, omits previous attempts for two rerun workflows, and excludes other repositories. No authenticated Blacksmith CLI was available. The official usage command below is the remaining financial authority.
Across 44 successful Affected Verification jobs, the median is 89s, but observed green jobs extend to 76m10s. These span revisions and cache states; they are not controlled benchmark repetitions. A blended median is a poor optimization objective. Report warm and forced-execution latency separately, including the tail.
Historical planning example, superseded as a requirement by Q2: at approximately 210 Behavior Contract runs/month, the full 3,000 allowance permits an average 14.3 normalized minutes/run; reserving 20% permits 11.4. The sampled mean is 33.6. The user rejected both a monthly design ceiling and reserve; neither is an acceptance gate.
The existing PR 221 right-sizes the aggregate runner and estimates $0.79/month savings. Even eliminating the entire sampled aggregate cost would recover less than 1% of the total. Its current failed run has not proven the success path on the smaller runner. PR 222 overlaps candidate validation/commit coverage; coordinate any later implementation with it.
Current system and candidate structure
Current graph, from .github/workflows/behavior-contract.yml and scripts/behavior-contract/run-ci-verification.mjs:
The graph already has affected selection, signed task caching, and cancellation of superseded PR work. These are existing capabilities, not new recommendations.
Candidate conceptual structure:
This is a model for task boundaries, not a proposal for another general-purpose proof registry. Reuse package scripts, Turbo, and existing evidence. Avoid test-title inventories or a second scheduler.
Evidence-backed opportunities
- Eliminate repeated irrelevant preparation. CLI update helper has 69 call sites and no default dependencyProcess replacement (apps/cli/src/update/run.test.ts:256–299). Some cases intentionally execute real installs; many protect other outcomes. Reuse immutable prepared scaffold/toolchain seeds with private copies for filesystem mutations. Keep actual fresh-install and lifecycle-script suppression cases fresh (:2403). Five authoring setup calls and four ordinary-refresh calls perform repeated full-update preparation (:1360). These are source counts, not measured runtime counts.
- Reduce historical replay, not merely its fixture self-test. September 22 correction: the fixture-only case took 1.226s; exact/resume took 701.025s. Readback-only skips retained publication artifacts but still reconstructs historical authority through isolated installs and CLI builds. The deeper trace below derives approximately 89 such builds across the two scenarios. Keep exact binding, resume, tamper, first-parent and provider readback invariants; Q9 accepts minimal real history in CI and full historical replay only on demand.
- Separate semantic breadth from production composition. behavioral-portfolio.test.ts contains 25 scaffold call sites and repeated built-process execution (:170). The full issue-215 benchmark runs inside it (:350), with real warm-update, readiness, warnings, and no-write assertions. It is valuable behavior, not a disposable stopwatch. The macOS witness invokes restricted scenarios on a different filesystem; it is not proven duplicate Linux evidence. Move broad combinations to nearer owned seams only with an invariant map, keeping representative built CLI paths and platform-specific safety.
- Make cache identity match actual capabilities. cache-identity.mjs:162–181 builds an all-host digest including browser, Docker, and executable identities. turbo.json:149 and :291–300 feed it to CLI types/tests. A browser change can invalidate nonbrowser proof. Partition by required capability only after proving the input closure; retain relevant runtime/platform/tool identities, trust authority, source dependencies, and deterministic environment.
- Audit broad build/task inputs. turbo.json:34–48 uses whole-workspace defaults for CLI build. A test edit can invalidate build; unrelated CLI content can invalidate the one coarse test task. Global inputs also include lint/format config. Exclude irrelevant inputs only after checking build copying and generation; do not trade correctness for attractive cache-hit counts. Existing CLI tests correctly depend on the CLI build (:292).
- Provision browser only for selected browser work. Chromium install is unconditional (workflow:62) and costs 17–18s in sampled jobs. Resolve selection before provision and coordinate host identity. Native Playwright guidance warns that browser cache restoration may cost as much as download; measure before adding another cache.
- Give root cache-policy proof one owner. run-ci-verification.mjs:101 and :124 request it in both static and affected jobs. Static lacks shared-cache setup and differs in environment/capabilities. Remove duplicated execution only while preserving the required proof. Keeping cheap static feedback parallel may still be worthwhile: observed static failures did not create material long-running waste in this sample.
- Reuse one PostgreSQL server within a suite. packages/db/src/baseline-authority-postgres.test.ts starts/removes four independent containers (:78, :149, :224, :295). Use separate databases and migration state on one suite-owned server if equivalent isolation is demonstrable. Do not delete migration/immutability/ledger scenarios. Cross-package server sharing is a larger ownership/concurrency decision.
- Consolidate wiki test runners. Four wiki files currently start Node, Bun, and Vitest sequentially. One owning runner can retain synchronization, projection, navigation, and route behavior with less ceremony. Probably a small saving beside CLI hotspots.
- Keep execution bounded without sharding. Organize proof by capability for ownership/cache reuse inside existing verification jobs. R5 rejects the original shard proposal. Historical subprocess starvation explains fileParallelism:false; adding parallelism is not a substitute for removing repeated installs/builds/setup. Monitor total process pressure on two CPUs.
Merely putting the three hot files on separate machines still leaves the 21m42s updater file as the lower bound before setup. Five equally balanced lanes over 43.5 serial minutes have an ideal 8.7-minute average, but setup, imbalance, and the 6.6-minute recovery case remain. This was historical scheduling arithmetic. R5 supersedes all numerical targets and shard experiments; it requires less actual work within existing job topology.
What to remove and what to protect
TDD's criterion is durable owned behavior through the nearest public seam. Do not replace real integration proof with mocks of the same behavior. Do not confuse runtime validation of untrusted data with static typing; owned public compatibility snapshots can be valuable.
Strong full-file removal candidates: three CLI catalog-registration files; CLI improve-mobile-frontend and verification-recovery-skills content tests; current SHA/tag snapshot in sync-skills-repo; release-guidance prose/history locks; API fake-service self-test; backoffice literal cache tags; parked web placeholder/source-copy tests. Partial culls include fixture self-validation, help/registration accounting, source transaction counting, byte-exact Philosophy prose, historical wiki existence/ingest flags, and UI class names. Backoffice auth-forwarding and cosmetic date formatting are medium-confidence candidates requiring consumer/UX confirmation.
Keep especially: user-authored-file preservation; no writes outside owned roots; partial failure/retry truthfulness; no lifecycle-script execution; no credential disclosure; live readiness despite cache hits; artifact tamper/authority rejection; exact historical release and resume; real database constraints and migrations; authentication/session lifecycle; critical browser failure recovery. These buy substantially more confidence than test count alone indicates.
For a merge/move/removal that protects behavior, record actor, authority, input class, seam, observable outcome, side effects, isolation, and retained proof's required gate. Static similarity is insufficient. For no-owned-invariant cases, state that explicitly. New behavior changes require actual RED before production edits; fixture refactoring must preserve semantic fingerprints. Full-file inventory follows.
Complete test inventory
All line references below point to the examined revision. Retain means meaningful evidence exists, not proof of zero duplication. Strengthen/move/merge preserve material behavior until equivalence is demonstrated. Dispositions and sensitivity are static candidates; no deletion or runtime reduction has been verified.
CLI: 68 files
Paths relative to apps/cli/src. All selected by Vitest src/**/*.test.ts through the CLI test task when affected. Generated dist copies are not selected. The old run.test-cases.test.ts exclusion is present in configuration but no such tracked source file exists at this revision.
| File | Candidate | Owned invariant / evidence |
|---|---|---|
| cli/behavioral-portfolio.test.ts | Strengthen | Built commands and filesystem outcomes; generic any-status smoke :233–249 is weak; retain failure/no-write/readiness :272–354 and benchmark :350. |
| cli/commit-gate-command.test.ts | Retain | Observed health before persisted receipt, no receipt on validation failure :82–109. |
| cli/ensure-command.test.ts | Retain | Rejection preserves bytes; migration idempotent; disabled policy persists :59–125. |
| cli/safety-invariants.test.ts | Retain | Confinement, no partial state, stream behavior, secret redaction :83–154. |
| cli/scaffold-subagent-alignment.test.ts | Retain | Generated roles preserve custom roles and planned specialist skills :138–178. |
| cli/update-command.test.ts | Strengthen | Keep JSON/progress/cache-clear; cull option existence :163–165 when behavior retained. |
| cli/update-presenter.test.ts | Merge | Preserve completedFiles/legacy changedFiles meaning in command-output proof :23–42. |
| content/finder-entrypoints-catalog-prompts.test.ts | Strengthen | Cull registration :23; preserve explicit-human-invocation guidance boundary :39. |
| content/handback.test.ts | Strengthen | Keep bounded handoff behavior :12–39; cull byte-exact prose, catalog, source routing locks. |
| content/improve-mobile-frontend.test.ts | Remove | Registry fields/membership only :8–14; no owned execution result. |
| content/scaffold-copy.project-verifier.test.ts | Merge | Preserve authored verifier and legacy handoff semantics :31–40 with generated-output proof. |
| content/verification-recovery-skills.test.ts | Remove | Registry paths/pack membership only :8–19. |
| content/wiki.test.ts | Strengthen | Config substring :6 should prove actual internal-file exclusion. |
| data/catalog/architect-pipeline-registration.test.ts | Remove | Exact catalog/membership only :8–23. |
| data/catalog/make-tsuite-registration.test.ts | Remove | Registry and extracted manifest registration only :30–51. |
| data/catalog/verification-registration.test.ts | Remove | Pack placement/order only :15–24. |
| data/hooks/format-edited-file.test.ts | Retain | Real shipped hook fixes, previews, diagnostics, failure and isolation :124, :402. |
| data/scripts/commit-gate-runner.test.ts | Retain | Real staged Git paths, worktrees, deleted files, owner cwd and aggregate failures. |
| data/scripts/harness-projection/authoring.test.ts | Retain | Allowed authoring scope depends on current proof, custom roles preserved :36–79. |
| data/scripts/harness-projection/core.test.ts | Strengthen | Gate contributions must produce no harness writes; prefer result over adapter-call check :15–60. |
| data/scripts/harness-projection/lint-feedback.test.ts | Retain | Owned finding/failure projection, clean silence, exhausted repair requests. |
| data/scripts/sync-subagents-validator.test.ts | Merge | Carry rejected input through delivered script to prove no projection. |
| features/commit-gate/index.test.ts | Retain | Unresolved until observed proof; preserve foreign hooks; invalidate changed config :49–148. |
| features/commit-gate/quality.test.ts | Retain | Command ownership and mutating formatter refusal across spellings :13–181. |
| features/commit-gate/setup.test.ts | Retain | Rootless/independent setup, live executables and hook integrity :154, :208, :461. |
| features/context-planning/compiler.test.ts | Retain | Portable identity, ownership, partial outcomes and typed incomplete contracts :77, :574. |
| features/operator-skill/lifecycle-target-delegation.test.ts | Retain | Stale authority cannot retry/install; preserve uncovered bindings and cancellation :446. |
| features/repository-check/application.test.ts | Retain | Independent health/drift survives dependency failure; dependent mutation blocked :91, :478. |
| features/scaffold-state/generation-inputs.test.ts | Retain | Filesystem input closure, hashes, unsupported files and symlink escape :100. |
| features/scaffold-state/portability.test.ts | Retain | Relocation preserves bytes but rechecks live hooks :196–199. |
| features/scaffold-state/reconcile.test.ts | Strengthen | Keep edit/deletion invalidation; replace reference identity/intermediate shapes :94, :262. |
| features/scaffold-update/validation-cache/cache.test.ts | Retain | Promotion, invalidation, bypass/failure no publication, optional storage :39, :162. |
| integrations/repository-detector.test.ts | Retain | Real manifests map to owned package-manager authority :63, :175. |
| integrations/skills-cli-target-delegation.test.ts | Retain | Fresh source authority, refusal before install, prompt preservation :422. |
| integrations/tool-management.test.ts | Retain | Nonmutating readiness, actual version evidence, redacted failures :75–83. |
| integrations/update-cache-filesystem.test.ts | Retain | Atomic process publication, corruption/recovery, private restore, confinement :63, :209. |
| integrations/update-cache-remote.test.ts | Retain | Owned signing/redirect/archive/outage policy and refusal before request :88, :603. |
| platform/commit-gate-capabilities.test.ts | Retain | Live hook/dependency/lock authority, foreign hook ownership, no false receipt :465. |
| platform/feature-application-operations.test.ts | Retain | Missing live tools reject cached validation, local unauthenticated cache :41–85. |
| platform/scoped-scaffold-operation.test.ts | Retain | Tool install suppresses lifecycle scripts; partial failure, opt-out, confinement :88–156. |
| runtime/config.test.ts | Retain | Secret/environment allowlist excludes unrelated variables :47. |
| runtime/scripts.test.ts | Retain/split | Changed-input refusal, candidate isolation, cleanup/cache/env/process failures :143, :2905. |
| runtime/validation-candidate.test.ts | Retain | Special/escaping files refused, contained aliases allowed :233. |
| scaffold/confined-root-alias.test.ts | Retain | Retarget/escape refusal leaves external bytes untouched :151–160. |
| scaffold/generated-gate-scoping.test.ts | Retain | Actual generated linter rejects staged invalid, ignores unstaged invalid :134–148. |
| scaffold/output-receipt-validation.test.ts | Retain | Malformed/failure receipt rejected before adoption :151; owned output :36. |
| scaffold/output-root-dependencies.test.ts | Retain | Planned/applied package agreement and conflicting ownership refusal :76–186. |
| scaffold/output-root-materialization.test.ts | Retain | Rootless placement and conflicts preserve inventory :263–272. |
| scaffold/output-wiki-plugin-alias.test.ts | Retain | Actual generated linter accepts owned alias and emits rule diagnostics :129. |
| scaffold/output.test.ts | Strengthen | Keep hook/role/adoption safety :928; replace object-identity memoization :57. |
| scaffold/project-verifier-preservation.test.ts | Retain | Arbitrary authored verifier bytes/absence survive generated updates :236–270. |
| scaffold/settings-selection.test.ts | Retain | Supported authority, read-only planning, independent output when quality unresolved :404. |
| scripts/promote-baseline-authority.test.ts | Retain | Redirect refusal and exact historical readback after stable advances :70, :188. |
| scripts/release-authority.test.ts | Retain | Wrong tree/squash source, duplicate and expired evidence refused :189, :486. |
| scripts/release-candidate.test.ts | Merge | Decoder roundtrip :9 should flow through actual authority acceptance. |
| scripts/release-changelog-selection.test.ts | Retain | Changed changelogs alone select baseline/npm/mixed/none :62–133. |
| scripts/release-convergence.test.ts | Retain | Stable product order, no-product no calls, immutable mismatch refusal :98–186. |
| scripts/release-dispatcher.test.ts | Retain | Preflight conflict/ambiguity refusal, prerelease metadata and archive cleanup :44–185. |
| scripts/release-guidance-contract.test.ts | Remove | Five prose/source/history snapshots :14–77; no executed release invariant. |
| scripts/release-impact-freshness.test.ts | Retain | Old unchanged notes cannot authorize later changed payload :35. |
| scripts/release-publication.test.ts | Retain/split | Tamper/exact identity, credentials, safe extraction, idempotency and readback :1650–1683. |
| scripts/release-recovery-publication.test.ts | Strengthen | Cull fixture-only :151; preserve exact/restart/tamper production proofs :294–340. |
| scripts/release-recovery.test.ts | Retain | First-parent history, notes/tree/version authority and strict recovery declaration :78–233. |
| scripts/release-verification-evidence.test.ts | Merge | Preserve forbidden release-state transport :9 in actual authority decision. |
| scripts/sync-skills-repo.test.ts | Remove | Current hardcoded source SHA/tag :10–13 only. |
| update/planned-wiki.test.ts | Retain | Read-only authored/managed/missing wiki states and authority :195–283. |
| update/run.test.ts | Strengthen/split | Preserve transaction/ownership/retry/no-write/install invariants; reduce repeated preparation. |
| update/wiki-alignment-owner.test.ts | Retain | Full update preserves authored/root rules and records truthful owner receipt :131–151. |
API: 29 files
Paths relative to apps/api/src. Workspace Vitest selects these when API is affected.
| File | Candidate | Owned invariant / evidence |
|---|---|---|
| auth-lifecycle-contract.test.ts | Merge | Resource getters' lazy init, absent config, races and cleanup :33; preserve auth-specific cases. |
| backoffice-access.test.ts | Retain | Invite activation/revocation/reinvite persists :155. |
| backoffice.test.ts | Retain | Triage, usage, remapping, historical identity and typed outages :186. |
| baseline-registry.test.ts | Retain | Durable/exact authority, digest/provenance integrity, credential-safe redirects :168. |
| baseline-release-inventory.test.ts | Retain | Immutable candidates, safe pagination and typed failures :53. |
| features/baseline-delivery/baseline-delivery.test.ts | Retain | Exact version and authority before download :39. |
| features/baseline-promotion/baseline-promotion.test.ts | Retain | CAS, idempotency, immutable evidence, rollback authority/history :45. |
| features/public-domain-root-contract.test.ts | Remove | Constructs and calls its own fake service :18; no owned implementation proof. |
| features/report-submission/report-submission.test.ts | Retain | Persist before provider, claim release, fallback metadata and duplicate policy :21. |
| features/typed-failure-contract.test.ts | Retain | Real submission preserves repository/provider failure context :21. |
| http-adapter-contract.test.ts | Retain | Credential-safe transport/pagination and interruption :288. |
| index.test.ts | Merge | Overlapping HTTP happy paths :559; retain distinct failures with public API owner. |
| integrations/baseline-artifact/baseline-artifact.test.ts | Retain | Verified bytes, redirect/auth/size bounds and interruption :44. |
| integrations/persistence/baseline-promotion-drizzle.test.ts | Retain | Durable transaction ambiguity and reconciliation :13. |
| integrations/persistence/report-repository.test.ts | Strengthen | Remove source transaction count :51; keep stable identity/recovery. |
| integrations/reports/provider-adapters.test.ts | Retain | Owned input/auth/interruption/failure mapping :35. |
| platform/http/baseline-download.test.ts | Retain | Binary response, typed invalid path/version and dispatch :27. |
| platform/http/baseline-promotion.test.ts | Retain | Publisher bearer boundary and failure mapping :62. |
| platform/runtime/lifecycle.test.ts | Merge | Preserve cleanup ordering, failures, late init disposal :7 in composition proof. |
| platform/runtime/resources.test.ts | Merge | Preserve concurrent acquisition/retry/terminal disposal :31 with lifecycle tests. |
| production-application-lifecycle-contract.test.ts | Retain | Application disposal before resources, failure aggregation and blocked init :7. |
| production-shutdown-contract.test.ts | Retain | Real process exactly-once shutdown, import inertness :82. |
| public-api-contract.test.ts | Retain | Owned HTTP compatibility, trusted identity, anonymous paths/outages/security :209. |
| reports.test.ts | Retain/split | Delivery recovery, canonical identity, duplicate/concurrent reports :552; transaction behavior :3348. |
| runtime/config-contract.test.ts | Retain | Runtime config rejection and secret redaction :41. |
| runtime/baseline-ledger-cutover.test.ts | Strengthen | Keep report/apply/retry/readback/token safety :134; cull discovery-only :213. |
| runtime/baseline-ledger-no-redeploy.runtime.test.ts | Retain | Real DB/runtime reconstruction maintains one linear authority :101. |
| runtime-product-contract.test.ts | Retain | Real process/DB/email/report/auth/invitation product flow :618. |
| telemetry.test.ts | Retain | Idempotency, trusted identity, monotonic activity and historical compatibility :208. |
Backoffice: 16 files
Paths relative to apps/backoffice. Vitest selects all except e2e, which belongs to test:browser. Playwright uses one worker and eight cases; server startup is a likely larger concern than browser test sharding.
| File | Candidate | Owned invariant / evidence |
|---|---|---|
| e2e/public-routes.test.ts | Retain | OTP/recovery, operator navigation, errors and persisted invitation :104; move output-dir assertion :93. |
| src/app/api/auth/[...path]/route.test.ts | Remove candidate | Mock forwarding/call count :25; first confirm real endpoint/auth proof. |
| src/features/operator-shell/shell.test.ts | Strengthen | Prove actual client serialization/signed-out outcome beyond prototype checks :7. |
| src/features/operator-workflows/actions.test.ts | Strengthen | Keep input/failure :58; visible freshness instead of exact invalidation-call lists. |
| src/modules/auth/login-path.test.ts | Retain | Reject external and recursive redirect destinations :6. |
| src/modules/auth/server.test.ts | Retain | Failed auth construction cleans DB, disposal idempotent :64. |
| src/modules/control-plane/cache-tags.test.ts | Remove | Literal helper strings :6, no consumer cache outcome. |
| src/modules/control-plane/client.test.ts | Retain | Cookie forwarding, failure mapping and transient refetch :33. |
| src/modules/presentation/date-time.test.ts | Remove candidate | Cosmetic/library formatting :6; verify timezone not critical workflow contract. |
| src/modules/runtime/config.test.ts | Retain | Injected config/defaults and redaction :34. |
| src/modules/runtime/shutdown.test.ts | Retain | Child process cleanup before termination :16. |
| src/public-routes.test.tsx | Strengthen | Keep auth/unavailable versus empty/mutations :76; cull cosmetic markers. |
| test-fixtures/browser/cleanup-order.test.ts | Retain | Exact owned process identity before signals/escalation :34. |
| test-fixtures/browser/process-capture.test.ts | Retain | Detached process identity and Portless capture :25. |
| test-fixtures/browser/readiness.test.ts | Strengthen | Real runner must await readiness; appended fake spawn event is insufficient :13. |
| test-fixtures/browser/workspace-bin.test.ts | Retain | Real filesystem resolves root-hoisted binary :11. |
Wiki and web: five files
| File | Candidate | Owned invariant / selection |
|---|---|---|
| apps/web/src/public-routes.test.tsx | Remove candidate | Parked placeholder and source title :26; explicitly permit zero behavior tests if accepted. |
| apps/wiki/scripts/source-page-tree.test.mjs | Merge | Two-route existence :13 into general navigation; explicitly Bun-selected. |
| apps/wiki/scripts/sync-content.test.mjs | Retain | Real projection preservation/pruning/read-only mode :14; cull prose/ingest flag lock :78; Node-selected. |
| apps/wiki/scripts/public-wiki-contract.test.ts | Merge | Real sync/parity/determinism :175; cull historical artifact matrix :204; Vitest-selected. |
| apps/wiki/src/public-routes.test.tsx | Retain | Redirects/docs/Markdown/LLM/missing route/OG behavior :182; assess mocked search echo; Vitest-selected. |
Shared packages: 21 files
Paths relative to packages. Workspace test task selects each when affected; config uses Bun, others Vitest.
| File | Candidate | Owned invariant / evidence |
|---|---|---|
| auth/src/auth-email-contract.test.ts | Retain | Real auth/DB/email/cookie lifecycle :449; cull source self-inspection :412 and assess provider replica. |
| auth/src/composition-contract.test.ts | Retain | Layer/config composition, release pool on failure :120. |
| auth/src/config-contract.test.ts | Retain | Decode/redaction/injected construction :67; source import boundary :152 is a separate architecture rule. |
| auth/src/index.test.ts | Merge | Keep owned OTP/cookie/Vercel policy :71; drop options-only duplication after real-session mapping. |
| auth/src/public-domain-root-contract.test.ts | Retain | Real owned email content and adapter mapping :10. |
| auth/src/typed-failure-contract.test.ts | Retain | Owned auth/persistence/email failure normalization :19. |
| config/public-config-contract.test.ts | Strengthen | Keep real compiler consumption :21; cull issue-181 literal config lock :59. |
| contract/src/api.test.ts | Merge | Runtime validation negatives :29; consolidate overlapping protocol routes/status cases. |
| contract/src/public-protocol-contract.test.ts | Retain | Owned public protocol snapshot/semantic canonicalization :251; compatibility exception. |
| db/src/baseline-authority-postgres.test.ts | Retain | Real migrations/immutable releases/rollback ledger :74; reuse server with isolated databases. |
| db/src/composition-contract.test.ts | Retain | Injected layer releases pool after connection rejection :34. |
| db/src/config-contract.test.ts | Retain | Validated explicit URL, secret-redacted startup failure :16. |
| db/src/index.test.ts | Retain | Explicit URL required despite ambient config :8. |
| db/src/postgres-contract.test.ts | Strengthen | Keep real DB constraints/migrations :329; assess removing deterministic provider imitation. |
| db/src/typed-failure-contract.test.ts | Retain | Owned connection/query/transaction failure mapping :23. |
| env/src/public-env-contract.test.ts | Retain | Runtime input/defaults/allowlist/redaction :124. |
| scaffold/src/context-plan.test.ts | Merge | Preserve runtime invalid scope/variant cases :64 in compatibility owner. |
| scaffold/src/harness-capability.test.ts | Merge | Preserve missing family/contradiction/failure semantics :18. |
| scaffold/src/index.test.ts | Merge | Keep v1/v2 negatives/compatibility :22; consolidate repetitive positives. |
| scaffold/src/public-scaffold-contract.test.ts | Retain | Runtime fixtures, legacy receipts, invalid variants :409; do not substitute TypeScript. |
| ui/src/public-ui-contract.test.tsx | Strengthen | Cull CSS/data-variant locks :39; retain/prove critical accessibility and interaction. |
Root and operator tooling: 12 files
R = ordinary root test:cache-policy; CW = separate cache witness workflow; U = no normal package/Turbo/workflow selector found. U does not mean useless or safe to delete. Deleting U files saves zero normal CI runtime.
| File | Gate | Candidate | Owned invariant / evidence |
|---|---|---|---|
| scripts/behavior-contract/affected-verification.test.ts | R | Strengthen | Preserve actual task routing :77; replace literal workflow recipe locks :36. |
| scripts/behavior-contract/cache-identity.test.ts | R | Retain | Real Turbo hashing/restoration across runtime/env/host contexts :119. |
| scripts/behavior-contract/cache-trust.test.ts | R | Retain | Credential export/redaction/signing and privilege boundaries :62. |
| scripts/behavior-contract/shared-cache-witness.test.ts | CW | Retain | Actual transfer attribution and fail-closed evidence :33. |
| scripts/behavior-contract/consumer-repositories.test.ts | U | Retain, decide gate | Installed-consumer provenance/cleanup/package authority :26. |
| scripts/behavior-contract/cutover-repeat.test.ts | U | Retain, decide gate | Repeat evidence, concurrency/cleanup/mutation detection :32. |
| scripts/behavior-contract/isolation.test.ts | U | Retain, decide gate | Resource ownership/symlink/race/signal safety and atomic manifests :47. |
| scripts/behavior-contract/packaged-runtime-product.test.ts | U | Retain, decide gate | Rebuild stale artifacts and reject stale source evidence :15. |
| scripts/behavior-contract/process-identity.test.ts | U | Retain, decide gate | GNU/BSD process identity, required fields and locale :7. |
| scripts/behavior-contract/validate.test.ts | U | Retain if supported | Reject invalid ownership/proof/waivers :10. |
| scripts/external-drift/npm-trusted-publisher-oidc-exchange.test.ts | U | Retain, decide gate | Trusted event/package token, masked environment/failures :10. |
| scripts/install-git-hooks.test.ts | U | Retain, decide gate | Idempotent local hook install and worktree delegation :25. |
Official runner guidance and cost experiments
Sources were read on 2026-09-21; marketing speed claims are not repository benchmarks.
| Source | Verified guidance | Application here |
|---|---|---|
| Blacksmith runner overview | 3,000 x64 2-vCPU minutes per organization; CPU-proportional consumption; ARM factor 0.625; no provider concurrency cap. | 4→8 vCPU must halve runtime for equal consumption. 4→2 saves only if slowdown is less than 2x. |
| Blacksmith metrics | CPU, memory, network and step timelines available. | Inspect the hot job before choosing workers/runner size; no speculative instrumentation project needed first. |
| Blacksmith native cache | Native Actions cache transparently accelerated with no additional charge; old Blacksmith forks archived; seven-day unused-entry eviction. | Use actions/cache, not outdated vendor fork advice. Task cache already exists; inspect misses first. |
| Blacksmith sticky disks | $0.50/GB/month; isolated clones of prior snapshot; default repository-wide commits, optional protected writes. | Not a zero-spend default. Do not confuse snapshots with shared live mutable disks. |
| setup-bun and Bun cache | setup-bun caches the executable; dependencies live in Bun install cache; Linux can hardlink cached packages. | Native dependency cache may help nested installs, but top-level install is already only a few seconds on sampled successful Blacksmith jobs. Measure before adding complexity. |
| Vitest v3 config and performance | fileParallelism:false forces one worker; shards distribute files, not cases. | Split giant capability groups first; two-worker script exists, but contention previously caused starvation. |
| Turbo caching | Correct task identity depends on actual declared inputs and outputs. | Independent capability tasks improve selective invalidation; no unsafe input exclusions. |
| GitHub concurrency | cancel-in-progress cancels running superseded work. | Already present for PRs. Main groups by SHA and does not cancel older main runs; changing this requires deciding which evidence must complete. |
| Blacksmith usage CLI | Usage can be broken down by runner, repository, workflow and job. | Use authenticated usage readback to replace timestamp estimates; do not assume GitHub billing rounding. |
| Playwright CI | Browser cache restoration may cost as much as download, and OS dependencies still need installation. | Prefer selecting required browser work before provisioning; benchmark cache alternatives. |
Cost model without assuming rounding:
normalized minutes = sum(job wall minutes × vCPU / 2 × architecture factor)
architecture factor: x64 = 1; ARM = 0.625
sharded work = useful work + repeated setup + transfer + aggregation
latency = longest dependency path + queue/scheduling timeFor illustration, four 11-minute 4-vCPU shards consume 88 normalized minutes before extra overhead, comparable to one 44-minute 4-vCPU job. They deliver feedback sooner but do not solve allowance pressure. Reducing work to four 6-minute shards would consume 48; the saving comes from less work, not the number of shards.
ARM may be cheaper if the same-size job takes less than 1.6x the x64 time. Native binaries, package-manager behavior, cache partitioning and required platform coverage must be verified before migration. Do not infer acceptance from public pricing alone.
Release and cache witnesses use GitHub-hosted runners, including macOS; they do not explain Blacksmith minutes. Moving them to Blacksmith would move consumption into this budget. Protected historical build/apply separation remains intentional; current publication eligibility does not require a prior PR Candidate Evidence lookup. Preserve actual current release rules rather than reintroducing historical assumptions.
Candidate experiment sequence and acceptance
- Restore a valid benchmark starting state: the latest hosted failure is frozen-lockfile mismatch, not test speed. Coordinate with open work; do not weaken --frozen-lockfile.
- Capture passing per-file/case durations and phase costs on one fixed revision/toolchain/runner: setup/copy/Git/install/build/command/cleanup, CPU/memory and subprocess pressure. Reuse existing runner metrics. A cached timing artifact must retain its original execution provenance and be labelled reused.
- Cull no-owned-invariant checks and consolidate redundant runner/setup work. For other changes, record semantic-lock mappings before edits. Confirm real cold-install, native filesystem, auth, database and release witnesses remain selected.
- Introduce capability-level execution/cache boundaries and immutable fixture seeds, each locally runnable. Validate task identity with forced execution, output deletion/restoration, relevant-input misses, irrelevant-input reuse, and live readiness where required.
- Use historical 4-vCPU results as context and measure comparable 2-vCPU configurations, as required by Q1. Do not add CLI shards; capability groups remain inside existing verification jobs with private mutable resources. Report both critical-path latency and summed runner work, including setup and retries; there is no accepted monthly budget gate.
- Only then consider more shards or ARM. Retain all expected shard identities, exactly-once selected-file assignment, signed output provenance, and aggregate fail-closed behavior before emitting tested-tree evidence. Do not shard the entire test:ci graph and rebuild everything on every runner.
Measure three distinct modes: task-result reads bypassed with warm dependency stores; clean dependency stores and task caches; normal warm reuse. Benchmark runs must not seed trusted cache evidence under misleading identities. Use several completed green repetitions; include failed/cancelled consumption separately and compare at the same revision. Initial repetitions give distributions, not a statistically robust p95; collect p95 over sustained operation.
Suggested acceptance candidates: warm feedback remains around the current two-minute workflow range; expensive misses improve through removed work, without a fixed percentage target; total normalized usage per comparable change falls; fixture failures and retries do not increase; complete proof remains required. Q10 removes the numerical target and sharding; latency and total work remain separate observations. No improvement is proven by this report.
September 22 technical deep dive
The user reopened requirements to understand concrete test types and pruning, with an aspirational 80% CI-time reduction. Source grounding is now child HEAD 9430c62496aed8b5f9a94c399d877b9f1aa7a39d, above PR #222 head cef8e31e. Three readonly research lanes traced updater tests, release reconstruction, and built CLI journeys. Their findings are consolidated here; no expensive tests or hosted workflows were launched.
The current inventory is 153 executable files, including 70 CLI files. The parent added features/project-settings/service.test.ts and features/scaffold-update/validation-plan.test.ts. The original inventory and timing sample remain historical: 151 total files, 68 CLI. New code anchors below refer to the current child; historical durations do not become current measurements.
Test value and execution cost are different axes
| Test kind | Concrete evidence | Pruning boundary |
|---|---|---|
| Registration, source text and prose pins | Seven obvious CLI full-file culls total 20ms in the historical sample | Delete when no owned capability is protected. Valuable cleanup, negligible latency saving. |
| Owned decisions and runtime compatibility | Context compiler: 21 cases / 208ms; commit-gate setup: 13 / 348ms; persisted invalid-input and legacy receipt cases | Keep cheap breadth through exported capabilities. Runtime decoding is not redundant TypeScript checking. |
| Application transactions with real files | update/run.test.ts: 60 / 1301.638s | Keep confinement, ownership, failure recovery and truthful receipt outcomes; remove irrelevant real installs and repeated setup. |
| Real adapters and installation | Script suppression, actual tool executability, hook installation/removal, subprocess cleanup | Retain representative real-boundary proof. A fake installer cannot prove installer safety. |
| Built CLI composition | cli/behavioral-portfolio.test.ts: 76 / 415.508s | Keep dispatch, output/exit semantics, one scaffold/update journey, cache wiring and failure/no-mutation witnesses. Move policy permutations down. |
| Historical release reconstruction | Recovery publication: 11 / 702.363s, almost entirely two cases | Preserve release protocol; decide explicitly whether routine CI must replay the old full monorepo builders. |
Deleting hundreds of cheap checks while preserving the same repeated installations would barely change latency. Conversely, a test may assert an important invariant at a needlessly expensive execution boundary. Such a test needs a cheaper proof, not an indiscriminate deletion.
Updater: repeated application preparation and real installation
Historical slow-case groups within update/run.test.ts:
| Group | Reported slow cases | Time |
|---|---|---|
| Scoped authoring receipt recovery | 16 | 679.974s |
| Candidate lint preview | 13 | 172.498s |
| Managed shared baseline prompt | 5 | 147.582s |
| Routine local quality installation | 6 | 142.077s |
| Converged projection proof recovery | 3 | 129.556s |
These are sums of explicitly printed slow cases, not the complete group case counts. Overall 47 printed slow updater cases explain 1298.719s of its 1301.638s. The largest authored-wiki/starter case took 83.652s.
test fixture
preparatory updates and authoring completion
runUpdateWithPreview
detect repository and compile desired context
materialize scaffold and run projection subprocess
compare state and plan reconciliation
injected lint preview (usually already fake)
apply each dependency action
write manifest
REAL package-manager install unless dependencyProcess supplied
restore manifest/locks on failure
completeLocalQualitySetup before receipt persistence
installation/tool-version/hook work as needed
verify outputs and publish receipt
further update/check calls for repair, refusal or repeatAnchors: apps/cli/src/update/run.test.ts:256 (runUpdateWithPreview), :231 (cleanPreview); update/run.ts:2328 (default process), :2420–2478 (per-dependency apply/install), :6477–6516 (quality setup); scaffold/output.ts:4536 (projection subprocess). A fake general tool runtime does not replace the separate dependency process. cleanPreview does not run candidate installation/lint and does not prevent live application installs.
Static repetition examples, not measured subprocess counts:
- Each of three receipt-recovery variants at
run.test.ts:950traverses five full updates and one preview. - Each current-handoff variant at
:1629traverses three updates and two previews. - The ready-validation fixture at
:309performs two preparatory updates plus per-scope authoring completion. It already substitutes installation at:327, so its remaining cost is not solved by removing npm. - The authored-wiki/starter mega-case at
:1925traverses twelve updates including preparation. Several calls merely classify one file state. - The real installation journey at
:2635traverses five updates; cold install and lifecycle-script suppression genuinely need the installer boundary, while follow-up policy breadth may use the owning setup capability.
Recommended seam map:
| Current assertion | Cheapest sufficient owner / retained integration |
|---|---|
| Which consumer/root must be validated | planUpdateValidation; minimal real-manifest discovery witness for optional dependencies |
| Catalog/manifest ownership and preservation | compareObservedState, scaffold reconciliation and real file application |
| Context identity and proof-to-action decisions | compileScaffoldContextPlan, assessScopedAuthoring; retain actual receipt persistence/race witness |
| Update response to installation success/failure | Real runUpdate transaction with controlled dependencyProcess; assert rollback, diagnostics and withheld proof |
| Hook/tool installation and script suppression | Real completeLocalQualitySetup boundary plus one full update cold-install witness |
| Full-tree convergence and authored-file preservation | A small number of real update/check/repeat journeys from independently copied prepared seeds |
Pure culls include plan-object identity assertions (:509) and a nominal explicit-check preservation test (:491) that only checks no issues/no application without proving commands survived. The latter can instead gain an actual compiled-command assertion if unique update wiring warrants one. Before/after receipt interruption, changed live inputs before publication, symlink/FIFO refusal and independent partial work are distinct safety invariants, not duplicate regressions.
Use controlled installed-state fixtures, not an unconditional no-op that invents healthy tools: subsequent checks may inspect executable versions, lock evidence and tool metadata. Keep private mutable roots. Copying a generated seed requires portability checks for absolute paths, hooks and symlinks; sharing writable roots or copying giant installed trees blindly is not the recommendation.
Some cheaper seams do not yet exist. Flat-route/hidden-placeholder behavior lives in private wikiSemanticState (run.ts:1710–1777). Supplying precomputed obligations to compareObservedState would bypass that behavior. Retain a read-only application witness or separately accept a coherent production refactor; do not call bypassed coverage equivalent.
Release recovery: a nested build farm inside two tests
Historical case durations: exact recovery/tamper 394.860s; partial/completed resume 306.165s; fixture self-test 1.226s. Removing only the fixture self-test cannot materially improve this hotspot.
recovery test → child runner → real historical Git fixture
createReleasePublicationPlan
readReleaseRecovery → reconstruct baseline target authority
each first-parent entry, sequentially
reconstruct base authority
reconstruct target authority
classify changed changelogs
selected non-readback product → build target again
validateReleasePublicationPlan
verify history, notes and artifacts
readback entries → reconstruct historical authority againEach default authority reconstruction creates a detached worktree, runs bun install --frozen-lockfile --ignore-scripts, builds that revision's CLI and baseline, assembles artifacts, calculates authority and deletes the worktree/output. Adjacent entries rebuild shared revisions; validation repeats work. Anchors: apps/cli/scripts/classify-release-impact.mjs:265–365; release-publication.mjs:978–983, :1735–1774, :2041; fixture apps/cli/src/scripts/release-recovery-publication.fixture.ts:77–153; runner release-recovery-publication.runner.ts:8–99.
| Static successful-path accounting | Authority builds | Product builds | Total |
|---|---|---|---|
| Exact plan: 17 real commits + synthetic current | 37 | 6 | 43 |
| Exact validation and tamper validations | 3 | 0 | 3 |
| Resume partial plan and validation | 37 | 5 | 42 |
| Resume completed plan and validation | 1 | 0 | 1 |
| Total | 78 | 11 | 89 |
This is a source-derived estimate of the successful fixture paths, not a traced installation count from the old log. Every CLI build also generates a baseline and runs built-command assertions; selected baseline products add nine further baseline generations. Warm dependency archives do not eliminate installation into each isolated tree.
Correction to the original report: readback-only is not build-free. It skips retained publication artifacts after authority reconstruction. Even the completed resume with no entries reconstructs one spent recovery target (release-publication.mjs:751–760). A test that reads frozen recovery declarations without tools does not prove the full planner or validator avoids builds.
Accepted Q9 direction: use a minimal real Git history with real first-parent relationships, local remote tags, current recovery declarations, actual planner/validator and tiny valid outputs to test binding, tampering and partial/completed resume. Keep the full historical replay as an explicit diagnostic. This consciously removes routine compatibility proof for the entire old monorepo/toolchain range; it is not a claim of identical coverage. The user selected full historical replay only on demand; an additional routine historical-builder witness is not required.
Production authority memoization or changelog-first selection could also reduce work, but changes production behavior and requires separately grounded TDD. It is not silently authorized as test-only cleanup. Existing release-publication.test.ts has 61 cases in 2.560s, illustrating the much cheaper semantic layer, although its temporary module with appended private exports deserves a seam cleanup rather than blanket endorsement.
Built CLI: broad policy matrices and a diagnostic benchmark
apps/cli/src/cli/behavioral-portfolio.test.ts:351 launches the issue-215 benchmark without a scenario argument. The default lifecycle runs ten sequential CLI commands: scaffold, adoption, cold, warm, no-op, source-only, managed drift, bypass, clear and human progress. The case took 50.567s historically. It is explicitly diagnostics-only, with no latency/speedup gate (scripts/behavior-contract/issue-215-update-benchmark.mjs:747–766). Several lifecycle samples assert only generic success; human progress is collected without a content/order assertion in this branch.
Move the full diagnostic benchmark out of the normal suite. Retain a focused real cold/warm composition witness. Reducing ten commands to scaffold/adoption/cold/warm removes six launches, not a proven 60% runtime. A valid converged seed could reduce it further, but its preparation must not hide the same work. Existing runtime/scripts.test.ts:1054 already proves non-empty warning replay, source invalidation with installation reuse, missing-tool refusal and uncacheable dynamic configuration through the actual candidate-preview/cache capability.
The generic 11-command matrix (portfolio:184–258, :458–471) accepts any string-valued status or error code and does not constrain process exit status. A newly broken command emitting a typed error can pass. Delete this weak shared matrix and its registry-completeness check; preserve meaningful command-specific refusal semantics in focused tests where required. That avoids twelve CLI invocations and eleven copied/Git-initialized fixtures without pretending generic output shape proved command success.
Further hotspots are narrow policies exercised through complete consumers:
- Missing/old/correct Oxlint catalog: 26.652s, 25.154s and 24.812s; three variants each run two full updates.
- Array workspace globs: 22.819s; incompatible Effect ownership: 22.142s; exact Effect/tsgo catalog versions: 19.820s.
- Prepare adoption/preservation: 17.695s; monorepo Effect prepare lifecycle: 16.657s.
- Four consumer dependency-section variants invoke full monorepo update to vary one manifest field.
Move broad manifest/prepare permutations to real plan/reconcile/file-application seams, retaining one executable witness per distinct composition boundary. scaffold/output-root-dependencies.test.ts:17 already consumes real detection, desired output, materialization and receipts; features/scaffold-state/reconcile.test.ts:249,276 protects catalog-before-dependency ordering. These do not automatically cover every preservation/conflict input: map the missing cases explicitly.
Keep focused built-artifact dispatch, aliases where packaging matters, exact public JSON/exit/stderr semantics, one scaffold/update composition, cold/warm cache wiring, controlled failure with unchanged consumer bytes, and meaningful delegation/transport behavior. Move inexpensive shell-grammar and direct candidate-preview tests out of a file whose initialization requires a built CLI. That also reduces their cache input closure.
Historical 80% arithmetic — superseded by Q10
The historical CLI run was 2612.83s. An 80% reduction means 522.566s, about 8m43s. Other files plus runner overhead consume 193.321s, leaving only 329.245s, about 5m29s, for the three hotspots if execution remains serial. Their combined 2419.509s must fall by 86.4% under that simplifying assumption. The single exact-recovery case already exceeds the entire hotspot allowance.
This is target arithmetic, not a speed prediction on 2-vCPU hardware. The old sample used different hardware/code and failed. Successful expensive workflow samples were roughly 43–46 minutes, implying about 8.6–9.2 minutes for an 80% workflow reduction; other uncached tasks could become the next critical path. Cold dependency stores, forced test execution with warm stores, and normal task-result reuse must be measured separately.
Reduce repeated real operations first, then schedule remaining independent proof. Shards shorten elapsed time but do not eliminate work and duplicate setup; workers within one 2-vCPU VM share that CPU/disk budget. Lower cache invalidation reduces how often work executes, not the cost of a genuine miss. Capability seams help both: a catalog policy test needs manifests/planner/reconciler inputs, while a built CLI smoke legitimately depends on packaged assets. Unrelated wiki prose should not invalidate either; shipped prompts remain inputs to consumers that use them.
Next evidence needed: phase timings for fixture preparation versus assertion work; actual install/build/process counts; one passing comparable 2-vCPU baseline and reduced portfolio; distinct warm/cold measurements; total runner work including retry/setup. This arithmetic is retained only as the rationale for the deeper investigation; Q10 removes the 80% target. No runtime improvement has been measured here.
Current decisions and remaining evidence
Q1–Q7 are accepted in the grill: only 2-vCPU workers, no automated macOS jobs, no monthly ceiling/reserve, cheapest meaningful proof, precise capability caches, PR/main verification and current publication authority retained, all eight unselected operator files reconciled. Existing trigger cancellation semantics remain. Manufactured tests need no replacement when they protect no owned behavior; any affected rule's verification route must remain truthful.
R5 closes shared understanding: Q10 supersedes Q8 with no numerical target and no sharding; Q11 accepts the consolidated proposals. Q9 retains minimal real recovery history in CI and full replay on demand. The compiled specification is authoritative for delivery. Preparation timings, resource pressure, portable fixture isolation and passing 2-vCPU benchmarks remain unmeasured. Blacksmith invoice readback can refine consumption reporting but is not a monthly-cap requirement.
Work log and rule evaluation
Read global/project agent settings before dispatch; inspected .agents/subagents/manifest.mjs and selected planning-discovery ownership/guidance. Four lanes stayed readonly; only coordinator authored this consolidated wiki artifact. No shared skills, source code, tests, workflows, provider settings or external messages changed. No full suite execution was claimed. GitHub issue/PRs were inspected as references only.
| Rules | Result | Current evidence |
|---|---|---|
| HI-WIKI-001 | Pass | Existing research route reused; report listed in owning meta.json. No new route surface or accepted spec invented. |
| HI-WIKI-003 | Pass | Page frontmatter, research links and route metadata follow current schema; owning content/public-contract checks recorded below. |
| HI-DOCS-001 | Pass | Candidate research stays in private wiki; implemented runbooks remain authoritative and unchanged. |
| HI-REPO-001; HI-API-001; HI-BACKOFFICE-001 | Pass for audit boundary | API/contracts/auth outcomes and ownership retained in recommendations; no implementation mutations. |
| HI-CLI-001; HI-CLI-003; HI-CLI-004 | Pass for audit boundary | CLI behavior/assets retain ownership, integrity and public-output evidence. |
| HI-WEB-001; HI-AUTH-001; HI-CONFIG-001; HI-CONTRACT-001; HI-DB-001; HI-ENV-001; HI-SCAFFOLD-001; HI-UI-001 | Pass for audit boundary | Owning surfaces and critical proof identified in inventory; replacement verification routes explicitly required before culls. |
| HI-REPO-002; HI-CLI-002 | Not applicable | No publication, product version, changelog or release-bearing change. |
| HI-REPO-003; HI-REPO-004 | Not applicable | No URL/default/callback or reusable-skill change. |
Validation: wiki check:content passed before and after synthesis; the owning public-wiki-contract suite passed all six tests. The six inventory tables contain exactly 68 + 29 + 16 + 5 + 21 + 12 = 151 file rows. Mutation-specific production validation and RED/GREEN are not applicable to this research-only change. This log does not claim all retained tests pass or any candidate deletion is semantically proven.
September 22 extension validation: wiki content check and all six public wiki contract tests pass; the four touched research/grill/log files were formatted. Three resumed readonly lanes supplied the updater, release reconstruction and built CLI findings. Coordinator consolidated them, corrected the older build/budget statements and persisted both R4 answers. No timing benchmark or test implementation ran.
Delivery planning refresh — September 22
Three resumed readonly lanes translated the accepted research into implementation kernels: updater preparation and controlled installation; minimal recovery history plus precise CI/cache wiring; and the remaining app/package/wiki portfolio. The current inventory is 153 executable test files after inherited parent changes. Existing public seams support the accepted test-only transformations without production release-planner changes. Each worker must return final selectors, actual input closure, capability needs, semantic fingerprints and proof; root configuration has one dependent owner. Independent plan review resolved T9 dependency ordering and assigned the Bun 1.4 diagnostic assertion to T3. This refresh adds execution detail, not new product decisions or measured savings.