Issue 217 CI Cost Grill Log
Issue 217 CI Cost Grill Log
Intake: 2026-09-21
User: "only have 2vcpu workers", accepts the preceding proposals, and requires aggressive test reduction and minimal cache invalidation. User explicitly invokes Requirements Grill.
Accepted direction:
- 2-vCPU worker constraint, with cross-platform scope to resolve in Q1.
- Aggressive removal of low-signal and redundant tests under the current TDD guidance.
- Preserve meaningful owned behavior through cheap public seams and representative real integration witnesses.
- Reduce repeated preparation and unnecessary cache invalidation before scaling concurrency.
- Optimize both total consumption and feedback latency; do not treat sharding alone as cost reduction.
This acceptance supersedes incompatible older whole-command-only/native-test prohibition for this optimization direction. It does not automatically accept every numerical hypothesis, open tradeoff, or unpublished implementation detail in the research report. Current publication policy remains authoritative; no prior-PR-receipt lookup requirement is reintroduced.
Round R1
Q1–Q6 were presented as one logical round. The user answered "agree with all", with explicit annotations taking precedence: remove the macOS integration test, impose no monthly minute ceiling, retain the cheapest meaningful proof, and minimize correct invalidation. Those answers are recorded below. Q7 resolves the two other macOS jobs.
Q1 — Runner and platform scope
Prerequisites: none.
Evidence anchors:
- .github/workflows/behavior-contract.yml:runs-on (three Blacksmith 4-vCPU jobs).
- .github/workflows/release.yml:runs-on (GitHub-hosted Ubuntu).
- .github/workflows/cache-reuse-witness.yml:consumer matrix and macos candidate; external-drift.yml:runs-on.
- Repository visibility readback: private. Official GitHub runner documentation lists standard private-repository Ubuntu at two cores and macOS ARM at three cores; no standard two-core macOS option.
Observed constraint: existing Ubuntu publication already satisfies two vCPUs and needs GitHub-hosted trusted publishing. The macOS witnesses cannot meet a literal two-core limit on their existing service.
Question: does "2-vCPU workers only" apply to Linux CI or literally every runner including macOS?
Recommendation: two-vCPU Linux with the existing macOS checks as an explicit exception, preserving actual cross-platform evidence. Alternative: strict two-vCPU scope and explicitly park hosted macOS proof.
Code consequence: determines workflow runner labels and whether cross-platform restoration, real macOS candidate behavior and external drift remain required. Merely setting two test workers does not satisfy a VM-size constraint.
Accepted answer: two-vCPU workers remain the requirement. Remove the macOS CLI integration test rather than retain the recommended platform exception for that test. The user's "can we just not do the integration macos test?" is accepted as that requested direction. Q7 separately asks whether the remaining cache-restoration and weekly filesystem jobs also leave hosted CI; no blanket macOS removal has been inferred.
Q2 — Harness cost allocation
Prerequisites: none.
Evidence anchors: research report September usage estimate; Blacksmith runner allowance documentation.
Observed constraint: 3,000 normalized minutes belongs to the whole organization; sampled Harness consumption is approximately 4,911 normalized minutes through September 21. Actual account usage is not read back.
Question: what share of the allowance is Harness's monthly design budget?
Recommendation: 2,400 minutes, leaving 600 reserve. Alternative: all 3,000 for Harness, or an explicit different allocation.
Code consequence: sets measured acceptance and capacity planning. A budget target is not permission to skip required verification after a monthly threshold.
Accepted answer: no maximum monthly minute ceiling and no allocation/reserve target. The user explicitly rejected the proposed 2,400-minute budget. Runtime and normalized consumption remain comparative optimization evidence, not enforced monthly gates. The earlier 3,000-minute allowance is financial context, not an acceptance ceiling.
Q3 — Test reduction completion
Prerequisites: accepted aggressive pruning direction.
Evidence anchors: complete 151-file research inventory; apps/cli/vitest.config.ts:test; three hotspot timing records.
Observed constraint: cheap low-signal files and expensive meaningful tests have very different costs. A deletion percentage does not predict either confidence or consumption.
Question: is pruning complete based on supported capability proof and CI cost/time, or must it also meet a minimum test-count reduction?
Recommendation: retain the cheapest meaningful proof of each supported capability/safety invariant, remove low-signal and redundant proof, and use CI cost/time for acceptance with no deletion quota.
Code consequence: controls retention criteria and whether a numeric test-count threshold becomes a delivery requirement. No redundant-test replacement is needed when no owned invariant exists.
Accepted answer: keep the cheapest meaningful proof of supported capabilities, public contracts and safety invariants. Aggressively remove low-signal and redundant tests. No minimum test-count reduction quota; do not manufacture replacement tests for checks with no owned invariant.
Q4 — Independently cached test groups
Prerequisites: accepted minimal correct invalidation direction.
Evidence anchors: turbo.json:@punks/cli#test; @punks/cli#build; globalDependencies; scripts/behavior-contract/cache-identity.mjs:verificationEnvironment.
Observed constraint: all CLI tests currently share one cache entry; broad source and host inputs make unrelated changes invalidate costly work.
Question: should independently cached groups follow capabilities or individual test files?
Recommendation: small capability groups, each with its actual consumed source, fixture and tool inputs. Per-file tasks may increase reuse but multiply task definitions/startup. A per-case cache or separate scheduler is not proposed.
Code consequence: defines task/module boundaries and the scope of input identity, while existing package ownership and signed Turbo persistence remain.
Accepted answer: small independently cached capability groups, with the lowest possible correct invalidation rate. Unrelated documentation must not invalidate product tasks. A changed shipped prompt/skill or other consumed artifact must invalidate its real consumers. Relevant transitive inputs, tool capabilities, trust and output identity remain in the cache contract.
Q5 — PR and main verification
Prerequisites: current publication authority grounded.
Evidence anchors: .github/workflows/behavior-contract.yml:on; .github/workflows/release.yml:release-evidence and production; root AGENTS.md:publication eligibility.
Observed constraint: verification runs on PRs and main today; current publication does not look up prior PR Candidate Evidence. Older closed grills describe a different model.
Question: retain PR and main tests with maximal correct reuse, or require PR-only tests and resolve main protection/authority separately?
Recommendation: retain both triggers during this optimization, maximizing reuse without a release-authority redesign. PR-only is an explicit requirements expansion, not a simple cache improvement.
Code consequence: determines when the required suite executes. A PR-only choice unblocks separate questions on protected merges, direct pushes, source-tree identity and publication authority. Cancellation policy follows after this decision.
Accepted answer: retain verification on both PRs and main, maximizing correct reuse. Keep current publication authority; do not introduce a PR Candidate Evidence lookup or redesign release protection. Existing trigger cancellation semantics remain until separately justified by implementation evidence.
Q6 — Unselected operator test scope
Prerequisites: completed file and selector inventory.
Evidence anchors: package.json:test/test:cache-policy; research report root/operator inventory; scripts/behavior-contract/isolation.test.ts and process-identity.test.ts.
Observed constraint: eight files have no normal package/Turbo/workflow selector, including supported cleanup/provenance safety candidates. Deleting them saves zero normal CI execution.
Question: reconcile these files as part of this effort, or leave them outside scope?
Recommendation: trace current consumers, retain unique proof of supported behavior in its appropriate affected task, and delete obsolete/duplicate proof. This can add some useful verification while much larger waste is removed elsewhere.
Code consequence: defines selected-test inventory completeness and ownership; no silent addition of an entire unreviewed suite.
Accepted answer: reconcile all eight unselected root/operator test files. Trace current consumers, keep unique proof of supported behavior in the appropriate affected task, and delete obsolete or redundant proof. Do not preserve dead tests solely because of historical issue names.
Round R2
The accepted answers leave one concrete platform decision. No numerical budget/latency target is required after Q2; measurement and exact concurrency belong to implementation validation.
Q7 — Remaining macOS jobs
Prerequisites: Q1, accepted removal of the macOS CLI integration test.
Evidence anchors: .github/workflows/cache-reuse-witness.yml:consumer macos matrix row and macos-candidate job; .github/workflows/external-drift.yml:macos-filesystem-capabilities.
Observed constraint: two other jobs still use three-core GitHub macOS runners: cross-platform cache restoration and weekly filesystem observation. Linux Blacksmith checks and GitHub-hosted publication are separate. GitHub-hosted Ubuntu already has two cores for this private repository; publication's npm trusted-publishing boundary remains intact.
Question: remove all automated macOS jobs, or only the CLI integration test and retain these two exceptions?
Recommendation: remove all three automated macOS jobs to satisfy the strict two-vCPU requirement. Retain meaningful Linux coverage and describe macOS as no longer automatically validated; this does not change supported runtime behavior.
Code consequence: determines whether the macOS matrix consumer and scheduled workflow also leave the selected CI graph. No Linux result may be relabelled as macOS proof.
Accepted answer: remove all automated macOS jobs. Linux verification remains; CI no longer claims actual macOS execution/restoration/filesystem validation. This resolves the platform exception without changing supported product behavior. Existing GitHub-hosted two-vCPU Linux publication retains its npm trust boundary.
Round R4 — technical pruning reopened
2026-09-22: the user stopped the final-confirmation path and requested deeper technical explanation of test types, pruning and hottest paths, ideally reaching an 80% CI-time reduction. This supersedes R3's empty-frontier/ready-for-confirmation state, not Q1–Q7. No specification or implementation begins. The consolidated research report's September 22 section contains the three readonly lanes' current-code findings.
Historical evidence distinguishes cheap low-signal tests (seven obvious CLI culls: 20ms) from expensive meaningful scenarios run at broad boundaries. The three CLI hotspots explain 93.1% of test-body time. Source tracing finds per-dependency real installation in updater tests, repeated preparatory updates, ten commands inside a diagnostic benchmark, and about 89 isolated installs/CLI builds across two historical recovery scenarios. Build counts are static successful-path estimates; timings come from a historical failed run, not current passing 2-vCPU execution.
Q8 — Meaning of the 80% target
Prerequisites: Q1, Q2, historical run and cache-state evidence.
Evidence anchors: research report:Actual latency and consumption / September 22 deep dive; detailed Actions job 105078205569; successful expensive runs 35185839490 and 35197493495.
Observed constraint: expensive CLI execution takes about 43.5 minutes in the detailed sample while warm workflows are about two minutes. Averaging them hides the slow path. Sharding changes latency and can increase summed runner work; narrower cache keys change execution frequency. Hardware/code differs from the required 2-vCPU target.
Question: for the aspirational 80% reduction, target uncached PR feedback with lower total runner work, total runner work primarily, or average PR time including cache hits?
Recommendation: target PR feedback when expensive CLI tests execute, require total runner work to fall too, and report warm reuse separately. Distinguish warm dependency stores from fully cold installation.
Code consequence: defines comparison/acceptance metrics without a monthly cap or test-count quota. It does not preselect shard count or promise a measured improvement.
Accepted answer: target 80% less uncached PR feedback time, with less total runner work. Warm cache-hit runs are measured separately. This is an accepted optimization target, not a claim of achievement or a monthly minute ceiling.
Q9 — Historical toolchain compatibility in routine CI
Prerequisites: Q3 and current release reconstruction trace.
Evidence anchors: apps/cli/scripts/release-publication.mjs:classifyEntry, createReleasePublicationPlan, validateReleasePublicationPlan; classify-release-impact.mjs:produceReleaseImpactAuthority; apps/cli/src/scripts/release-recovery-publication.runner.ts:exact/resume.
Observed constraint: two cases replay 17 historical commits plus a synthetic current entry, with repeated base/head authority reconstruction and validation. Source accounting estimates 89 isolated installs/CLI builds. Readback-only avoids retained publication artifacts but not historical authority builds. Existing fast tests cover many release semantics, but not this entire old builder compatibility surface.
Question: use minimal real history in CI and full replay on demand, retain one additional historical-builder witness in CI, or keep the entire full replay mandatory?
Recommendation: use actual planner/validator over a minimal real Git history with tiny valid outputs for recovery binding, tamper rejection and partial/completed resume; keep full historical replay as an explicit diagnostic. Real commits, first-parent relationships, remote tags and release authority remain tested. Routine execution of every old builder is consciously relinquished.
Code consequence: determines the required historical compatibility proof. Production planner memoization or metadata-first classification is separate behavior-changing work, not silently included in test-fixture pruning.
Accepted answer: minimal real history in CI; full historical replay only on demand. Preserve recovery authority, tamper rejection, first-parent and resume behavior through the actual planner/validator and real Git/artifact boundaries. Continuous execution of all those historical monorepo builders is no longer required.
R4 glossary review distinguishes feedback latency, summed runner work, task-result reuse and historical-builder compatibility without publishing new canonical terms. Q8 and Q9 are both answered. The requested technical explanation and discussion continue; this is not final approval to create a spec or implement. Further owned-seam decisions can follow answers or new evidence; no final shared-understanding approval is requested in this round.
Round R5 — consolidate and close
2026-09-22: the user said to forget the hard 80% cap, make CI as fast as possible without sharding, remove the identified hot-path inefficiencies, bring the proposals together, and confirmed agreement. This explicitly closes shared understanding with the corrections below. No repeated confirmation is required.
Q10 — Remove the numerical target and exclude sharding
Prerequisites: Q8 and the technical hotspot walkthrough.
Evidence anchors: research report:September 22 technical deep dive; apps/cli/vitest.config.ts:fileParallelism; current workflow topology.
Observed constraint: elapsed latency, summed runner work and task reuse differ. Shards can shorten latency by duplicating setup and increasing cost; the identified waste exists inside executed tests.
Question resolved by direct user instruction: retain a hard 80% threshold or optimize the identified waste without sharding?
Accepted answer: no percentage speedup threshold or fixed CI runtime ceiling. Optimize as far as the identified inefficiencies permit, using only 2-vCPU workers and no test sharding. Preserve the aim of lower uncached feedback time and lower total runner work; measure warm cache reuse separately. No monthly cap or deletion quota. This supersedes Q8's numerical target and earlier shard experiments, including those in the initial research report.
Code consequence: no shard matrix or new duplicate jobs distributing test subsets. Capability-sized cache tasks remain valid inside the existing verification topology. Performance work must remove redundant operations and invalidation; progress is measured, not judged against a promised percentage.
Q11 — Accept the complete efficiency proposal
Prerequisites: Q3, Q4, Q6, Q9, Q10 and the completed technical walkthrough.
Evidence anchors: update/run.test.ts:runUpdateWithPreview; update/run.ts:applyDependency; release-publication.mjs:classifyEntry; cli/behavioral-portfolio.test.ts; research report complete inventory and evidence-backed opportunities.
Observed constraint: the top three files account for 93.1% of historical CLI test-body time. Important invariants coexist with incidental package installs, repeated generated fixtures, broad executable permutations and historical reconstruction. Obvious registration/prose culls alone save only milliseconds.
Question resolved by direct user instruction: consolidate and accept the recommended pruning and efficiency changes?
Accepted answer: yes. Keep supported behavior through the cheapest meaningful proof. Remove irrelevant real installs from filesystem/receipt scenarios; share immutable preparation with private mutable state; reduce repetitive full-update traversals; move catalog/prepare/policy breadth to owned seams; remove weak command/registration/source/prose tests; keep focused built-command and real-adapter witnesses; replace routine historical replay with minimal real history; move diagnostic benchmarks to on-demand use. Apply the accepted cache/provisioning/root-proof/database/wiki-runner efficiencies from the report, preserving their stated correctness/isolation conditions. Reconcile the full inventory and eight unselected operator files.
Code consequence: one coherent optimization scope spans CLI test fixtures/seams, CI/cache selection, and redundant proof elsewhere. No essential safety guard, current publication authority, or supported product behavior is deleted merely because its test was hot. Small testability refactors must preserve behavior; production release-planner redesign remains parked.
Final consistency: existing glossary terms retain their identities. Behavioral Test includes exported application/adapter outcomes; CLI Behavioral Test is its built-process subset. PR/main verification and current publication authority supersede older contrary glossary axioms. All active branches are grounded and closed at 100%; the frontier is empty and shared understanding is confirmed. Compile the accepted scope into one specification above PR #222.
Delivery stack instruction
Subsequent explicit Create Spec instruction, 2026-09-22: the user requested a refreshed rebase onto #222's branch and closure of that PR before compiling. Fetched parent head 4d4b7b71622ee90228143af57f995f0f6de4bff8; rebase reports up to date and ancestry proves inclusion. Closed #222 without merge and retained its branch. #223 remains open with base team/stefan/cli-update-quick-fixes. This changes provider status, not the accepted CI/test requirements or a claim that parent code landed in main. The existing spec identity is reused and its dependency/base evidence is refreshed.
2026-09-22 resume: restacked PR #223 onto updated parent head cef8e31e8a4fc7c8ac57ca7439b7590683c77a0e. The parent's implemented validation/generation/commit behavior does not change Q1–Q7 requirements. Its two new CLI test files bring current inventory to 153 (70 CLI); retain the original 151-file audit as historical evidence and reassess the added proofs during implementation. Final glossary consistency pass is complete and no design question remains unanswered. R3 presents the persisted shared understanding for explicit closure before canonical glossary synthesis and specification.
User requests main ← PR #222 ← this work. The child branch is team/stefan/issue-217-ci-cost, to be rebased onto the current head of team/stefan/cli-update-quick-fixes and opened as a draft with that branch as its base. PR #222 remains independently owned; do not rewrite its branch or absorb its edits into this PR's diff.
Only requirements/research artifacts are currently authored. The stacked PR will not claim test pruning or CI optimization has shipped. Historical measurements remain tied to 7e147b3300c391db979b87e20c4159a5a03d12cb and require refreshed validation after the parent implementation.
Glossary and consistency record
Stack checkpoint: rebased child branch onto PR #222 head 4d61376a and preserved both sides of the research navigation/wiki log conflicts. The initial commit attempt was blocked by 497 inherited wiki lint errors; the generated bundle accounts for 273 diagnostic locations and other existing source/configuration accounts for the rest. No source or lint policy is modified by this task.
Handback resolution, 2026-09-21: the user answered "yes go" to the explicit request for a docs-only hook bypass and dependent draft PR. This authorizes a command-scoped exception for the research/requirements commit; repository hook configuration and future implementation gates remain unchanged. It is not approval to claim failing baseline checks passed or to repair unrelated source.
The ordinary branch push then failed in the existing pre-push hook: release:candidate rejects its obsolete --base argument. The same docs-only checkpoint exception applies to the push needed for the authorized draft PR; record the failure without changing the hook or claiming release-candidate success.
After rebase, six public wiki contract tests still pass. Full content validation reports two stale specification metadata files inherited unchanged from PR #222. This child leaves both parent-owned files intact and records the additional baseline failure; it does not claim the full wiki check passes.
After Q7, the design frontier is empty. Domain Modeling rechecked the complete glossary: two-vCPU execution, zero macOS CI claims, cheapest supported proof, minimum correct invalidation and no monthly cap are consistent. Final shared-understanding confirmation is pending against the saved artifacts; no implementation or canonical glossary publication is claimed yet.
Before R1, Domain Modeling used the status glossary as working persistence and checked the existing CI Verification and Publication glossary. Applicable High-Signal Test, Behavioral Test, CLI Behavioral Test, Safety Invariant, Affected Verification and Stable Aggregate Check meanings are preserved. Historical full-command-only and publication-receipt axioms conflict with current accepted direction/current code and are flagged explicitly. No canonical glossary is rewritten before closure.
Brainstorm was completed after technical grounding under the newly accepted two-vCPU constraint. Its unresolved choices form Q1–Q6; later numerical targets and trigger consequences retain explicit dependencies. Show Me's design-tree view is derived from this record and carries no additional authority.
Delivery steering — Bun version
2026-09-22: user annotated the current CI Bun 1.3.5 pin with “lets make it 1.4 btw if we can”. Accepted bounded amendment: align current repository/CI pins to local Bun 1.4.0 after frozen-install, build and affected-test compatibility proof. No unrelated dependency upgrade or automatic rewrite of historical-version fixtures is authorized. Local-first verification remains required; no new grilling round is needed for this explicit instruction.