Harness Intelligence Wiki
Grilling

CI Test Suite Pruning Grill Log

CI Test Suite Pruning Grill Log

Source

  • Research: [[content/docs/project/research/ci-suite-reduction-and-npm-release-research-report]]
  • User direction: retain only high-signal behavioral tests and remove useless regression, fixture, inventory, and implementation-shape tests.

Suite objective

Q1

Prerequisites:

  • none

Question: What kind of automated tests belong in the target suite?

Accepted answer:

  • Retain only high-signal tests that prove a named product capability, public contract, or critical safety invariant.
  • Prefer behavioral tests through the nearest public seam.
  • A retained test must survive an internal implementation rewrite when observable behavior is unchanged.

Q2

Prerequisites:

  • Q1

Question: What happens to tests that do not meet the retained-test standard?

Accepted answer:

  • Delete them.
  • Do not keep low-signal tests in a nightly graveyard, quarantine, legacy suite, or historical regression pack.
  • A test earns retention through current signal, not through age or the existence of an old bug.

Implementation-shape boundary

Q3

Prerequisites:

  • Q1

Question: Do tests that inspect source strings, import ordering, line counts, test names, exact call spelling, or release-version fixtures belong in the target suite?

Accepted answer:

  • No.
  • Delete implementation-shape tests and their meta-test inventories.
  • If an underlying rule remains important, it must be represented by a behavioral test or one purpose-built static enforcement mechanism, subject to Q5.

Round R1 Frontier

Round R1 is complete. Q4 through Q7 are answered.

CLI safety seam

Q4

Prerequisites:

  • Q1

Question: Which CLI safety tests may use a lower native seam?

Accepted answer:

  • None.
  • A retained CLI safety test must execute the complete hi or hint command.
  • Delete lower-level native safety tests even when they make symlink, race, permission, rollback, or interrupted-write fault injection easier.
  • Safety behavior remains eligible only when it can be proved through the complete public CLI command.

Architecture enforcement

Q5

Prerequisites:

  • Q3

Question: How are valid architecture constraints enforced after source-shape tests are deleted?

Accepted answer:

  • Prefer behavioral outcomes, compilation, and package boundaries.
  • When those cannot express a genuine architecture invariant, enforce it with one purpose-built Oxlint, AST, or build rule outside the test suite.
  • Do not replace deleted source-string tests with another collection of spelling checks.

Deletion evidence

Q6

Prerequisites:

  • Q1
  • Q2

Question: What evidence is required before deleting or replacing a test?

Accepted answer:

  • After a test is classified as low-signal under the accepted standard, delete it directly.
  • Do not require a unique-witness audit, replacement test, mutation proof, or RED/GREEN cycle before deletion.
  • Pure test deletion has tdd_status: not-required because it does not change production behavior.
  • Any later production behavior change still follows ordinary vertical RED/GREEN TDD.

Portfolio sufficiency

Q7

Prerequisites:

  • Q1

Question: What proves that the final test portfolio is sufficient?

Accepted answer:

  • Use a capability-and-safety inventory.
  • Map important capabilities, public contracts, and accepted safety invariants to their strongest retained Behavioral Tests.
  • Do not require a line-coverage percentage, fixed test count, or historical regression count.
  • Treat the 40–50% reduction as a minimum target, not a stopping rule.

Round R2 Frontier

Round R2 is complete. Q8 through Q10 are answered.

Public command inventory

Q8

Prerequisites:

  • Q1
  • Q4
  • Q7

Question: Which public commands must have a full-command Behavioral Test?

Accepted answer:

  • Keep one primary full-command Behavioral Test for every displayed command: check, ensure, init, scaffold, update, operator, report, deprecated skills, tools, and upgrade.
  • Add another full-command test only when it proves a distinct public outcome.
  • Do not create a scenario matrix for every command.

Executable and output modes

Q9

Prerequisites:

  • Q1
  • Q4

Question: How are hi, hint, human output, and JSON output covered without duplication?

Accepted answer:

  • Use hi throughout the primary Behavioral Test suite.
  • Prove the hint executable alias once for the whole product.
  • Where JSON is documented, parse the JSON and assert its stable meaning.
  • Assert human text only for the root command atlas and one representative failure.
  • Do not duplicate the command suite across executable aliases or output modes.

External provider boundary

Q10

Prerequisites:

  • Q1
  • Q4

Question: How may full-command tests control external providers?

Accepted answer:

  • Execute the real built or packaged CLI process against a real temporary repository.
  • Run a local loopback HTTP substitute for the external provider and select it through public configuration such as DP_CONTROL_PLANE_URL.
  • Let the CLI use its production Effect composition and HTTP adapter. Do not inject internal module mocks into the spawned process.
  • Do not contact live providers from CI.

Verified existing architecture:

  • Focused tests already inject Effect services, configuration, and HttpClient through Layers.
  • The production composition root provides FetchHttpClient.layer and decodes DP_CONTROL_PLANE_URL into the CLI control-plane configuration.
  • Existing tests already use loopback HTTP servers and real spawned CLI processes in separate places.
  • The retained full-command provider witnesses must combine those seams: spawn hi, point it at the loopback provider through public configuration, and assert the public result.

Round R3 Frontier

Round R3 is complete. Q11 through Q13 are answered.

Focused test eligibility

Q11

Prerequisites:

  • Q1
  • Q4

Question: May focused unit tests remain below the complete CLI process?

Accepted answer:

  • Retain focused tests only for dense pure domain rules or typed protocol decoding where several meaningful cases would make full-command tests wasteful or opaque.
  • Exercise an exported domain or protocol API without internal mocks.
  • Delete focused orchestration, presentation, source-shape, call-order, fixture-churn, and historical regression tests.

Catastrophic safety inventory

Q12

Prerequisites:

  • Q1
  • Q4
  • Q7

Question: Which catastrophic safety outcomes need extra full-command witnesses?

Accepted answer:

  • Keep one product-wide full-command witness proving repository containment and no writes outside the selected root.
  • Keep one proving a failed write leaves no partial managed state.
  • Keep one proving failures redact secrets.
  • Use the command that best exposes each outcome. Do not repeat the safety matrix for every write command.
  • If an outcome cannot be reproduced through the public CLI, Q4 excludes the lower-level native witness.

Built and packaged artifacts

Q13

Prerequisites:

  • Q4
  • Q8
  • Q9

Question: Do command witnesses execute built output or an installed package tarball?

Accepted answer:

  • Execute the primary command suite against built dist output.
  • Keep one installed-tarball smoke proving the published package surface, both hi and hint, and the reported package version.
  • Do not install the tarball separately for every command witness.

Turbo cache constraint

New user direction:

  • Heavily use the Turbo task graph and share cache reuse through every CI lane.

Verified current gaps:

  • Trusted internal pull-request and protected jobs configure signed Turborepo Remote Cache through GitHub OIDC.
  • Fork pull-request jobs do not configure shared Turbo cache.
  • Root test and test:ci run multiple suites directly before delegating only part of their work to Turbo.
  • @punks/cli#build is explicitly cache: false, despite being repeated across CI jobs.
  • Several CLI verification scripts wrap Turbo tasks, while other workflow steps build or test outside the graph.

Security boundary:

  • Untrusted fork code must not receive authority to write artifacts into a trusted cache namespace.
  • GitHub Actions cache can safely restore default-branch entries to fork pull requests while keeping writes scoped to the pull-request merge ref.
  • The exact trusted-versus-fork cache design remains open in Q14.

Round R4 Frontier

Round R4 is complete. Q14 through Q16 are answered.

Cache trust boundary

Q14

Prerequisites:

  • Q7

Question: How does cache sharing include untrusted fork pull requests safely?

Accepted answer:

  • Use one signed Turborepo Remote Cache namespace across trusted internal pull requests, main, and release verification when task inputs are identical.
  • Let fork pull requests restore the default-branch GitHub Actions .turbo cache and write only to their isolated pull-request cache scope.
  • Do not give untrusted fork code authority to write trusted remote-cache artifacts.
  • Remove artificial namespace and environment differences from deterministic task hashes unless they change observable task behavior.

Turbo graph ownership

Q15

Prerequisites:

  • Q7
  • Q11

Question: Must every retained check participate directly in the Turbo task graph?

Accepted answer:

  • Yes.
  • Define every deterministic build, lint, typecheck, test, and package check as a package-owned Turbo task.
  • Root package scripts only delegate with turbo run.
  • CI does not invoke Vitest, Bun tests, or custom verification suites directly outside the Turbo graph.
  • Workflow-only setup, aggregation, artifact transfer, and external publication may remain outside the graph when they are not repository tasks.

Cache opt-outs

Q16

Prerequisites:

  • Q7
  • Q13

Question: Which retained tasks may opt out of caching?

Accepted answer:

  • Cache every deterministic retained task, including the CLI build and full-command Behavioral Tests.
  • Declare precise task inputs, outputs, environment hashes, and dependency edges so cache reuse remains correct.
  • Permit cache: false only for external publication, state mutation, or a genuinely nondeterministic operation.
  • A historical preference for freshness or an old regression does not justify bypassing cache.

Requirements frontier

The frontier is empty. Shared-understanding confirmation remains before specification.

Detailed workflow fan-out, affected-package selection, stable aggregate checks, and release triggering remain in the separate CI Workflow Topology grill. This grill owns the retained-test portfolio and the rule that every retained deterministic check is graph-owned and cached.

Cross-domain reconciliation

The CI Workflow Topology grill consumes Q14 through Q16 directly: affected selection runs through package-owned Turbo tasks; deterministic retained checks are cacheable; identical trusted tasks share signed remote cache; fork writes remain isolated.

Topology's former "retained native witnesses" wording does not preserve any lower-level CLI safety suite. Q4 and Q12 remain authoritative: only full-command CLI safety witnesses survive, and schedules may contain only named external-drift witnesses that pull-request verification cannot establish.

Q17

Prerequisites:

  • Q13
  • Q16

Question: Does any installed-tarball execution smoke remain?

Accepted answer:

  • Supersede Q13's installed-tarball smoke and delete it.
  • Keep the affected pull-request command suite against built dist.
  • Let publication perform only the non-test npm pack --json --dry-run integrity precondition accepted by the topology grill.
  • Do not execute the packed tarball before publication.

Integrated consistency pass

  • The retained portfolio runs only on affected pull-request revisions and executes CLI Behavioral Tests against built dist.
  • Publication consumes exact-tree Candidate Evidence and performs package integrity, mutation, reconciliation, and readback without executing a test suite or installed tarball.
  • No lower-level CLI native safety matrix survives in pull-request, release, or scheduled workflows.
  • Every deterministic retained repository check is a cacheable Turbo task with an explicit owner and correct trusted-versus-fork cache authority.
  • The cross-domain frontier is empty. Shared-understanding confirmation is pending.

Shared-understanding confirmation

  • The user confirmed the integrated test-portfolio and workflow model on 2026-08-14.
  • All pruning branches are closed at 100%.
  • The accepted requirements are ready for specification compilation with the CI Workflow Topology grill.

On this page