Harness Intelligence Wiki
Research

Issue 217: Remaining test repetition

Remaining test repetition

Authority and method

The user requested another aggressive pruning pass after the first implementation, explicitly invoking Brainstorm and Parallel Research. The accepted direction is the cheapest meaningful proof, two-vCPU workers, no sharding, no numerical target, and minimum correct cache invalidation. The user also explicitly authorized the previously requested one-time commit/pre-push hook bypass. That waiver does not change repository hooks or turn failed inherited lint into passing evidence.

Four readonly lanes inspected the implemented suite at e41938f0: updater/scaffold; remaining CLI; release/root tooling; and API/apps/packages/wiki. They inspected assertions, production seams and routine selectors. No lane ran tests or changed files during discovery. The coordinator owns this consolidated report. A follow-up on the expensive CLI policy portfolio resolved the second-largest hotspot before its implementation task was assigned.

Sources below refer to that implementation commit, before this second pruning pass. Findings are static equivalence arguments unless a local result is explicitly identified. Test-only deletions need no fabricated product RED. Subsequent implementation and focused GREEN belong in the linked spec's plan, portfolio and validation record.

System model and measured concentration

The verification system connects an owned behavior to its proof, the proof to its real inputs and execution capabilities, and the resulting evidence to the aggregate gate. Repeating a complete installation or publication for every combination of independent input classes increases work without necessarily proving another behavior. Removing cheap input enumeration while retaining repeated full journeys attacks the wrong layer.

The complete local CLI run after the first pass passed 863 cases in 63 files in 1,530.919 seconds. Its largest files were:

FileCasesLocal body time
src/update/run.test.ts65898.888 s
src/cli/portfolio-policy.test.ts52208.108 s
src/update/wiki-alignment-owner.test.ts172.037 s
src/scaffold/project-verifier-preservation.test.ts362.144 s
src/cli/behavioral-portfolio.test.ts654.335 s
src/scaffold/output.test.ts2542.764 s
src/scaffold/generated-gate-scoping.test.ts133.900 s
src/runtime/scripts.test.ts12130.344 s

Evidence: /tmp/hi-217-local-proof/cli-complete-final.log and .json, summarized durably in Validation. These are observations from Linux x86_64 on an AMD Ryzen 5 3600 with affinity restricted to two logical CPUs, not an Apple M3 Pro or a Blacksmith runner. Two logical CPUs are not proof of equivalent physical capacity, filesystem or network latency. No hosted percentage improvement follows from this run. The updater's previous local 2,937.77-second result is a local before/after comparison only. Per-case timings can include first-use seed preparation; they cannot simply be summed as predicted deletion savings.

The two largest files account for about 72% of total measured CLI wall time. The 64-case quality-command suite takes about 12 ms; deleting its supported syntax cases would reduce the count substantially while barely affecting feedback time. Source: recorded full-run log and apps/cli/src/features/commit-gate/quality.test.ts.

Lane A: updater and scaffold

Source: updater tests, scaffold output, verifier preservation.

CandidateWork removedSurviving evidence and limitation
Pure handoff-render assertion at verifier-preservation line 148One caseBoth retained real scaffold/update verifier journeys already check that rendered instruction and the actual preserved verifier.
Scaffold stale/tampered/missing handoff matrix at output line 955Two complete preparation rowsRetain stale-to-current publication, completed authoring, exact bytes and receipt hash. Updater lines 1690 onward retain tampered and missing repair; initial scaffold creates the absent output. This deliberately stops repeating every damage class through both entrypoints.
Healthy updater handoff row at line 1690One scenarioMove exact byte preservation into the healthy repeat after tampered repair. Retain missing creation separately. The recorded 28.571-second row includes seed preparation that moves to another row.
Cold and converged projection-proof matrices at lines 918 and 988Two cold rowsRetain one cold-publication witness and all three missing/corrupt/stale receipt-only repairs. Move manifest-hash and exact repeated-receipt assertions to the retained repairs. Each invalidity remains tested; cold-publication multiplied by every invalidity is reduced.
Valid unchanged-authoring row at line 2303One scenarioMove accepted-proof, empty-pending-set and truthful operation assertions into the existing successful completed-authoring recovery witness. Retain empty, missing and retargeted ownership failures.
Empty and whitespace-only wiki route previews at line 1989One repeated traversal if direct empty-input proof existsRetain whitespace through the complete updater. Delete the empty traversal only after locating the selected direct predicate proof.
Both preview and check for each converged invalidityTwo repeated traversalsKeep both interfaces for one invalidity and distribute the other invalidities between them. Retain all three actual repairs. This reduces interface-by-invalidity combinations, not invalidity classes.

The candidate case reduction is seven. The two removed scaffold rows cost 9.673 and 9.804 seconds in the observed run; the three cold rows cost 18.108, 18.794 and 19.106 seconds. These are evidence of repeated expensive work, not a hosted forecast. All six receipt-publication validation cases, nine confinement cases, four distinct ordinary-refresh cases, real install/script-suppression/hook behavior and both real verifier journeys remain. They protect different outcomes rather than permutations of one outcome.

Lane B: remaining CLI

Source: runtime scripts, validation cache, runtime composition.

  • At runtime line 1372, six input-invalidation rows each repeat four identical failure/retry/corrupt-envelope probes. Keep six cold-plus-changed pairs and one recovery sequence: 36 candidate-preview calls become 16. This is a work reduction; extracting a named recovery test may increase the test count by one.
  • Remove the call-order-only preparation case at line 3296. The retained preparation-mutation case at line 2630 reads the copied manifest and proves the changed command reaches lock refresh, install, patch and lint. It checks the useful effect as well as order.
  • Remove the smaller cache-identity case at cache line 52 after moving its unique closedInputs: false refusal into the full dependency matrix at line 80.
  • Remove the cache reuse case at line 127 after strengthening the existing production-runtime composition witness with the exact installation-kind payload. The old case's title claims reuse across roots, but both wrappers use the same root.

These last three candidates remove three cases. The cache files take milliseconds; the runtime file takes 30.344 seconds for all 121 cases. No candidate-specific timing was recorded. The four catalog consumer sections remain independently observable: combining them in one repository could let successful dev-dependency discovery hide broken optional or peer discovery. Numeric lint severity and ordinary staged symlink refusal also retain their own observations.

Policy portfolio follow-up

The policy portfolio contains 16 cheap helper/catalog cases and 36 application cases. Its helper still invokes real runScaffold and runUpdate (test-fixtures/portfolio/application.ts:217 onward). Static requests total 26 scaffold, seven check-update and 25 write-update calls; seed reuse means these are not 58 guaranteed complete executions.

Remove the exact Oxlint catalog leaf case at line 2182 (retained root-dependency application and catalog-consumer proof) and the missing-object catalog row at line 2398 (the stronger missing-array catalog repair at line 1313 checks preserved unrelated fields, narrow ownership and actual idempotence). Keep old-object repair, but remove its repeated update: the array-form repair already owns that repeat observation. The removed two cases cost 7.224 and 15.606 seconds in the recorded run. This reduces representation-by-entrypoint combinations, which is recorded rather than presented as exhaustive proof.

The three prepare-receipt cases at lines 736, 819 and 970 cover old segment present, desired segment present and both present. The initial recommendation to move them to exported reconciliation was rejected during implementation: that seam consumes normalized observations and cannot prove private receipt-aware command composition or published receipt effects (update/run.ts:2614–2648,3208–3224). Constructing those observations in tests would reproduce the policy being tested. All three actual mutation cases remain. This is an evidence-based correction to the initial static equivalence inference.

Eight further declaration/refusal cases can use the existing public planning seam: whitespace prepare (583), incompatible tuple (1049), duplicate declarations (1896), two pnpm catalog rows (2148), unknown protocol (2536), old Oxlint without React (2563), and incompatible runtime ownership (2590). Preserve exact declared values, typed failures and unchanged input bytes. Retained React/Effect materialization and actual mutation-refusal tests continue to prove application. No new production export is justified merely to make the tests convenient. With the three prepare-receipt cases retained, the bounded cuts and these migrations can avoid about 12 static full-application requests, subject to worker verification; net time remains unmeasured until local execution.

Lane C: release and root tooling

Source: publication, recovery integration, authority, classification, root cache identity.

Remove three publication cases: exact checkout/no Candidate Evidence at line 173 is subsumed by the no-release/no-side-effects case at line 699; nullish filtering at line 236 checks private array assembly already exercised through the real recovery planner; npm-only exact readback at line 1591 is covered by the real planner/validator integration at recovery line 129. Remove the direct compatibility-tamper row at recovery line 79 because that integration already mutates compatibility and observes actual refusal.

The four publication classifier rows at line 192 use a fake classifier that reproduces the policy. Production classifier tests already cover all four product selections. Keep one successful product publication boundary witness and remove three copied-policy rows. Consolidate the two authority selection cases at lines 481 and 500 while preserving older-valid/expired selection and a separate duplicate refusal observation. These changes remove eight cases, with little expected runtime impact: whole publication and recovery files took 4.251 and 4.911 seconds respectively.

Move the three command-membership assertions from root cache identity line 253 into the existing topology dry-run at ci-topology.test.ts:78, then delete the duplicate test. This removes one real Turbo process. Preserve the independently reviewed fixture-input perturbation repair in that same topology file.

Retain the eight distinct recovery refusal dimensions, baseline-only versus mixed readback, startup/ENOENT/nonzero/provider outage distinctions, process ownership races, root-versus-hoisted dependency layouts, OIDC authority and aggregate wiring. Similar names do not make these equivalent authority boundaries.

Lane D: API, apps and packages

Source: wiki sync, wiki public contract, auth policy, auth composition, backoffice routes, browser routes.

Four residual duplicates can be removed: wiki's check-only/no-write case (the real public contract fingerprints the entire tree); auth's sign-in policy row and successful Resend forwarding case (the configured OTP composition checks the complete same payload); and the mocked protected-route unauthorized redirect (the required browser journey checks exact redirect URLs and usable controls, including /access). Keep non-sign-in policy rows, provider rejection, forbidden and unavailable route outcomes. The auth and backoffice files took 7 ms and 30 ms respectively; these cuts improve signal and maintenance, not the runtime bottleneck.

Do not delete real invalid-OTP, secure-cookie, sign-out/session-revocation or database constraint proof as though API happy-path tests replace it. They do not. The sole shared UI case checks owned accessibility forwarding beyond the login browser's observations. API typed failures preserve diagnostic causes, a different outcome from releasing resources after failure. The corresponding sources are packages/auth/src/auth-email-contract.test.ts:332, packages/db/src/postgres-contract.test.ts:293, packages/ui/src/public-ui-contract.test.tsx:36, and apps/api/src/features/typed-failure-contract.test.ts:22. Removing these would sacrifice unique coverage for little measured savings.

Synthesis and execution decision

The user has authorized aggressive pruning. Proceed with duplicate deletion, assertion migration and representative cross-product reduction under the existing accepted behavioral guarantees. Record the reduced combinations explicitly; do not call them exhaustive coverage. Preserve every identified invalidity/ownership class at its owning real boundary. Implement through disjoint scoped tasks and run the changed retained witnesses locally.

The four lanes agree that test count is a poor cost proxy. No claim is made that a large fraction of the remaining suite is useless. The strongest residual cuts remove repeated full traversals and candidate processes. Larger removal of unique filesystem, recovery, installation, auth, SQL or browser proof remains a product-risk choice, not a demonstrated equivalence. No numerical count or speedup promise is introduced.

Cache reuse stays bounded by correctness. The retained review found that updater inputs omitted test-fixtures/monorepo-root/**; its applied repair now makes actual consumed fixture changes invalidate updater while unrelated build evidence remains reusable. Thirty-two local cache-policy tests and root static/format checks passed for that repair. Removing this required input to improve hit rate would reuse invalid evidence.

Implementation disposition, exact counts and focused execution results are authoritative in Test portfolio, Plan, and Validation. Hosted performance remains unmeasured for this revision; the next ordinary uncached run can supply that observation without commissioning repeated benchmark workflows.

On this page