Issue 215: safe and faster update research
Issue 215: safe and faster update research
Scope and decision status
The user requested analysis and brainstorming before implementation. They identified the main performance complaint as the delay between submitting the command and its first output. Findings below support candidate solutions; this report does not approve requirements or authorize a particular redesign.
Source revision: 1ee7f8b72f8fe116ad74aa15923238d3b985c751. Environment: Linux, Node 24.19.0, Bun 1.4.0. Issue evidence includes separate Linux and macOS reproductions; their environments must remain distinct.
Three readonly research lanes investigated staging/confinement, lint/bootstrap, and scaffold/planner consistency. The coordinator inspected presentation, baseline resolution, prior learnings, local data volume, and repository-detection timing. No application code was changed.
Evidence and confidence
| Finding | Evidence | Confidence |
|---|---|---|
| Ignored sockets and FIFOs break candidate copying. | apps/cli/src/runtime/scripts.ts:166-177 copies recursively, excluding only Git metadata and dependencies. Running the same copy options on isolated ignored socket and FIFO fixtures produces Node's ERR_INTERNAL_ASSERTION: Unreachable code. | Reproduced on Linux. |
| Staging failures are misclassified. | The broad catch at runtime/scripts.ts:1003-1007 reports config-load for preparation failures. | Confirmed in source. |
| Candidate containment compares incompatible root spellings. | runtime/scripts.ts:606 compares a lexical candidate root with a canonical child. A temporary symlink-root fixture rejects a contained directory; canonicalizing both sides accepts it. | Reproduced with an alias root on Linux; actual Darwin execution remains unverified. |
| Plain-text Oxlint failure output loses its cause. | JSON parsing failure at runtime/scripts.ts:835-844 retains stderr or the parser exception, discarding stdout. | Confirmed in source; matches reported stdout beginning Failed to .... |
| Quiet bootstrap loses subprocess diagnostics. | integrations/tool-management.ts:270-275 passes stdio: "ignore"; best-effort handling at :560-580 can retain only the resulting error message/stack. | Confirmed in source. |
| Scaffold and check disagree about which packages get lint configuration. | scaffold/output.ts:1939-1956 uses different Python exclusion predicates depending on planOnly. Scaffold materialization and check planning select different modes. | Confirmed source divergence; exact macOS consumer fixture has not been reproduced. |
| Planned paths can appear on a failed update. | update/run.ts:6439 returns computed allChanges as changedFiles, while :6354-6398 separately attributes planned/completed operations and derives applied. | Confirmed in source; a listed path does not establish a write. |
| Intermediate update output is suppressed by the production adapter. | platform/feature-application-operations.ts:346-362 resolves baseline and invokes update with emitOutput: false. presentation/present.ts:209-220 publishes the returned operation result after execution. | Confirmed in source. Installer-inherited output may still appear in interactive mode. |
The fixture dependency failure is less certain. Candidate install ownership reads package.json workspaces and treats selected directories outside those roots as separate installs (runtime/scripts.ts:558-631). Selection includes managed dependency workspaces and package-backed commit-gate scopes (update/run.ts:5845-5859). This can explain a fixture losing access to workspace/catalog resolution, but the exact failing macOS fixture and its package-manager ownership remain necessary to prove the cause. The ownership resolver also lacks pnpm-workspace.yaml handling.
Performance evidence
Feedback latency
The production update adapter disables intermediate output. Even if enabled, the first update event is late: update/run.ts:5637, after detection, context compilation, preliminary desired-state generation, wiki alignment, possible second detection/context compilation, final planning and staging.
Sources: update/run.ts:4073,4153,4385,4472,4516,4537,4625,4635,4711,4978,5637; platform/feature-application-operations.ts:346-362.
Consequently, silence can cover work much later than startup. A progress message improves responsiveness but does not establish a runtime improvement. JSON stdout should remain one complete result; early human progress belongs on an appropriate separate channel.
Concrete filesystem volume
On 2026-09-16:
du -sh /home/stefan/repos/harness-intelligence/.devpunks/delivery
52G
du -sh --exclude=.git --exclude=node_modules /home/stefan/repos/harness-intelligence/.devpunks/delivery
15G
du -sh /home/stefan/repos/harness-intelligence/apps/wiki
16M
du -sh /home/stefan/repos/harness-intelligence/docs
448KThe delivery tree totals 52 GB; approximately 15 GB remains after matching the candidate copier's existing .git/node_modules exclusions. This is a concrete potential volume multiplier and exposes the special-file crash. These are disk-usage measurements, not copied-byte counters; they do not establish how much the failed invocation copied before aborting.
Wiki alignment also copies the wiki and docs (update/run.ts:1248-1280,1346-1352), before candidate lint copies the live and staged trees (runtime/scripts.ts:517-518). Avoiding irrelevant traversal/copying addresses both correctness and cost.
Small measured component
Repository detection was called three times per root using the source in the main checkout and its installed Effect v4 dependencies. Its detector file SHA-256 exactly matched this research checkout (dc79a4fece51e327c3b8a1b326bf0ebe25379f0b38b1f8066ae195c81009d0cb).
| Root | Three elapsed samples | Detected manifests |
|---|---|---|
| Research worktree | 67, 46, 40 ms | 13 |
| Main Harness checkout | 47, 51, 48 ms | 13 |
| Local collective-intelligence checkout | 1292, 850, 769 ms | 27 |
These are component measurements, not command benchmarks or evidence about the separate macOS checkout. The measurements do not justify prioritizing Harness detector micro-optimization over staging and repeated setup.
Other source-confirmed repeated work
- Preliminary/final planning and materialization call through the output pipeline more than once; wiki package election can legitimately alter later inputs. Reuse must respect those dependencies.
- Tool bootstrap runs on an apply request even when
shouldApplyis false (update/run.ts:5748-5794). - An installed agent-browser can still execute its install command unless the existing system-browser readiness check passes (
integrations/tool-management.ts:518-548;data/catalog/tools.ts:3-22). Its own install command may short-circuit; its current duration was not measured. - Candidate validation performs frozen isolated installation; changed manifests or missing Bun lock add a lockfile-only step (
runtime/scripts.ts:652-682,988-1000). Install, patch and lint targets run sequentially. - Lint buffer overflow can trigger a second complete lint execution (
runtime/scripts.ts:743). Its agent-format fallback also needs a nonzero-exit/empty-findings regression (:746-775). - Warm baseline loading still hashes the archive, extracts a verification copy and compares content digests (
baseline/resolve.ts:385-427). This is integrity work; any optimization must preserve verification. - Remote baseline metadata resolution precedes update execution and has a five-second default timeout (
baseline/resolve.ts:55,544-546). No claim is made that it dominated this user's run.
Measurement blockers
The installed CLI's read-only hi update --check --baseline bundled failed immediately with:
Bundled baseline asset missing: frameworks/nestjs/nestjs-best-practices/.gitignoreIts full-command timing is therefore unusable for this issue. Local captured stdout/stderr: /tmp/hi-215-profile-ikfeh7ur/ (ephemeral diagnostic files).
Direct source imports in this dependency-free worktree initially resolved an incompatible Effect installation, failing on Schema.TaggedErrorClass / Schema.Literals. Detection measurements used the main checkout's explicit installed dependency paths. No dependency installation or packaged-asset repair was performed for this analysis.
Candidate solution
One final plan, one candidate, one attributable result
A small shared lifecycle should own these steps:
- Report the active phase immediately in human mode; record timings and work counters.
- Resolve baseline, repository ownership and package-manager topology.
- Compile one final desired plan after legitimate package election, and retain its immutable artifacts.
- If current owned bytes, receipts and required readiness already satisfy that plan, finish through a verified no-op path.
- Otherwise construct one isolated candidate from required inputs and the planned overlay; validate it.
- Apply the validated plan through the existing confined filesystem boundary, then report completed writes and remaining failures precisely.
This is a candidate internal design, not a proposal for a new public command family. Scaffold, check and update should share selection/ownership decisions; dry-run versus write should change execution, not which lint scopes exist.
Safe, smaller staging
Define candidate inputs separately from copy execution. Inputs must include actual lint/type-resolution source, transitive configuration, package-manager/workspace files, planned artifacts, required ignored/generated inputs, and project-owned verifier references. Respect intended deletions as well as additions.
Use Git-aware exclusions where applicable, with explicit required-input inclusion and a non-Git policy. Ignore status is not ownership authority. A tracked-only snapshot would omit legitimate local work; discarding every ignored path would risk verifier references and required configuration.
Before copying, inspect entry kinds without following arbitrary symlinks. Skip irrelevant sockets, FIFOs and transient trees. A required unsupported entry should produce an actionable staging error. Resolve canonical containment consistently and preserve accepted project-owned file bytes, modes and symlink semantics.
Prune excluded directories before traversal. Merely refusing to copy their leaf files still wastes time walking large evidence trees. Reflink/copy-on-write support may help retained regular files, but must have a portable fallback and cannot replace input selection. Writable hardlinks would violate candidate isolation.
Make failures and writes legible
Retain phase, subprocess command/cwd, exit code and bounded stdout/stderr. Handle plain text as a useful failure diagnostic when JSON decoding fails. Quiet mode should capture subprocess output.
The public result should distinguish planned paths, completed writes, validation failures and tool readiness failures. Existing operation-state evidence is useful; improve its presentation and audit direct-write attribution before adding duplicate status fields. Changing changedFiles semantics or replacing it requires an explicit compatibility decision.
Optimize repeated work after selection is correct
- Reuse final compiled artifacts and hashes within the invocation rather than recomputing an equivalent plan.
- Check tool/browser readiness before installing; separate readiness evidence from a repeated installer invocation.
- Reuse dependency downloads immediately where existing package-manager caches permit. Reuse prepared installations only with complete keys and isolated mutation: lockfile, package/workspace/catalog inputs, package-manager/runtime/platform identity, patches and lifecycle inputs.
- Bound concurrency across genuinely independent install/lint roots after package ownership is established. Shared installs, Effect patches and apply ordering cannot run indiscriminately in parallel.
- Consider archive-verification optimization only after measuring it and preserving the existing integrity contract.
Persistent lint-result caching, a background daemon and weaker validation are unnecessary first steps.
Options and tradeoffs
| Option | Benefit | Limitation |
|---|---|---|
| Narrow patch to copy filters, path normalization and errors | Quickly removes proven crashes and diagnostics loss. | Can leave repeated planning, install cost and planner divergence. |
| Shared final plan plus required-input staging | Addresses ownership consistency, irrelevant copies and duplicated generation together. | Requires careful input completeness and scaffold/check/update compatibility tests. |
| Cache and parallelize the current pipeline first | Can improve some repeated runs. | Preserves unnecessary work and divergent selection; invalidation and shared-state risks complicate proof. |
Recommended candidate: deliver the narrow correctness fixes and phase measurement first, then introduce shared-plan reuse/selective staging through the existing seams. Use measured costs to decide whether dependency-install caching or concurrency is worth the additional state.
Proposed validation and performance contract
Correctness cases:
- Ignored sockets, FIFOs, virtualenv external/dangling links and large evidence trees do not enter the candidate when irrelevant.
- Required source/config inputs cannot silently disappear because of ignore status.
- Root path aliases work; escaping selected roots remain rejected.
- Mixed Python/JS scaffold then check agrees on lint configuration and selection bytes; authoring-pending findings remain separate.
- Workspace/catalog fixtures are assigned to their real resolution owner or receive a precise unsupported-layout error.
- Project-owned verify-behavior references survive byte-for-byte, including ignored/untracked references and CRLF content.
- Oxlint plain-text failures and installer failures retain diagnostics in JSON mode.
- A validation failure reports no completed managed writes; no-op and partial failures retain truthful attribution.
- Existing confined apply, rollback and receipt-publication safeguards remain in force.
Performance cases:
- Measure first progress event, time to plan and total duration separately.
- Record phase duration, files/bytes staged, generation passes, install roots, subprocess count and cache reuse.
- Compare warm no-op, small managed update and cold dependency-changing update.
- Include a large irrelevant evidence tree and a mixed-language monorepo.
- Measure bundled/local baseline separately from network resolution.
- Report repeated-sample medians/tails only after collecting sufficient runs.
Candidate UX target: first human progress within 250 ms on the benchmark host. Candidate work budgets: zero evidence-tree files staged, one final materialization, no installer invocation when unchanged tool readiness is proven. Total-time targets should be chosen after a valid before measurement; no end-to-end speedup has yet been demonstrated.
Unresolved decisions
- Exact total-time targets and cold/warm benchmark environments.
- Compatibility strategy for clarifying legacy
changedFilesin JSON. - Progress defaults for non-interactive/JSON callers; preserve single-document stdout.
- Complete required-input policy for ignored files and repositories without Git.
- Whether prepared-install caching is needed after copy/planning reduction.
- Exact package ownership of the failing macOS fixture.
Prior constraints consulted
.agents/notes/2026-08-04-scaffold-parent-identity.md: retain operation-root identity and stable-parent confinement, including rollback and publication..agents/notes/2026-07-21-linux-ci-ownership-readiness-and-json.md: keep subprocess output from corrupting JSON.- Research on issues 194, 197, 203/207 and 181: preserve prior receipt evidence, project ownership, scaffold/update distinctions, and full candidate validation.
- Existing verifier preservation code/tests:
scaffold/output.ts:3669-3718,4059-4083;scaffold/project-verifier-preservation.test.ts:173-265.
Runtime reproduction receipts
Using the current copy options (recursive, verbatim symlinks, only .git/node_modules excluded), isolated fixtures under an ignored .delivery/ directory yielded:
socket ERR_INTERNAL_ASSERTION Unreachable code
fifo ERR_INTERNAL_ASSERTION Unreachable codeEach fixture used a temporary directory, a Unix socket from Node's net server or a FIFO from mkfifo, and was cleaned after execution.
For a temporary symlink alias -> real containing candidate/app, comparison using the same containment predicate yielded:
{"lexicalWithin":true,"canonicalWithin":false,"correctCanonicalWithin":true}These confirm bounded mechanisms. They are not a successful full issue reproduction or an implemented fix.