Delivery Phase Flow Optimization Research
Delivery Phase Flow Optimization Research
Executive answer
Research is now sufficient to define prompt-level requirements for a faster delivery flow. It is not sufficient to claim a specific token, latency, or cost reduction.
The clearest supported optimization ground is repeated orchestration around the native coding harness:
- repeated loading and serialization of authority;
- repeated model resumptions;
- broad worker briefs;
- fan-out, waits, and follow-up coordination;
- whole-wave release barriers;
- repeated review preparation and routing;
- complete review reruns after ordinary repairs.
The six focused root tasks processed 394.13 million input tokens. Of those, 383.46 million, or 97.3%, were reported as cached input. They also produced 1,934 shell calls, 578 explicit waits, 103 child spawns, 217 child follow-up or message operations, and 17 compactions. Across every discoverable actor in the same six task trees, observed input reached at least 2.410 billion tokens. These are processed-context observations, not unique context or billing figures.
The corpus supports prioritizing fewer resumptions, smaller context packets, dependency-local dispatch, restrained fan-out, bounded waiting, and one default full review. It does not establish that any one mechanism caused the totals.
The target remains prompt, skill, context, delegation, and tool policy. Codex, Claude, Cursor, OpenCode, or another user-selected native harness performs the work. Harness Intelligence does not need a Delivery executor, deterministic delivery runtime, scheduler service, journal, watcher, outbox, lock service, or cache service.
The two accepted specifications remain separate leaf authorities:
- Project Verification defines reusable project-owned executable knowledge beneath verify-behavior and keeps Verification separate from readonly Code Review (apps/wiki/content/docs/project/specs/cli/project-verification-capability/SPEC.md:19-44, 80-150, 178-228).
- Dynamic Implement Spec Execution Policy defines dependency-local task release, shared-current-branch workers with disjoint scopes, Context Pointers, retained assurance gates, early draft-PR visibility, and one default Full Code Review Pass (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:18-33, 35-116, 139-170).
They belong in one delivery-optimization change set because the execution policy decides when work and review run, while Project Verification supplies the reusable behavioral proof used inside implementation. Flattening them would erase their distinct decision trails and authority boundaries.
System boundary
Operator
The operator is a coding agent inside the native harness selected by the user. It reads repository instructions and accepted artifacts, invokes skills, delegates through the native harness, operates tools, and returns evidence.
Harness Intelligence responsibility
Harness Intelligence supplies:
- phase and skill prompts;
- context-selection rules;
- authority and evidence pointers;
- routing and stop conditions;
- native-harness delegation policy;
- worker input and output contracts;
- review and verification boundaries.
It does not execute a second delivery engine beside the native harness.
Protected constraints
- delivery-phase remains human-invoked (.agents/skills/delivery-phase/SKILL.md:1-20).
- Direct one-phase or HITL invocation stops at its requested boundary (.agents/skills/delivery-phase/SKILL.md:22-27, 59-64).
- Applicable TDD, architecture, runtime, UI, provider-readback, Verification, and final-acceptance gates retain authority (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:24-33, 83-92).
- Verification remains owned by implement-spec through verify-behavior. review-phase remains readonly Code Review (apps/wiki/content/docs/project/specs/cli/project-verification-capability/SPEC.md:72-78, 134-137).
- Review report retention and returned-route validation remain separate proof boundaries (.agents/skills/review-phase/phases/retain-report.md:35-75; .agents/skills/review-phase/phases/return-route.md:22-61).
- Changed requirements, weaker gates, or substantial redesign still stop at handback.
Explicitly excluded
The following proposals from the earlier report are superseded:
- deterministic Delivery runtime;
- Delivery executor;
- compiled delivery kernel;
- receipt or event journal;
- scheduler service;
- provider watcher or outbox;
- lease or lock service;
- evidence cache service;
- any infrastructure required to make phase prompts work.
Existing repository artifacts and provider records remain the durable authorities. A compact context packet is a model-created view for the current task, not a new source of truth.
Research method and lane coverage
Four readonly lanes ran in parallel and the coordinator reconciled them against current repository sources.
- Authority reconciliation compared both temporary handoffs, both closed grills, both compiled specifications, superseding decisions, source branch histories, and the consolidated branch.
- Skill-chain reconstruction traced delivery-phase, implement-spec, verify-behavior, review-phase, debugging, docs ingest, and closeout.
- Historical task tracing inspected selected Codex JSONL records, including root and child actors, token snapshots, tools, waits, delegation, and compactions.
- Architecture challenge applied the brainstorm lenses of intake, state, control, feedback, recovery, and handoff while preserving the prompt-only boundary.
Source snapshot:
- Harness base: origin/main at b9bb213dc490753020501ef81f42978d1caee89e.
- Canonical shared skills: wearedevpunks-skills main at 584ea4bddd5bcfcf780baf9fb67ff736ba7d0626.
- Project Verification handoff: /tmp/harness-verification-handoff.chgr3Z/project-verifier-handoff.md.
- Implement Spec policy handoff: /private/tmp/implement-spec-execution-policy-handoff-2026-09-04.md.
- Skill composition visualization: /Users/stefan/.codex/visualizations/2026/09/03/01a069b4-ee54-71a3-b2b0-5517dbfa77b7/agnostic-skill-composition-graph.html.
Temporary handoffs are discovery aids. The retained specifications, grill histories, repository commits, and provider readbacks are authority.
Historical task measurement protocol
Protocol identity: 2026-09-04/v1.
Selection
The focused corpus intentionally includes:
- the PR 182 brainstorm and prompt delivery;
- the Implement Spec execution-policy requirements and specification task;
- the Project Verification requirements and specification task;
- issue 193, a small change explicitly handled without delivery-phase;
- issue 194, a larger Harness repair;
- a full delivery of the Evaluation Suite in Collective Intelligence.
This is a purposive, heterogeneous corpus. It is useful for finding recurring amplification patterns. It is not a statistically representative sample.
Actor identity
- The first session_meta.payload.session_id groups root and child actors.
- The actor identity is session_meta.payload.id.
- The root actor has id equal to session_id.
- Root-only and all-actor totals are reported separately.
Token reduction
total_token_usage is cumulative within an accounting segment. Summing every snapshot would overcount.
For each actor and each token field:
- The first snapshot contributes its value.
- A later value greater than or equal to the previous value contributes only the nonnegative delta.
- A decrease begins a reset segment and contributes the current value.
- Actor totals are summed only after this per-actor reduction.
The report keeps input, cached input, cache-write input, output, and reasoning output separate. Derived noncached input equals input minus cached input. Cache-write input was zero in these six focused root records.
Operation counts
- Unique tool operations are keyed by call_id.
- Shell calls are exec operations.
- Explicit waits combine wait and wait_agent operations.
- Spawns are spawn_agent operations.
- Follow-ups combine followup_task and send_message operations.
- Compactions count only outer JSONL records whose type is compacted.
- Root active time is the sum of task_complete.duration_ms. It includes tool and wait activity inside the root turn and is not pure model compute.
Known accounting limits
- Input tokens include cached and replayed prompt context. They are not unique bytes and are not billing evidence.
- Root active time is not first-to-last elapsed time.
- Parent waits and child execution overlap; they cannot be added as serial active work.
- Five Evaluation Suite child logs were not discoverable despite 50 recorded spawns, so that all-actor total is a lower bound.
- The logs do not contain trustworthy durable markers for every delivery and review phase transition. Review-only token attribution is therefore unknown.
- Encrypted or missing child content proves activity occurred, not what the child did.
- Different tasks used different scopes, tools, retries, provider delays, and continuations.
Focused corpus results
Root activity and tokens
| Task | Root active h | Input M | Cached % | Derived noncached M | Output K | Reasoning K |
|---|---|---|---|---|---|---|
| PR 182 brainstorm and prompting | 4.12 | 90.08 | 96.7 | 2.93 | 152.64 | 67.63 |
| Implement Spec policy | 0.35 | 9.45 | 90.8 | 0.87 | 66.10 | 24.98 |
| Project Verification | 0.89 | 30.43 | 94.7 | 1.60 | 128.82 | 44.80 |
| Issue 193 quick fix | 0.81 | 19.97 | 98.6 | 0.28 | 31.05 | 10.80 |
| Issue 194 repair | 3.16 | 114.38 | 97.7 | 2.64 | 152.56 | 64.44 |
| Evaluation full delivery | 4.98 | 129.81 | 98.2 | 2.34 | 261.93 | 66.54 |
| Total | 14.31 | 394.13 | 97.3 | 10.67 | 793.09 | 279.18 |
Root active time is an activity envelope, not model-only work.
Root orchestration operations
| Task | Shell | Explicit waits | Spawns | Follow-ups | Compactions |
|---|---|---|---|---|---|
| PR 182 brainstorm and prompting | 487 | 108 | 12 | 30 | 4 |
| Implement Spec policy | 69 | 0 | 0 | 0 | 1 |
| Project Verification | 171 | 10 | 5 | 4 | 2 |
| Issue 193 quick fix | 115 | 14 | 6 | 16 | 0 |
| Issue 194 repair | 572 | 147 | 30 | 60 | 4 |
| Evaluation full delivery | 520 | 299 | 50 | 107 | 6 |
| Total | 1,934 | 578 | 103 | 217 | 17 |
All discoverable actors
| Task tree | Actor logs found | Observed input M | Cached % | Derived noncached M |
|---|---|---|---|---|
| PR 182 brainstorm and prompting | 13 | 140.18 | 96.9 | 4.34 |
| Implement Spec policy | 1 | 9.45 | 90.8 | 0.87 |
| Project Verification | 6 | 42.47 | 93.5 | 2.74 |
| Issue 193 quick fix | 7 | 36.89 | 97.4 | 0.97 |
| Issue 194 repair | 31 | 275.62 | 97.0 | 8.24 |
| Evaluation full delivery | 46 | at least 1,905.72 | 98.2 | at least 34.10 |
| Total | 104 | at least 2,410.32 | 97.9 | at least 51.26 |
All-actor output was 5.49 million tokens and reasoning output was 1.55 million tokens. These totals remain activity observations, not quality or cost measures.
Root source records
| Task | Root task identity | Root JSONL |
|---|---|---|
| PR 182 | 01a05d00-904b-7ca2-a483-c608b8ed901c | /Users/stefan/.codex/archived_sessions/rollout-2026-09-01T14-45-13-01a05d00-904b-7ca2-a483-c608b8ed901c.jsonl |
| Implement Spec policy | 01a05d01-0229-7ed0-a0c6-7f39e4079608 | /Users/stefan/.codex/archived_sessions/rollout-2026-09-01T14-45-42-01a05d01-0229-7ed0-a0c6-7f39e4079608.jsonl |
| Project Verification | 01a05d00-02f7-73a2-8024-ae3b472b319a | /Users/stefan/.codex/archived_sessions/rollout-2026-09-01T14-44-37-01a05d00-02f7-73a2-8024-ae3b472b319a.jsonl |
| Issue 193 | 01a06797-12b3-70f2-abb2-c3b419e204a5 | /Users/stefan/.codex/sessions/2026/09/03/rollout-2026-09-03T16-05-49-01a06797-12b3-70f2-abb2-c3b419e204a5.jsonl |
| Issue 194 | 01a0680b-a308-7c83-b12f-5127aeb8c8c6 | /Users/stefan/.codex/sessions/2026/09/03/rollout-2026-09-03T18-13-08-01a0680b-a308-7c83-b12f-5127aeb8c8c6.jsonl |
| Evaluation full delivery | 01a014ed-3703-78b3-9c94-f96f891a2492 | /Users/stefan/.codex/sessions/2026/08/18/rollout-2026-08-18T14-51-25-01a014ed-3703-78b3-9c94-f96f891a2492.jsonl |
What the corpus supports
Fact: repeated context carriage is large
Cached input represents 97.3% of focused root input and 97.9% of all observed actor input. This does not make cached tokens equivalent to fresh reasoning or equal in cost. It does show that the harness repeatedly processed substantial already-seen context.
Consequence: prompt size, resumption frequency, copied worker context, and fan-out all deserve optimization.
Fact: child fan-out magnifies processed context
The full Evaluation delivery root observed 129.81 million input tokens. Its root plus 45 discoverable children observed at least 1.906 billion. Issue 194 grew from 114.38 million at the root to 275.62 million across 31 actors.
Consequence: parent-only token totals understate the size of heavily delegated flows.
Unknown: how much of the child context was required to preserve correctness. The corpus does not establish an optimal agent count.
Fact: router re-entry is not the only amplifier
Issue 193 was intentionally handled without delivery-phase, yet its root still recorded 19.97 million input tokens, six spawns, 16 follow-ups, 14 waits, and 115 shell calls.
Consequence: optimizing only delivery-phase routing would leave global instructions, skill loading, delegation, validation, and coordination costs.
Fact: current source has repeated orchestration boundaries
delivery-phase loads one phase and full delivery re-enters routing after every phase (.agents/skills/delivery-phase/SKILL.md:13-20). review-phase reloads its router on every invocation or resume (.agents/skills/review-phase/SKILL.md:14-23).
The common successful review path contains:
- wrapper plus router;
- target and source preparation;
- one autoreview call;
- five normative review axes;
- parent candidate verification and one report;
- retention and committed-byte verification;
- returned-route revalidation.
The wrapper, router, prepare, run, retain, and return files contain 3,758 words before their referenced contracts and invoked lenses. The complete review-phase Markdown tree contains 7,824 words. Static source size is only a context proxy; the logs do not show exactly which bytes each model call received.
Consequence: review has a concrete prompt-level opportunity to prepare once, carry one frozen packet, run required semantic work once, and avoid reconstructing unchanged authority between gates.
Fact: accepted requirements already remove two known barriers
The Implement Spec policy accepts:
- release each dependent when its own prerequisites and Task Gate pass, without waiting for unrelated running work (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:67-74);
- one default Full Code Review Pass, focused repair validation, affected Verification reruns, and a second full pass only for accepted high-risk repairs (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:101-112).
These are requirements authority, not hypotheses.
Current composition and amplification points
The implemented high-level path is:
human
|
v
delivery-phase router
+-> requirements or spec repair
+-> backlog projection
+-> create-plan
+-> implement-spec
| +-> worker waves
| +-> task and architecture gates
| +-> verify-behavior
| +-> final acceptance
+-> review-phase
| +-> prepare frozen target
| +-> autoreview once
| +-> five normative axes
| +-> parent aggregation
| +-> retained report proof
| +-> returned route proof
+-> repair or debugging
+-> docs ingest
+-> closeoutProject Verification adds the accepted implementation-time branch:
implement-spec
-> verify-behavior
-> selected project verifier references
+-> Uncovered Surface
| -> create-verification-skill
| -> Reference Smoke Proof
| -> resume original scenario
+-> Uncovered Behavior
| -> update-verification-skill
| -> prove drive path
| -> resume original scenario
+-> product-failure
-> debugging-phaseThis branch follows the single verify-behavior entrypoint and keeps app mechanics project-owned (apps/wiki/content/docs/project/specs/cli/project-verification-capability/SPEC.md:82-99, 112-150, 204-228).
Amplification point: broad worker payloads
The current worker brief repeats goals, dependencies, paths, related tasks, description, acceptance, TDD, architecture, runtime, UI, risk, skill guidance, and output obligations (.agents/skills/implement-spec/references/parallel-worker-brief.md:5-93).
Candidate: carry only a small task kernel plus resolvable pointers. This aligns with accepted AC-007 and AC-008 (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:79-82).
Amplification point: whole-wave coordination
The current orchestration builds a ready wave, launches every unblocked task, waits for workers and shared-artifact review, then computes the next wave (.agents/skills/implement-spec/references/parallel-orchestration.md:60-105).
Accepted leaf-spec direction: the native harness recomputes the eligible frontier after each Task Gate. An unrelated slow task does not hold a ready dependent.
Amplification point: shared Markdown writes
Workers currently update their plan entries, while the parent rereads and reconciles shared plan and implementation-note state (.agents/skills/implement-spec/references/parallel-worker-brief.md:95-111).
Candidate: workers return concise deltas and evidence pointers. The parent alone updates shared plan and notes. This is a prompt ownership rule, not a write service.
Amplification point: repeated waiting
The focused roots recorded 578 explicit waits and 217 child follow-up or message operations. The full delivery root alone recorded 299 waits and 107 follow-ups.
Candidate: perform one bounded wait over all known children, continue any independent parent work, and issue another wait only after a meaningful state change or when no independent work remains. Do not narrate unchanged state.
Tradeoff: fewer checks may increase terminal-state detection latency. The native harness wait primitive and experiment results should determine the bound.
Prompt-level target
Full-delivery control loop
An explicitly requested full delivery should use this model-owned loop inside the selected native harness:
explicit full delivery
-> cold route once
-> assemble compact Delivery Context Packet
-> execute selected phase
-> receive compact Transition
-> update the packet
-> continue to the next in-bounds phase
-> stop only at:
HITL requested by the user
handback
missing access or approval
contradictory or stale authority
terminal failure
closeoutDirect human invocation of one phase still stops after that phase. A cold route runs again after a new task, cold resume, changed bounds, contradictory authority, or an unresolvable stale packet.
This is prompt behavior within the current task. It adds no executor and no new durable state.
Delivery Context Packet
The packet should contain only:
- goal and accepted-bounds identity;
- authoritative spec, plan, notes, and handoff pointers;
- branch, base, target, and current diff identity;
- current phase and next action;
- active task or review identity;
- evidence pointers and freshness facts needed by the next action;
- exact blocker, if any;
- stop condition.
It should not copy full specifications, plans, skills, reports, or handoff prose. The agent resolves a pointer only when the selected action needs it.
Transition
Every phase exit should return:
- outcome;
- authority or evidence created;
- changed facts;
- invalidated evidence;
- next eligible phase;
- exact blocker or stop reason;
- pointers needed by the next phase.
The Transition is conversational structured output. Existing artifacts remain the authority.
Native-harness task dispatch
- Recompute readiness after every Task Gate.
- Dispatch every truly independent, disjoint ready task supported by the native harness.
- Do not spawn duplicate investigators for the same question unless deliberate cross-checking is required.
- Do not spawn for work cheaper than transferring, supervising, and validating its context.
- Keep overlapping write scopes serial.
- Preserve Architecture Checkpoints as hard barriers for affected work.
- Let independent eligible work continue after a failed Task Gate.
No universal fan-out number is yet supported. The dependency graph and measured coordination cost should bound fan-out.
Worker input and output
A worker receives:
- task identity and intended outcome;
- dependencies and owned paths;
- acceptance references;
- exact checks and required gates;
- applicable skill names;
- pointers to source detail.
A worker returns:
- changed paths and concise delta;
- acceptance and check outcomes;
- evidence pointers;
- blocker or risk;
- any context invalidation.
The parent validates the delta and is the sole writer of shared plan and implementation-note summaries.
Handoff
A resumable handoff carries goal, branch/base, current phase, next action, artifact pointers, relevant hashes or provider identities, and blocker. It does not restate long rationale already retained in a spec, plan, report, or grill.
The current handoff skill already prefers pointers over copied content (.agents/skills/handoff/SKILL.md:7-13). The optimization is to enforce that contract in phase and worker handoffs.
Review-phase optimization
Proven current boundaries
Review preparation normalizes and freezes one target and governing source set before semantic review (.agents/skills/review-phase/phases/prepare-review.md:29-49).
One review run:
- invokes autoreview exactly once;
- runs Standards, skill adherence, architecture, simplify, and Spec exactly once against the same frozen snapshot;
- lets the parent verify candidates and build one report (.agents/skills/review-phase/phases/run-review.md:20-74).
Retention independently validates identity, freshness, uniqueness, allowed paths, committed bytes, and retained-ref containment. Retryable retention reuses the same local report without rerunning lenses (.agents/skills/review-phase/phases/retain-report.md:35-75).
Return routing independently revalidates the retained pass and derives one route without executing repair (.agents/skills/review-phase/phases/return-route.md:22-61).
These proof boundaries are valuable and remain.
Prompt-only optimized review
One review epoch should:
- Prepare and freeze one Review Packet once.
- Pass only the frozen target identity, source pointers, scope, and required output contract to review actors.
- Invoke autoreview once as advisory input.
- Run the five normative axes exactly once, concurrently when the native harness supports it.
- Require each actor to return only candidates and evidence pointers.
- Let one parent verify, deduplicate, aggregate, and derive the report.
- Retain and validate the same report without repeating semantic review.
- Return the validated route.
After an ordinary accepted repair:
- run Focused Repair Validation;
- rerun only Verification scenarios invalidated by the repair;
- do not automatically run a second Full Code Review Pass.
Run one second Full Code Review Pass only when the accepted repair changes:
- architecture;
- security or authorization;
- a public contract;
- runtime or deployment topology;
- accepted scope.
Do not exceed two Full Code Review Passes in one delivery goal without explicit human direction. These limits are already accepted in AC-019 through AC-023 (apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:101-112).
Review candidate that remains unproven
Combining the five normative axes into one size-adaptive integrated model call could reduce repeated target reads. The current corpus cannot show that it preserves cross-file defect recall, severity, route correctness, or skill adherence. Keep the existing five axes exactly once and parallelize them until a controlled seeded-defect experiment establishes an integrated alternative.
Brainstorm lens synthesis
Intake
Fact: delivery routing may inspect spec, backlog, plan, notes, diff, retained review, handoff, validation, docs, provider, and PR state (.agents/skills/delivery-phase/SKILL.md:13-20; .agents/skills/delivery-phase/phases/router.md:6-66).
Consequence: cold entry needs broad discovery; every warm transition does not.
Candidate: distinguish cold route from warm transition and retain only the compact Delivery Context Packet inside the current task.
Unresolved tradeoff: packet staleness. Any contradiction or changed accepted bounds must force cold reconstruction.
State
Fact: specifications, plans, implementation notes, retained review reports, handoffs, Git, PRs, provider records, and validation evidence already divide authority.
Consequence: adding another durable authority would increase reconciliation.
Candidate: use a disposable model-held projection of those sources for the current task.
Unresolved tradeoff: define the minimum freshness check before trusting a warm transition.
Control
Fact: current worker waves delay newly eligible dependents until wave reconciliation. The accepted spec replaces that barrier with dependency-local release.
Consequence: the native harness can start useful work sooner without a separate scheduler.
Candidate: require the parent prompt to recompute eligibility after each Task Gate and enforce disjoint scopes.
Unresolved tradeoff: exact fan-out policy. More concurrency can increase context and coordination faster than it reduces latency.
Feedback
Fact: the corpus contains many waits and follow-ups, including tasks with no delivery-phase.
Consequence: conversational progress reporting and frequent polling are global costs.
Candidate: communicate only transitions, blockers, decisions, and material evidence. Use one bounded native-harness wait over known children.
Unresolved tradeoff: wait duration versus responsiveness.
Recovery
Fact: review retention already distinguishes retryable retention from semantic review and forbids rerunning lenses for a fresh reusable report (.agents/skills/review-phase/phases/retain-report.md:71-79).
Consequence: recovery should resume the failed operation, not restart its whole phase chain.
Candidate: every failure response names failed step, reusable evidence, invalidated evidence, retry action, and stop condition.
Unresolved tradeoff: logs lack reliable phase markers for historical recovery cost attribution.
Handoff
Fact: the two temporary handoffs contain useful authority pointers but also repeat substantial narrative and stale local-state observations.
Consequence: a new agent must still rediscover live Git and provider state.
Candidate: hand off stable rationale by pointer and include only live identity, current phase, next action, and explicit unknowns.
Unresolved tradeoff: overly terse handoffs may omit the reason for an unusual boundary. That rationale must live in the linked authoritative artifact.
Requirements-ready priorities
Priority 0: proposed requirements frontier
This list combines accepted leaf-spec constraints with grounded umbrella candidates. The umbrella items remain candidates until the requirements grill accepts them.
- Keep the entire solution prompt/skill/context/delegation-level.
- Preserve the user-selected native harness as the executor.
- Distinguish cold route from warm in-task transition.
- Define Delivery Context Packet and Transition contracts.
- Implement the accepted dependency-local Execution Frontier as parent prompt policy.
- Replace copied worker briefs with minimum task kernels and Context Pointers.
- Make the parent the sole updater of shared plan and implementation-note summaries.
- Use one bounded wait over known children and avoid unchanged-state narration.
- Preserve one default Full Code Review Pass and accepted risk-triggered second pass.
- Preserve Verification, review retention, and returned-route boundaries.
Priority 1: controlled experiments before changing semantics
- Determine when delegation pays for its transfer and coordination overhead.
- Determine safe default fan-out by task shape.
- Compare copied briefs with pointer-based packets.
- Measure cold-route versus warm-transition context and resumption costs.
- Compare five independent review axes with an integrated review on seeded defects.
- Test whether any cross-task evidence reuse can be safe under exact identity.
Not requirements candidates
- Delivery runtime or executor;
- scheduler, watcher, outbox, lock, journal, or cache infrastructure;
- removal of accepted assurance gates;
- merger agents or per-task worker branches;
- moving Verification into review-phase;
- making docs ingest own Project Verifier Feature Maps.
Controlled experiment contract
Numerical savings remain unknown until paired baseline and prompt-variant runs use the same goal, accepted artifacts, branch state, model, reasoning setting, tools, and injected delays.
Experiment 1: small and medium delivery
Compare current prompts with the Context Packet and Transition flow. Measure:
- input, cached input, derived noncached input, output, and reasoning tokens;
- model resumptions and compactions;
- bytes and files loaded;
- shell, provider, wait, spawn, and follow-up operations;
- root active and first-to-last elapsed time separately;
- identical acceptance and assurance outcomes.
Experiment 2: skewed dependency graph
Use one fast prerequisite chain and one unrelated delayed task. Verify the dependent starts after its own Task Gate, then measure dispatch latency and idle time.
Experiment 3: worker context
Run identical tasks with full serialized briefs and pointer-based packets. Compare transferred bytes, resolved bytes, token fields, follow-ups, compactions, and acceptance-evidence completeness.
Experiment 4: review
Use identical frozen small, medium, large sparse, and large connected bundles with seeded Standards, skill, architecture, simplicity, Spec, security, and cross-file defects.
Compare:
- current review orchestration;
- optimized five-axis-once orchestration;
- integrated review only as an experimental variant.
Require equal or better finding recall, severity, route correctness, report completeness, target freshness, and retained-report validity.
Experiment 5: failure and resume
Inject a failed Task Gate, stale source, child failure, retention retry, and handoff resume. Require:
- independent work continues where allowed;
- no duplicate mutation;
- no overlapping writes;
- no stale evidence accepted;
- semantic review does not rerun for retention-only failure;
- final state matches the baseline.
No fixed percentage target is justified yet. Establish repeated-run distributions first. Faster execution never compensates for a failed assurance invariant.
Decision ledger
Accepted authority
- Dependency-local release after each task's Task Gate.
- Current-branch workers with disjoint Active Write Scopes.
- Minimum inline task kernel plus Context Pointers.
- Applicable assurance gates remain authoritative.
- Architecture Checkpoints remain cumulative hard barriers.
- One early draft PR after meaningful pushed state.
- implement-spec completes Verification before delivery-owned Code Review.
- One default Full Code Review Pass with focused repair validation and risk-triggered escalation.
- One verify-behavior entrypoint with project-owned progressively disclosed Project Verifier content.
- review-phase remains readonly and separate from Verification.
Grounded candidates for requirements
- Cold-route and warm-transition distinction.
- Delivery Context Packet and phase Transition output.
- Parent-only shared summary updates.
- Bounded native-harness wait policy.
- Minimal worker deltas and evidence pointers.
- One frozen Review Packet reused through one review epoch.
- Explicit failure response naming reusable and invalidated evidence.
Superseded or rejected
- Runtime, executor, kernel, journal, watcher, outbox, lock, and cache-service proposals.
- Treating router re-entry as the sole or proven dominant cause.
- Converting cached input into unique context, billing, or cost.
- Replacing Verification with Code Review or Code Review with Verification.
- Automatic full-review repetition after ordinary repairs.
- Flattening the two accepted leaf specifications.
Unknown
- Exact token, latency, and monetary savings.
- Review-only historical token consumption.
- Optimal fan-out or worker-size threshold.
- Safe integrated-review threshold and defect-recall parity.
- Best wait duration for each native harness.
- Whether cross-task evidence reuse is worth its identity checks.
- Provider and model differences.
Branch authority consolidation
A literal merge or rebase of both source tips would import unrelated changes:
- origin/team/stefan/verification-skill ends at 4fb0d683 and also contains an unrelated autoreview change;
- origin/team/stefan/scheduled-implement-spec contains e08c3efc, an unrelated backlog-provider configuration change.
The consolidated branch is based on origin/main and carries only the relevant authority commits:
| Authority | Original commit | Consolidated commit |
|---|---|---|
| Project Verification requirements closure | a8e02a26 | 252989f7 |
| Project Verification compiled specification | 90999aa4 | bea5bf99 |
| Project Verification ownership correction | 120c99c8 | b6f60a17 |
| Dynamic Implement Spec Execution Policy | 1f40cb9e | 24a723bf |
PR 182 is closed and unmerged. PR 196 is the draft consolidated pull request on team/stefan/delivery-phase-optimization. The current report supersedes its earlier runtime-oriented research revision.
Conflicts and uncertainty
- Historical activity is observational. It supports prioritization and hypotheses, not causality.
- The tasks differ in scope, repository, tools, model behavior, provider latency, user pauses, retry count, and child topology.
- Static skill size is a context proxy. Actual prompt assembly and cache behavior are not fully observable from source.
- The review contract proves five axes are requested; it does not expose a reliable phase-level token ledger.
- Heavy caching may make repeated tokens cheaper, but it does not remove model resumptions, orchestration calls, or context-window pressure.
- Fewer agents can reduce coordination but can also remove useful independent checking. Controlled quality comparisons are required.
Report validation
- Targeted Oxfmt: passed.
- apps/wiki content synchronization check: passed.
- apps/wiki content check: passed.
- Git diff whitespace check: passed.
Synthesized conclusion
There is now a clear requirements frontier.
The delivery flow should become one compact prompt-controlled loop inside the native harness: cold route once, carry a small current context packet, execute the selected phase, consume a structured transition, and continue until a real stop. Implementation should release safe dependents as soon as their own gates pass. Workers should receive pointers and return deltas. Review should freeze once, run its required semantic axes once, aggregate once, retain once, and avoid a second full pass after ordinary repairs.
The research does not authorize skill implementation, backlog mutation, release, merge, or deployment. The next action is a bounded requirements grill that converts Priority 0 into an umbrella prompt-level specification while referencing both accepted leaf specs. Experiments should remain acceptance evidence for numerical targets and any semantic review-topology change.