Harness Intelligence Wiki
Grilling

Delivery Phase Flow Optimization Grill Log

Delivery Phase Flow Optimization Grill Log

Scope

Define one prompt-level delivery flow that composes the retained Project Verification and Dynamic Implement Spec Execution Policy specifications, reduces latency and context overhead across the full delivery lifecycle, and preserves every accepted assurance authority.

Evidence Base

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md
  • apps/wiki/content/docs/project/specs/cli/project-verification-capability/SPEC.md
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md
  • /tmp/harness-verification-handoff.chgr3Z/project-verifier-handoff.md
  • /private/tmp/implement-spec-execution-policy-handoff-2026-09-04.md
  • Current delivery-phase, create-plan, implement-spec, verify-behavior, review-phase, handoff, and related reference files.

Inherited Accepted Authority

  • The user-selected native harness executes. Harness Intelligence supplies prompts, skills, context selection, routing, delegation rules, tool policy, and evidence contracts.
  • No Delivery runtime, executor, scheduler, watcher, outbox, lock, journal, cache service, or new infrastructure is in scope.
  • The two retained leaf specifications remain separate authorities and will be composed by one umbrella specification.
  • Preserve dependency-local Execution Frontier release, Task Gates, current-branch workers, disjoint Active Write Scopes, Context Pointers, Architecture Checkpoints, applicable assurance gates, and early draft pull-request visibility.
  • implement-spec owns Verification through verify-behavior and the Project Verifier. review-phase remains readonly Code Review.
  • Delivery runs one Full Code Review Pass by default. Ordinary repairs receive Focused Repair Validation and affected Verification reruns. A second Full Code Review Pass is limited to accepted high-risk triggers; more than two needs explicit human direction.

Brainstorm Synthesis

The current system repeatedly reconstructs a wide authority surface, transfers broad worker briefs, waits and follows up often, reconciles shared Markdown writes, and reloads review state across several gates. The smallest coherent prompt-level change is to carry compact disposable projections through one active native-harness task, keep durable artifacts authoritative, move shared summary writes to the parent, make waits state-driven, and reuse one frozen review input within one review pass. These remain candidates until accepted below.

In-Task Continuity

Q1

Prerequisites:

  • none

Evidence anchor:

  • .agents/skills/delivery-phase/SKILL.md:Quick Start
  • .agents/skills/delivery-phase/phases/router.md:Inputs
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Full-delivery control loop

Observed constraint:

  • Full Delivery loads one phase at a time and re-enters routing. Existing durable artifacts already divide authority. The research found repeated context carriage but does not justify another durable state system.

Question: Should Full Delivery carry one disposable current-state projection between in-task phases and cold-route only when the task starts or resumes, accepted bounds change, relevant authority changes, or freshness cannot be proven?

Recommendation:

  • Yes.

Why:

  • It removes repeated broad reconstruction while preserving existing authorities and the prompt-only boundary.

Code consequence:

  • delivery-phase and selected phase exits need a transient projection and explicit invalidation contract. No new file, service, or execution engine owns it.

Accepted answer:

  • Yes. Full Delivery carries one disposable current-state projection between in-task phases.
  • It is never a durable authority.
  • A new task, resumed handoff, changed accepted bounds, changed relevant authority, inconsistent branch or target identity, or unprovable freshness forces a cold route.

Shared Artifact Ownership

Q2

Prerequisites:

  • none

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel-worker-brief.md:Scope rule
  • .agents/skills/implement-spec/references/parallel-orchestration.md:Step 5

Observed constraint:

  • Workers currently update plan entries while the parent rereads, validates, and reconciles PLAN.md and IMPLEMENTATION-NOTES.md. Concurrent shared summaries add collision and reconstruction cost.

Question: Should the parent be the sole writer of shared plan and implementation-note summaries while workers write only their Active Write Scope and return concise deltas and evidence pointers?

Recommendation:

  • Yes.

Why:

  • One summary owner removes concurrent Markdown mutation and lets workers return smaller results without changing product-code ownership.

Code consequence:

  • Worker briefs and output contracts stop assigning shared plan or notes writes. Parent reconciliation becomes the single authoritative summary update.

Accepted answer:

  • Yes. The parent is the sole writer of shared PLAN.md and IMPLEMENTATION-NOTES.md summaries.
  • Workers edit only their Active Write Scope and return concise deltas and evidence pointers.

Waiting And Feedback

Q3

Prerequisites:

  • none

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Root orchestration operations
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Amplification point: repeated waiting

Observed constraint:

  • Six focused roots recorded 578 explicit waits and 217 child follow-up or message operations. The corpus does not identify one optimal timeout across native harnesses.

Question: Should waiting be material-state driven: continue independent parent work first, otherwise issue one bounded native-harness wait over all known children, wait again only after a material change or timeout, and never narrate unchanged state?

Recommendation:

  • Yes. Leave the exact timeout to the native harness and controlled tests.

Why:

  • This cuts polling and conversation noise without inventing a cross-harness timing constant.

Code consequence:

  • Orchestration prompts need a shared definition of material child events, bounded wait behavior, and terminal-state handling.

Accepted answer:

  • The parent orchestrates; implementation edits are delegated to scoped workers.
  • The parent processes one completed task at a time, runs its Task Gate, recomputes the Execution Frontier, and immediately dispatches newly eligible dependents.
  • The parent never waits for an unrelated worker wave to complete.
  • When no orchestration action is ready, one bounded native-harness wait may cover active workers and must wake on the first task completion or attention event.
  • Do not narrate unchanged state. Exact timeout remains native-harness-specific.

Review Input And Semantic Topology

Q4

Prerequisites:

  • none

Evidence anchor:

  • .agents/skills/review-phase/phases/prepare-review.md:Bounded Action
  • .agents/skills/review-phase/phases/run-review.md:Inputs
  • .agents/skills/review-phase/phases/retain-report.md:Invariants
  • .agents/skills/review-phase/phases/return-route.md

Observed constraint:

  • Review already freezes one target and governing source set, but the prompt graph reloads related state across prepare, semantic review, retention, and returned-route gates.

Question: Should one frozen disposable review input projection be prepared once per Full Code Review Pass and reused through semantic review, aggregation, report retention, and returned-route validation, while retention and routing still recompute identity and freshness?

Recommendation:

  • Yes.

Why:

  • It removes unchanged input reconstruction while preserving separate retention and route proof boundaries.

Code consequence:

  • Review gates share one immutable input contract. Any bounds, target, or governing-source mismatch invalidates it and returns to preparation.

Accepted answer:

  • Yes. One frozen disposable review input projection is prepared for one Full Code Review Pass and reused through semantic review, aggregation, retention, and returned-route validation.
  • Retention and routing retain their independent identity and freshness checks.

Q5

Prerequisites:

  • none

Evidence anchor:

  • .agents/skills/review-phase/phases/run-review.md:Bounded Action
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Review candidate remains unproven

Observed constraint:

  • Current review invokes autoreview once and five distinct normative axes once. Combining the axes may reduce repeated reads, but current evidence cannot prove equal defect recall, severity, route correctness, or skill-adherence coverage.

Question: Should the first implementation preserve the five distinct normative axes, run them once and concurrently when supported, and keep one integrated size-adaptive semantic reviewer experimental until seeded-defect parity is proven?

Recommendation:

  • Yes.

Why:

  • It captures proven orchestration savings now and makes the more aggressive semantic change earn its safety case.

Code consequence:

  • Initial review changes optimize input sharing, briefs, concurrency, aggregation, and pass count. Integrated review remains behind a controlled experiment and later accepted requirement.

Accepted answer:

  • Yes. Keep the five normative review axes distinct, run each exactly once and concurrently when supported, and invoke autoreview once.
  • Keep an integrated size-adaptive semantic reviewer experimental until controlled seeded-defect comparison proves parity.

Fan-Out Policy

Q6

Prerequisites:

  • none

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel.md:Contract
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-001..AC-006
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:All discoverable actors

Observed constraint:

  • Safe tasks should release immediately, but fan-out magnifies context and coordination. Native harness capacity differs, and the evidence does not support one universal worker count.

Question: Should fan-out be eligibility- and native-capacity-based, with no universal numeric cap, so every eligible worker-worthy task may start when slots and disjoint Active Write Scopes permit?

Recommendation:

  • Yes.

Why:

  • A global number would either idle capable harnesses or overload smaller ones. The task graph and native capacity are the useful controls.

Code consequence:

  • Delivery and planning prompts need a worker-worthiness rule and must query current native capacity rather than encode one fixed fan-out number.

Accepted answer:

  • Yes. Fan-out follows task eligibility, native-harness capacity, disjoint Active Write Scopes, and worker-worthiness.
  • Do not encode one universal numeric worker cap.

Performance Proof

Q7

Prerequisites:

  • none

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Controlled experiment contract
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Conflicts and uncertainty

Observed constraint:

  • Historical tasks show large processed context and orchestration activity, but they differ too much to establish causal savings or a defensible fixed percentage target.

Question: Should delivery optimization require paired baseline and variant evidence with unchanged assurance outcomes, while deferring any fixed token or latency percentage target until repeated-run distributions exist?

Recommendation:

  • Yes.

Why:

  • It makes speed claims falsifiable without inventing a threshold from observational data.

Code consequence:

  • The umbrella specification must define controlled fixtures, measures, comparison identity, and non-regression gates; numerical targets remain provisional.

Accepted answer:

  • Yes. Require paired baseline and optimized measurements with identical assurance outcomes.
  • Do not set a fixed token or latency percentage until repeated-run distributions support one.

Compact Delivery Contracts

Q8

Prerequisites:

  • Q1

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Delivery Context Packet
  • .agents/skills/delivery-phase/phases/router.md:Inputs
  • .agents/skills/delivery-phase/references/phase-handoff.md:Base Shape

Observed constraint:

  • The router needs enough current state to select one next phase, but copying specifications, plans, reports, skill bodies, and handoff prose recreates the cost this change targets.

Question: Should Delivery Context Packet become the canonical term for the disposable current-task projection, containing only goal and accepted-bounds identity, authoritative artifact pointers, branch/base/target identity, current phase and next action, active Task or review identity, relevant evidence freshness, exact blocker, and stop condition?

Recommendation:

  • Yes. Exclude copied artifact bodies, full reports, skill text, command logs, and retained rationale.

Why:

  • This is the smallest state set that can route safely and resolve detailed context only when the selected action needs it.

Code consequence:

  • delivery-phase defines one canonical packet schema used only inside the active native-harness task.

Accepted answer:

  • Delivery Context Packet is the canonical disposable current-task projection.
  • It contains only goal and accepted-bounds identity, authoritative artifact pointers, branch/base/target identity, current phase and next action, active Task or review identity, relevant evidence freshness, exact blocker, and stop condition.
  • It excludes copied artifact bodies, full reports, skill text, command logs, and retained rationale.

Q9

Prerequisites:

  • Q1

Evidence anchor:

  • .agents/skills/delivery-phase/phases/router.md:Inputs
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:State

Observed constraint:

  • A disposable projection saves work only while its routing facts remain fresh. Trusting stale bounds, target, branch, or evidence could select the wrong phase.

Question: Should the projection remain valid only inside the same active task and delivery-goal identity while accepted-bounds hash, next-action authority pointers, branch/base/target identity, and required evidence freshness still match?

Recommendation:

  • Yes. Rebuild on a new task or handoff resume, changed mode or bounds, changed required authority, inconsistent Git target, stale evidence, or any unresolved contradiction. A normal phase change alone does not rebuild it.

Why:

  • These checks protect routing authority without forcing every phase to reconstruct unrelated state.

Code consequence:

  • Warm continuation receives an explicit validity predicate; failed predicate returns to cold routing.

Accepted answer:

  • Yes. Warm continuation requires the same active task and delivery-goal identity, matching accepted bounds, matching next-action authority pointers, consistent branch/base/target identity, and fresh required evidence.
  • A new task or handoff resume, changed mode or bounds, changed required authority, inconsistent Git target, stale evidence, or contradiction forces cold routing.
  • A normal phase change alone does not force cold routing.

Q10

Prerequisites:

  • Q1
  • Q3

Evidence anchor:

  • .agents/skills/delivery-phase/references/phase-handoff.md:Base Shape
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Transition

Observed constraint:

  • Phase exits currently leave varying completion prose. The next route needs changed facts and invalidations, not a repeated account of the whole phase.

Question: Should Phase Result become the canonical compact phase-exit contract containing phase, outcome, authority or evidence created, changed facts, invalidated evidence, next eligible phase, exact blocker or stop reason, and only the Context Pointers needed next?

Recommendation:

  • Yes. Keep detailed rationale in its authoritative artifact and reference it.

Why:

  • A uniform result gives the parent one material event that can update the Delivery Context Packet without another broad scan.

Code consequence:

  • Every delivery phase returns the same minimum result envelope while retaining its current phase-specific durable evidence.

Accepted answer:

  • Phase Result is the canonical compact phase-exit contract.
  • It contains phase, outcome, authority or evidence created, changed facts, invalidated evidence, next eligible phase, exact blocker or stop reason, and only the Context Pointers needed next.
  • Detailed rationale remains in its authoritative artifact.

Granular Task Flow

Q11

Prerequisites:

  • Q2
  • Q3

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel-worker-brief.md:Worker output contract
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-004..AC-008

Observed constraint:

  • The parent needs enough worker evidence to run a Task Gate. Full logs, copied requirements, and worker-authored shared summaries are not required for that decision.

Question: Should Task Result become the canonical worker return contract containing task identity, changed paths and concise delta, acceptance outcomes, RED/GREEN and required validation evidence pointers, required skill-evidence pointers, architecture or public-seam deltas, invalidated context, and an exact blocker or risk?

Recommendation:

  • Yes. Exclude copied source text, full command logs, and shared PLAN.md or IMPLEMENTATION-NOTES.md edits.

Why:

  • This retains every Task Gate input while minimizing transfer and giving the parent a predictable validation surface.

Code consequence:

  • Worker briefs and all worker roles share one result envelope; the parent rejects incomplete results instead of requesting broad narrative follow-ups.

Accepted answer:

  • Task Result is the canonical compact worker return contract.
  • It contains task identity, changed paths and concise delta, acceptance outcomes, RED/GREEN and required validation evidence pointers, required skill-evidence pointers, architecture or public-seam deltas, invalidated context, and an exact blocker or risk.
  • It excludes copied source text, full command logs, and shared plan or implementation-note edits.

Q12

Prerequisites:

  • Q1
  • Q2
  • Q3

Evidence anchor:

  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-001..AC-004
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-024..AC-025

Observed constraint:

  • Dependency-local release is already accepted. Shared-artifact reconciliation remains authoritative but does not need to delay a safe dependent.

Question: After one Task Result passes its Task Gate, should the parent release its Active Write Scope, recompute and dispatch the newly eligible frontier immediately, then reconcile shared summaries while those workers run, with reconciliation mandatory before an applicable Architecture Checkpoint or finalization?

Recommendation:

  • Yes.

Why:

  • This implements the corrected task-granular flow and removes the whole-wave barrier without losing durable records.

Code consequence:

  • Dispatch moves ahead of shared-summary writes on the critical path; checkpoints and finalization still require complete reconciliation.

Accepted answer:

  • Yes. After a Task Result passes its Task Gate, the parent releases its Active Write Scope, recomputes the Execution Frontier, and dispatches newly eligible Tasks immediately.
  • The parent reconciles shared summaries while those workers run.
  • Reconciliation remains mandatory before an applicable Architecture Checkpoint or finalization.

Q13

Prerequisites:

  • Q2
  • Q6

Evidence anchor:

  • .agents/skills/create-plan/references/plan-schema.md:Task contract
  • .agents/skills/implement-spec/references/parallel.md:Contract

Observed constraint:

  • Native capacity can only improve throughput when each delegated Task has a real ownership and acceptance boundary. Very small worker tasks can cost more to transfer and validate than they save.

Question: Should Worker-worthy Task mean an atomic, independently ownable plan outcome with an exclusive Active Write Scope, explicit acceptance boundary, Task Gate checks, and enough cohesive implementation work to justify delegation, with smaller operations folded into the nearest such Task during planning?

Recommendation:

  • Yes. The parent still delegates every implementation edit; it does not absorb micro-edits during execution.

Why:

  • This controls coordination cost before fan-out without imposing a universal task-size or worker-count constant.

Code consequence:

  • create-plan must consolidate microtasks before implement-spec; the execution frontier contains only worker-worthy Tasks.

Accepted answer:

  • Worker-worthy Task is an atomic, independently ownable plan outcome with an exclusive Active Write Scope, explicit acceptance boundary, Task Gate checks, and enough cohesive implementation work to justify delegation.
  • create-plan folds smaller operations into the nearest Worker-worthy Task.
  • The parent delegates every implementation edit and never absorbs micro-edits during execution.

Frozen Review Contract

Q14

Prerequisites:

  • Q4

Evidence anchor:

  • .agents/skills/review-phase/phases/prepare-review.md:Completion Evidence
  • .agents/skills/review-phase/phases/run-review.md:Inputs

Observed constraint:

  • Each review actor needs the same frozen identity and scope but only its own lens-specific source detail.

Question: Should Review Packet become the canonical frozen input for one Full Code Review Pass, containing mode, lineage, run identity and ordinal, accepted-bounds identity and hash, normalized target locator and snapshot hash, governing-source pointers and hashes, applicable Spec, Standards, scoped guidance, plan, implementation-note and Verification evidence pointers, and the report and routing output contracts?

Recommendation:

  • Yes. Do not embed full diffs, source files, skill bodies, reports, or validation logs; actors resolve only their lens-specific pointers.

Why:

  • Every actor receives identical review authority while repeated source bodies stay out of orchestration prompts.

Code consequence:

  • Prepare Review creates one immutable packet; autoreview, five normative axes, aggregation, retention, and return routing consume its identity.

Accepted answer:

  • Review Packet is the canonical frozen input for one Full Code Review Pass.
  • It contains mode, lineage, run identity and ordinal, accepted-bounds identity and hash, normalized target locator and snapshot hash, governing-source pointers and hashes, applicable Spec, Standards, scoped guidance, plan, implementation-note and Verification evidence pointers, and report and routing output contracts.
  • It excludes full diffs, source files, skill bodies, reports, and validation logs. Each actor resolves only lens-specific pointers.

Q15

Prerequisites:

  • Q4
  • Q5

Evidence anchor:

  • .agents/skills/review-phase/phases/run-review.md:Declared Exits
  • .agents/skills/review-phase/phases/retain-report.md:Declared Exits
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-019..AC-023

Observed constraint:

  • Some failures make semantic findings stale; report retention and provider failures do not when the frozen identities still match.

Question: Should accepted-bounds, target snapshot, governing-source set, or required skill-obligation changes invalidate the Review Packet and semantic run, while retention, push, ref-approval, or returned-route failures reuse the existing report and packet when hashes still match?

Recommendation:

  • Yes. Ordinary repairs use Focused Repair Validation. Only accepted high-risk repair triggers create a second Review Packet and Full Code Review Pass.

Why:

  • This retries the failed operation instead of paying for semantic review again.

Code consequence:

  • Review recovery gains explicit semantic-invalidating and retention-only failure classes.

Accepted answer:

  • Accepted-bounds, target snapshot, governing-source set, or required skill-obligation changes invalidate the Review Packet and semantic run.
  • Retention, push, ref-approval, and returned-route failures reuse the existing report and Review Packet when their hashes still match.
  • Ordinary repairs use Focused Repair Validation. Only accepted high-risk repair triggers create a second Review Packet and Full Code Review Pass.

Performance Acceptance

Q16

Prerequisites:

  • Q7

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Controlled experiment contract
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Known accounting limits

Observed constraint:

  • Wall time, fresh input, cached processed context, model resumptions, and coordination operations can move in different directions. One fixed percentage is not supported yet.

Question: Should performance use a Pareto acceptance rule: assurance outcomes must match, at least one primary efficiency measure must improve beyond observed run-to-run noise, and no other primary measure may regress beyond that noise?

Recommendation:

  • Yes. Treat end-to-end time, noncached input, total processed input, resumptions and compactions, bytes loaded, and wait, spawn, follow-up, shell, and provider operations as separately reported primary measures.

Why:

  • This accepts genuine improvements without hiding a major cost increase behind one faster number.

Code consequence:

  • Any meaningful speed-versus-token regression remains an explicit human tradeoff instead of an automatic pass.

Accepted answer:

  • Use an assurance-constrained Pareto rule.
  • Assurance outcomes must match, at least one primary efficiency measure must improve beyond observed run-to-run noise, and no other primary measure may regress beyond that noise.
  • Report end-to-end time, noncached input, total processed input, resumptions and compactions, bytes loaded, and wait, spawn, follow-up, shell, and provider operations separately.
  • A meaningful speed-versus-token regression requires an explicit human decision.

Q17

Prerequisites:

  • Q5
  • Q7

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 1
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 2
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 3
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 4
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 5

Observed constraint:

  • A single happy-path run cannot expose dependency skew, context-transfer cost, review recall, or recovery regressions.

Question: Should the required comparison suite cover small and medium full delivery, a skewed dependency graph, full versus pointer-based worker context, small and large frozen reviews with seeded defects, failure and resume cases, and both covered and Uncovered Behavior Project Verifier paths?

Recommendation:

  • Yes.

Why:

  • Together these fixtures exercise every optimization boundary and both retained leaf specifications.

Code consequence:

  • The umbrella specification defines fixture outcomes and measurement fields; it does not require new runtime infrastructure.

Accepted answer:

  • The controlled comparison covers small and medium full delivery, a skewed dependency graph, full versus pointer-based worker context, small and large frozen reviews with seeded defects, failure and resume cases, and both covered and Uncovered Behavior Project Verifier paths.
  • The fixture suite uses the native harness and existing evidence surfaces; it creates no runtime infrastructure.

Pointer And Result Authority

Q18

Prerequisites:

  • Q8
  • Q11
  • Q14

Evidence anchor:

  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-007..AC-008
  • apps/wiki/content/docs/project/specs/cli/project-verification-capability/SPEC.md:AC-033
  • .agents/skills/review-phase/phases/prepare-review.md:Completion Evidence

Observed constraint:

  • Delivery Context Packet, Task Result, and Review Packet all depend on Context Pointers, but a path without authority identity or freshness cannot prevent stale context use.

Question: Should every Context Pointer contain authority kind, stable locator, optional section or symbol selector, content or state identity, and the freshness rule the consumer must check?

Recommendation:

  • Yes. The locator may be a repository path, immutable commit and path, provider identity, evidence locator, or another existing authoritative handle. Do not embed the pointed content.

Why:

  • One small pointer shape supports progressive disclosure while making stale or ambiguous references detectable.

Code consequence:

  • Plans, packets, results, handoffs, and review inputs use the same Context Pointer contract.

Accepted answer:

  • Every Context Pointer contains authority kind, stable locator, optional section or symbol selector, content or state identity, and the freshness rule its consumer must check.
  • The locator uses an existing authoritative handle and never embeds the pointed content.

Q19

Prerequisites:

  • Q10

Evidence anchor:

  • .agents/skills/delivery-phase/references/phase-handoff.md:Base Shape
  • .agents/skills/delivery-phase/phases/router.md:Route

Observed constraint:

  • Phase-specific internal states are detailed and useful, but the delivery parent needs a small common outcome set to choose continue, retry, handback, or stop.

Question: Should every Phase Result use exactly one common outcome: complete, blocked, failed, skipped, or human_steering_required, while phase-specific state remains behind a Context Pointer?

Recommendation:

  • Yes. blocked names the missing condition and resume action; failed means the current bounded attempt cannot continue; skipped requires an exact no-op reason.

Why:

  • A five-value envelope is enough for routing and prevents every phase from copying its internal state graph into the parent prompt.

Code consequence:

  • The Delivery Context Packet consumes one common outcome while selected phases retain their existing detailed state contracts.

Accepted answer:

  • Every Phase Result uses exactly one common outcome: complete, blocked, failed, skipped, or human_steering_required.
  • blocked names the missing condition and resume action; failed means the current bounded attempt cannot continue; skipped requires an exact no-op reason.
  • Phase-specific internal state remains behind a Context Pointer.

Q20

Prerequisites:

  • Q11
  • Q12

Evidence anchor:

  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-004
  • .agents/skills/implement-spec/references/parallel-orchestration.md:Completion rule

Observed constraint:

  • A worker supplies evidence but cannot authoritatively declare its own Task Gate passed. The parent owns validation and dependent release.

Question: Should a Task Result report only ready_for_gate or blocked, while the parent alone records the Task Gate outcome as passed, repair_required, or blocked?

Recommendation:

  • Yes. A missing or malformed Task Result is blocked; it never becomes an inferred pass.

Why:

  • This keeps task acceptance with the orchestrator and separates worker completion from safe downstream consumption.

Code consequence:

  • Dependents enter the Execution Frontier only after a parent-recorded passed Task Gate.

Accepted answer:

  • A Task Result reports only ready_for_gate or blocked.
  • The parent alone records the Task Gate outcome as passed, repair_required, or blocked.
  • A missing or malformed Task Result is blocked; it never becomes an inferred pass.

Recovery And Cross-Task Handoff

Q21

Prerequisites:

  • Q8
  • Q9
  • Q10

Evidence anchor:

  • .agents/skills/handoff/SKILL.md
  • .agents/skills/delivery-phase/references/phase-handoff.md:Base Shape
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Handoff

Observed constraint:

  • A Delivery Context Packet is intentionally task-local. A new task still needs enough durable identity to reconstruct current state without repeated narrative.

Question: Should Delivery Handoff contain only delivery-goal and bounds identity, branch/base/target identity, current phase and next action, stable Context Pointers with identities, relevant provider identities, exact blocker, explicit unknowns, and at most one sentence explaining an unusual boundary with a pointer to its rationale?

Recommendation:

  • Yes. Never persist the Delivery Context Packet itself. The receiving task always cold-routes from the handoff and current authorities.

Why:

  • This preserves resumability without turning transient state or stale rationale into a second source of truth.

Code consequence:

  • $handoff and delivery phase handoffs share one compact cross-task contract.

Accepted answer:

  • Delivery Handoff contains only delivery-goal and bounds identity, branch/base/target identity, current phase and next action, stable Context Pointers with identities, relevant provider identities, exact blocker, explicit unknowns, and at most one sentence explaining an unusual boundary with a pointer to its rationale.
  • Never persist the Delivery Context Packet itself. The receiving task always cold-routes from the handoff and current authorities.

Q22

Prerequisites:

  • Q11
  • Q12
  • Q13

Evidence anchor:

  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-003
  • .agents/skills/implement-spec/references/parallel.md:Loop rule
  • .agents/skills/implement-spec/references/parallel-worker-brief.md:Worker output contract

Observed constraint:

  • A failed or incomplete task may leave partial changes in the shared current branch. Releasing the same paths immediately could create overlapping ownership or hide the actual blocker.

Question: When a Task Result is blocked or its Task Gate requires repair, should only its dependent chain stop, while independent disjoint Tasks continue and its Active Write Scope stays unavailable until the parent either validates cleanup or assigns exactly one scoped repair worker?

Recommendation:

  • Yes. The repair brief points to the original Task Result and only the missing gate or defect; the parent still makes no implementation edit.

Why:

  • This contains failure to the affected chain and preserves shared-branch write isolation.

Code consequence:

  • Task recovery becomes scope-local and idempotent instead of restarting a wave or stopping the entire execution frontier.

Accepted answer:

  • A blocked or repair-required Task stops only its dependent chain. Independent disjoint Tasks continue.
  • Its Active Write Scope remains unavailable until the parent validates cleanup or assigns exactly one scoped repair worker.
  • The repair brief points to the original Task Result and only the missing gate or defect. The parent makes no implementation edit.

Review Actor Contract

Q23

Prerequisites:

  • Q14
  • Q15

Evidence anchor:

  • .agents/skills/review-phase/phases/run-review.md:Actor-Like Gate Boundary
  • .agents/skills/review-phase/phases/run-review.md:Bounded Action

Observed constraint:

  • Review actors need a compact return that the parent can verify and deduplicate. They do not own the retained report or final routing decision.

Question: Should Review Lens Result contain lens identity, Review Packet identity, explicit clean or candidate findings, and for each candidate only location, evidence pointer, impact, proposed severity, proposed return route, and uncertainty?

Recommendation:

  • Yes. The parent verifies candidates, deduplicates them, assigns stable finding IDs, confirms severity and route, and alone assembles the report.

Why:

  • This removes repeated target summaries and prevents advisory actors from becoming competing review authorities.

Code consequence:

  • autoreview and all five normative axes return one common candidate envelope, with advisory versus normative provenance preserved.

Accepted answer:

  • Review Lens Result contains lens identity, Review Packet identity, explicit clean or candidate findings, and for each candidate only location, evidence pointer, impact, proposed severity, proposed return route, and uncertainty.
  • The parent verifies and deduplicates candidates, assigns stable finding IDs, confirms severity and route, and alone assembles the report.
  • Advisory versus normative provenance remains explicit.

Frontier Priority

Q24

Prerequisites:

  • Q12
  • Q13

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel-orchestration.md:Step 3
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-001..AC-003

Observed constraint:

  • Native capacity may be smaller than the eligible Execution Frontier. Plan order alone can delay the task that unlocks the most downstream work or exposes a risky assumption.

Question: When eligible Worker-worthy Tasks exceed native capacity, should the parent first run any task whose result could invalidate substantial downstream work, then prefer the longest remaining dependency chain, then the task unlocking more dependents, with plan order as the stable tie-break?

Recommendation:

  • Yes.

Why:

  • This discovers expensive wrong assumptions early and otherwise shortens the critical path without a scheduler service.

Code consequence:

  • implement-spec uses one prompt-level priority rule when it cannot launch the complete eligible frontier.

Accepted answer:

  • When eligible Worker-worthy Tasks exceed native capacity, first run any Task whose result could invalidate substantial downstream work.
  • Then prefer the longest remaining dependency chain, then the Task unlocking more dependents, with plan order as the stable tie-break.

Measurement Reliability

Q25

Prerequisites:

  • Q16
  • Q17

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Controlled experiment contract
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Conflicts and uncertainty

Observed constraint:

  • One paired run cannot distinguish a prompt improvement from ordinary model and tool variance, while unlimited repetitions would consume the resources the change aims to save.

Question: Should each fixture run as three matched baseline/variant pairs, extend to five only when the first three disagree on the Pareto result, and return inconclusive after five rather than sampling until a preferred result appears?

Recommendation:

  • Yes.

Why:

  • This gives bounded repeated evidence and prevents cherry-picking or an unbounded benchmark cost.

Code consequence:

  • The benchmark report retains every pair, comparison identity, distribution, and final passed, failed, or inconclusive outcome.

Accepted answer:

  • Each fixture runs as three matched baseline and optimized pairs.
  • Extend to five only when the first three disagree on the Pareto result.
  • After five, return inconclusive rather than sampling until a preferred result appears.
  • Retain every pair, comparison identity, distribution, and final passed, failed, or inconclusive outcome.

Q26

Prerequisites:

  • Q5
  • Q17

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Experiment 4
  • .agents/skills/review-phase/phases/run-review.md:Completion Evidence

Observed constraint:

  • Faster review is unacceptable if it loses high-impact findings, weakens a mandatory axis, misroutes repair, or accepts a stale target.

Question: Should review parity require no missed seeded critical or high finding, equal or better per-axis recall, severity and route correctness, complete lens outcomes, unchanged target freshness, and valid retained-report evidence across the matched runs?

Recommendation:

  • Yes. Any critical or high miss fails the optimized review regardless of token or time savings.

Why:

  • This makes the review optimization prove the exact safety properties its parallelism and context reduction could threaten.

Code consequence:

  • The five-axis optimization must pass parity before becoming the default; an integrated reviewer still requires a later accepted requirement even if its experiment passes.

Accepted answer:

  • Review parity requires no missed seeded critical or high finding, equal or better per-axis recall, severity and route correctness, complete lens outcomes, unchanged target freshness, and valid retained-report evidence across matched runs.
  • Any critical or high miss fails the optimized review regardless of token or time savings.
  • An integrated reviewer still requires a later accepted requirement even if its experiment passes.

Pointer Resolution

Q27

Prerequisites:

  • Q18

Evidence anchor:

  • .agents/skills/delivery-phase/phases/router.md:Inputs
  • .agents/skills/delivery-phase/references/phase-handoff.md:Base Shape
  • apps/wiki/content/docs/project/specs/cli/implement-spec-execution-policy/SPEC.md:AC-008

Observed constraint:

  • A missing, stale, malformed, or ambiguous Context Pointer cannot safely supply authority to its consumer. Broad rediscovery or inferred content would erase the pointer's trust boundary.

Question: Should the consumer stop before acting, let the parent perform one bounded refresh from the named authority, update and cold-route only when one current identity is proven, and otherwise return blocked with the exact pointer and missing proof?

Recommendation:

  • Yes. Never invent pointed content, silently use a near match, or fall back to copying a broad authority set.

Why:

  • This makes progressive disclosure fail closed without turning every pointer miss into an unbounded research pass.

Code consequence:

  • All packet, result, worker, review, and handoff consumers share one pointer-resolution failure contract.

Accepted answer:

  • A consumer stops before acting when a Context Pointer is missing, stale, malformed, or ambiguous.
  • The parent may perform one bounded refresh from the named authority. It updates the pointer and cold-routes only when that refresh proves exactly one current identity.
  • Otherwise the result is blocked and names the pointer plus the missing proof. The consumer never invents pointed content, selects a near match, or broad-loads authorities as fallback.

Entry-Mode Compatibility

Q28

Prerequisites:

  • Q9
  • Q19
  • Q21

Evidence anchor:

  • .agents/skills/delivery-phase/SKILL.md:Entry Modes
  • .agents/skills/delivery-phase/SKILL.md:Stop Conditions
  • .agents/skills/review-phase/SKILL.md:Bootstrap

Observed constraint:

  • Full Delivery is the only mode authorized to continue across phases. Direct phase, HITL checkpoint, standalone review, resume, and closeout requests retain narrower user-selected boundaries.

Question: Should only Full Delivery use the continuing Delivery Context Packet loop, while direct one-phase and standalone invocations cold-route their own bounded input, emit one Phase Result, and stop at the requested boundary?

Recommendation:

  • Yes. A same-task Full Delivery continuation may stay warm when Q9 remains true; every cross-task handoff resume cold-routes.

Why:

  • This improves end-to-end delivery without silently expanding authority granted by a narrow invocation.

Code consequence:

  • Existing public entrypoints and explicit HITL stops remain compatible with the optimized loop.

Accepted answer:

  • Only Full Delivery uses the continuing Delivery Context Packet loop.
  • A direct phase, HITL checkpoint, standalone review, resume, or closeout invocation cold-routes its bounded input, emits one Phase Result, and stops at the requested boundary.
  • A same-task Full Delivery continuation may stay warm while Q9 remains true. Every cross-task handoff resume cold-routes.

Q29

Prerequisites:

  • Q6
  • Q14
  • Q23
  • Q24

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel.md:Contract
  • .agents/skills/review-phase/phases/run-review.md:Bounded Action

Observed constraint:

  • Native harnesses expose different concurrency. The accepted model still separates the parent orchestrator from implementation workers and preserves five distinct review axes.

Question: When only one worker slot exists, should Tasks and review axes run sequentially through the same packets and gates, while a harness with no subagent capability returns blocked instead of letting the parent silently absorb implementation edits?

Recommendation:

  • Yes. Sequential capacity changes throughput, not ownership or assurance.

Why:

  • This preserves the accepted execution model across harnesses without pretending unavailable concurrency exists.

Code consequence:

  • Capacity one is supported; capacity zero for required delegated implementation is an explicit capability blocker.

Accepted answer:

  • Capacity one runs Tasks and review axes sequentially through the same packets and gates.
  • Capacity changes throughput, not parent ownership, worker ownership, or assurance.
  • A harness with no subagent capability returns blocked when delegated implementation is required. The parent never absorbs implementation edits.

Contract Rollout And Legacy Artifacts

Q30

Prerequisites:

  • Q18
  • Q19
  • Q20
  • Q21
  • Q23

Evidence anchor:

  • AGENTS.md:Shared skills source of truth
  • .agents/skills/delivery-phase/SKILL.md:Phase Files
  • .agents/skills/implement-spec/SKILL.md:Advanced features
  • .agents/skills/review-phase/SKILL.md:Bootstrap

Observed constraint:

  • Delivery Context Packet, Phase Result, Task Result, Review Packet, Review Lens Result, and Delivery Handoff cross several skills. A partial rollout would create mixed vocabulary and incompatible expectations.

Question: Should canonical shared-source changes to delivery-phase, create-plan, implement-spec, review-phase, handoff, their required references, tests, and operator documentation ship as one compatible contract change, then sync to Harness with the exact source receipt in this single Harness pull request?

Recommendation:

  • Yes. Project Verifier implementation remains its own leaf module and plan Tasks but lands in the same Harness delivery change set.

Why:

  • One synchronized contract avoids an old producer feeding a new consumer or the reverse while preserving leaf-spec ownership.

Code consequence:

  • Implementation planning may use dependency-local Tasks, but final contract activation is atomic at the shared-skill version boundary.

Accepted answer:

  • Roll out the shared contracts as one compatible canonical-source change across delivery-phase, create-plan, implement-spec, review-phase, handoff, required references, tests, and operator documentation.
  • Push the canonical Devpunks source first, then sync its exact receipt into this single Harness pull request.
  • Project Verifier remains a separately owned leaf module and plan scope but lands in the same Harness delivery change set.

Q31

Prerequisites:

  • Q8
  • Q10
  • Q11
  • Q14
  • Q21

Evidence anchor:

  • .agents/skills/delivery-phase/references/artifact-state.md
  • .agents/skills/review-phase/phases/router.md:Inputs Authority

Observed constraint:

  • Existing plans, notes, handoffs, and retained review reports predate the new compact contracts. Rewriting historical artifacts would add migration cost and could invalidate retained evidence.

Question: Should new writes use the compact contracts while a cold route derives them from valid legacy artifacts without bulk rewriting, preserves valid retained review reports and ordinals, and returns blocked when required authority cannot be reconstructed?

Recommendation:

  • Yes.

Why:

  • This gives forward compatibility without a repository-wide migration or loss of existing evidence.

Code consequence:

  • Compatibility is a read-time normalization boundary; historical authoritative bytes remain unchanged.

Accepted answer:

  • New writes use the compact contracts.
  • A cold route normalizes valid legacy plans, notes, handoffs, and retained review reports at read time without rewriting historical authoritative bytes.
  • Preserve valid retained review identities and ordinals. Return blocked when required authority cannot be reconstructed.

Bounded Repair

Q32

Prerequisites:

  • Q20
  • Q22

Evidence anchor:

  • .agents/skills/implement-spec/references/parallel.md:Loop rule
  • .agents/skills/delivery-phase/phases/implement.md:Rules

Observed constraint:

  • A fixed task-retry count is not supported, but repeated generic follow-ups with no new evidence consume context and cannot change the result.

Question: Should another scoped repair be allowed only when the latest Task Result or Task Gate supplies new actionable evidence, while the same blocker without new evidence, a required scope redesign, or a weaker gate returns human_steering_required rather than looping?

Recommendation:

  • Yes. Do not impose one universal retry count; require evidence progress on every retry.

Why:

  • This preserves useful autonomous repair while terminating stagnant conversational loops.

Code consequence:

  • Repair progression is evidence-based and idempotent, with $handback retaining its existing authority boundary.

Accepted answer:

  • Permit another scoped repair only when the latest Task Result or Task Gate supplies new actionable evidence.
  • The same blocker without new evidence, a required scope redesign, or a weaker gate returns human_steering_required through $handback.
  • Do not impose one universal retry count. Every retry must prove evidence progress.

Optimization Admission

Q33

Prerequisites:

  • Q25
  • Q26

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Controlled experiment contract

Observed constraint:

  • An inconclusive benchmark proves neither regression nor the required efficiency improvement.

Question: Should failed or inconclusive performance evidence prevent the optimized contracts from becoming the default unless the human explicitly accepts the measured tradeoff?

Recommendation:

  • Yes. Never label or activate an unproven default optimization automatically.

Why:

  • The change exists to improve delivery measurably; absence of proof must remain visible.

Code consequence:

  • Default activation requires a passing benchmark or an explicit retained human exception.

Accepted answer:

  • A failed or inconclusive benchmark prevents the optimized contracts from becoming the default.
  • Default activation requires passing performance evidence or an explicit retained human exception accepting the measured tradeoff.
  • Never label or activate an unproven optimization automatically.

Q34

Prerequisites:

  • Q8
  • Q10
  • Q11
  • Q14
  • Q23

Evidence anchor:

  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Prompt-level target
  • apps/wiki/content/docs/project/research/delivery-phase-flow-optimization-research-report.md:Known accounting limits

Observed constraint:

  • A universal token limit is not grounded and could truncate required evidence, but an unconstrained narrative envelope would recreate the current context cost.

Question: Should compactness be enforced by required fields, pointer-only authority, forbidden copied payloads, and concise-delta output rather than one universal token or word limit, with later benchmarks allowed to justify numeric budgets?

Recommendation:

  • Yes.

Why:

  • Structural bounds are testable now and preserve required evidence across differently sized Tasks and reviews.

Code consequence:

  • Contract tests reject embedded authority bodies and unbounded narrative fields; no arbitrary truncation threshold ships initially.

Accepted answer:

  • Enforce compactness through required fields, pointer-only authority, forbidden copied payloads, and concise-delta output.
  • Contract tests reject embedded authority bodies and unbounded narrative fields.
  • Do not ship a universal token or word limit. Later benchmark evidence may justify numeric budgets.

Accepted Round R5 — 2026-09-05

This compact record captures the accepted Astra re-analysis. It does not reopen the retained leaf decision trails or the existing AC-044..AC-048 experiment policy.

Q35 — Trigger-specific instruction surfaces

Accepted: load project facts and hard boundaries broadly, workflow contracts required by the current action selectively, and optional techniques only when their trigger applies. Record selected surface identities and freshness in disposable context.

Q36 — Bounded derived context

Accepted: permit a bounded source-attributed excerpt or rationale when it avoids a predictable lookup. Carry source identity and freshness and mark it non-authoritative. Full bodies, broad payload dumps, and unsourced excerpts remain invalid.

Q37 — Diagnosis before repair

Accepted: bounded authorized hypothesis-discriminating diagnosis may produce new actionable evidence before another repair. Stagnant repeated repair, redesign, changed requirements, missing access, human decisions, and weaker-gate requests hand back through $handback.

Q38 — Planning read and runtime conflicts

Accepted: the swarm-planner primitive within create-plan, the planning step of delivery-phase, records read dependencies, shared runtime resources, and evidence-stability inputs alongside write scopes. Owning execution and evidence gates recheck affected inputs; no lock or execution service is introduced.

Q39 — Stronger verification composition

Accepted: compose the Project Verifier leaf contract for falsifiable claims, important negative or failure checks, matching proof reuse where identity and conditions permit, and targeted obsolete Feature Map retirement or supersession.

Q40 — Review completeness and snapshot reuse

Accepted: Review Lens Results use clean, findings, or incomplete; partial coverage is never clean, and an incomplete attempt may retry only through bounded useful authorized diagnosis or the existing review route. Semantic changes during an active pass restart review. Accepted repairs after a completed pass follow focused validation or the risk-triggered second pass. Same-snapshot complete evidence may be reused only when identity, scope, and freshness match.

Q41 — Delegation boundary

Accepted: the parent always delegates implementation edits. Read-dependency and evidence-stability detail belongs to the swarm-planner primitive within create-plan, the planning step of delivery-phase.

Q42 — Experiments declined

Declined: additional experiments comparing delegation topology, hybrid excerpts, reduced instruction surfaces, or verifier maintenance cost. AC-044..AC-048 and the assurance-constrained admission policy remain unchanged. Further review-phase optimization ideas remain candidates for a separate later grill.

Accepted Round R6 — 2026-09-05

Accepted review optimization topology. Historical Q5 remains append-only and is superseded by this round.

Q43 — Primary and challenger

Replace five separate lens executions with one comprehensive primary reviewer covering Standards, skill adherence, architecture, simplicity, and Spec obligations, plus one independent risk-focused challenger. autoreview fills the primary role and is no longer an extra advisory scan.

Q44 — Shared factual preparation

Prepare target, rules, evidence, and relevant dependencies once, without persuasion narrative or target rediscovery. Existing helpers perform mechanical schema, identity, evidence-cardinality, serialization, and routing checks.

Q45 — Parent adjudication

The parent groups duplicates before investigation, verifies every underlying issue and distinct claim, preserves primary/challenger provenance, and owns IDs, severity, routes, and retained report assembly. No automatic parent discovery pass follows.

Q46 — Risk coverage

The challenger runs in parallel when capacity permits and may split independent risk areas for larger changes without a fixed numeric cap. Primary coverage retains every obligation, including security.

Q47 — Compatibility and gates

Existing incomplete-with-candidates semantics, one-pass focused repair, maximum-two-pass policy, and AC-044..AC-048 experiment/admission criteria remain unchanged. Legacy five-lens reports normalize at read time without rewriting history.

On this page