Harness Intelligence Wiki
SpecsCLIScaffold, Managed Skill, and Operator Follow-ups

Scaffold, Managed Skill, and Operator Follow-ups Plan

Plan: Scaffold, Managed Skill, and Operator Follow-ups

Initial situation

The child branch starts at live PR #123 head fb8c216f, not its obsolete pre-rebase head 435af1ec. Research commit 82c8cfe2 proves the five defects and is retained independently. The compiled spec is retained at 02b84eac with a stable GitHub blob URL. The initial control-plane check was unavailable; the closeout retry resolves authority and reports broad inherited projection/settings drift. This branch applies no active-projection rewrite.

GitHub issues #122, #124, #125, and #126 are the product-facing backlog. The operator marker defect has no provider item and remains linked through this spec and child PR. No backlog mutation is planned: the issue set already exists and the work request does not ask to reshape provider hierarchy.

Problem and solution shape

  • Replace #122's verbose accepted prompt-authoring obligations with compact repository-invariant and exact-trigger pointer requirements, then change its renderer and public contracts together.
  • Repair #124, #125, and #126 in the shared skills repository first, with tests, push an immutable source ref, then synchronize Harness from that exact ref.
  • Align operator host-marker detection and safe environment forwarding with supported Codex markers while retaining one explicit Skills CLI target.
  • Refresh projections, fixtures, docs, and baseline evidence only through owning generators and supported commands.

The operator-skill feature remains a deep module: callers provide an action; its runtime boundary resolves one active agent and the Skills CLI adapter receives one explicit binding. Marker resolution stays at that boundary. Prompt rendering remains behind renderPromptSpecMarkdown; compactness is verified through the rendered interface rather than private helpers.

Resolved decision ledger

DecisionStateEvidence
Child base is live PR #123 headLockedGitHub reports fb8c216f; old 435af1ec is on the pre-rebase lineage.
#122 output becomes concise pointer-style guidanceLockedUser report plus retained Collective Intelligence before/after evidence.
#125 fixes canonical Markdown with |LockedGFM/Oxfmt semantic reproduction.
Shared skills remain source authorityLockedRoot repository guidance and user convention.
All AI-context Markdown applies writing-for-agentsLockedUser steering on 2026-08-12.
Operator keeps exact --agent argvLockedSkills CLI no-agent behavior broadens mutation scope.
Recognized Codex markers suppress Harness promptLockedCODEX_CI reproduction and installed Skills CLI detection contract.
Parent merges before childLockedUser stack instruction.
Production publication remains dormantParkedExplicit approval/environment gates remain outside scope.

Constraints and assumptions

  • Preserve unrelated worktree changes; all implementation uses clean scoped worktrees.
  • Shared-skill edits activate writing-for-agents, are pushed first, and are synchronized by exact immutable ref.
  • Every AI-context Markdown change activates and applies writing-for-agents. This includes skills, phase routers, prompt specs, handoffs, AGENTS.md, CLAUDE.md, and agent-facing runbook instructions. Completion requires an explicit skill-evidence record for each task that changes one of these surfaces.
  • Behavior changes are test-first. Generated-only projection changes are validated through source contracts and readback.
  • hi check --json is read-only evidence: the recovered check reports broad inherited active-projection/settings drift, which stays outside this branch.
  • The child PR base remains the parent branch until PR #123 merges.

Research and codebase findings

  • apps/cli/src/content/scaffold-copy.ts owns scoped prompt text; focused tests currently require complete skill-semantic derivation.
  • Shared delivery router accepts a state that its backlog child phase rejects.
  • The canonical Effect table is malformed before Oxfmt expands it.
  • Docs onboarding is correct in Harness packaged bytes but not shared main, released projection, or active managed projection.
  • Harness marker resolution and safe runtime environment forwarding are narrower than Skills CLI 1.5.22; the adapter's exact-agent argv itself is correct.

Dependency graph

T1 ──┐
     ├── T5 ──┐
T4 ──┘        │
T2 ───────────┼── T6 ── T7
T3 ───────────┘

Tasks

T1: Repair canonical shared skills for #124, #125, and #126

  • depends_on: []
  • location: /Users/stefan/Desktop/repos/wearedevpunks-skills/skills, /Users/stefan/Desktop/repos/wearedevpunks-skills/tests
  • owned_paths: [skills/phases/delivery-phase/**, skills/frameworks/effect/effect-recoverable-actions/**, skills/agnostic/docs/docs-onboarding/**, tests/wayfinder-lifecycle.contract.test.mjs, tests/effect-recoverable-actions.contract.test.mjs, tests/docs-onboarding.contract.test.mjs]
  • wave_boundary: W1
  • description: Create a clean shared-skills branch from current origin/main; add public contract REDs; make retention/blob verification part of delivery spec completeness, escape the Effect table pipe, and use hi init; validate and push the source head plus a unique immutable sync tag at that exact commit.
  • validation: Shared contracts prove unretained spec routing, semantic three-column table rendering after Oxfmt, and supported onboarding command; pushed branch and immutable tag resolve to the same intended commit.
  • status: Completed
  • log: Shared branch team/stefan/scaffold-skill-followups was pushed at 408746ad5fdf10d75513cd1b63c71c6f22fa7fb7; immutable tag sync/scaffold-skill-followups-408746ad5fdf resolves to the same commit. Focused contracts pass 34/34 and the full shared suite passes 104/104.
  • files edited/created: skills/phases/delivery-phase/phases/router.md, skills/phases/delivery-phase/phases/spec.md, skills/frameworks/effect/effect-recoverable-actions/references/strategy-matrix.md, skills/agnostic/docs/docs-onboarding/SKILL.md, tests/wayfinder-lifecycle.contract.test.mjs, tests/effect-recoverable-actions.contract.test.mjs, tests/docs-onboarding.contract.test.mjs
  • backlog_item_id: #124, #125, #126
  • backlog_item_url: https://github.com/wearedevpunks/harness-intelligence/issues/124, https://github.com/wearedevpunks/harness-intelligence/issues/125, https://github.com/wearedevpunks/harness-intelligence/issues/126
  • relation_mode: body-links
  • assigned_skills: [writing-for-agents, tdd, codebase-design, simplify]
  • implementation_skill_guidance:
    • skill: writing-for-agents applicable_behavior: Keep routing and onboarding instructions compact, positive, and singular at canonical source.
    • skill: tdd applicable_behavior: Capture one failing public contract per defect before each minimal source fix.
    • skill: codebase-design applicable_behavior: Preserve each skill file as the authoritative interface and keep projection mechanics outside it.
    • skill: simplify applicable_behavior: Review only the touched source/test delta and remove redundant wording without changing accepted behavior.
  • tdd_status: required
  • tdd_target: A local unretained/unverified spec routes to spec repair, formatted strategy Markdown has three cells, and docs onboarding includes hi init while excluding hi scaffold init.
  • red_command: node --test tests/wayfinder-lifecycle.contract.test.mjs tests/effect-recoverable-actions.contract.test.mjs tests/docs-onboarding.contract.test.mjs
  • expected_red_failure: Missing retention still routes to backlog; the table parser sees a split mode expression; docs onboarding still contains hi scaffold init.
  • green_command: node --test tests/wayfinder-lifecycle.contract.test.mjs tests/effect-recoverable-actions.contract.test.mjs tests/docs-onboarding.contract.test.mjs
  • reason_not_testable:
  • red_evidence: Focused contracts failed for missing remote authority routing, four semantic Effect-table cells, and absent hi init.
  • green_evidence: node --test tests/wayfinder-lifecycle.contract.test.mjs tests/effect-recoverable-actions.contract.test.mjs tests/docs-onboarding.contract.test.mjs passed 34/34; node --test tests/*.test.mjs passed 104/104.
  • codebase_design_notes: Canonical skill Markdown is the interface; sync is the adapter to consumer repositories.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T2: Compact scoped prompt authoring for #122

  • depends_on: []
  • location: apps/wiki/content/docs/project/specs/cli/scoped-agent-guidance-authoring-contract/SPEC.md, apps/cli/src/content/scaffold-copy.ts, focused content/scaffold tests
  • owned_paths: [apps/wiki/content/docs/project/specs/cli/scoped-agent-guidance-authoring-contract/SPEC.md, apps/cli/src/content/scaffold-copy.ts, apps/cli/src/content/content.test.ts, apps/cli/src/scaffold/format-summary.ts, apps/cli/src/scaffold/run.test.ts]
  • wave_boundary: W1
  • description: Supersede verbose acceptance language, write RED public-output expectations, and make scoped specs require concise repo invariants and terse exact-trigger skill pointers while retaining IDs, Source Guide, mirrors, and validation seams.
  • validation: Focused renderer/scaffold tests pass; the fixed 15-scope fixture preserves semantic assertions and totals at most 700 generated authoring-instruction lines.
  • status: Completed
  • log: Replaced complete skill-semantic restatement and generic evidence prose with repository invariants and exact-trigger pointers. The fixed 15-scope fixture measures 690 lines against the 700-line limit.
  • files edited/created: apps/wiki/content/docs/project/specs/cli/scoped-agent-guidance-authoring-contract/SPEC.md, apps/cli/src/content/scaffold-copy.ts, apps/cli/src/content/content.test.ts, apps/cli/src/scaffold/format-summary.ts, apps/cli/src/scaffold/run.test.ts
  • backlog_item_id: #122
  • backlog_item_url: https://github.com/wearedevpunks/harness-intelligence/issues/122
  • relation_mode: body-links
  • assigned_skills: [writing-for-agents, tdd, codebase-design, quality-types, simplify]
  • implementation_skill_guidance:
    • skill: writing-for-agents applicable_behavior: Reduce context load by pointing to installed skills with exact triggers and keeping only repository-specific invariants inline.
    • skill: tdd applicable_behavior: Test rendered public Markdown before changing the renderer, one semantic obligation at a time.
    • skill: codebase-design applicable_behavior: Keep one renderer seam and prevent policy duplication across renderer, handoff, and summary surfaces.
    • skill: quality-types applicable_behavior: Preserve existing typed prompt-spec inputs; do not introduce parallel descriptive state derivable from selected skills.
    • skill: simplify applicable_behavior: Remove repeated generic authoring prose and keep tests focused on durable outcomes.
  • tdd_status: required
  • tdd_target: A fixed 15-scope rendered fixture totals at most 700 authoring-instruction lines while preserving exact selected IDs, terse triggers, Source Guide pointers, mirror rules, and validation seams.
  • red_command: bun run --cwd apps/cli test src/content/content.test.ts src/scaffold/run.test.ts
  • expected_red_failure: New concise contract assertions fail because output still requires complete skill-derived prose and detailed generic sections.
  • green_command: bun run --cwd apps/cli test src/content/content.test.ts src/scaffold/run.test.ts
  • reason_not_testable:
  • red_evidence: Public rendered-Markdown contract first lacked repository-specific invariants; the first minimal renderer measured 705 lines; the next assertion exposed duplicated handoff policy.
  • green_evidence: bun run --cwd apps/cli test src/content/content.test.ts src/scaffold/run.test.ts passed 51/51; bun run --cwd apps/cli check-types, Oxfmt, and git diff --check passed.
  • codebase_design_notes: renderPromptSpecMarkdown remains the interface/test surface; summary text points to it instead of duplicating it.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T3: Detect supported Codex operator markers without prompting

  • depends_on: []
  • location: apps/cli/src/integrations/operator-skill-context.ts, runtime config, operator public-output tests, Skills CLI adapter tests
  • owned_paths: [apps/cli/src/integrations/operator-skill-context.ts, apps/cli/src/integrations/operator-skill-context.test.ts, apps/cli/src/runtime/config.ts, apps/cli/src/runtime/config-contract.test.ts, apps/cli/src/cli/public-output-contract.test.ts, apps/cli/src/integrations/skills-cli.test.ts]
  • wave_boundary: W1
  • description: Add CODEX_CI and CODEX_SANDBOX to safe host detection/forwarding; prove HI_OPERATOR_AGENT precedence; retain interactive unknown-host prompting and noninteractive fail-closed behavior; preserve exact child target across inspection, install, update, migration removal, and guarded retries.
  • validation: Resolver, config, public-output/PTY, retry, and exact argv tests pass for both Codex markers and every operator action; explicit override and unknown-host behavior remain correct.
  • status: Completed
  • log: Added one typed marker map for CODEX_CI and CODEX_SANDBOX, reused it for safe child-environment forwarding, retained override precedence and fail-closed unknown-host behavior, and kept exactly one explicit agent on every lifecycle/retry invocation.
  • files edited/created: apps/cli/src/integrations/operator-skill-context.ts, apps/cli/src/integrations/operator-skill-context.test.ts, apps/cli/src/runtime/config.ts, apps/cli/src/runtime/config-contract.test.ts, apps/cli/src/cli/public-output-contract.test.ts, apps/cli/src/integrations/skills-cli.test.ts
  • backlog_item_id: not_applicable
  • backlog_item_url: not_applicable
  • relation_mode: body-links
  • assigned_skills: [tdd, codebase-design, quality-types, effect, simplify]
  • implementation_skill_guidance:
    • skill: tdd applicable_behavior: Capture the packaged public prompt regression under CODEX_CI before changing marker resolution.
    • skill: codebase-design applicable_behavior: Keep host detection at the runtime boundary and exact-agent mutation behind the Skills CLI adapter.
    • skill: quality-types applicable_behavior: Derive marker keys from one typed mapping where practical and keep invalid states unrepresentable.
    • skill: effect applicable_behavior: Preserve typed InvalidOperatorContext failure flow and injected application-operation test seam.
    • skill: simplify applicable_behavior: Avoid a second independent Codex marker list if the existing mapping can own resolution and forwarding cleanly.
  • tdd_status: required
  • tdd_target: hi operator status|install|update|migrate under CODEX_CI=1 and CODEX_SANDBOX=1 reaches execution without prompting; explicit override wins; unknown interactive hosts prompt; unknown noninteractive hosts fail closed; every child/retry path carries one exact --agent identity.
  • red_command: bun run --cwd apps/cli test src/integrations/operator-skill-context.test.ts src/runtime/config-contract.test.ts src/cli/public-output-contract.test.ts src/integrations/skills-cli.test.ts
  • expected_red_failure: Codex marker resolution is undefined or packaged output contains Enter the Skills CLI agent name:.
  • green_command: bun run --cwd apps/cli test src/integrations/operator-skill-context.test.ts src/runtime/config-contract.test.ts src/cli/public-output-contract.test.ts src/integrations/skills-cli.test.ts
  • reason_not_testable:
  • red_evidence: Packaged CODEX_CI operator contract failed because stdout contained Enter the Skills CLI agent name:.
  • green_evidence: Resolver/configuration/adapter suites passed 55/55; packaged marker, unknown-host, lifecycle, and retry contracts passed 8/8; typecheck, build, Oxfmt, and git diff --check passed.
  • codebase_design_notes: Agent-marker resolver is the interface; runtime config forwards its safe inputs; Skills CLI remains one explicit adapter.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T4: Add pre-sync baseline and managed-consumer semantic contracts

  • depends_on: []
  • location: apps/cli/src/baseline/baseline-release-scripts.native.test.ts
  • owned_paths: [apps/cli/src/baseline/baseline-release-scripts.native.test.ts]
  • wave_boundary: W1
  • description: Before packaged skill synchronization, add semantic archive and isolated managed-consumer assertions for the Effect three-column row and hi init onboarding bytes; capture the real failure against current packaged/projection inputs.
  • validation: The focused test fails for the named malformed table or stale onboarding projection before T5 changes packaged sources.
  • status: Completed
  • log: Candidate archive and isolated managed consumer share one semantic readback helper. The focused target failed before sync because both distributions split the Effect expression into four cells; onboarding command semantics already passed.
  • files edited/created: apps/cli/src/baseline/baseline-release-scripts.native.test.ts
  • backlog_item_id: #125, #126
  • backlog_item_url: https://github.com/wearedevpunks/harness-intelligence/issues/125, https://github.com/wearedevpunks/harness-intelligence/issues/126
  • relation_mode: body-links
  • assigned_skills: [tdd, codebase-design, simplify]
  • implementation_skill_guidance:
    • skill: tdd applicable_behavior: Capture the semantic archive/consumer failure before synchronized source bytes change.
    • skill: codebase-design applicable_behavior: Test the baseline archive and managed consumer as public distribution interfaces, not copier internals.
    • skill: simplify applicable_behavior: Add one focused semantic helper rather than duplicating archive setup.
  • tdd_status: required
  • tdd_target: Candidate baseline extraction and isolated managed update preserve a three-column Effect expression and docs onboarding hi init bytes.
  • red_command: bun run --cwd apps/cli test src/baseline/baseline-release-scripts.native.test.ts
  • expected_red_failure: Current packaged/projection inputs split the Effect mode expression or retain hi scaffold init.
  • green_command: bun run --cwd apps/cli test src/baseline/baseline-release-scripts.native.test.ts
  • reason_not_testable:
  • red_evidence: Targeted semantic test exited 1 with 1 failed/6 skipped; archive and consumer both split "validate" | "either" into separate cells instead of the expected three-column row.
  • green_evidence: After T5 exact-ref sync, bun run --cwd apps/cli test src/baseline/baseline-release-scripts.native.test.ts passed, including candidate archive and isolated managed-consumer semantic readback.
  • codebase_design_notes: Baseline archive and installed managed skill are the distribution interfaces.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T5: Synchronize exact shared-skill source into packaged contracts

  • depends_on: [T1, T4]
  • location: Harness packaged skill trees and exact-ref sync receipt
  • owned_paths: [apps/cli/skills/**, apps/cli/.devpunks-cache/skills-sync.json]
  • wave_boundary: W2
  • description: Run bun run sync:skills from T1's immutable pushed tag and accept only affected packaged sources plus the receipt. Do not rewrite active repository projections while control-plane baseline authority is unavailable; preserve .devpunks/pre-existing-skills snapshots.
  • validation: Requested tag and receipt commit equal T1's recorded immutable SHA; canonical and packaged bytes match; T4 becomes GREEN; no unrelated skill drift enters the branch.
  • status: Completed
  • log: The sync script qualifies unprefixed refs as branches, so the approved exact-tag invocation used refs/tags/sync/scaffold-skill-followups-408746ad5fdf. The gitignored receipt resolves 408746ad5fdf10d75513cd1b63c71c6f22fa7fb7; whole-tree canonical/package comparison passes; active projections and preserved snapshots are unchanged.
  • files edited/created: apps/cli/skills/frameworks/effect/effect-recoverable-actions/references/strategy-matrix.md, apps/cli/skills/phases/delivery-phase/phases/router.md, apps/cli/skills/phases/delivery-phase/phases/spec.md; gitignored apps/cli/.devpunks-cache/skills-sync.json
  • backlog_item_id: #124, #125, #126
  • backlog_item_url: issue URLs above
  • relation_mode: body-links
  • assigned_skills: [repo-asset-management, tdd, simplify]
  • implementation_skill_guidance:
    • skill: repo-asset-management applicable_behavior: Treat packaged and managed files as generator-owned assets; refresh from canonical source and verify ownership metadata.
    • skill: tdd applicable_behavior: Reuse T1 contracts and add projection/readback assertions before accepting generated changes.
    • skill: simplify applicable_behavior: Reject unrelated sync drift and keep the generated delta limited to affected assets and receipts.
  • tdd_status: not_applicable
  • tdd_target: Exact-ref generated asset synchronization and readback.
  • red_command: not_applicable
  • expected_red_failure: not_applicable
  • green_command: HI_SKILLS_REPOSITORY_REF=<immutable-sync-tag> bun run sync:skills && test "$(jq -r .commit apps/cli/.devpunks-cache/skills-sync.json)" = "<immutable-shared-head>" && bun run --cwd apps/cli fixtures:check:managed-assets && bun run --cwd apps/cli test src/baseline/baseline-release-scripts.native.test.ts
  • reason_not_testable: Generated asset synchronization; behavior REDs belong to T1 and readback validation proves the projection.
  • red_evidence:
  • green_evidence: Full tree diff -qr passed, four touched canonical file hashes match packaged bytes, T4 baseline test passed, sync-focused tests passed 9/9, and git diff --check passed. Managed fixture check is intentionally deferred with exactly five expected hashes: three synced files and two T2 prompt outputs.
  • codebase_design_notes: The sync command is the only adapter between shared source and packaged skill assets.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T6: Update operator/scaffold docs and distribution evidence

  • depends_on: [T2, T3, T5]
  • location: apps/cli/README.md, docs/README.md, docs/runbooks/hi-cli-scaffolding.md, spec implementation notes, baseline tests/changelogs
  • owned_paths: [apps/cli/README.md, docs/README.md, docs/runbooks/hi-cli-scaffolding.md, apps/wiki/content/docs/project/runbooks/hi-cli-scaffolding.md, apps/cli/test-fixtures/public-output/managed-assets.json, apps/cli/test-fixtures/public-output/context-outcome.json, apps/cli/src/data/bundled-baseline-identity.generated.ts, CHANGELOG.md, BASELINE_CHANGELOG.md]
  • wave_boundary: W3
  • description: Document compact prompt continuation and supported marker detection, record source/projection lineage and T4/T5 evidence, and prepare the baseline changelog without activating publication.
  • validation: Docs match implemented behavior; current archive/isolated-consumer test passes; wiki/spec checks pass.
  • status: Completed
  • log: Reconciled stable prompt/operator behavior across CLI and runbook docs, split executable and baseline Unreleased ledger entries, refreshed managed/context fixtures and bundled identity through supported generators, and synchronized the wiki runbook projection. Temporary candidate lineage was pruned from durable docs and retained in release/implementation evidence.
  • files edited/created: CHANGELOG.md, BASELINE_CHANGELOG.md, apps/cli/README.md, docs/README.md, docs/runbooks/hi-cli-scaffolding.md, apps/wiki/content/docs/project/runbooks/hi-cli-scaffolding.md, apps/cli/test-fixtures/public-output/managed-assets.json, apps/cli/test-fixtures/public-output/context-outcome.json, apps/cli/src/data/bundled-baseline-identity.generated.ts
  • backlog_item_id: #122, #124, #125, #126
  • backlog_item_url: issue URLs above
  • relation_mode: body-links
  • assigned_skills: [writing-for-agents, tdd, codebase-design, simplify]
  • implementation_skill_guidance:
    • skill: writing-for-agents applicable_behavior: Describe the operator and prompt contracts as concise stable behavior, not duplicated command/source inventories.
    • skill: tdd applicable_behavior: Add semantic archive assertions before accepting baseline distribution bytes.
    • skill: codebase-design applicable_behavior: Document the marker-resolution boundary and source-to-projection seam explicitly.
    • skill: simplify applicable_behavior: Reconcile existing docs rather than layering duplicate explanations.
  • tdd_status: not_applicable
  • tdd_target: Documentation and implementation-evidence reconciliation.
  • red_command: not_applicable
  • expected_red_failure: not_applicable
  • green_command: bun run --cwd apps/cli test src/baseline/baseline-release-scripts.native.test.ts && bun run wiki:check
  • reason_not_testable: Docs and changelog reconciliation; T4/T5 own behavior RED/GREEN.
  • red_evidence:
  • green_evidence: CLI check including managed fixture passed; prompt/scaffold passed 51/51; operator/config/adapter passed 55/55; candidate archive/consumer passed; wiki content, Oxfmt, changelog parsing, and git diff --check passed. Full baseline file passed 6/7; the unrelated publisher fixture cannot spawn git under its mocked PATH, while the required semantic target passes independently.
  • codebase_design_notes: Baseline archive is the distribution interface; semantic readback tests it without coupling to copier internals.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

T7: Run acceptance, review, and stacked PR handoff

  • depends_on: [T6]
  • location: repository-wide touched-surface validation, PLAN/IMPLEMENTATION-NOTES, GitHub child PR
  • owned_paths: [apps/wiki/content/docs/project/specs/cli/scaffold-skill-operator-followups/PLAN.md, apps/wiki/content/docs/project/specs/cli/scaffold-skill-operator-followups/IMPLEMENTATION-NOTES.md]
  • wave_boundary: W4
  • description: Run focused-to-outer validation, simplify/review the total diff, record real RED/GREEN and blockers, push child, open draft stacked PR against the parent branch, and verify stack sync --dry-run. Do not merge or publish.
  • validation: All available gates pass or are classified with evidence; PR base/head and ancestry are correct; control-plane check remains explicitly pending if authority is unavailable.
  • status: Completed
  • log: Repository check passed 12/12 workspaces. The outer test gate passed root contracts, non-CLI workspaces, runtime validation, wiki production build, and one 23-test update shard; the runner then sent SIGTERM to the other three long update shards without a reported assertion failure. Frozen standards and spec reviews pass after two pointer/frontmatter repairs. Draft PR #127 targets PR #123's branch; stack sync previewed the inferred parent-first topology and stack sync --apply reconciled it.
  • files edited/created: apps/wiki/content/docs/project/specs/cli/scaffold-skill-operator-followups/PLAN.md, apps/wiki/content/docs/project/specs/cli/scaffold-skill-operator-followups/IMPLEMENTATION-NOTES.md
  • backlog_item_id: not_applicable
  • backlog_item_url: not_applicable
  • relation_mode: body-links
  • assigned_skills: [review, simplify, stack]
  • implementation_skill_guidance:
    • skill: review applicable_behavior: Review the accepted diff and test evidence without expanding into unrelated repository debt.
    • skill: simplify applicable_behavior: Remove unnecessary complexity introduced by this branch before final validation.
    • skill: stack applicable_behavior: Preserve parent-first topology, verify the dry-run, and keep the child based on the parent until merge.
  • tdd_status: not_applicable
  • tdd_target: Acceptance and stack topology evidence.
  • red_command: not_applicable
  • expected_red_failure: not_applicable
  • green_command: bun run check && bun run test && stack sync --dry-run
  • reason_not_testable: Validation and delivery orchestration task; behavior RED/GREEN occurs in owning tasks.
  • red_evidence:
  • green_evidence: bun run check passed 12/12; all touched focused suites passed; final standards and spec reviews passed; stack status shows main -> #123 -> #127, and applied sync reports #127 rebased onto the parent branch.
  • codebase_design_notes: Review tests public seams; no new interface is introduced.
  • review_mode: cli
  • runtime_validation: not_required
  • runtime_target: not_applicable
  • runtime_evidence: not_applicable
  • runtime_cleanup: not_applicable

Parallel execution waves

WaveTasksStart condition
W1T1, T2, T3, T4Immediately; write scopes are disjoint.
W2T5T1 and T4 complete; immutable shared tag exists and distribution RED is captured.
W3T6T2, T3, and T5 complete.
W4T7T6 complete.

Testing strategy

Each behavior task uses vertical RED/GREEN cycles through public contracts. Shared skill tests prove instructions and semantic Markdown. Harness content tests prove rendered prompt output. Operator tests prove packaged terminal behavior and exact adapter argv. Distribution tests inspect the built archive semantically. Validation then expands to typecheck, managed fixtures, wiki content, build, and repository checks. Control-plane scaffold authority is retried last and may remain an external blocker.

Risks and mitigations

  • Shared source drift: branch from fetched shared origin/main, push first, pin exact SHA, inspect sync diff before accepting it.
  • #122 overcompaction: assert semantic invariants and representative output, not only line count.
  • Operator over-broadening: retain exact --agent argv and existing retry identity.
  • Formatter-only false confidence: parse/render the table after formatting and archive extraction.
  • Baseline authority unavailable: distinguish local build/readback from remote authority and never claim drift or publication readiness without the supported check.
  • Stack drift: verify child merge-base and PR base before push and after stack dry-run.

Backlog sync

Skipped. Existing GitHub issues #122, #124, #125, and #126 already express the product-facing work. This planning run is not authorized to create or reshape provider hierarchy; the child PR will link all four issues and the retained spec.

Unresolved questions

  • Exact baseline release identifier remains intentionally unset until all implementation and remote authority gates pass. This does not block coding.
  • Whether #122 should close the prior implemented spec or mark it superseded is a documentation bookkeeping choice resolved during T2 from existing repo convention; it does not change accepted product behavior.

On this page