Commit Graph

10 Commits

Author SHA1 Message Date
drew eb1d81828f feat(controller): implementer_retrigger_ci — agent re-triggers CI on ci-infra-failure
A CI job hard-killed by OOM / pod eviction produces no verdict — nothing
in the diff to fix. The implementer correctly emits
outcome=ci-infra-failure, but that alone only shuffles workflow state;
CI never re-runs, so every retry reads the same dead run and the
workflow loops until the _MAX_CI_INFRA_FAILURE backstop STUCKs it
(observed live on PR #39 in run-3).

New ``implementer_retrigger_ci(owner, repo, pr_branch)`` MCP tool on the
implementer response-builder lets the agent kick a fresh CI run. Forgejo
15.x has no Actions rerun API, so it reuses the controller's existing
mechanism — ``ci_rerun.trigger_ci_rerun_via_empty_commit`` — an empty
commit that advances the PR head SHA, which is the unambiguous
fresh-state signal for CI (and a stale review).

- ci_rerun.py is loaded standalone (importlib by file path, registered
  in sys.modules before exec so its @dataclass resolves) — the
  response-builder MCP must not pull the heavy tools.controller.master
  package.
- _retrigger_ci resolves the Forgejo base URL + token from env
  (FORGEJO_URL / FORGEJO_API_BASE, FORGEJO_TOKEN / GITEA_TOKEN) and
  never raises.
- Once per session: a second implementer_retrigger_ci call is refused
  benignly so a looping agent cannot pile junk commits on the PR.
- The implementer prompt's ci-infra-failure block now instructs the
  agent to call the tool before emitting the outcome.

10 new tests (test_mcp_builders.py, test_worker_prompts.py); full
controller suite (1192) green.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 00:22:46 -04:00
drew f76568a871 feat(controller): ci-infra-failure implementer outcome → bounded rerun
The implementer-side counterpart to the `indeterminate` verdict. When
the implementer IS dispatched onto a CI failure and finds the log
carries no verdict (a hard-kill / OOM — nothing in the diff to fix),
it can now emit `outcome=ci-infra-failure` instead of being forced to
`blocked` → STUCK.

`ci-infra-failure` forbids commits/files/blockers (no-work invariant,
like `noop`) and fires `implementer_ci_infra_failure`, routing
IMPLEMENTING → DISCOVERED so the CI-freshness gate reruns CI under its
bounded RERUN_BUDGET. Backstopped by `_MAX_CI_INFRA_FAILURE=4` so a
mis-classification cannot loop the gate forever.

Wired through: the V1 contract enum, the implementer MCP builder
(outcome value + no-work invariant), the state machine (event +
IMPLEMENTING→DISCOVERED transition), the outcome mapper (+ per-workflow
cap), tick's `_count_prior_ci_infra_failure`, and the implementer
prompt — which now surfaces the outcome whenever the CI summary shows
a failing gate, with guidance to use it ONLY when the log genuinely
shows no verdict.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 22:55:28 -04:00
drew 0db0a15dad feat(controller): RUN_CI_LOCAL verdict source, ci-not-ready outcome, escalation hardening
Adds RUN_CI_LOCAL — an on-demand local-CI verdict source for when the
cluster's Forgejo CI is broken — plus robustness fixes, the telemetry
Live-tab rewrite, and PR-level cost attribution.

Controller:
- RUN_CI_LOCAL: the master swaps its Forgejo CI callbacks for local
  `forgejo-runner exec` runs (tools/run-ci-full-local.sh + local_ci.py).
  Async per-(owner,repo,SHA) on-disk job cache; preflights the
  forgejo-runner binary + Docker daemon at startup (fail loud, not a
  red verdict on every PR); GCs finished run dirs + per-run actcache.
- ci-not-ready implementer outcome + implementer_ci_not_ready event:
  an implementer that runs before the on-demand verdict exists parks
  in AWAITING_CI instead of dead-ending at STUCK; capped against
  ci_red ping-pong.
- ci_poll_exhaustion skips its sweep while a local CI run is in
  flight, so AWAITING_CI workflows queued behind on-demand CI are not
  STUCK'd by the remote-CI-sized timeout.
- Escalate the workflow after repeated worker-internal-error at a
  tier instead of retrying to pickup-exhaustion -> STUCK.
- forgejo_http: normalise Forgejo's per-gate `status` key to `state`
  so failing gates are actually counted (they previously all read as
  pending).
- Per-tier worker timeouts bumped +15 min; a timed-out attempt's
  dirty-worktree residue is preserved on auto-scratch/pr-<N> before
  the next attempt's reset.
- Implementer agents now verify only the CI-flagged gate(s) via a
  targeted re-run rather than the full local battery before claiming
  resolved/noop. Re-running the whole suite CI will run anyway was the
  #1 cause of implementer timeouts; CI remains the real gate and
  re-dispatches the implementer on red.

Telemetry:
- Live tab rebuilt on /api/live (controller DB run state + the live
  OpenCode session forest) after the live_log_writer sidecar was
  retired with the legacy dispatchers.
- Durable per-attempt input/output payloads surfaced in the Live
  drill-down, archived-session detail, and Workflows timeline.
- PR-level cost attribution: worker session tags carry -pr-<n>;
  backfill_llm_activity_pr.py repairs rows written before the fix.

Shared:
- tools/controller/session_tag.py — one canonical controller-tag
  parser shared by the telemetry server and the backfill.

Tests: new coverage for local_ci (state machine, log parsing,
_summarize_run, GC, in-flight probe, preflight), the ci-not-ready
path, the escalation/ci-not-ready SQL counters, the ci_poll
in-flight skip, and CI-status payload parsing across both sources.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 11:49:05 -04:00
drew 0bc734c020 style: ruff format the controller-state-machine branch (288 files)
Applies `ruff format` to the accumulated formatting debt on this branch.
Formatting-only — no behavioral changes. Required for CI/lint's format
gate (`nox -s format -- --check`), which the branch was failing on 288
tracked files that drifted from ruff's canonical style.

In-progress WIP files are intentionally excluded so this commit stays a
clean formatting-only diff.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 00:09:17 -04:00
drew a6986008ee feat(controller): trial-5 batch — dispute path, merge pipeline split, conflict-resolver hardening
T5-1   reviewer feedback rendered full-body to the implementer
T5-4/9 implementer dispute path — dispute-at-any-tier with per-tier cap,
       OPERATOR_ATTENTION state on stalemate, pr-review-worker-dispute agent
T5-5   reviewer BLOCKING ISSUE EVIDENCE RULE + 5-step validation
T5-7   merge step split into a singleton process — impl/review masters write
       APPROVED and stop; merge_drive owns APPROVED -> MERGING -> MERGED
T5-10  merge process is fully deterministic; base conflicts bounce to the
       controller's CONFLICT_RESOLVING (LLM); conflict_drive sidecar retired
T5-11  implementer fast success path — verified-clean outcome so a no-op
       after conflict resolution doesn't force busywork
T5-12  conflict-resolver permissions fixed across all paths (/tmp/** glob)
T5-13  conflict-resolver PR-intent prehydration (title/body/comments)

Adds tools/_controller_db_bridge.py so merge_drive reads the controller DB
directly (Option B), plus APPROVED + OPERATOR_ATTENTION states, the
dispute/verified-clean events, and the V1 contract fields backing them.
Reviewer model: baseline -> sonnet, dispute -> opus.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 17:54:07 -04:00
drew 46764b841b fix(controller): batch S — 11 fixes from trial-4 finding + 2nd code-review pass
Trial-4 ran the controller past ANALYZING for the first time. One
critical regression surfaced live (T4-1 — scheduler dup-attempt
enqueue), and a parallel agent-driven code review found 8 more bugs
across "races + error paths + agent quality" classes. Batch S
addresses 11 of those.

LIVE-OBSERVED REGRESSION (trial-4 2026-05-19 00:49):

T4-1 — Scheduler enqueues duplicate estimator after workflow advanced
File: tools/controller/master/scheduler.py
The "already pending" check filtered on
``status IN ('pending', 'in_progress')`` — missing the brief window
where the prior attempt is ``status='complete'`` but tick hasn't yet
processed its outcome. Result: scheduler enqueues a 2nd estimator/etc.;
when its outcome fires from a now-advanced state, IllegalTransition
→ STUCK. Hit wf=1 in trial-4. Fix: also skip when an unprocessed
``complete`` attempt exists (finished_at > w.last_transition_at).

E-5 — IllegalTransition over-aggressive STUCKing
File: tools/controller/master/tick.py
Companion fix to T4-1. Even if T4-1 escapes in some other path (or
a worker delays writing outcome past tick), a stale-outcome
(workflow already advanced via parallel path like ci_status_poll or
reconciliation) shouldn't STUCK. New ``_is_stale_role_outcome``
helper recognizes "this role's outcome arrived after the workflow
moved past its origin state" → consume the attempt, bump
last_transition_at, continue. Only genuine state corruption → STUCK.

E-3 — Corrupted output_payload silently wedges workflow
File: tools/controller/master/tick.py
Pre-fix _decode_output_payload swallowed json.JSONDecodeError →
mapper returned None → tick bumped last_transition_at but workflow
never moved. Operator had no signal. Now raises
``CorruptedOutputPayload`` → tick routes to STUCK with reason.

E-1 — Contract-violation routed to STUCK on first attempt
File: tools/controller/master/outcomes.py + tick.py
v9 spec promised retry-once-with-corrective-prompt for
contract-violation; the column ``strict_parse_retries`` existed but
nothing read/incremented it. _map_failed_outcome now takes
``prior_contract_violations`` count (queried in tick.py); STUCKs
only when count ≥ _CONTRACT_VIOLATION_RETRY_LIMIT (2). First two
violations re-enqueue.

R-1 — Reconciliation flipped workflow state mid-attempt
File: tools/controller/master/reconciliation.py
Worker holding a lock + heartbeating; reconciliation flipped current_state
to MERGED/ABANDONED based on Forgejo; tick.py then skipped the
worker's eventual write (terminal-state exclusion). Worker's output
lost. Fix: _apply_transition first checks for in_progress attempts
on the same workflow and defers if any exist.

R-2 + R-8 — _apply_transition lacks current_state guard
Files: reconciliation.py + merging.py
Same pattern as ci_status_poll's existing TOCTOU defense. UPDATE now
filters ``WHERE current_state = :from_state``; on rowcount=0, skip
the event row. Prevents racing ticks from over-writing each other.

R-4 — Reaper UPDATE didn't re-check heartbeat freshness
File: tools/controller/reaper.py
A worker's healthy heartbeat between reaper's SELECT and UPDATE
would be silently overwritten; the worker's later _write_outcome
(filtered on locked_by_instance) returned rowcount=0 → output lost.
UPDATE now includes the same freshness filter as the SELECT, so
fresh heartbeats protect the row.

E-8 — merging_retry_count not reset on STUCK/abandoned paths
File: tools/controller/master/merging.py
Pre-fix, only 200/409/422 paths reset the counter. 403 (branch
protection), 404 (externally closed → ABANDONED), retry-exhausted
(STUCK) leaked stale counts. If operator unsticks a STUCK workflow
back through MERGING, the stale count made it STUCK again sooner
than expected. All terminal-state-changing paths now reset.

A-1 — commit_shas validation accepted any ≥7-char string
File: tools/controller/mcp/implementer_builder.py
Tightened to ``re.fullmatch(r"[0-9a-f]{7,40}")``. Pre-fix an agent
could pass any 7+ char string; head_sha_advanced accepted the
hallucination; CI poll then 404'd on the fake SHA forever (until
2h ci_poll_exhaustion).

A-2 — merging-409 → IMPLEMENTING(tier=NULL) trap
File: tools/controller/master/merging.py
On 409 the handler routes to IMPLEMENTING(tier=tier_last_succeeded);
if that's NULL (e.g., metadata-only → REVIEWING → approve → 409 path
where implementer_pushed never fired), the MCP rejects tier=NULL →
contract-violation → STUCK. Fix: default to current_tier when
tier_last_succeeded is NULL.

A-9 — Reviewer cross-field check: verdict ↔ suggested_next_action
File: tools/controller/mcp/reviewer_builder.py
Agent could set verdict=approve + suggested_next_action=abandon; the
master fired reviewer_approve regardless. New _VERDICT_ACTION_COMPAT
map enforces compatible pairs at finalize.

TESTS:

- New file ``test_batch_s_fixes.py`` with 13 regression tests, one
  per fix class.
- Updated test_master_outcomes.py for the new contract-violation
  retry behavior.
- Updated test_master_prefetch.py fixture to bump last_transition_at
  past the seeded attempts (the new T4-1 filter would otherwise
  correctly identify the pre-seeded attempts as unprocessed).

Total: 819 → 832 tests, 0 regressions.

DEFERRED to a follow-up batch (per PENDING_FIXES.md):
- 8 MEDIUM items (error UX, dead fields, signal-loss in error paths)
- 5 LOW items (agent quality polish)
- 4 CONFIRMED-CLEAN (no fix needed)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 21:15:26 -04:00
drew 2ad0958cb7 fix(controller): batch O — trial-2 fix: interpolate real attempt_id into prompts
Trial run-2 surfaced the actual blocker: the response-builder MCPs
shared their module-global ``_STATE`` across OpenCode sessions (they're
registered as ``type: local`` in opencode.json, so OpenCode spawns one
subprocess per OpenCode-server lifetime, not per session). The CA2/PD5
cross-session reset I added in the M-batch keyed on
``identity.attempt_id`` — but the prompt template literally read:

  1. ``estimator_start(attempt_id=..., workflow_id=..., ...)``

The ``...`` were placeholder ellipses, not interpolated. So agents
guessed ``attempt_id=1`` every single time. The cross-session reset
compared ``1 == 1``, decided "same attempt — no reset", and the prior
attempt's ``finalized=True`` persisted forever. Every estimator
session after the first failed with ``"response already finalized;
no further mutations allowed"`` on every ``_set_*`` call → no
finalize → 30s worker timeout → workflow STUCK.

Observed in trial run-2: 2 successes (attempts 1, 2), then 15
consecutive failures (attempts 3–17 across all 6 workflows) before
the pickup_guard would have STUCK every workflow.

Two fixes (defense in depth):

1. **Unconditional reset on _start** in all 5 builders. We can't
   distinguish "agent retry in same session" from "new session reusing
   this MCP" reliably — the observable signature is identical. Just
   reset whenever ``_STATE.started`` is True; ``reset_for_new_attempt``
   logs a WARN if the prior state was in-flight so abandoned attempts
   are still visible to ops. The agent's last ``_start`` always wins.

2. **Interpolate real attempt_id + workflow_id + pr_number + head_sha
   into all 5 prompts' Output contract sections**. The agent_runner
   now injects ``attempt_id`` into ``input_payload`` (matching the
   existing ``workspace_dir`` injection pattern). Each ``build_*_prompt``
   reads ``input_payload.get("attempt_id")`` and bakes the concrete
   value into the MCP-call signature shown to the agent, plus a
   ``PASS THESE EXACT VALUES`` directive.

Also strengthened each prompt's ``_finalize`` line with
``**You MUST call this tool — without it the controller times out.**``
so the agent understands the contract is hard, not optional.

Updated tests:
- ``test_double_start_refused`` → ``test_double_start_resets_silently``:
  pins the new permissive-reset behavior + asserts the WARN log fires.
- ``test_implementer_rejects_intra_session_double_start`` updated for
  same reason; now asserts the second _start succeeds + last-wins
  semantics (used_tier == new tier).
- NEW ``test_prompt_interpolates_real_attempt_id`` parametrized over
  all 5 roles: asserts ``attempt_id=42`` + ``workflow_id=7`` appear
  literally in the rendered prompt AND that ``attempt_id=...`` /
  ``workflow_id=...`` placeholders DO NOT.

Total: 795 → 800 tests, 0 regressions.

The model-override files (.md + .txt) are still in place from the
prior turn — that change is orthogonal to this bug fix; OpenCode
ignores the dispatcher's pass-through per the README, and the actual
generation has been on haiku the whole time. The .md frontmatter
change to sonnet will only take effect on the next OpenCode restart.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 19:11:17 -04:00
drew 84da774212 fix(controller): batch M — post-Phase-1m adversarial review fixes
Three rounds of adversarial review (Chief Architect / Principal Dev /
Senior Test Engineer) on commits 3ca794be7..db12f45ac surfaced ~35
issues. This commit addresses 25+ across criticals, highs, and
mediums, and adds 40 new tests covering the changes plus key gaps the
review identified.

CRITICALS (M1):
- CA1: stale {role}_output.json from a prior attempt on the same
  per-PR workspace was readable as "fresh" output of the new attempt.
  agent_runner now unlinks the MCP-canonical path AND every fallback
  path BEFORE the session runs.
- CA2/PD5: opencode.json-registered MCP subprocesses persist across
  OpenCode sessions, but BuilderState was module-singleton. Added
  reset_for_new_attempt() + cross-session detection (compare
  identity.attempt_id) to every *_start; force-resets with WARN if
  prior attempt was interrupted (timeout / lost lock).
- PD3: inline-JSON callback could overwrite an MCP-written canonical
  V1 file with adapted-from-prose garbage. Callback now inspects
  existing files and skips when V1 is already present.
- PD4: FORGEJO_URL = .rstrip("/api/v1") is a character-set strip —
  catastrophic for hosts whose path contains /v1 in the middle.
  Replaced with explicit endswith()-based suffix strip.
- CA10: clone URL embedded $FORGEJO_TOKEN, persisted into
  .git/config where any agent could cat it. Token now sourced via
  local credential.helper at clone-time, URL kept clean.
- CA12: state.finalized was set BEFORE the file write, so disk-full
  / OSError left the agent unable to retry finalize. Reordered.

HIGHS (M2):
- CA3/PD12: output_path validation (NUL-byte rejection, must be
  absolute, parent-not-file check) in finalize_and_emit.
- CA6: ci_status_poll SELECT only considered implementer attempts;
  conflict_resolver also pushes commits. SQL now unions both roles.
- PD9: ci_status_poll could advance on a stale "resolved" SHA from a
  blocked attempt (whose head_sha_after == head_sha_before). Added
  outcome='resolved' filter.
- CA8: cancelled/stale CI states mapped to ci_red_retry_same_tier,
  burning pickup_count on healthy PRs. Both now wait (treated as
  operator/system action, not failure). timed_out stays red.
- TE9: unknown Forgejo CI states now WARN-log instead of silently
  being treated as pending — operators see new state strings.
- PD8: ci_status_poll event_type strings standardized to match the
  state-machine event names (ci_green / ci_red_retry_same_tier)
  instead of legacy ci-green / ci-red.
- CA7: inline-JSON callback now checks lost_lock_check BEFORE write
  so a file isn't staged after lock loss.
- PD10: atomic .tmp + os.replace writes in both MCP finalize and
  inline callback so the poller never sees a half-written file.
- PD16: inline_output_callback exceptions now re-raise as WorkerError
  instead of being silently logged (root cause was buried 30s later
  in a canonical-output timeout).
- CA9: WorkerConfig manual rebuild on --max-concurrent/--poll-interval
  silently dropped new fields. Use dataclasses.replace, matching
  round-4 P5 fix in master/__main__.py.

MEDIUMS (M3) — legacy_adapter quality upgrades:
- PD1: unrecognized confidence values now WARN instead of silently
  defaulting to "medium" — surfaces agent prompt drift.
- PD2: estimator recommended_tier clamped to {0,1,2} so an out-of-
  range int doesn't bypass the adapter's whole purpose.
- PD7: reviewer blocking_issues list-of-strings coerced into the
  list-of-BlockingIssue-dict shape strict_parse requires.
- PD13: conflict_resolver prompt defaults tier=1 + warns instead of
  raising; the scheduler always sets it but defends against drift.
- PD14: summarizer summary < 50 chars padded with a clear marker so
  strict_parse accepts it (and the truncation is visible).
- PD15: implementer blockers capped at 4096 chars each so a buggy
  agent can't blow up audit log / DB column.
- PD17: launch script accepts either FORGEJO_TOKEN or GITEA_TOKEN
  with a clear error if both are unset.
- PD22: conflict_resolver adapter accepts singular commit_sha
  fallback, matching implementer.
- CA4: every adapter invocation logs role + payload key fingerprint
  so operators can measure agent-migration progress.
- estimator + summarizer now have explicit _start tools (the prompts
  already referenced them; previously absent → first call would fail).

TESTS (M4) — added 40 tests in test_post_review_fixes.py:
- Cross-session MCP state reset (implementer + reviewer + estimator
  + summarizer; intra-session double-start still rejected).
- finalize_and_emit output_path precedence (arg > env > stdout),
  parent-dir creation, rejection of relative/NUL paths, failed-write
  leaves state retryable.
- legacy_adapter quality: tier clamping, blocker cap, non-string
  commit warning, blocking_issues string coercion, conflict_resolver
  full roundtrip + non-resolved head clearing, summarizer padding,
  confidence warning, V1-passthrough no-log.
- opencode.json registration parity: every MCP the prompts name is
  registered with the correct module path.
- Per-role prompts mention {role}_output.json (canonical poller path)
  + the "DO NOT emit chat-JSON" directive.
- FORGEJO_URL suffix-strip parametrized table.
- agent_runner stale-file cleanup: prior-attempt file is unlinked
  before a new session can read it as phantom output.

Also updated 2 pre-existing tests for the CA8 / PD8 / PD13 behavior
changes (cancelled→wait, event_type renaming, conflict_resolver
default-tier warning).

Total: 741 → 781 tests, 0 regressions.

DEFERRED (M5 follow-up — non-trial-blocking):
- CA5: head_sha verification via git cat-file (requires subprocess).
- CA11: discovery_interval_s wall-time cadence (vs iteration count).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 18:26:06 -04:00
drew db12f45acb feat(controller): Phase 1m — wire response-builder MCPs into OpenCode (Option A)
The trial-path expedient (legacy_adapter harvesting chat-JSON) is now
the SAFETY NET; the architecture-intended path (agents call the
controller's response-builder MCPs, which validate against V1 +
write canonical JSON) is now wired end-to-end.

Four pieces:

1. MCP `finalize` accepts `output_path` argument
   ``mcp/_builder_base.finalize_and_emit`` adds ``output_path: str | None``
   kwarg. Precedence: explicit arg > ``CONTROLLER_CANONICAL_OUTPUT_PATH``
   env > stdout. This solves the per-attempt path problem that
   blocked opencode.json registration (the env var is static; the
   per-attempt path comes from the prompt, the agent passes it as a
   tool arg). Each role's finalize updated:
   - estimator_finalize(output_path)
   - implementer_finalize(output_path)
   - reviewer_finalize(output_path)
   - conflict_finalize(output_path)
   - summarizer_finalize(output_path)
   Creates parent directory if missing (so the controller doesn't
   need to pre-create). Returns ``wrote_to`` in the ok dict so tests
   can pin the path.

2. opencode.json registers the 5 controller MCPs
   ``.opencode/opencode.json`` adds:
   - implementer-response-builder
   - reviewer-response-builder
   - estimator-response-builder
   - conflict-resolver-response-builder
   - summarizer-response-builder
   Each spawned via ``python -m tools.controller.mcp.{role}_builder``
   with PYTHONPATH=/repo-root so the controller imports resolve.

3. Controller prompt builder injects the EXACT tool sequence
   ``worker/prompts.py`` rewrites each role's "Output contract"
   section. The old "PREFERRED MCP / FALLBACK file-write" instruction
   becomes a single REQUIRED contract: numbered tool calls (``X_start``,
   ``X_set_*``, ``X_finalize(output_path=...)``) with the explicit
   per-attempt path baked in. Explicit "DO NOT emit a JSON object in
   your final chat message" instruction to override the legacy
   contract baked into the agent system prompts.

4. Agent permission whitelists include the new MCPs
   - ``.opencode/agents/task-implementor.md`` (source for tier-{0,1,2,min}
     variants — regenerated via tools/sync_tier_models.py)
   - ``.opencode/agents/pr-review-worker.md``
   - ``.opencode/agents/estimator-implementation.md``
   - ``.opencode/agents/conflict-resolver-worker.md``
   Each adds the matching ``"{role}*": allow`` pattern.

Safety net preserved:
The legacy_adapter (commit bcc59d38a) is KEPT as a fallback path.
If an agent ignores the new instruction + emits chat-JSON anyway,
``opencode_session.py:inline_output_callback`` harvests it +
``legacy_adapter`` normalizes to V1 + writes to the canonical out
path. Both paths produce valid V1 → ``strict_parse`` succeeds. The
MCP path is now the WORKING preferred path; the chat-JSON harvest
is the safety net.

Total: 738 controller tests pass (+0 net — wiring change, no new
behavior tests), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 18:10:48 -04:00
drew eebb5718a8 feat(controller): Phase 1a — 5 response-builder MCP servers
Per-attempt MCP subprocesses that enforce V1 contract invariants at
construction time. Worker LLM calls builder tools incrementally; the
MCP validates each call against the schema + cross-field invariants;
`{role}_finalize()` emits canonical Pydantic-validated JSON to stdout
for the worker controller to read (Phase 1c). Defense-in-depth: the
controller strict-parses whatever finalize emits.

Builders shipped:

- reviewer_builder: 9 tools. Auto-acks all CISummary gates as passed
  at start; reviewer only calls record_gate to discuss specifics.
  reviewer_override_gate requires ≥20-char justification. Approve
  with any failed gate is refused with an actionable error pointing
  at the override path. Request-changes requires ≥1 blocking issue.
  Verdict-vs-blocking-issues invariant checked at finalize.

- implementer_builder: 7 tools. Outcome-specific finalize invariants:
  resolved → ≥1 commit + ≥1 file; blocked → ≥1 blocker; noop → no
  commits/files/blockers.

- estimator_builder: 4 tools. Lightweight; requires
  recommended_tier + confidence + reasoning at finalize. Reasoning
  capped at 2048 chars.

- conflict_resolver_builder: 8 tools. outcome='resolved' requires
  new_head_sha + ≥1 commit + ≥1 file. resolution_strategy enum-checked.

- summarizer_builder: 3 tools. Summary 50-2000 chars (enforced at
  MCP layer and Pydantic).

Shared infrastructure:

- _builder_base.py: BuilderState dataclass + invariant guard helpers
  (require_started / require_not_finalized) + audit-record-with-summary
  + finalize_and_emit (validates against Pydantic model class,
  emits canonical JSON to stdout, marks finalized).

41 builder tests in test_mcp_builders.py (happy paths + every
invariant + outcome-specific paths + audit summarization + JSON
round-trip through Pydantic strict-parse). Plus the existing 62
Phase-0 tests. 103 controller tests total. Full auto_agents suite
(2465 tests) still passes.
2026-05-18 11:09:40 -04:00