Commit Graph

31 Commits

Author SHA1 Message Date
drew a6986008ee feat(controller): trial-5 batch — dispute path, merge pipeline split, conflict-resolver hardening
T5-1   reviewer feedback rendered full-body to the implementer
T5-4/9 implementer dispute path — dispute-at-any-tier with per-tier cap,
       OPERATOR_ATTENTION state on stalemate, pr-review-worker-dispute agent
T5-5   reviewer BLOCKING ISSUE EVIDENCE RULE + 5-step validation
T5-7   merge step split into a singleton process — impl/review masters write
       APPROVED and stop; merge_drive owns APPROVED -> MERGING -> MERGED
T5-10  merge process is fully deterministic; base conflicts bounce to the
       controller's CONFLICT_RESOLVING (LLM); conflict_drive sidecar retired
T5-11  implementer fast success path — verified-clean outcome so a no-op
       after conflict resolution doesn't force busywork
T5-12  conflict-resolver permissions fixed across all paths (/tmp/** glob)
T5-13  conflict-resolver PR-intent prehydration (title/body/comments)

Adds tools/_controller_db_bridge.py so merge_drive reads the controller DB
directly (Option B), plus APPROVED + OPERATOR_ATTENTION states, the
dispute/verified-clean events, and the V1 contract fields backing them.
Reviewer model: baseline -> sonnet, dispute -> opus.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 17:54:07 -04:00
drew ab8a5e5bfb fix(controller): T4-4 — route issue workflows to STUCK when role has no prefetch
Trial-4 observed wf=6 (entity #34, kind='issue') promoted ANALYZING→
IMPLEMENTING by the estimator, but build_implementer_input raises
``ValueError("implementer prefetch needs kind='pr', got 'issue'")``.
The scheduler logged WARNING and retried every ~5s forever — log spam
+ workflow never reached a terminal state. Pre-fix:

  2026-05-18 21:13:43 WARNING ... prefetch failed ...
  2026-05-18 21:13:50 WARNING ... prefetch failed ...
  2026-05-18 21:13:55 WARNING ... prefetch failed ...
  (50+ identical lines over the trial)

Fix: before calling prefetch, scheduler checks the workflow's kind
column. If kind='issue' and role in
{implementer, reviewer, conflict_resolver}, the workflow is routed
to STUCK with reason='issue-not-supported (T4-4)'. Operator-visible
controller_events row tags this as 'issue-not-supported'.

Future work: IssueImplementerOutputV1 contract already exists in
contracts/v1.py:362; a follow-up batch can add ``build_issue_
implementer_input`` + remove this short-circuit. For the trial we
just need issues to stop wedging.

Test: TestIssueWorkflowRoutedToStuck in test_batch_s_fixes.py.

833 controller tests pass (was 832; +1).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 21:20:17 -04:00
drew 46764b841b fix(controller): batch S — 11 fixes from trial-4 finding + 2nd code-review pass
Trial-4 ran the controller past ANALYZING for the first time. One
critical regression surfaced live (T4-1 — scheduler dup-attempt
enqueue), and a parallel agent-driven code review found 8 more bugs
across "races + error paths + agent quality" classes. Batch S
addresses 11 of those.

LIVE-OBSERVED REGRESSION (trial-4 2026-05-19 00:49):

T4-1 — Scheduler enqueues duplicate estimator after workflow advanced
File: tools/controller/master/scheduler.py
The "already pending" check filtered on
``status IN ('pending', 'in_progress')`` — missing the brief window
where the prior attempt is ``status='complete'`` but tick hasn't yet
processed its outcome. Result: scheduler enqueues a 2nd estimator/etc.;
when its outcome fires from a now-advanced state, IllegalTransition
→ STUCK. Hit wf=1 in trial-4. Fix: also skip when an unprocessed
``complete`` attempt exists (finished_at > w.last_transition_at).

E-5 — IllegalTransition over-aggressive STUCKing
File: tools/controller/master/tick.py
Companion fix to T4-1. Even if T4-1 escapes in some other path (or
a worker delays writing outcome past tick), a stale-outcome
(workflow already advanced via parallel path like ci_status_poll or
reconciliation) shouldn't STUCK. New ``_is_stale_role_outcome``
helper recognizes "this role's outcome arrived after the workflow
moved past its origin state" → consume the attempt, bump
last_transition_at, continue. Only genuine state corruption → STUCK.

E-3 — Corrupted output_payload silently wedges workflow
File: tools/controller/master/tick.py
Pre-fix _decode_output_payload swallowed json.JSONDecodeError →
mapper returned None → tick bumped last_transition_at but workflow
never moved. Operator had no signal. Now raises
``CorruptedOutputPayload`` → tick routes to STUCK with reason.

E-1 — Contract-violation routed to STUCK on first attempt
File: tools/controller/master/outcomes.py + tick.py
v9 spec promised retry-once-with-corrective-prompt for
contract-violation; the column ``strict_parse_retries`` existed but
nothing read/incremented it. _map_failed_outcome now takes
``prior_contract_violations`` count (queried in tick.py); STUCKs
only when count ≥ _CONTRACT_VIOLATION_RETRY_LIMIT (2). First two
violations re-enqueue.

R-1 — Reconciliation flipped workflow state mid-attempt
File: tools/controller/master/reconciliation.py
Worker holding a lock + heartbeating; reconciliation flipped current_state
to MERGED/ABANDONED based on Forgejo; tick.py then skipped the
worker's eventual write (terminal-state exclusion). Worker's output
lost. Fix: _apply_transition first checks for in_progress attempts
on the same workflow and defers if any exist.

R-2 + R-8 — _apply_transition lacks current_state guard
Files: reconciliation.py + merging.py
Same pattern as ci_status_poll's existing TOCTOU defense. UPDATE now
filters ``WHERE current_state = :from_state``; on rowcount=0, skip
the event row. Prevents racing ticks from over-writing each other.

R-4 — Reaper UPDATE didn't re-check heartbeat freshness
File: tools/controller/reaper.py
A worker's healthy heartbeat between reaper's SELECT and UPDATE
would be silently overwritten; the worker's later _write_outcome
(filtered on locked_by_instance) returned rowcount=0 → output lost.
UPDATE now includes the same freshness filter as the SELECT, so
fresh heartbeats protect the row.

E-8 — merging_retry_count not reset on STUCK/abandoned paths
File: tools/controller/master/merging.py
Pre-fix, only 200/409/422 paths reset the counter. 403 (branch
protection), 404 (externally closed → ABANDONED), retry-exhausted
(STUCK) leaked stale counts. If operator unsticks a STUCK workflow
back through MERGING, the stale count made it STUCK again sooner
than expected. All terminal-state-changing paths now reset.

A-1 — commit_shas validation accepted any ≥7-char string
File: tools/controller/mcp/implementer_builder.py
Tightened to ``re.fullmatch(r"[0-9a-f]{7,40}")``. Pre-fix an agent
could pass any 7+ char string; head_sha_advanced accepted the
hallucination; CI poll then 404'd on the fake SHA forever (until
2h ci_poll_exhaustion).

A-2 — merging-409 → IMPLEMENTING(tier=NULL) trap
File: tools/controller/master/merging.py
On 409 the handler routes to IMPLEMENTING(tier=tier_last_succeeded);
if that's NULL (e.g., metadata-only → REVIEWING → approve → 409 path
where implementer_pushed never fired), the MCP rejects tier=NULL →
contract-violation → STUCK. Fix: default to current_tier when
tier_last_succeeded is NULL.

A-9 — Reviewer cross-field check: verdict ↔ suggested_next_action
File: tools/controller/mcp/reviewer_builder.py
Agent could set verdict=approve + suggested_next_action=abandon; the
master fired reviewer_approve regardless. New _VERDICT_ACTION_COMPAT
map enforces compatible pairs at finalize.

TESTS:

- New file ``test_batch_s_fixes.py`` with 13 regression tests, one
  per fix class.
- Updated test_master_outcomes.py for the new contract-violation
  retry behavior.
- Updated test_master_prefetch.py fixture to bump last_transition_at
  past the seeded attempts (the new T4-1 filter would otherwise
  correctly identify the pre-seeded attempts as unprocessed).

Total: 819 → 832 tests, 0 regressions.

DEFERRED to a follow-up batch (per PENDING_FIXES.md):
- 8 MEDIUM items (error UX, dead fields, signal-loss in error paths)
- 5 LOW items (agent quality polish)
- 4 CONFIRMED-CLEAN (no fix needed)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 21:15:26 -04:00
drew 3e23853ffa fix(controller): batch R — wire 5 V1-contract fields the controller silently dropped
Trial run-3 (2026-05-19) surfaced the first instance of a broader bug
class: V1 contract fields existed and agents emitted them, but no
controller code wired them into state transitions. An adversarial
"walk the happy path" code review found 4 more, all listed below.

The class shape: a V1 field is "Required iff X" by contract docstring,
the worker emits it correctly, but the master reads the wrong field
(or doesn't read it at all), so a critical state transition silently
no-ops or drops to the wrong default.

FIX #0 — outcome-mapper early-return (committed earlier in this
session) — moved role dispatch before the ``outcome is None`` guard
so estimator+reviewer+summarizer (V1 contracts without an ``outcome``
field) are correctly handled. Without this fix, all estimator
attempts in trial run-3 completed successfully then were silently
discarded, stranding all 6 workflows in ANALYZING.

FIX #1 — current_tier never written from estimator's recommended_tier
File: tools/controller/master/tick.py
The ANALYZING→IMPLEMENTING UPDATE wrote only current_state /
last_transition_at / entered_state_at. recommended_tier from the
estimator payload was never extracted, so every PR ran at the
workflow's creation-time tier (typically 0) regardless of what the
estimator recommended — the entire tier-escalation ladder was
informational-only. Fix: per-event ``extra_set`` clauses; on
``estimator_done`` / ``estimator_metadata_only`` events, set
``current_tier = :rec_tier`` from the payload (with 0..2 validation).
Tests: TestEstimatorRecommendedTierWritten (3 cases).

FIX #2 — approved_at_sha never passed to merge callback
File: tools/controller/master/merging.py, forgejo_http.py
ReviewerOutputV1.approved_at_sha is the exact SHA the reviewer
signed off on. Pre-fix the MergeCallback signature was
``(owner, repo, pr_number)`` — Forgejo merged whatever HEAD currently
was. Race condition: a concurrent push (operator or another driver)
between approval and merge would silently merge unapproved code.
Fix: extended signature to ``(owner, repo, pr_number, approved_at_sha)``;
SQL SELECT now pulls the latest reviewer attempt's output_payload as
a subquery; merge_pr forwards it to Forgejo as ``head_commit_id``
(Forgejo refuses with 409 if HEAD has advanced). Defensive: still
merges when approved_at_sha is None but logs a WARNING. Tests:
TestApprovedAtShaPassedToMerge (2 cases).

FIX #3 — tier_last_succeeded column had ZERO writers
File: tools/controller/master/tick.py
The schema column existed; the merging.py 409-conflict path read it
to recover the last-known-good tier; but NOTHING ever wrote to it.
Every workflow's tier_last_succeeded was permanently NULL → the
409-recovery path transitioned to IMPLEMENTING(tier=NULL) → scheduler
silently coerced to tier 0. Fix: on ``implementer_pushed`` event,
``UPDATE workflows SET tier_last_succeeded = current_tier``. Tests:
TestTierLastSucceededWritten.

FIX #4 — outcome column NULL for estimator/reviewer/summarizer
File: tools/controller/worker/runner.py
``workflow_attempts.outcome`` is the operator-facing audit column.
Pre-fix the runner extracted ``output_payload.get("outcome")``
blindly — works for implementer/conflict_resolver but those three
roles have no ``outcome`` field. Result: ``SELECT … WHERE outcome IS
NOT NULL`` audit queries silently missed every estimator/reviewer/
summarizer attempt. Fix: new ``_derive_outcome_for_audit(role, payload)``
helper synthesizes meaningful per-role values:
  - implementer/conflict_resolver: payload['outcome'] (unchanged)
  - reviewer: payload['verdict']
  - estimator: 'metadata-only' OR f'tier-{recommended_tier}'
  - summarizer: 'summarized'
Tests: TestOutcomeAuditColumn (parametrized 8 cases).

FIX #5 — conflict_resolver new_head_sha never preferred
File: tools/controller/worker/runner.py
ConflictResolverOutputV1.new_head_sha is "Required iff outcome='resolved'"
(the canonical post-rebase branch tip). Pre-fix runner.py used
``commit_shas[-1]`` for head_sha_after — works for normal git rebase
--continue but wrong for resolvers that did force-pushed merge commits
where the last commit SHA ≠ the branch tip. CI status poll would then
poll the wrong SHA. Fix: when role=='conflict_resolver', prefer
``new_head_sha`` over commits[-1]. Tests:
TestConflictResolverNewHeadShaUsed (2 cases).

ALSO updated existing tests that papered over the original bug:
- test_master_outcomes.py: estimator tests used to inject a fake
  ``"outcome": "(implicit)"`` field; now use real V1 shape (no outcome).
  Reviewer tests now use ``verdict`` (the real V1 field) not ``outcome``.
- test_master_tick.py reviewer tests: same `verdict` switch.
- test_master_merging.py: updated all 13 ``lambda o, r, n: ...``
  merge-callback stubs to the new 4-arg signature.

CONFIRMED-CLEAN (no fix needed) by the same code review:
- outcomes.py post-fix-#0
- prefetch.py field reads
- prompts.py field accesses
- ci_status_poll.py role+outcome filter
The above were verified to handle all 5 V1 contract shapes correctly.

Total: 802 → 819 controller tests, 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 20:40:31 -04:00
drew 84da774212 fix(controller): batch M — post-Phase-1m adversarial review fixes
Three rounds of adversarial review (Chief Architect / Principal Dev /
Senior Test Engineer) on commits 3ca794be7..db12f45ac surfaced ~35
issues. This commit addresses 25+ across criticals, highs, and
mediums, and adds 40 new tests covering the changes plus key gaps the
review identified.

CRITICALS (M1):
- CA1: stale {role}_output.json from a prior attempt on the same
  per-PR workspace was readable as "fresh" output of the new attempt.
  agent_runner now unlinks the MCP-canonical path AND every fallback
  path BEFORE the session runs.
- CA2/PD5: opencode.json-registered MCP subprocesses persist across
  OpenCode sessions, but BuilderState was module-singleton. Added
  reset_for_new_attempt() + cross-session detection (compare
  identity.attempt_id) to every *_start; force-resets with WARN if
  prior attempt was interrupted (timeout / lost lock).
- PD3: inline-JSON callback could overwrite an MCP-written canonical
  V1 file with adapted-from-prose garbage. Callback now inspects
  existing files and skips when V1 is already present.
- PD4: FORGEJO_URL = .rstrip("/api/v1") is a character-set strip —
  catastrophic for hosts whose path contains /v1 in the middle.
  Replaced with explicit endswith()-based suffix strip.
- CA10: clone URL embedded $FORGEJO_TOKEN, persisted into
  .git/config where any agent could cat it. Token now sourced via
  local credential.helper at clone-time, URL kept clean.
- CA12: state.finalized was set BEFORE the file write, so disk-full
  / OSError left the agent unable to retry finalize. Reordered.

HIGHS (M2):
- CA3/PD12: output_path validation (NUL-byte rejection, must be
  absolute, parent-not-file check) in finalize_and_emit.
- CA6: ci_status_poll SELECT only considered implementer attempts;
  conflict_resolver also pushes commits. SQL now unions both roles.
- PD9: ci_status_poll could advance on a stale "resolved" SHA from a
  blocked attempt (whose head_sha_after == head_sha_before). Added
  outcome='resolved' filter.
- CA8: cancelled/stale CI states mapped to ci_red_retry_same_tier,
  burning pickup_count on healthy PRs. Both now wait (treated as
  operator/system action, not failure). timed_out stays red.
- TE9: unknown Forgejo CI states now WARN-log instead of silently
  being treated as pending — operators see new state strings.
- PD8: ci_status_poll event_type strings standardized to match the
  state-machine event names (ci_green / ci_red_retry_same_tier)
  instead of legacy ci-green / ci-red.
- CA7: inline-JSON callback now checks lost_lock_check BEFORE write
  so a file isn't staged after lock loss.
- PD10: atomic .tmp + os.replace writes in both MCP finalize and
  inline callback so the poller never sees a half-written file.
- PD16: inline_output_callback exceptions now re-raise as WorkerError
  instead of being silently logged (root cause was buried 30s later
  in a canonical-output timeout).
- CA9: WorkerConfig manual rebuild on --max-concurrent/--poll-interval
  silently dropped new fields. Use dataclasses.replace, matching
  round-4 P5 fix in master/__main__.py.

MEDIUMS (M3) — legacy_adapter quality upgrades:
- PD1: unrecognized confidence values now WARN instead of silently
  defaulting to "medium" — surfaces agent prompt drift.
- PD2: estimator recommended_tier clamped to {0,1,2} so an out-of-
  range int doesn't bypass the adapter's whole purpose.
- PD7: reviewer blocking_issues list-of-strings coerced into the
  list-of-BlockingIssue-dict shape strict_parse requires.
- PD13: conflict_resolver prompt defaults tier=1 + warns instead of
  raising; the scheduler always sets it but defends against drift.
- PD14: summarizer summary < 50 chars padded with a clear marker so
  strict_parse accepts it (and the truncation is visible).
- PD15: implementer blockers capped at 4096 chars each so a buggy
  agent can't blow up audit log / DB column.
- PD17: launch script accepts either FORGEJO_TOKEN or GITEA_TOKEN
  with a clear error if both are unset.
- PD22: conflict_resolver adapter accepts singular commit_sha
  fallback, matching implementer.
- CA4: every adapter invocation logs role + payload key fingerprint
  so operators can measure agent-migration progress.
- estimator + summarizer now have explicit _start tools (the prompts
  already referenced them; previously absent → first call would fail).

TESTS (M4) — added 40 tests in test_post_review_fixes.py:
- Cross-session MCP state reset (implementer + reviewer + estimator
  + summarizer; intra-session double-start still rejected).
- finalize_and_emit output_path precedence (arg > env > stdout),
  parent-dir creation, rejection of relative/NUL paths, failed-write
  leaves state retryable.
- legacy_adapter quality: tier clamping, blocker cap, non-string
  commit warning, blocking_issues string coercion, conflict_resolver
  full roundtrip + non-resolved head clearing, summarizer padding,
  confidence warning, V1-passthrough no-log.
- opencode.json registration parity: every MCP the prompts name is
  registered with the correct module path.
- Per-role prompts mention {role}_output.json (canonical poller path)
  + the "DO NOT emit chat-JSON" directive.
- FORGEJO_URL suffix-strip parametrized table.
- agent_runner stale-file cleanup: prior-attempt file is unlinked
  before a new session can read it as phantom output.

Also updated 2 pre-existing tests for the CA8 / PD8 / PD13 behavior
changes (cancelled→wait, event_type renaming, conflict_resolver
default-tier warning).

Total: 741 → 781 tests, 0 regressions.

DEFERRED (M5 follow-up — non-trial-blocking):
- CA5: head_sha verification via git cat-file (requires subprocess).
- CA11: discovery_interval_s wall-time cadence (vs iteration count).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 18:26:06 -04:00
drew 3ca794be75 feat(controller): autonomous CI status polling — closes the last trial gap
The Phase 2 trial previously required operator-intervention SQL to
advance workflows from AWAITING_CI → REVIEWING (no automated CI
status polling). This commit wires the missing tick so the trial
runs end-to-end without manual help.

Components:
- ``master/forgejo_http.py``: new ``get_ci_status`` callback wraps
  Forgejo's ``/commits/{sha}/status`` combined-status endpoint;
  added to ``ForgejoCallbacks``.
- ``master/ci_status_poll.py`` (NEW): ``run_ci_status_poll_tick``
  scans AWAITING_CI workflows, fetches CI status keyed on the
  latest implementer attempt's ``head_sha_after``, and applies
  state transitions via ``apply_event``. TOCTOU-defended UPDATE
  (``WHERE current_state='AWAITING_CI'``) + per-row exception
  isolation.
- ``master/loop.py``: new ``ci_status_poll_args=(owner, repo,
  get_ci_status)`` kwarg + ``ci_status_poll_interval_s`` config
  (default 60s) + ``MasterTickReport.ci_status_poll`` field.
- ``master/__main__.py``: threads ``callbacks.get_ci_status`` into
  the loop.

State mapping (Forgejo combined-status state → event):
- success / neutral / skipped / warning → ci_green → REVIEWING
- failure / error / cancelled / timed_out / stale →
  ci_red_retry_same_tier → IMPLEMENTING
- pending / queued / in_progress / action_required → no-op (wait)
- None / unknown / fetch failure → no-op (transient)

The ``ci_polling_exhausted`` timeout (default 2h) remains as the
safety net for CI that genuinely never reports.

Tests (+14 in test_master_ci_status_poll.py):
- Happy paths (success→green, failure→red, pending→wait)
- Error paths (callback raises; workflow without head_sha)
- Event row shape (event_type='ci-green'/'ci-red', reason payload)
- Extended state mapping (cancelled, neutral, in_progress)
- Other-repo isolation
- LoopIntegration end-to-end via master_main_loop with safety timer

RUNBOOK updated: removed the manual SQL workaround; added the
autonomous CI poll's tunables.

Total: 726 controller tests pass (+14 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 17:37:04 -04:00
drew f57d9f9478 fix(controller): batch L — round-4 trial-blockers (A1, P1, P3, P4, P5, T5)
Round-4 adversarial review found 5 trial-blockers + 1 silent-debt
item the post-round-3 deep pass missed. All fixed.

A1 — pre-clone the workspace so the agent has a worktree to operate on
``worker/__main__.py``: the agent_runner closure now constructs a
``PerPRWorkspace`` from input_payload.owner/repo/pr_number + the
FORGEJO_URL+FORGEJO_TOKEN env vars. Pre-flight:
- ``workspace.ensure_present()`` creates the dir skeleton.
- ``workspace.clone_if_absent()`` clones the repo into
  ``{workspace_dir}/worktree/`` if not already present (idempotent).
- ``workspace.fetch_and_validate(head_sha, head_ref)`` refreshes +
  verifies the workspace is at the expected head. ``StaleInputError``
  → ``WorkerError(outcome='stale-input')`` so the master re-prefetches
  without burning a pickup. ``RuntimeError`` → ``worker-internal-error``.
Previously the agent saw an empty workspace_dir + had no repo.

P1 — partial-write defense in the canonical-output poller
``worker/agent_runner.py:_wait_for_canonical_output`` now polls each
path with a two-pass quiescence check (size stable + content parses
as JSON) before returning. Partial writes (agent crashed mid-flush)
are skipped + the polling loop continues. The previous
``f.read().strip()`` returned partial JSON which then tripped
``ContractValidationError`` → ``worker-internal-error`` with no
record of WHICH path; now logs source path on every read.

P3 — TOCTOU defense in promote_discovered
``master/promote.py``: the UPDATE now filters
``current_state='DISCOVERED'``. If a concurrent reconciliation
moved the row off DISCOVERED between SELECT and UPDATE, rowcount=0
+ we skip the event-row write. No duplicate audit entry; no
overwriting a pause-by-label-removal.

P4 — explicit tuple-length validation in reconciliation_args + discovery_args
``master/loop.py``: previously a 6-tuple silently fell into the
``else`` 4-tuple unpack, raised ValueError("too many values"), got
swallowed by the per-iter ``except Exception``, and reconciliation
silently died forever. Now: ``elif n == 4`` + ``else: raise TypeError``.
The TypeError still hits the per-iter except (so the loop doesn't
crash) but ``logger.exception`` surfaces the actionable message in
journald. Operator sees "reconciliation_args must be a 4- or 5-tuple;
got length 6" instead of zero indication.

P5 — --tick-interval CLI flag preserves other config fields
``master/__main__.py``: replaced the manual ``MasterConfig(...)``
rebuild (which dropped reconciliation/ci_poll/discovery intervals)
with ``dataclasses.replace(cfg_loop, tick_interval_s=args.tick_interval)``.
Operators who pass --tick-interval no longer silently revert the
other intervals to defaults.

T5 — scheduler._commit_escalation uses safe_json_dumps
``master/scheduler.py``: the escalation event row's payload was the
only call site that bypassed safe_json_dumps. Now consistent — a
future contributor adding a datetime/Decimal field won't trip raw
json.dumps at runtime.

Tests (+4 net):
- ``test_worker_agent_runner.py::test_partial_write_not_read``: pins
  P1 (truncated fallback file + valid MCP output → MCP wins).
- ``test_master_promote.py::test_toctou_state_change_between_select_and_update``:
  pins P3 (steal state via monkey-patch → no double-promotion, no
  extra event row).
- ``test_master_loop.py::test_reconciliation_args_wrong_length_logs_not_silent``:
  pins P4 (6-tuple → logged error, not silent forever).
- ``test_entry_points.py::test_tick_interval_flag_preserves_other_cfg_fields``:
  pins P5 (env-set non-default intervals survive --tick-interval).

Total: 711 controller tests pass (+4 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 17:23:35 -04:00
drew 251eeb21ff fix(controller): more pipeline run-blockers — merging tick, periodic discovery, worker create_all, systemd ordering
Continuing the round-3 deep-pass cleanup. Three more run-blockers
+ one robustness fix.

RB5 — MERGING handler never invoked from master loop:
``run_merging_tick`` was exported by the master package but no caller
fired it. Workflows that transition to MERGING (via reviewer
approval) would sit there indefinitely with no Forgejo merge call.

Fix:
- ``master/loop.py`` accepts a ``merging_args=(owner, repo,
  merge_callback)`` kwarg. When set, the tick fires every iteration
  (cheap if no workflows in MERGING).
- ``MasterTickReport`` gains ``merging: MergingHandlerReport | None``.
- ``master/__main__.py`` wires it from the Forgejo callback bundle.

RB6 — periodic discovery never fires:
``run_discovery`` was only called at startup via
``run_startup_backfill`` + the ``--discovery-only-once`` smoke flag.
PRs created after master startup would not be discovered until the
master restarted.

Fix:
- ``master/loop.py`` accepts ``discovery_args=(owner, repo, list_prs,
  list_issues)`` or the 5-tuple with kwargs. Periodic tick on its
  own cadence (``CONTROLLER_DISCOVERY_INTERVAL_S``, default 30s).
- ``MasterTickReport`` gains ``discovery: DiscoveryReport | None``.
- ``master/__main__.py`` wires it + threads ``require_opt_in_label``
  through.

RB-robust — worker calls create_all defensively:
Master is normally responsible for schema creation (workers run
After= it via systemd ordering). But if the worker is started in
isolation (test / local dev / unit ordering broken), it'd crash on
the first query against missing tables.

Fix:
- ``worker/__main__.py`` calls ``create_all(engine)`` after
  ``build_engine``. ``create_all`` is idempotent (CREATE TABLE IF
  NOT EXISTS); safe to call from both master + worker.
- ``cleveragents-controller-worker@.service`` adds
  ``After=cleveragents-controller-master.service`` +
  ``Wants=cleveragents-controller-master.service`` so systemd
  enforces the start ordering in production.

Total: 703 controller tests pass (no test changes; all new wiring
is exercised by master_main_loop tests via the new kwargs).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 16:47:10 -04:00
drew febb352618 fix(controller): pipeline run-blockers — promoter, scheduler, owner/repo, workspace_dir patch
Round-3 deep pass identified four issues that would have prevented an
actual end-to-end pipeline run:

RB1 — DISCOVERED → ANALYZING never fired in production:
The state machine defines (DISCOVERED, discovery_picked_up) →
ANALYZING but NO production code fires the event. Workflows
created by discovery would sit in DISCOVERED forever.

Fix:
- New ``master/promote.py``: ``run_promote_discovered_tick`` scans
  for DISCOVERED workflows + fires ``discovery_picked_up`` via
  apply_event (state-machine invariants stay enforced) + emits a
  ``discovery-promoted`` controller_events row per transition.
- Composes with the master loop's other ticks; runs every iteration
  (cheap — typically 0-1 row).

RB2 — scheduler.schedule_next_attempts never called from master loop:
The scheduler was exported by the master package but never invoked.
It creates the ``workflow_attempts`` rows that workers dequeue —
without it, workers would have nothing to pick up.

Fix:
- ``master/loop.py`` now accepts a ``prefetch: PrefetchCallback``
  kwarg. When provided, the loop runs promote_discovered + scheduler
  every iteration after tick/reaper/reconciliation.
- ``MasterTickReport`` gains ``promote_discovered`` and ``scheduler``
  optional fields so on_iteration callbacks see both.
- ``master/__main__.py`` builds a ``PrefetchDataCallbacks`` from the
  Forgejo callback bundle and constructs the production
  ``make_prefetch_callback(engine, callbacks)`` — wires through to
  the loop's new prefetch kwarg.

RB3 — owner / repo missing from V1 input contracts:
The implementer / reviewer / estimator / conflict-resolver V1 inputs
had pr_number but not owner/repo. The OpenCode agent would have
had no way to know which Forgejo repo to clone — it would have had
to derive owner/repo from process env, coupling the worker to a
single repo.

Fix:
- ``contracts/v1.py``: added ``owner: str`` and ``repo: str``
  (min_length=1) to ImplementerInputV1, ReviewerInputV1,
  EstimatorInputV1, ConflictResolverInputV1.
- ``master/prefetch.py``: builders populate owner/repo from the
  Workflow row (already known at prefetch time).
- Existing test fixtures in ``test_contracts_v1.py`` updated.

RB4 — input_payload.workspace_dir placeholder reached the agent:
Prefetch wrote ``workspace_dir = "<worker-injected>"`` as a
placeholder; the worker never patched it before invoking the
OpenCode session. The prompt builder rendered the literal
placeholder string into the agent's prompt — the agent had no idea
where to clone.

Fix:
- ``worker/agent_runner.py``: patches input_payload.workspace_dir
  with the real path immediately before calling run_opencode_session.
  Uses a shallow copy so the caller's dict isn't side-effected.
- ``worker/__main__.py``: workspace_dir naming convention is now
  ``pr-{owner}-{repo}-{pr_number}`` (matches workspace.py's
  PerPRWorkspace convention) so the janitor's pr-* glob + the
  agent's expected workspace location agree. Falls back to
  ``pr-attempt-{N}`` for legacy input_payloads missing owner/repo.

Tests:
- ``test_master_promote.py`` (NEW, +7 tests):
  - empty DB no-op
  - single workflow promoted
  - multiple promoted in one tick
  - only DISCOVERED targeted (non-DISCOVERED untouched)
  - controller_events row emitted with correct shape
  - idempotent after first promotion
  - LoopIntegration end-to-end: DISCOVERED → ANALYZING → pending
    estimator attempt visible in workflow_attempts (pins the entire
    previously-broken pipeline from discovery to enqueue)

Total: 703 controller tests pass (+7 net), 0 regressions.

Without these four fixes, the pipeline would have looked alive in
unit tests but produced zero work in a real deployment.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 16:43:36 -04:00
drew 9b6b64f0d3 fix(controller): batch I — round-3 correctness items (R2, R7, R9)
R2 — STUCK short-circuits the opt-in label gate:
Round-2's batch G shipped reconciliation-ordering with merged/closed
winning over label removal. But _decide_transition returns
("STUCK", "pr-not-found-on-forgejo") for 404s — that STUCK was
falling through to the label gate. Operator removing the label on a
deleted PR could PAUSE it forever.

Fix: extend the short-circuit set in master/reconciliation.py to
include STUCK alongside MERGED/ABANDONED. All Forgejo-terminal
transitions now bypass the label gate.

Test: test_pr_404_takes_priority_over_label_removal pins the
contract end-to-end (404 + no opt-in label → STUCK, not PAUSED).

R7 — defense-in-depth for corrupt workflow state in ci_poll:
ci_poll.py only caught IllegalTransitionError from apply_event, but
apply_event raises ValueError for states not in KNOWN_STATES (DB
row corruption, unknown-state guard miss). A single bad row would
have aborted the whole tick.

Fix: broaden the exception handler to (IllegalTransitionError,
ValueError). One bad row is skipped; valid rows still STUCK.

Test: test_unknown_state_skips_row_doesnt_crash_tick monkey-patches
apply_event to raise ValueError once + verifies the tick processes
the other workflow normally.

R9 — distinct event_types per reconciliation reason:
Pause / resume / external-merge / external-close / external-issue-
close / pr-not-found-on-forgejo all used event_type='reconciliation'.
Operators querying controller_events for "what happened" could
only distinguish via JSON-payload LIKE queries — dialect-specific
(SQLite LIKE vs Postgres ::jsonb->>).

Fix: master/reconciliation.py introduces _REASON_TO_EVENT_TYPE
mapping:
- opt-in-label-removed       → 'label-pause'
- opt-in-label-restored      → 'label-resume'
- externally-merged          → 'external-merge'
- externally-closed-not-merged → 'external-close'
- issue-closed-externally    → 'external-issue-close'
- pr-not-found-on-forgejo    → 'external-pr-deleted'
Unknown reasons fall back to 'reconciliation' so future contributors
adding a new reason still emit a well-formed row.

Both _apply_transition and _apply_transition_with_pre_pause now
derive event_type via _event_type_for(reason).

Tests updated (filter by new event_type per case):
- test_master_reconciliation.py::TestEventRows refactored:
  - test_transition_emits_external_merge_event
  - test_closed_pr_emits_external_close_event (NEW)
  - test_pr_404_emits_external_pr_deleted_event (NEW)
  - test_consistent_workflow_no_event widened to all 7 event_types
- test_label_gate.py: pause/resume tests filter by 'label-pause' /
  'label-resume' respectively + assert the event_type matches.

Total: 693 controller tests pass (+2 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 16:32:23 -04:00
drew c4a7f0027d fix(controller): batch H — round-2 important items (N5-N10)
Six items from the round-2 adversarial review's "important" list.

N5 — restricted JSON encoder replaces ``default=str``:
``default=str`` silently stringified custom objects, sets, and bytes
to ``"<MyObj at 0x...>"`` — masking worker output bugs. Now uses a
restricted encoder (``tools/controller/_json_safe.safe_json_dumps``)
with an allowlist:
- datetime / date → ISO-8601 string
- Decimal → str (preserves precision)
- UUID → canonical string
- Path → str
- set / frozenset → sorted list (best-effort)
- Everything else → TypeError (a worker output regression surfaces
  loudly instead of writing garbage to the DB)

Replaces ``json.dumps(..., default=str)`` at:
- ``worker/runner.py`` (terminal-state UPDATE write)
- ``master/scheduler.py`` (input_payload INSERT + UPDATE)
Tests: ``test_json_safe.py`` (+15 tests) — every allowlisted type +
rejected types (custom class, bytes, complex) + nested structures
+ kwargs forwarding.

N6 — strict parser check now runs BEFORE backfill + loop:
Previously the strict-parser exit could run AFTER backfill (the test
``test_strict_parser_coverage_blocks_startup_with_stubs`` passed
because backfill's ``fail_if_called`` AssertionError was swallowed
by the bare ``except Exception``, then strict exited 2). The test
asserted the right outcome via the wrong path.

Fix:
- ``master/__main__.py``: parser-coverage check moved to immediately
  after engine creation, BEFORE backfill + loop. Strict-mode failure
  exits 2 without wasting a Forgejo round-trip + without dependent
  code paths firing.
- Test refactored: count-based assertions on backfill and loop call
  counts (0 each) instead of fail_if_called. Catches regressions
  where the strict check moves back below either.

N7 — _to_aware_datetime unit-tested in isolation:
Previously exercised only via end-to-end comment-filter test. New
``TestToAwareDatetime`` (+12 tests) covers: None, empty string, Z
suffix, +00:00 offset, microseconds preserved, naive datetime →
UTC, malformed string → None, partial string → None, unsupported
types → None, timezone abbreviations → None, equality across
Z + offset forms (the regression the helper exists to defend).

N8 — runtime=None lazy-import branch tested:
Previously all 24 HTTP factory tests injected a fake runtime; the
production path (``build_callbacks(cfg=None)`` → sys.path injection
+ lazy import of ``tools._claim_runtime``) was untested.
``TestBuildCallbacksDefaultRuntime`` (+2 tests): verifies the import
succeeds + every ForgejoCallbacks attribute is callable; verifies
the import is idempotent (second call doesn't crash on sys.path
re-insert).

N9 — resume event-row emission asserted:
Round-1's batch B added the PAUSE event-row test
(``test_label_removed_emits_event_with_reason``) but not RESUME.
``test_label_restored_emits_event_with_reason`` pins the resume's
controller_events shape (from_state=PAUSED, to_state=<prior>,
reason="opt-in-label-restored", source="reconciliation") so
operators auditing the timeline see both pause + resume.

N10 — concurrent SQLite dequeue test:
The dequeue docstring claims SQLite BEGIN-DEFERRED concurrent
dequeues "retry via busy_timeout (5s)" — but no test verified.
``TestSQLiteConcurrentDequeue::test_two_threads_racing_one_wins``
spawns two threads, both attempt dequeue simultaneously via a
threading.Barrier. Asserts: neither thread raises SQLITE_BUSY-
without-retry, exactly one acquires the attempt, the loser sees
no_pending_eligible (winner committed first).

Total: 691 controller tests pass (+31 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 16:15:01 -04:00
drew 147e3403c1 fix(controller): batch G — show-stoppers from round-2 review (N1–N4)
Four items the round-2 adversarial review flagged as ship-blockers.

N1 — PID-reuse defense is now WIRED in production:
Round 1's batch D shipped ``subprocess_starttime`` in the sidecar +
janitor checks against it, but NO production code wrote sidecars.
The defense was unwired; tests passed against a code path that
production never invoked.

Fix:
- ``worker/agent_runner.py`` accepts ``workspace_dir`` and
  ``opencode_server_url`` kwargs. When ``workspace_dir`` is set, it
  writes a sidecar (``{workspace_dir}/worker.session``) immediately
  after MCP spawn capturing the real PID + starttime from
  ``/proc/{pid}/stat`` field 22. Removes it on attempt completion.
- ``worker/__main__.py`` builds the per-attempt workspace dir
  (``{workspace_root}/pr-attempt-{N}/``) and threads it through the
  agent_runner closure with the OpenCode URL. The naming convention
  is picked up by the janitor's ``pr-*`` glob; when per-PR shared
  workspaces ship (Phase 1k++ follow-up), it changes to
  ``pr-{owner}-{repo}-{N}``.
- Test: ``TestSidecarWiring`` (+2 tests) verifies the sidecar appears
  during the attempt, carries the right PID + starttime + instance,
  and is cleaned up post-attempt.

N2 — AWAITING_CI escape event firing is now WIRED in production:
Round 1's batch D shipped ``ci_polling_exhausted`` /
``ci_flake_retries_exhausted`` in TRANSITIONS, but NO production code
emitted them. Workflows could still hang in AWAITING_CI forever.

Fix:
- New ``master/ci_poll.py``: ``run_ci_poll_exhaustion_tick`` scans
  workflows whose ``entered_state_at`` is older than
  ``CONTROLLER_AWAITING_CI_TIMEOUT_S`` (default 7200s) and fires
  ``ci_polling_exhausted`` via ``apply_event`` → STUCK + emits a
  ``ci_poll_exhausted`` controller_events row with the threshold
  payload.
- ``master/loop.py`` integrates the new tick on its own cadence
  (``ci_poll_exhaustion_interval_s`` env, default 300s). Composes
  with the existing master loop. ``MasterTickReport`` gains
  ``ci_poll_exhaustion: CIPollExhaustionReport | None``.
- Tests: ``test_master_ci_poll.py`` (+7 tests) — happy path, fresh
  workflow stays untouched, only AWAITING_CI is targeted (other
  long-lived non-terminal states ignored), event row shape pinned,
  default threshold matches the documented 2h, end-to-end loop
  integration (master_main_loop drives the exhaustion +
  workflow → STUCK without operator intervention).
- Dialect-portable SQL (Postgres interval, SQLite julianday).
- Handles SQLite returning TIMESTAMP as str from text() queries
  (no .isoformat() on str).

N3 — externally-merged/closed PRs now win over label removal:
Round-1's PAUSE-on-label-removed shipped, but reconciliation
checked the label gate BEFORE checking merged/closed. Operators
removing the opt-in label on an already-merged PR would PAUSE the
workflow forever — never transitioning to MERGED.

Fix:
- ``master/reconciliation.py:_reconcile_one`` re-ordered:
  1. Check terminal-state mappings (merged/closed) FIRST — apply
     immediately if they fire.
  2. THEN the opt-in label gate (pause/resume).
  3. Fall through to "consistent" otherwise.
- Tests: ``test_externally_merged_takes_priority_over_label_removal``
  + ``test_externally_closed_takes_priority_over_label_removal``
  pin the contract. Both seed an IMPLEMENTING workflow + Forgejo
  reporting "merged/closed AND no opt-in label" → workflow
  transitions to MERGED/ABANDONED (not PAUSED) + pre_pause_state
  stays None.

N4 — graceful handling of empty env vars:
``int(os.environ.get("CONTROLLER_FORGEJO_REQUEST_TIMEOUT_S", "30"))``
crashes with non-actionable ``int('') ValueError`` if the operator
sets the env to empty/whitespace (common when sourcing a partially-
edited /etc/cleveragents/master.env file).

Fix:
- ``master/forgejo_cfg.py:_env_int(name, default)`` — empty or
  whitespace-only values fall back to the documented default; only
  non-numeric values still raise (with a clear message naming the
  variable).
- Tests: ``test_empty_env_value_falls_back_to_default`` +
  ``test_whitespace_only_env_falls_back`` + updated
  ``test_malformed_env_raises_value_error`` to match the new
  "not a valid integer" wording.

Total: 660 controller tests pass (+13 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 16:11:27 -04:00
drew abff38a274 refactor(controller): batch F — ControllerForgejoConfig + parser strict mode (items 16, 17)
Final two items from the consolidated adversarial-review punch list.

ITEM 16 — ControllerForgejoConfig:
Previously ``master/__main__.py:build_cfg_stub`` reached into
``tools/_mcp_common.ForgejoCfg`` via sys.path injection and mutated
``cfg.owner`` / ``cfg.repo`` after construction. That inverted the
dependency direction (the controller is the new system; it shouldn't
reach into legacy pipeline modules) and tied controller deployments
to whatever schema ForgejoCfg happened to have.

Replaced with ``master/forgejo_cfg.py``: a dataclass owning exactly
the fields ``_claim_runtime`` reads (token, request_timeout_s,
api_retries, claim_ttl_seconds) plus the controller's own
(owner, repo). ``from_environment(owner, repo)`` reads the same env
vars the legacy ForgejoCfg used (FORGEJO_TOKEN, CONTROLLER_FORGEJO_*
tunables) so operators don't have to relearn anything.

``build_cfg_stub`` is now a 1-line delegate to ``from_environment``;
no sys.path mutation, no cross-package import.

ITEM 17 — parser coverage validator + strict mode:
Master startup now calls ``validate_parser_coverage()`` (already
exposed by Phase 1j's parser registry) and logs the gap loudly:

  WARNING CI parser coverage: 3/10 real (7 stub: ['bandit', 'build',
  'radon', 'robot_framework', 'semgrep', 'slipcover', 'vulture']).
  Stub-parser gates fall back to raw_log_excerpt; implementers see
  the log but not structured findings.

Operators who want the strict plan-v10 "STUCK on unknown tool"
behaviour set ``CONTROLLER_STRICT_PARSER_COVERAGE=1``; master then
refuses to start (exit 2) until every EXPECTED_PARSER has a real
implementation. Default off: ship-stubs is the migration-friendly
path; strict is the after-everything-is-implemented gate.

Tests:
- test_forgejo_cfg.py (+13 tests): dataclass shape, env resolution
  (defaults, FORGEJO_TOKEN preferred over CONTROLLER_FORGEJO_TOKEN,
  malformed env raises), no-sys-path-injection AST audit + returns
  the typed dataclass.
- test_entry_points.py (+2 tests): strict-mode blocks startup +
  default mode only warns.

Total: 647 controller tests pass (+12 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:40:45 -04:00
drew d71046b9a0 fix(controller): batch D — PID-reuse, AWAITING_CI escape, flake bound, scheduler skip
Four safety items from the consolidated adversarial-review punch list.

ITEM 8 — PID-reuse hazard in janitor:
The janitor SIGKILL'd whatever process happened to live at the
sidecar's recorded subprocess_pid. Between sidecar write and janitor
sweep, the OS can reuse the PID for an unrelated process; the janitor
was killing innocents under fork-heavy workloads.

Fix:
- ``session_sidecar.py``: added ``subprocess_starttime`` field
  (Optional[int]) + ``read_proc_starttime(pid)`` helper that reads
  ``/proc/{pid}/stat`` field 22 (clock ticks since boot — monotonic
  for a (boot, pid) pair).
- ``WorkerSession.from_dict`` filters unknown keys so forward + back
  compat with sidecars from earlier/later versions is preserved.
- ``janitor._pid_alive`` and ``_kill_with_grace`` accept
  ``expected_starttime``; on mismatch they short-circuit and DON'T
  signal the impostor.
- ``_kill_with_grace`` return semantics tightened: True iff a signal
  was actually delivered (False for "PID gone" / "PID reused"). The
  ``sessions_killed`` counter now reflects real kills.

ITEM 9 — AWAITING_CI escape from infinite poll:
Previously AWAITING_CI could only exit via ``ci_green`` /
``ci_red_*`` / ``ci_flake_retry`` — if CI hangs forever (runner
outage, broken integration, etc.) the workflow had no controller-
driven STUCK path; only operator_unstick could rescue it.

Fix: new ``ci_polling_exhausted`` event → STUCK. The master's
AWAITING_CI poll handler is the natural place to emit it once a
threshold passes (deferred to a follow-up — Phase 1k+ ships the
event in the table; the timer fires it).

ITEM 10 — ci_flake_retry was unbounded:
The ``ci_flake_retry`` self-loop on AWAITING_CI had no encoded
ceiling. Pathological flaky CI could loop forever (the docstring
said "retry once per gate" but nothing enforced it).

Fix:
- New ``workflows.ci_flake_retries_remaining`` column (server_default
  '1', default 1 — operators tune via ``CONTROLLER_CI_FLAKE_RETRIES``
  at startup or via direct UPDATE).
- New ``ci_flake_retries_exhausted`` event → ESCALATING. Master
  decrements the column on each ci_flake_retry; at 0 the next CI
  failure routes through ci_red_* (regular path) or this new
  event (escalates if the operator wants a hard ceiling).

ITEM 11 — scheduler now skips PAUSED workflows:
Without this, the scheduler could enqueue a fresh attempt for a
PAUSED workflow between two reconciliation ticks (race: label
removed at T+0, reconciliation runs at T+300, scheduler ticks at
T+30 with stale DB state). The window is at most one attempt of
worker work.

Fix: ``schedule_next_attempts`` SQL now lists only
{ANALYZING, IMPLEMENTING, REVIEWING, CONFLICT_RESOLVING, ESCALATING}
explicitly; PAUSED is excluded by absence. Reconciliation owns the
PAUSED → resume transition; scheduler doesn't touch it.

Schema additions:
- ``workflows.ci_flake_retries_remaining`` (INTEGER NOT NULL DEFAULT 1)
- ``workflows.awaiting_ci_started_at`` (TIMESTAMP NULL) — for the
  poll-exhaustion timer (timer impl deferred; column is staged).

Tests:
- TestReadProcStarttime — 3 tests (Linux skip-guard) for the
  /proc/pid/stat parser (self-pid > 0, missing pid is None,
  invalid pid is None).
- TestJanitor::test_pid_reuse_defended_via_starttime — pins the
  contract end-to-end (real subprocess + fabricated wrong starttime
  → janitor doesn't signal).
- TestPhase1kPlusTransitions — 5 tests pinning the new events +
  proving the load-bearing invariants still pass.
- test_scheduler_skips_paused_workflows — pins item 11.

Total: 603 controller tests pass (+10 net), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:30:45 -04:00
drew a1c6646a64 fix(controller): batch C — placeholder patching + pickup_count semantics
Items 3 + 4 from the consolidated adversarial-review punch list.

ITEM 3 — placeholders no longer poison the audit trail:
The prefetch (master/prefetch.py) writes input_payload with
``attempt_id=0`` and ``attempt_number=1`` as placeholders because
the autoincrement PK isn't known until after INSERT. Previously
those values stayed in the DB forever — post-mortem queries against
``workflow_attempts.input_payload`` would show ``attempt_id=0``
and operators would chase ghosts.

Fix: ``master/scheduler.py:_insert_pending_attempt`` now patches both
fields with their real values:
- attempt_number: patched BEFORE the INSERT (we compute it as MAX+1).
- attempt_id: patched via a follow-up UPDATE after INSERT (we need
  the autoincrement first). One extra UPDATE per attempt; cheap
  compared to forever-incorrect audit trail.

Test: ``test_scheduler_patches_attempt_id_and_number_into_payload``
asserts the stored payload carries the real values, not the
placeholders.

ITEM 4 — pickup_count tracks REAPS, not dequeues:
Previously the dequeue path bumped ``pickup_count = pickup_count + 1``
on every successful pickup. With ``MAX_PICKUPS=3`` (default), 3
crashed-mid-attempt workers would STUCK the workflow — but that's
the wrong semantic. A worker that successfully picks an attempt
and runs it should NOT burn a pickup. Only failures (stale-heartbeat
reset by the reaper) should count toward the exhaustion limit.

Fix:
- ``db/dequeue.py`` (both postgres + sqlite paths): removed the
  ``pickup_count = pickup_count + 1`` UPDATE. Dequeue is a healthy
  pickup; doesn't bump.
- ``reaper.py``: added ``pickup_count = pickup_count + 1`` to the
  reset UPDATE. Each reap = one failed pickup.
- Docstrings updated to reflect the new semantics in both files.

Tests:
- Updated existing assertions in ``test_db_dequeue.py`` and
  ``test_reaper_and_pickup_guard.py`` to reflect: dequeue keeps
  pickup_count; reaper bumps it.
- ``TestPickupCountSemantics``: 2 new tests pin the contract end-to-end
  — N healthy dequeues stay at 0; alternating dequeue→reap→dequeue
  walks pickup_count up by 1 per reap.

Impact: a worker pool that crashes 3 times mid-attempt now needs
3 REAPS (not 3 dequeues) to STUCK the workflow. With default
TTL=600s + reaper_interval=60s, that's 30+ minutes of repeated
mid-attempt failure before STUCK — appropriately conservative.

Total: 593 controller tests pass (+3 new), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:24:42 -04:00
drew feeec3e9c7 fix(controller): batch B — PAUSED state for label-gate pause/resume (item 2)
Adversarial review flagged: removing the opt-in label transitions a
live workflow to ABANDONED — but ABANDONED is TERMINAL with only
``operator_unstick`` re-entry → DISCOVERED, losing all prior
controller_events continuity. Operators removing the label to "pause"
a long-running PR will be surprised it restarted from scratch.

Fix: introduce a non-terminal ``PAUSED`` state.

State machine changes (``tools/controller/state_machine.py``):
- ``PAUSED`` added to KNOWN_STATES (non-terminal — has exits via
  ``opt_in_label_restored`` and ``operator_unstick``).
- ``opt_in_label_removed`` / ``opt_in_label_restored`` events
  documented in EVENTS but NOT listed per-state in TRANSITIONS —
  they're out-of-band master-driven events written directly by
  reconciliation. Listing them per-state breaks the
  per-state-event-set invariants (ESCALATING / CONFLICT_RESOLVING).
- ``(PAUSED, operator_unstick) → DISCOVERED`` for the escape hatch.

Schema change (``tools/controller/db/models.py``):
- ``workflows.pre_pause_state: Mapped[str | None]`` column captures
  the resume target. Master writes it on pause; clears it on resume.

Reconciliation logic (``tools/controller/master/reconciliation.py``):
- New ``_apply_transition_with_pre_pause`` helper writes both
  ``current_state`` and ``pre_pause_state`` atomically + emits the
  ``reconciliation`` event row.
- On label removal (current != PAUSED): captures pre_pause_state,
  transitions to PAUSED.
- On label restoration (current == PAUSED): reads pre_pause_state
  (fallback DISCOVERED for legacy NULL data), transitions back,
  clears pre_pause_state.
- PAUSED workflows are now SCANNED by reconciliation (not just
  non-terminals) so we can detect label-restored.

Behaviour matrix:
| current | label    | result                                   |
|---------|----------|------------------------------------------|
| any !=  | absent   | → PAUSED, pre_pause_state = current     |
| PAUSED  | present  | → pre_pause_state (or DISCOVERED)       |
| PAUSED  | absent   | stays PAUSED (no transition)            |
| any !=  | present  | regular state checks (no-op for label)  |

Tests (test_label_gate.py refactor + 3 new tests):
- test_label_removed_pauses_workflow (was: abandons)
- test_label_restored_resumes_from_pre_pause_state (new)
- test_paused_workflow_without_label_stays_paused (new)
- test_resume_fallback_when_pre_pause_state_missing (new — legacy
  data without the new column)
- Existing event-reason test still passes (reason string unchanged).

Total: 590 controller tests pass, 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:20:38 -04:00
drew 6ba1926c52 fix(controller): batch A — datetime json, Forgejo state map, ISO compare, dequeue docs
Four narrow bug fixes flagged by adversarial code review (items 1, 5,
6, 7 from the consolidated critique).

ITEM 1 — datetime → json.dumps crash (silent write-after-work failure):
- ``worker/runner.py:274`` and ``master/scheduler.py:283`` now pass
  ``default=str`` to ``json.dumps`` so nested datetime fields (e.g.
  CISummary.observed_at) serialize without raising.
- Before this fix: a worker would do its real work, then crash on
  the terminal-state UPDATE with TypeError, get recorded as
  ``worker-internal-error``, and the output payload would be lost.
- Test: TestDatetimeSerializationSafety in test_master_ci_summarize
  + test_scheduler_handles_datetime_in_input_payload in
  test_master_prefetch (both pin the regression — the with-default
  test passes, the without-default test asserts the TypeError so
  future maintainers see the failure mode).

ITEM 5 — Forgejo state mapping completeness:
- Extended ``_FORGEJO_STATE_TO_GATE_STATUS`` in
  ``master/ci_summarize.py`` to cover ``cancelled``, ``timed_out``,
  ``action_required``, ``queued``, ``in_progress``, ``neutral``,
  ``skipped``, ``stale`` — states observed across Forgejo / Gitea /
  GH-mirror that previously collapsed to ``pending``, telling the
  implementer "CI is still running" when really a job was cancelled.
- ``cancelled`` / ``timed_out`` / ``action_required`` / ``stale``
  now map to ``error`` (the gate failed).
- ``queued`` / ``in_progress`` stay ``pending`` (still running).
- ``neutral`` / ``skipped`` → ``passed``/``skipped`` (informational).
- Test: TestExtendedForgejoStates — 6 tests covering each new state.

ITEM 6 — lexicographic ISO comparison drops/dupes comments:
- ``master/prefetch.py:_comment_bodies_since`` and ``_iso`` replaced
  with ``_to_aware_datetime`` + datetime comparison. Forgejo emits
  ``2026-05-18T12:00:00Z``; Python's ``datetime.isoformat()`` emits
  ``2026-05-18T12:00:00+00:00`` — a string compare gives 'Z' (0x5A)
  vs '+' (0x2B) which silently misorders timestamps.
- Now parses via ``datetime.fromisoformat`` (with Z → +00:00
  rewrite), defaults naive timestamps to UTC, and compares as
  ``datetime``.
- Test: test_comments_filter_handles_z_suffix_vs_offset_form pins
  the regression.

ITEM 7 — false BEGIN IMMEDIATE claim in dequeue docstring:
- The dequeue docstring claimed ``BEGIN IMMEDIATE`` was applied by
  session_scope; it wasn't. Attempted a global ``begin``-event
  listener that conflicted with StaticPool's shared-connection
  model (test_prefetch_callback_works_in_scheduler broke).
- Reverted to a documentation fix: SQLite stays on default
  BEGIN DEFERRED (the SQLITE_BUSY retry via busy_timeout=5s is
  acceptable for single-host dev) + the docs make MULTI-MACHINE
  REQUIRES POSTGRES explicit at three call sites
  (db/session.py, db/dequeue.py, RUNBOOK.md was already updated
  in Phase 1l). Postgres has FOR UPDATE SKIP LOCKED which is what
  production actually uses.

Tests: 586 controller tests pass (+10 new), 0 regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 15:17:38 -04:00
drew 72c272504c feat(controller): Phase 1k — controller-managed opt-in label gate
Migration safety mechanism: the controller only manages PRs and
issues carrying a configurable opt-in label (default
``controller-managed``). Operators opt PRs in for parallel-run
trials, can pause management mid-flight by removing the label, and
gradually roll out without exposing the controller to PRs that
human reviewers are actively driving.

Components:
- tools/controller/master/label_gate.py — single source of truth for
  the configured label name + pure predicates/filters over Forgejo
  PR/issue dicts.
  - ``opt_in_label_name()`` reads ``CONTROLLER_OPT_IN_LABEL`` env
    (default 'controller-managed'); empty/whitespace falls back.
  - ``has_opt_in_label(entity, name)`` defensively handles every
    degenerate shape (non-dict entity, non-list labels, non-dict
    label entries, missing name field).
  - ``filter_by_opt_in_label`` / ``count_filtered`` for callers.

Wired through:
- discovery.run_discovery + backfill.run_startup_backfill +
  reconciliation.run_reconciliation_tick each accept
  ``opt_in_label`` and ``require_opt_in_label`` kwargs.
- Function defaults are ``require_opt_in_label=False`` for API
  back-compat (existing 30+ discovery/backfill/recon tests work
  without changes).
- __main__.py defaults to ``--no-opt-in-label`` OFF (gate ENABLED in
  production); add ``--no-opt-in-label`` to bypass.
- DiscoveryReport gains a ``label_filtered_out`` counter.

Reconciliation behavior:
- When opt_in_label is configured AND the Forgejo response carries a
  ``labels`` field AND the opt-in label is NOT present, the workflow
  transitions to ABANDONED with reason ``opt-in-label-removed`` +
  emits a controller_events 'reconciliation' row.
- Partial Forgejo responses (no ``labels`` field) skip the label
  check — never ABANDON on incomplete data.

Master loop extension:
- ``reconciliation_args`` now accepts an optional 5th element — a
  kwargs dict threaded through to ``run_reconciliation_tick``.
  __main__.py uses this to pass ``require_opt_in_label`` per the CLI
  flag. 4-tuple back-compat preserved.

Tests (+29 in test_label_gate.py, 0 regressions across 569 tests):
- Predicate edge cases (every degenerate shape returns False)
- Env-var resolution (default, override, empty, whitespace)
- filter/count helpers
- Discovery + backfill: kept/filtered counts, gate disabled,
  explicit label overrides env
- Reconciliation: label removed → ABANDONED, label present →
  no-op, partial response → no-op, gate disabled → bypass, event
  row records reason

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 14:45:22 -04:00
drew 3e269ce011 feat(controller): Phase 1j — deterministic CI summarizer + priority parsers
Replaces "ci_summary=None / failing_gates=[]" placeholders from Phase
1h with a real summarizer that maps Forgejo combined-status →
CISummary V1 dict by running per-tool deterministic parsers on each
failing gate's log.

Priority parsers shipped (cover lint/format/typecheck/unit_tests, the
4 most-failed gates):
- ruff      — F+E codes from `nox -s lint`; Would-reformat lines from
              `nox -s format`. Aggregates to single error_class when
              all findings share one code, else RuffMixed.
- pyright   — error/warning/information diagnostics; rule name pulled
              from trailing `(reportName)` parens. Abs-path
              normalization strips container prefixes.
- behave    — failing scenarios (file:line + name), AssertionError
              extraction. Feature/scenario summary line aggregation.

Stub parsers for not-yet-shipped tools (robot_framework, slipcover,
bandit, semgrep, vulture, radon, build): return a structured
CIFailure with error_class="parser-pending-{name}" + the raw log
excerpt. Operators see the failure; implementer still has log
context. Phase 1j+ replaces stubs with real parsers without changing
the gate-→-session map.

Components:
- _base.py        — ParserResult dataclass + select_log_excerpt()
                    (tail-N-lines smart selection within 16KB cap)
- _stub.py        — make_stub(name) factory for pending tools
- _registry.py    — resolve(parser_name) + resolve_for_nox_session()
                    + validate_parser_coverage()
- master/ci_summarize.py — summarize_ci_status(head_sha, status,
                    log_fetcher) orchestrator. Handles:
                    - composite multi: gates → CIFailure.composite_findings
                    - log_fetcher returning None → log-fetch-failed
                    - log_fetcher raising → caught + log-fetch-failed
                    - Unknown gate context → NoParserAvailable
                    - Forgejo state=None → unknown summary
                    - Parameterized matrix gates ("unit_tests-3.13")
                      → base session name resolution

Tests (+46 across 2 new files, 0 regressions):
- Per-parser canonical + empty + garbage input
- Registry resolution (real vs stub), coverage validator
- Summarizer V1 contract round-trip
- Composite security_scan composite_findings shape
- Error paths (None status, raising fetcher, unknown gate)
- Parser version aggregation across mixed gates

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 14:40:16 -04:00
drew dfdfbf762b feat(controller): Phase 1h — prefetch callbacks (V1 input assembly)
Per plan v9, the master assembles the worker input_payload at
attempt-enqueue time so the worker dequeues a ready-to-use payload
with no extra Forgejo I/O of its own.

This phase ships per-role V1-input builders + a factory matching the
scheduler's PrefetchCallback protocol:

- build_implementer_input  → ImplementerInputV1 shape (head_sha,
  head_ref, base_branch, active_reviews, pr_comments_since_last_attempt,
  prior_attempts, diff_summary)
- build_reviewer_input     → ReviewerInputV1 shape (full_diff,
  prior_implementer_attempts, implementer_claim, prior_reviews)
- build_estimator_input    → EstimatorInputV1 shape (pr_title, pr_body,
  diff_summary) — works for both PR and issue kinds
- build_conflict_resolver_input → ConflictResolverInputV1 shape with
  conflicted_files=[] stub (worker patches via git rebase)
- make_prefetch_callback(engine, callbacks) → routes by role; returns
  (payload, "V1") matching the scheduler's PrefetchCallback signature

Forgejo HTTP wiring adds four new callbacks (get_pr_details,
get_pr_diff, list_pr_reviews, list_pr_comments) plumbed through
ForgejoCallbacks.

Worker-side patches (post-dequeue, pre-validation):
- attempt_id, attempt_number (known from dequeue)
- workspace_dir (worker filesystem path)
- wallclock_budget_s (worker config)

What this phase DOES NOT yet produce:
- ci_summary / failing_gates — Phase 1j (deterministic CI summarizer)
- Issue-kind estimator's title/body — needs list_issue_details
  callback (defer to future phase)
- conflict_resolver's actual conflicted_files — needs worker-side git
  rebase + conflict-parse pass

Tests (+26 in test_master_prefetch.py, 0 regressions):
- Per-role shape validation + V1 contract parse after worker patches
- Prior-attempts merge (verbatim cap=3, oldest-first, total count)
- Active-reviews projection (filters invalid states/missing user)
- pr_comments_since_last_attempt filtering by finished_at
- Factory routes by role; unknown role raises
- Scheduler integration end-to-end (real prefetch → real INSERT)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 14:33:07 -04:00
drew b76c5f05e7 feat(controller): wire reconciliation into master main loop
Composes reconciliation as the 4th tick layer at its own cadence.

tools/controller/master/loop.py:

- MasterConfig gains reconciliation_interval_s (default 300s per
  plan v9).
- MasterTickReport gains reconciliation: ReconciliationReport | None.
- master_main_loop gains reconciliation_args parameter — tuple of
  (owner, repo, get_pr_state_cb, get_issue_state_cb). When provided,
  runs run_reconciliation_tick every reconciliation_interval_s.
  When None, reconciliation is disabled (useful for tests + one-shot
  modes).
- Reconciliation exception is caught + logged; master keeps running.
- Iteration log line now includes reconciled=N.

tools/controller/master/__main__.py:

- Passes reconciliation_args from ForgejoCallbacks (built earlier
  in the entry point) so the production master automatically runs
  reconciliation against the configured (owner, repo).

3 new tests in test_master_loop.py:
- reconciliation_fires_when_configured: workflow with externally-
  merged state → reconciliation transitions to MERGED.
- reconciliation_skipped_when_args_none: workflows untouched +
  no reconciliation reports.
- reconciliation_exception_doesnt_break_loop: per-row fetch failures
  don't crash the master.

Total: 433 controller tests; full auto_agents suite 2795 pass.
2026-05-18 14:23:42 -04:00
drew 7d1dfb8635 feat(controller): Phase 1g — periodic reconciliation tick
Catches externally-merged / externally-closed PRs that the controller
didn't directly merge (operator clicked the merge button in Forgejo's
UI; collaborator closed a PR while controller was waiting). Inverse
of discovery: discovery ADDS new entities; reconciliation re-syncs
KNOWN ones.

tools/controller/master/reconciliation.py:

- run_reconciliation_tick(engine, owner, repo, get_pr_state,
  get_issue_state=None):
  - Scans all non-terminal workflows for (owner, repo).
  - Calls get_pr_state / get_issue_state per workflow.
  - Decision table:
    - PR merged=True → MERGED (reason='externally-merged')
    - PR state=closed not merged → ABANDONED ('externally-closed-
      not-merged')
    - PR state=open → consistent (no transition)
    - PR not found (404) → STUCK ('pr-not-found-on-forgejo')
    - Issue state=closed → ABANDONED ('issue-closed-externally')
  - Per-row failures isolated: one Forgejo flake doesn't kill the
    whole sweep. Failed fetches recorded as ReconciliationAction
    with reason='fetch-failed' (workflow untouched).

- Pure decision function (_decide_transition) separated from SQL
  writes (_apply_transition) for testability.

tools/controller/master/forgejo_http.py:

- Added get_pr_state + get_issue_state callbacks to ForgejoCallbacks.
- HTTP shape: 404 → None (workflow → STUCK); non-200/404 → raise
  (workflow recorded as fetch-failed, not silently STUCK'd).

15 new tests in test_master_reconciliation.py:
- basics (empty DB, terminal workflows skipped, other-repo skipped)
- PR state mappings (open=consistent, merged → MERGED, closed →
  ABANDONED, not-found → STUCK)
- issue state mappings (open=consistent, closed → ABANDONED,
  no-callback → consistent)
- fetch failures (raised exception → fetch-failed, workflow untouched)
- event rows (reconciliation event emitted on transition; none on
  consistent)
- ReconciliationAction dataclass shape

Total: 430 controller tests; full auto_agents suite 2792 pass.
2026-05-18 14:20:10 -04:00
drew ba2e9472bc feat(controller): Phase 1f — master startup backfill
When the master starts (first deploy or after a long outage), it
needs to learn about existing open PRs/issues that weren't created
via discovery-tick-during-uptime. Backfill = discovery + a one-time
marker so subsequent restarts know "this isn't the first time."

tools/controller/master/backfill.py:

- run_startup_backfill(engine, owner, repo, list_prs, list_issues):
  - Calls run_discovery (already idempotent — skips existing entities)
  - Records a 'controller-backfill-complete' marker in controller_events
    associated with the first new workflow OR an existing workflow OR
    skipped if Forgejo is truly empty (no FK target)
  - Returns BackfillReport{first_time, discovery}.

- has_backfill_run(engine, owner, repo): existence-check on the marker
  by parsing controller_events.payload. Multi-tenant isolated — a
  marker for (owner_a, repo_a) doesn't satisfy a check for
  (owner_b, repo_b).

- The marker is informational; the dedup is provided by discovery's
  unique-constraint skip. The marker exists so operators can answer
  "has backfill ever run for this repo?" in one SQL query.

Wired into master __main__:
- Runs AFTER engine/create_all + Forgejo callback wiring, BEFORE
  master_main_loop.
- Try/except wrapped so Forgejo flake at startup doesn't prevent
  the main loop from running (discovery tick will retry).
- New --skip-backfill flag for tests + warm restarts.

8 new tests in test_master_backfill.py:
- has_backfill_run: no marker → False; multi-tenant isolation
  (different owner OR different repo → False).
- First-time backfill creates workflows + marker; empty Forgejo
  skips marker (no FK target).
- Second run reports first_time=False; picks up newly-appeared PRs
  + emits a second marker.
- Multi-tenant (owner-a, repo-a) and (owner-b, repo-b) both get
  their own marker.
- New workflows are in DISCOVERED state.

Total: 415 controller tests; full auto_agents suite 2777 pass.
2026-05-18 14:16:23 -04:00
drew 7b1b68b784 feat(controller): Phase 1e — master + worker __main__ entry points
Runnable as `python -m tools.controller.{master,worker}` for systemd
deployment. Wires the production callbacks (Forgejo HTTP +
OpenCode session adapter + agent_runner) into the previously-shipped
main loops.

tools/controller/master/__main__.py:

- Args: --owner, --repo (required); --opencode-url, --log-level,
  --tick-interval, --discovery-only-once (smoke flag).
- Reads CLEVERAGENTS_DB_URL (returns 2 if unset — operator-actionable
  error visible in systemd logs).
- build_cfg_stub reuses _mcp_common.ForgejoCfg → PAT rotation +
  Forgejo URL env vars work without a new config module.
- SIGTERM/SIGINT → stop_event → master_main_loop drains + exits.
- --discovery-only-once: runs one discovery sweep + exits. Useful for
  initial backfill or smoke testing the Forgejo callback wiring
  without committing to the long-running loop.

tools/controller/worker/__main__.py:

- Args: --opencode-url, --roles (default all 5; comma-separated),
  --max-concurrent, --poll-interval, --log-level, --no-startup-janitor.
- Empty --roles after parsing → exit 2.
- Startup sequence:
  1. Run orphan-workspace sweep (unless --no-startup-janitor)
  2. Build OpenCode session adapter via wire_opencode_session
  3. Wrap production_agent_runner with the session injected
  4. Wire SIGTERM/SIGINT → stop_event
  5. Enter worker_main_loop

9 new tests in test_entry_points.py:
- Arg parsing: master --owner/--repo required; both DB-URL-missing
  returns 2; worker empty-roles returns 2.
- --help works via subprocess for both (no SystemExit issue).
- Master --discovery-only-once smoke: stubs build_callbacks + cfg
  builder; verifies exit 0 + the discovery call chain runs.
- Worker --no-startup-janitor: stubs sweep + loop; verifies sweep
  is skipped when flag set, runs when not.

Total: 407 controller tests; full auto_agents suite 2769 pass.

Controller v1 is now end-to-end deployable as two systemd units
(one per machine for each instance). The remaining work for a real
production rollout is:
- Real prefetch callbacks for the scheduler (assembling V1 input
  payloads from Forgejo data)
- Backfill at master startup (existing-PRs → DISCOVERED rows)
- Reconciliation tick (DB ↔ Forgejo sync)
- Per-role prompt templates (the V1 prompt stub works; specialised
  per-role prompts will land as agents are migrated)
2026-05-18 14:12:56 -04:00
drew ae940f4564 feat(controller): Phase 1d-3c-4 — HTTP adapter wiring callbacks to _claim_runtime
The production glue between the controller's callback protocols
(discovery / forgejo_writes / merging) and the existing Forgejo HTTP
client in _claim_runtime. Single build_callbacks(cfg) factory; tests
use the same module with a fake runtime stub.

tools/controller/master/forgejo_http.py:

- ForgejoCallbacks dataclass bundling every callback the controller
  needs: list_prs, list_issues, list_comments, post_comment,
  get_labels, add_label, remove_label, merge_pr.

- build_callbacks(cfg, runtime=None): wires each callback as a thin
  closure over runtime.get/post/delete. Production omits ``runtime``
  to use the real _claim_runtime module; tests inject a fake.

- Forgejo path conventions match the API:
    GET /repos/{o}/{r}/pulls?state=open
    GET /repos/{o}/{r}/issues?state=open&type=issues  (excludes PRs)
    GET/POST /repos/{o}/{r}/issues/{n}/comments
    GET/POST /repos/{o}/{r}/issues/{n}/labels
    DELETE   /repos/{o}/{r}/issues/{n}/labels/{name}  (URL-encoded)
    POST     /repos/{o}/{r}/pulls/{n}/merge  (body: {"Do": "merge"})

- Robust response handling:
  - list endpoints: non-200 → empty list; non-list body → empty;
    non-dict items filtered out.
  - post_comment: 200/201 ok; other → RuntimeError.
  - add_label / remove_label: 200/201/204 → True; remove-404 → True
    (label already gone = goal achieved); else False.
  - merge_pr: returns normalized MergeResponse. On 404, fetches the
    PR's actual state (merged=True → pr_state='merged'; state='closed'
    → 'closed'; PR fetch failure or non-200 → leave pr_state=None so
    the merging handler defaults to ABANDONED conservatively).
  - Any callback exception → synthetic 503 so the merging handler's
    retry logic kicks in cleanly.

23 new tests in test_master_forgejo_http.py:
- list_prs (path format, non-200 → empty, non-list body → empty,
  filters non-dict items)
- list_issues (type=issues filter)
- list_comments (path format)
- post_comment (201, 200, non-2xx raises, non-dict body)
- labels (get / add 201/500 / remove 204/404/URL-encoded)
- merge (200, 409, 500, callback-raises-as-503, 404+merged,
  404+closed, 404+pull-fetch-failure)

Plus a bug fix surfaced by the test_404_with_pull_fetch_failure test:
the PR-state-fetch branch was returning 'open' on a 500 response;
now correctly checks status==200 before inspecting the body.

Total: 366 controller tests; full auto_agents suite 2728 pass.
2026-05-18 14:00:11 -04:00
drew 94821d2702 feat(controller): Phase 1d-3c-3 — MERGING handler (6-response-shape table)
The master's per-tick handler for workflows in MERGING. Per plan v6+v9
non-blocking retry: state stays MERGING across ticks, retry counter +
backoff tracked in workflows.merging_retry_count and
merging_retry_next_attempt_at.

tools/controller/master/merging.py:

- MergeResponse: normalized {status_code, pr_state, error_message}
  the callback returns. status_code drives the 6-response table:
    200          → MERGED
    409          → IMPLEMENTING(tier_last_succeeded) with reason=
                    'post-approval-base-conflict'
    403          → STUCK ('branch-protection')
    422          → AWAITING_CI ('ci-required-status-missing')
    5xx (any)    → stay MERGING; bump retry_count + schedule
                    next_attempt_at = now + 2^retry_count seconds
                    (capped at 60s); STUCK at retry_count >= 5
    404 + pr_state='merged'  → MERGED ('externally-merged')
    404 + pr_state='closed'  → ABANDONED ('externally-closed')
    404 + no pr_state        → ABANDONED (conservative default)

- run_merging_tick(engine, merge, owner, repo) — sweeps workflows
  WHERE current_state='MERGING' AND kind='pr' AND owner+repo match
  AND (next_attempt_at IS NULL OR next_attempt_at <= now). Per row:
  call merge → map response → transition (or schedule retry) + emit
  controller_events row.

- Callback failure (callback raises) wrapped as a synthetic 503 so
  the retry logic kicks in cleanly.

- Counter resets on non-retry responses (409 / 422) — keeps backoff
  fresh for future retries.

- MAX_MERGE_RETRIES = 5; MAX_BACKOFF_S = 60.0.

14 new tests across 7 classes:
- 200 happy path + event row
- 409 → IMPLEMENTING(tier_last_succeeded)
- 403 → STUCK
- 422 → AWAITING_CI + retry counter reset
- 404 paths (merged + closed + no-state default)
- 5xx retry (counter bumped + backoff scheduled +
  in-window-skipped + max-retries-STUCK + callback-raises-as-5xx)
- Owner/repo filter (other repo's MERGING untouched)
- kind='pr' filter (issues never processed even if mis-seeded)

Total: 343 controller tests; full auto_agents suite 2705 pass.
2026-05-18 13:56:26 -04:00
drew 3e3796d918 feat(controller): Phase 1d-3c-2 — Forgejo write helpers (status + labels)
Idempotent status comment posting via fingerprint markers + per-label
no-op-aware adjust. Production wires the callbacks to existing
_status_comments + _claim_runtime helpers; tests use fakes.

tools/controller/master/forgejo_writes.py:

- compute_fingerprint(workflow_id, event_kind, content_key) →
  16-char SHA256 prefix. Deterministic; same inputs → same fingerprint.
  Different (workflow_id OR event_kind OR content_key) → different fp.

- build_marker(fp) → HTML comment "<!-- controller:fingerprint:abc -->".
  Searchable + invisible in Forgejo's rendered UI.

- comment_has_fingerprint(body, fp) → bool. Used by post_status_comment
  for the dedup check before posting.

- post_status_comment(): full idempotency protocol —
  1. compute fingerprint
  2. list_comments callback → check existing for marker
  3. if found → return duplicate (no post)
  4. else post via post_comment callback
  Failure modes:
  - list_comments raises → fall through to post (conservative;
    fingerprint match on next sweep catches the duplicate)
  - post_comment raises → return failed; caller retries

- adjust_labels(add, remove): one-shot get_labels + per-label
  add/remove. add-when-already-present → no-op; remove-when-absent
  → no-op. Per-label failure isolated. get_labels failure marks
  all requested actions failed (caller retries).

19 new tests:
- fingerprint helpers (5: deterministic, distinct inputs, marker
  format, body match, empty body no-match)
- post_status_comment (5: new post, duplicate skip, list-failure
  fall-through, post failure, distinct event_kinds get distinct fps)
- adjust_labels (9: empty no-op, add-when-absent, add no-op,
  remove-when-present, remove no-op, get failure, per-label
  isolation with mixed failures, returning-False failure)

Total: 329 controller tests; full auto_agents suite 2691 pass.

Phase 1d-3c-3 (MERGING handler) + Phase 1c-3 (real OpenCode + MCP
spawn) remain.
2026-05-18 13:53:09 -04:00
drew e8045d4b9b feat(controller): Phase 1d-3c-1 — discovery sweep + worker test deflake
Discovery polls Forgejo for open PRs/issues and inserts a fresh
DISCOVERED workflow for any entity the controller hasn't seen yet.
Same callback-injection pattern as the scheduler's prefetch: tests
provide synthetic ListPRsCallback / ListIssuesCallback; production
wires them to the existing _review_fetch helpers in a follow-up.

tools/controller/master/discovery.py:

- run_discovery(engine, owner, repo, list_prs, list_issues=None) —
  idempotent. Skips existing entities via a one-shot
  SELECT(kind, entity_number) WHERE owner+repo membership check.
  Inserts a DISCOVERED workflow + a controller_events 'discovered'
  row per new entity. Per-callback try/except: a PR-list failure
  doesn't block issue discovery, and vice versa.
- Defensive coerce_int: rejects bools (True is an int subclass —
  bug class to avoid). Accepts string digits.
- Returns DiscoveryReport(prs_seen, issues_seen, new_workflows,
  existing_skipped, new_entities[]).

12 new tests:
- basics (empty, PRs only, issues only, both PRs+issues)
- idempotency (2nd sweep skips; PR #42 + issue #42 coexist;
  other (owner, repo) isolated)
- malformed input (non-int + bool skipped; string digit accepted)
- callback failures (PR raises → still process issues; issues
  callback optional)
- event row creation per new workflow

Plus a worker test deflake:
test_stolen_lock_between_agent_return_and_write was asserting the
specific 'lost-lock-at-write' outcome, but under full-suite CPU
contention the heartbeat thread can fire between the agent's return
and _write_outcome — taking the 'lost-lock' (via lost_lock_event)
path instead. Both are valid for this scenario; loosened the
assertion to accept either.

Total: 310 controller tests; full auto_agents suite 2672 pass.
2026-05-18 13:49:30 -04:00
drew 337af855f3 feat(controller): Phase 1d-3b — per-workflow scheduler + static escalation
When the state machine transitions a workflow into a state that
needs a worker (ANALYZING / IMPLEMENTING / REVIEWING /
CONFLICT_RESOLVING), the scheduler enqueues a fresh pending
workflow_attempts row with the input_payload prefetched.

tools/controller/master/scheduler.py:

- schedule_next_attempts(engine, prefetch=...) — finds workflows
  in worker-needing states without a matching pending/in_progress
  attempt; enqueues one fresh attempt per. Skips:
  - workflows already with pending/in_progress attempt for the role
  - ESCALATING at MAX_TIER → resolved to ABANDONED (no enqueue)
  - prefetch raised → per-workflow isolated failure

- Static escalation policy resolved inline: ESCALATING + current_tier
  → IMPLEMENTING(min(tier+1, MAX_TIER)) OR ABANDONED. Both produce
  controller_events transition rows with the appropriate v9 event
  name (escalate_next_tier_available / escalate_max_tier_exhausted).

- State→role table: ANALYZING → estimator, IMPLEMENTING → implementer,
  REVIEWING → reviewer, CONFLICT_RESOLVING → conflict_resolver.

- PrefetchCallback is parameterized; production wires it to the
  Forgejo prefetch (Phase 1d-3c), tests inject a fake.

- attempt_number monotonically increments via COALESCE(MAX, 0) + 1.

- DEFERRED to Phase 1d-3c: actual Forgejo prefetch implementation
  (this commit ships the scheduler skeleton + the prefetch callback
  contract).

20 new tests:
- state→role parametrization (4 states × matching role)
- no-double-enqueue (3 cases: pending blocks, in_progress blocks,
  different-role doesn't block)
- escalation (4 cases: tier 0→1, 1→2, MAX→ABANDONED, event row content)
- prefetch raises → per-workflow skip
- non-schedulable states parametrized (7 cases: DISCOVERED,
  AWAITING_CI, MERGING, MERGED, ABANDONED, STUCK, CREATED_PR)
- attempt_number monotonicity (uses MAX+1)

Total: 298 controller tests; full auto_agents suite 2660 pass.
2026-05-18 13:44:38 -04:00
drew 130bbbf3c5 feat(controller): Phase 1d-3a — master main loop (composite tick)
Long-running master orchestrator composing the deterministic
per-iteration work shipped in Phase 1d-1/1d-2. One iteration =
state-machine tick + reaper + pickup guard, in that order
(reasoning in module docstring).

tools/controller/master/loop.py:

- MasterConfig: tick_interval_s (default 30s), reaper_interval_s
  (default 60s), pickup_guard_max_pickups (default 3 per v6).
- run_master_iteration(engine): single synchronous iteration —
  composable for tests + master_main_loop.
- master_main_loop(engine): runs run_master_iteration on a loop until
  stop_event fires. Reaper runs less often than tick (every
  reaper_interval_s, not every tick_interval_s). on_iteration
  callback gets each MasterTickReport (used by tests + future
  structured logging).

7 new tests:
- run_master_iteration: empty DB no-op; tick advances state
  (blocked → STUCK); reaper + pickup guard compose (stale
  in_progress at pickup limit → reaped → STUCK in one iteration).
- master_main_loop: runs until stop; empty loop exits cleanly;
  on_iteration exception doesn't break loop; reaper runs less
  frequently than tick.

Phase 1d-3+ remains for: discovery (Forgejo poll), per-workflow
scheduling (enqueue next workflow_attempts after transition),
Forgejo writes, MERGING state's merge call, reconciliation,
backfill, operator CLI.

Total: 278 controller tests; full auto_agents suite 2640 pass.
2026-05-18 13:41:27 -04:00
drew b519ebd98f feat(controller): Phase 1d-2 — outcome mapper + master tick handler
The master-side state machine driver. When a worker writes a complete
(or failed) attempt, the tick handler picks it up next pass, maps the
outcome to a state machine event, and applies the transition.

tools/controller/master/:

- outcomes.py: map_outcome_to_event() — pure function. Inputs:
  role, current_state, output_payload, status, head_sha_advanced,
  attempts_remaining_at_tier, conflict_count_at_current_tier.
  Returns EventMapResult(event_name|None, reason).
  - Implementer: resolved+pushed → implementer_pushed;
    resolved+NO push → implementer_competence_failure (worker lied);
    rebase-failed / competence-failure / blocked / noop all routed.
  - Reviewer: approve → reviewer_approve; request-changes → retry vs
    escalate based on attempts_remaining_at_tier; abstain →
    reviewer_abstain; comment → no transition (operator review needed).
  - Estimator: is_metadata_only → estimator_metadata_only; else
    estimator_done.
  - Conflict resolver: counts at current tier — 1st → first; 2nd →
    second_same_tier (escalate); 3rd+ → three_plus (STUCK).
  - Summarizer: doesn't drive state transitions.
  - status='failed' policy: contract-violation → pickup_exhausted;
    other failed outcomes (worker-internal-error / stale-input /
    ttl-insufficient-for-retry / git-clone-failed / lost-lock / etc.)
    → no event (master re-enqueues per the v9 table).
  - status='reaped' → no event (pickup guard handled).

- tick.py: run_tick() — one master tick. Queries
  workflow_attempts WHERE status IN ('complete', 'failed') AND
  finished_at > workflow.last_transition_at (heuristic for "not yet
  processed"). Per row:
  1. Validate current_state ∈ KNOWN_STATES (per v6 unknown-state
     guard); if not, transition workflow → STUCK with reason='unknown-state'.
  2. Map outcome → event via outcomes.map_outcome_to_event.
  3. If no event, bump workflow.last_transition_at so we don't
     re-process forever.
  4. Apply event via state_machine.apply_event; IllegalTransition →
     STUCK with reason='illegal-event'.
  5. Commit transition (workflows.current_state + entered_state_at +
     last_transition_at) AND insert controller_events row.

  Returns TickReport with attempts_processed + transitions_applied +
  transitions_log.

42 new tests:
- outcomes (22): status='failed' policy (3 paths) + each role's
  happy paths + edge cases (unknown outcome / missing field /
  None payload / unknown role).
- tick (20): implementer transitions (resolved+push → AWAITING_CI,
  resolved-no-push → ESCALATING, rebase → CONFLICT_RESOLVING,
  blocked → STUCK), reviewer transitions (approve → MERGING,
  request-changes → IMPLEMENTING), event row content, no
  reprocessing on second tick, unmapped attempts bump
  last_transition_at, terminal workflows skipped (4 parametrized),
  unknown current_state → STUCK.

Total: 271 controller tests; full auto_agents suite 2633 pass.
2026-05-18 13:38:28 -04:00