Files
cleveragents-core/tools/controller/pickup_guard.py
T
drew 976817fa22 feat(controller): Phase 1d-1 — state machine + reaper + pickup guard
The deterministic spine of the master controller. State machine is
pure data with 6 load-bearing invariants enforced via property tests.
Reaper resets stale-heartbeat workflow_attempts to pending. Pickup
guard transitions workflows to STUCK when an attempt has been
re-pended too many times without success.

tools/controller/state_machine.py:

- KNOWN_STATES = 12; TERMINAL_STATES = {MERGED, ABANDONED, STUCK,
  CREATED_PR}. STUCK's only allowed exit is the operator-driven
  operator_unstick event (back to DISCOVERED).
- 32 TRANSITIONS entries covering DISCOVERED → ANALYZING →
  IMPLEMENTING ↔ AWAITING_CI / CONFLICT_RESOLVING / ESCALATING →
  REVIEWING → MERGING → MERGED. Plus pickup_exhausted exits from
  IMPLEMENTING/CONFLICT_RESOLVING/REVIEWING.
- 27 named events with descriptions. apply_event() lookup raises
  IllegalTransitionError (lists legal events from current state)
  or ValueError on unknown state (per v6 unknown-state guard).
- 6 LOAD-BEARING invariants for v1 (per v9 simplification):
  1. no_path_implementing_to_reviewing_skips_ci (Hard Rule #1
     constructional fix for the no-mans-land race)
  2. terminal_states_have_no_exits (only STUCK→operator_unstick OK)
  3. tier_monotonic_non_decreasing
  4. every_pr_workflow_includes_reviewing
  5. conflict_resolving_bounded (1st→IMPLEMENTING, 2nd→ESCALATING,
     3rd→STUCK; structurally encoded)
  6. escalation_deterministic
- reachable_from() honors cycles (DISCOVERED ∈ reachable(DISCOVERED)
  via STUCK→operator_unstick path; AWAITING_CI self-loops via
  ci_flake_retry).

tools/controller/reaper.py:

- reap_stale_attempts(): SELECT in_progress attempts whose
  lock_heartbeat_at + lock_ttl_seconds < NOW (per-row TTL respects
  per-role differences — estimator 180s, reviewer 720s, tier-2
  implementer 2160s). UPDATEs status='pending', clears lock columns,
  preserves pickup_count (the pickup guard handles that). Inserts
  controller_events row with reason='lock-ttl-expired' per reap.
- Dialect-portable: Postgres uses interval arithmetic; SQLite uses
  julianday(). Same logic either way.

tools/controller/pickup_guard.py:

- transition_exhausted_to_stuck(): finds attempts with status='pending'
  AND pickup_count >= MAX_PICKUPS (default 3 per v6 blocker fix)
  AND workflow not already terminal. Transitions workflow → STUCK,
  marks attempt as 'reaped', inserts controller_events with
  reason='attempt-pickup-exhausted' + pickup_count + max_pickups.

45 new tests:
- state_machine: basic shape (states partition, every transition uses
  known states + defined events), apply_event success/error paths,
  events_from + reachable_from helpers (including cycle awareness),
  per-invariant zero-violations against the live table, per-invariant
  monkeypatch-violations to prove the checks catch the bug class they
  claim to, parametrised sanity check "every non-terminal can reach
  some terminal".
- reaper: empty DB / fresh heartbeat / stale heartbeat reaped /
  per-row TTL respected / event row created / only-in-progress
  reaped / multiple stale attempts.
- pickup guard: empty DB / below limit / at limit / in-progress not
  checked / terminal workflow skipped / event payload content /
  default max_pickups matches v6.

Total: 229 controller tests; full auto_agents suite 2591 pass.
2026-05-18 13:34:28 -04:00

131 lines
4.7 KiB
Python

"""Master-side pickup-exhaustion guard.
Per plan v6 blocker fix: an attempt that's been dequeued
``MAX_PICKUPS`` times without success means something is structurally
wrong (worker bug, persistent OpenCode failure, infrastructure
issue). The master detects these and transitions the workflow to
STUCK with ``reason='attempt-pickup-exhausted'``.
The dequeue helper (``db/dequeue.py``) already enforces the
``pickup_count < MAX_PICKUPS`` filter — so exhausted attempts are
NEVER picked up. This guard handles the secondary case: an attempt
that JUST hit the limit after a failed run; the workflow needs to
transition to STUCK so it stops being re-enqueued.
Runs on the master's tick loop. Cheap query; safe to run every tick.
"""
from __future__ import annotations
import json
import logging
import os
from dataclasses import dataclass, field
from datetime import datetime, timezone
from sqlalchemy import text
from sqlalchemy.engine import Engine
from .db.session import session_scope
logger = logging.getLogger(__name__)
DEFAULT_MAX_PICKUPS = int(os.environ.get("CONTROLLER_MAX_ATTEMPT_PICKUPS", "3"))
@dataclass
class PickupGuardReport:
"""Per-tick summary: which workflows got transitioned to STUCK."""
workflows_stuck: int = 0
stuck_workflows: list[tuple[int, int]] = field(default_factory=list)
# Each entry: (workflow_id, attempt_id_that_exhausted)
def transition_exhausted_to_stuck(
engine: Engine, *, max_pickups: int = DEFAULT_MAX_PICKUPS,
) -> PickupGuardReport:
"""Find attempts at or past MAX_PICKUPS that are also currently
pending (i.e., the last run failed and re-pended), and transition
their workflow to STUCK.
"At or past MAX_PICKUPS" means ``pickup_count >= max_pickups``
AND ``status = 'pending'``. The dequeue helper won't pick these
up (its filter is ``pickup_count < max_pickups``), so they'd loop
in the pending pool forever without this guard.
"""
report = PickupGuardReport()
now = datetime.now(timezone.utc)
with session_scope(engine) as session:
# Find candidate attempts + their workflows. A workflow in
# an already-terminal state doesn't need re-transitioning.
rows = session.execute(
text(
"SELECT a.attempt_id, a.workflow_id, a.pickup_count, "
" w.current_state "
"FROM workflow_attempts a "
"JOIN workflows w ON w.workflow_id = a.workflow_id "
"WHERE a.status = 'pending' "
" AND a.pickup_count >= :limit "
" AND w.current_state NOT IN ('MERGED', 'ABANDONED', 'STUCK', 'CREATED_PR')"
),
{"limit": max_pickups},
).all()
if not rows:
return report
for r in rows:
# Transition workflow → STUCK; mark the attempt as reaped
# (so the pending pool is clean).
session.execute(
text(
"UPDATE workflows SET "
" current_state = 'STUCK', "
" last_transition_at = :now, "
" entered_state_at = :now "
"WHERE workflow_id = :wf_id"
),
{"now": now, "wf_id": r.workflow_id},
)
session.execute(
text(
"UPDATE workflow_attempts SET status = 'reaped' "
"WHERE attempt_id = :aid"
),
{"aid": r.attempt_id},
)
session.execute(
text(
"INSERT INTO controller_events "
"(workflow_id, ts, event_type, from_state, to_state, "
" attempt_id, payload, forgejo_write_pending, replay_attempts) "
"VALUES (:wf_id, :ts, 'transition', :from_state, 'STUCK', "
" :aid, :payload, 0, 0)"
),
{
"wf_id": r.workflow_id,
"ts": now,
"from_state": r.current_state,
"aid": r.attempt_id,
"payload": json.dumps({
"reason": "attempt-pickup-exhausted",
"pickup_count": r.pickup_count,
"max_pickups": max_pickups,
}),
},
)
report.workflows_stuck += 1
report.stuck_workflows.append((r.workflow_id, r.attempt_id))
if report.workflows_stuck:
logger.warning(
"pickup guard: %d workflow(s) transitioned to STUCK "
"(pickup-exhausted): %s",
report.workflows_stuck, report.stuck_workflows,
)
return report
__all__ = ["DEFAULT_MAX_PICKUPS", "PickupGuardReport", "transition_exhausted_to_stuck"]