976817fa22
The deterministic spine of the master controller. State machine is
pure data with 6 load-bearing invariants enforced via property tests.
Reaper resets stale-heartbeat workflow_attempts to pending. Pickup
guard transitions workflows to STUCK when an attempt has been
re-pended too many times without success.
tools/controller/state_machine.py:
- KNOWN_STATES = 12; TERMINAL_STATES = {MERGED, ABANDONED, STUCK,
CREATED_PR}. STUCK's only allowed exit is the operator-driven
operator_unstick event (back to DISCOVERED).
- 32 TRANSITIONS entries covering DISCOVERED → ANALYZING →
IMPLEMENTING ↔ AWAITING_CI / CONFLICT_RESOLVING / ESCALATING →
REVIEWING → MERGING → MERGED. Plus pickup_exhausted exits from
IMPLEMENTING/CONFLICT_RESOLVING/REVIEWING.
- 27 named events with descriptions. apply_event() lookup raises
IllegalTransitionError (lists legal events from current state)
or ValueError on unknown state (per v6 unknown-state guard).
- 6 LOAD-BEARING invariants for v1 (per v9 simplification):
1. no_path_implementing_to_reviewing_skips_ci (Hard Rule #1
constructional fix for the no-mans-land race)
2. terminal_states_have_no_exits (only STUCK→operator_unstick OK)
3. tier_monotonic_non_decreasing
4. every_pr_workflow_includes_reviewing
5. conflict_resolving_bounded (1st→IMPLEMENTING, 2nd→ESCALATING,
3rd→STUCK; structurally encoded)
6. escalation_deterministic
- reachable_from() honors cycles (DISCOVERED ∈ reachable(DISCOVERED)
via STUCK→operator_unstick path; AWAITING_CI self-loops via
ci_flake_retry).
tools/controller/reaper.py:
- reap_stale_attempts(): SELECT in_progress attempts whose
lock_heartbeat_at + lock_ttl_seconds < NOW (per-row TTL respects
per-role differences — estimator 180s, reviewer 720s, tier-2
implementer 2160s). UPDATEs status='pending', clears lock columns,
preserves pickup_count (the pickup guard handles that). Inserts
controller_events row with reason='lock-ttl-expired' per reap.
- Dialect-portable: Postgres uses interval arithmetic; SQLite uses
julianday(). Same logic either way.
tools/controller/pickup_guard.py:
- transition_exhausted_to_stuck(): finds attempts with status='pending'
AND pickup_count >= MAX_PICKUPS (default 3 per v6 blocker fix)
AND workflow not already terminal. Transitions workflow → STUCK,
marks attempt as 'reaped', inserts controller_events with
reason='attempt-pickup-exhausted' + pickup_count + max_pickups.
45 new tests:
- state_machine: basic shape (states partition, every transition uses
known states + defined events), apply_event success/error paths,
events_from + reachable_from helpers (including cycle awareness),
per-invariant zero-violations against the live table, per-invariant
monkeypatch-violations to prove the checks catch the bug class they
claim to, parametrised sanity check "every non-terminal can reach
some terminal".
- reaper: empty DB / fresh heartbeat / stale heartbeat reaped /
per-row TTL respected / event row created / only-in-progress
reaped / multiple stale attempts.
- pickup guard: empty DB / below limit / at limit / in-progress not
checked / terminal workflow skipped / event payload content /
default max_pickups matches v6.
Total: 229 controller tests; full auto_agents suite 2591 pass.
131 lines
4.7 KiB
Python
131 lines
4.7 KiB
Python
"""Master-side pickup-exhaustion guard.
|
|
|
|
Per plan v6 blocker fix: an attempt that's been dequeued
|
|
``MAX_PICKUPS`` times without success means something is structurally
|
|
wrong (worker bug, persistent OpenCode failure, infrastructure
|
|
issue). The master detects these and transitions the workflow to
|
|
STUCK with ``reason='attempt-pickup-exhausted'``.
|
|
|
|
The dequeue helper (``db/dequeue.py``) already enforces the
|
|
``pickup_count < MAX_PICKUPS`` filter — so exhausted attempts are
|
|
NEVER picked up. This guard handles the secondary case: an attempt
|
|
that JUST hit the limit after a failed run; the workflow needs to
|
|
transition to STUCK so it stops being re-enqueued.
|
|
|
|
Runs on the master's tick loop. Cheap query; safe to run every tick.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import logging
|
|
import os
|
|
from dataclasses import dataclass, field
|
|
from datetime import datetime, timezone
|
|
|
|
from sqlalchemy import text
|
|
from sqlalchemy.engine import Engine
|
|
|
|
from .db.session import session_scope
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
DEFAULT_MAX_PICKUPS = int(os.environ.get("CONTROLLER_MAX_ATTEMPT_PICKUPS", "3"))
|
|
|
|
|
|
@dataclass
|
|
class PickupGuardReport:
|
|
"""Per-tick summary: which workflows got transitioned to STUCK."""
|
|
|
|
workflows_stuck: int = 0
|
|
stuck_workflows: list[tuple[int, int]] = field(default_factory=list)
|
|
# Each entry: (workflow_id, attempt_id_that_exhausted)
|
|
|
|
|
|
def transition_exhausted_to_stuck(
|
|
engine: Engine, *, max_pickups: int = DEFAULT_MAX_PICKUPS,
|
|
) -> PickupGuardReport:
|
|
"""Find attempts at or past MAX_PICKUPS that are also currently
|
|
pending (i.e., the last run failed and re-pended), and transition
|
|
their workflow to STUCK.
|
|
|
|
"At or past MAX_PICKUPS" means ``pickup_count >= max_pickups``
|
|
AND ``status = 'pending'``. The dequeue helper won't pick these
|
|
up (its filter is ``pickup_count < max_pickups``), so they'd loop
|
|
in the pending pool forever without this guard.
|
|
"""
|
|
report = PickupGuardReport()
|
|
now = datetime.now(timezone.utc)
|
|
with session_scope(engine) as session:
|
|
# Find candidate attempts + their workflows. A workflow in
|
|
# an already-terminal state doesn't need re-transitioning.
|
|
rows = session.execute(
|
|
text(
|
|
"SELECT a.attempt_id, a.workflow_id, a.pickup_count, "
|
|
" w.current_state "
|
|
"FROM workflow_attempts a "
|
|
"JOIN workflows w ON w.workflow_id = a.workflow_id "
|
|
"WHERE a.status = 'pending' "
|
|
" AND a.pickup_count >= :limit "
|
|
" AND w.current_state NOT IN ('MERGED', 'ABANDONED', 'STUCK', 'CREATED_PR')"
|
|
),
|
|
{"limit": max_pickups},
|
|
).all()
|
|
|
|
if not rows:
|
|
return report
|
|
|
|
for r in rows:
|
|
# Transition workflow → STUCK; mark the attempt as reaped
|
|
# (so the pending pool is clean).
|
|
session.execute(
|
|
text(
|
|
"UPDATE workflows SET "
|
|
" current_state = 'STUCK', "
|
|
" last_transition_at = :now, "
|
|
" entered_state_at = :now "
|
|
"WHERE workflow_id = :wf_id"
|
|
),
|
|
{"now": now, "wf_id": r.workflow_id},
|
|
)
|
|
session.execute(
|
|
text(
|
|
"UPDATE workflow_attempts SET status = 'reaped' "
|
|
"WHERE attempt_id = :aid"
|
|
),
|
|
{"aid": r.attempt_id},
|
|
)
|
|
session.execute(
|
|
text(
|
|
"INSERT INTO controller_events "
|
|
"(workflow_id, ts, event_type, from_state, to_state, "
|
|
" attempt_id, payload, forgejo_write_pending, replay_attempts) "
|
|
"VALUES (:wf_id, :ts, 'transition', :from_state, 'STUCK', "
|
|
" :aid, :payload, 0, 0)"
|
|
),
|
|
{
|
|
"wf_id": r.workflow_id,
|
|
"ts": now,
|
|
"from_state": r.current_state,
|
|
"aid": r.attempt_id,
|
|
"payload": json.dumps({
|
|
"reason": "attempt-pickup-exhausted",
|
|
"pickup_count": r.pickup_count,
|
|
"max_pickups": max_pickups,
|
|
}),
|
|
},
|
|
)
|
|
report.workflows_stuck += 1
|
|
report.stuck_workflows.append((r.workflow_id, r.attempt_id))
|
|
|
|
if report.workflows_stuck:
|
|
logger.warning(
|
|
"pickup guard: %d workflow(s) transitioned to STUCK "
|
|
"(pickup-exhausted): %s",
|
|
report.workflows_stuck, report.stuck_workflows,
|
|
)
|
|
return report
|
|
|
|
|
|
__all__ = ["DEFAULT_MAX_PICKUPS", "PickupGuardReport", "transition_exhausted_to_stuck"]
|