"""Blocking OpenCode worker dispatcher. Implements ``run_worker_blocking`` per ``docs/development/conflict-drive-plan.md`` § 6.2. The single entry point spins up an OpenCode session, dispatches an agent prompt, polls until the session is idle, performs tolerant JSON extraction on the final assistant message, and tears the session down — all from a deterministic Python driver with a watchdog timeout. The expected caller is :mod:`tools.conflict_drive`, but the surface is agent-agnostic and can be reused by other deterministic drivers that need a single-shot LLM worker. Model resolution ---------------- At session creation, :func:`run_session_blocking` resolves the model for the dispatched agent via :func:`_resolve_role_model`, which reads from ``.opencode/models/.txt`` (falling back to ``.opencode/models/default.txt``). The resolved ``{providerID, id}`` is passed to ``POST /session`` for **observability only** — empirically OpenCode does NOT consume this field for generation; it generates from its startup-cached ``opencode.json`` ``agent..model`` (see the footgun note in ``.opencode/models/README.md``). A model swap therefore requires editing ``opencode.json`` (or its ``{file:...}`` target) and restarting OpenCode — not just a ``.txt`` edit. If neither ``.txt`` file exists, the ``model`` field is simply omitted from the POST. Outcomes -------- The worker returns a :class:`WorkerResult` whose ``outcome`` is one of: - ``resolved`` — every conflict was semantically resolved. - ``unresolvable`` — the rebase finished with conflict markers committed; the driver will detect them via ``git diff --check`` and escalate via ``auto/needs-implementer``. - ``rebase-failed`` — the worker could not progress the rebase at all. - ``timeout`` — the watchdog wallclock fired (default 900 s). - ``transport-error`` — OpenCode HTTP 5xx, network reset, or any unexpected I/O error during the dispatch lifecycle. The driver maps these onto the failure-class table in plan § 5.1; ``verification-fail`` and ``lease-violation`` are driver-side classifications, not worker outcomes. """ from __future__ import annotations import datetime as _dt import json import logging import os import re import sys import time import urllib.error import urllib.request from collections.abc import Callable, Iterable from dataclasses import dataclass, field from pathlib import Path from typing import Any, Literal # Match the sibling-loader pattern used elsewhere in ``tools/`` # (see ``_dispatch_runtime.py`` for the canonical comment). Both # direct invocation (``python tools/...``) and pytest's # ``spec_from_file_location`` loader work after this idempotent insert. _TOOLS_DIR = str(Path(__file__).resolve().parent) if _TOOLS_DIR not in sys.path: sys.path.insert(0, _TOOLS_DIR) from _loader import ( # noqa: E402 type: ignore[import-not-found] load_sibling as _load_sibling, ) # Mid-flight pattern-check helper for the doom-loop / retry-cascade # soft-threshold probe. Loaded lazily via the sibling loader so this # module's import surface stays minimal and the helper can be unit # tested in isolation. _patterns = _load_sibling("_session_pattern_checks", "_session_pattern_checks.py") logger = logging.getLogger("opencode_worker") Outcome = Literal[ "resolved", "unresolvable", "rebase-failed", "timeout", "transport-error", ] @dataclass class WorkerResult: """Typed return value from :func:`run_worker_blocking`. ``raw_response`` carries the verbatim final-assistant message text purely for telemetry / debug; the driver never parses it directly (the typed ``outcome`` and ``files_touched`` fields above are the contract). ``error_kind`` is populated only for ``timeout`` / ``transport- error`` outcomes (and is ``None`` for ``resolved`` / ``unresolvable`` / ``rebase-failed``); it lets dispatchers route on the cause of a failure (``transport-create-session`` vs ``transport-poll`` vs ``startup-grace-elapsed``) without parsing free-text from ``raw_response``. """ outcome: Outcome files_touched: list[str] = field(default_factory=list) wallclock_seconds: float = 0.0 session_id: str = "" raw_response: str = "" # Typed as ``str | None`` rather than ``ErrorKind | None`` because # the field is also written from the conflict-driver wrappers # (``run_worker_blocking``) which surface workers' own # outcome-derived strings ("worker-failed", etc.) that are NOT # part of the closed :data:`ErrorKind` set. Loosening to ``str`` # at the dataclass level keeps both sources flowing through one # field without a cross-module type alias. Cycle-log consumers # treat the field as an opaque label and pivot on string equality. error_kind: str | None = None SessionStatus = Literal["completed", "timeout", "transport-error"] # Stable identifiers a dispatcher cycle log can pivot on. Matches every # ``transport-error`` / ``timeout`` return site in # :func:`run_session_blocking` so a SQL filter on ``error_kind`` always # returns a useful row count. # # :data:`SessionStatus` describes the *lifecycle* of a session # (``completed`` / ``timeout`` / ``transport-error``); :data:`ErrorKind` # describes the *cause* of a non-completed session in finer-grained # terms. The two are related but distinct: every ``transport-error`` # session has a ``transport-*`` or ``startup-grace-elapsed`` # error_kind; every ``timeout`` session has ``watchdog-timeout``; # ``completed`` sessions have ``error_kind=None``. ErrorKind = Literal[ "transport-create-session", "transport-create-session-malformed", "transport-prompt-async", "transport-poll-status", "transport-message-fetch", "startup-grace-elapsed", "watchdog-timeout", # Mid-flight doom-loop / hang-pattern early-abort kinds (G4 # harvest 2026-05-15). Emitted by the soft-threshold probe that # runs past ~30% of ``timeout_seconds`` when # ``WORKER_DOOM_LOOP_ABORT_ENABLED=1``. The status stays # ``timeout`` (the session was forcibly ended) but the specific # cause lets an operator distinguish a real hang from "the model # just needed more time." Detection logic lives in # ``_session_pattern_checks``. "doom-loop", "retry-cascade", "permission-deadlock", "empty-reasoning-loop", ] @dataclass class SessionResult: """Outcome-agnostic result from running a one-shot OpenCode session. Workers that do not emit a fixed JSON schema (review, implementation) use this directly; the conflict driver keeps using :func:`run_worker_blocking`, which classifies the parsed JSON into a domain-specific :class:`WorkerResult.outcome`. ``status`` describes only the *session* lifecycle: did the session finish normally (``completed``), hit the watchdog (``timeout``), or fail at the OpenCode HTTP layer (``transport-error``). ``parsed_json`` is the last parseable JSON object found in the final assistant message, or ``None`` if the worker did not emit one — which is the normal case for review / implementation workers. ``error_kind`` is ``None`` on ``completed``; on ``timeout`` it is ``watchdog-timeout``; on ``transport-error`` it identifies the specific lifecycle phase that failed so dispatchers can route on cause without parsing ``raw_response``. """ status: SessionStatus wallclock_seconds: float = 0.0 session_id: str = "" raw_response: str = "" parsed_json: dict[str, Any] | None = None # Typed as ``str | None`` rather than ``ErrorKind | None`` for # symmetry with :class:`WorkerResult.error_kind` — keeps the two # dataclasses aligned and avoids forcing every cross-module # consumer to import :data:`ErrorKind`. Production writers # (``run_session_blocking``) populate it only with values from # :data:`ErrorKind`; the looser type is purely for downstream # convenience. error_kind: str | None = None # Deepest BFS distance observed in the subagent task-tool tree, # captured by :func:`_archive_subagent_tree` from OpenCode's # ``/session`` walk just before the root session is DELETEd. # ``0`` for a flat session (no subagents), ``N > 0`` for a chain # of depth N, and ``None`` when the walk did not run (archive # disabled, transport error, or wrapper session never created). # Phase 4 telemetry's ``subagent_max_depth`` is sourced from this # field; the flatten-decision framework reads it directly. Kept # on ``SessionResult`` (not buried inside ``parsed_json``) # because not every worker emits structured JSON, but every # worker has a subagent tree by virtue of being an OpenCode # session. subagent_max_depth: int | None = None # ─── HTTP layer (kept intentionally minimal — no shared retries) ────────── # # We deliberately do NOT reuse ``_claim_runtime``'s state-change retry # policy: OpenCode's writes are not idempotent (POST /session creates a # new session every call) and a transient retry would multiply session # leaks. Each request gets one shot; any error short-circuits to # ``transport-error``. # Tuple of exceptions every dispatch step catches. ``urllib`` raises # ``HTTPError`` (subclass of ``URLError``) on non-2xx, so this covers # both transport-level failures (DNS, connection refused) and OpenCode # 5xx / 4xx. _TRANSPORT_EXC: tuple[type[BaseException], ...] = ( urllib.error.URLError, TimeoutError, OSError, ) def _request( method: str, url: str, *, body: Any | None = None, timeout: int = 30, ) -> Any: """Issue a single HTTP request to the OpenCode server. Returns the parsed JSON body on 2xx; raises ``urllib.error.HTTPError`` (a subclass of ``URLError``) on non-2xx, and lets all other transport-level exceptions propagate. Callers catch :data:`_TRANSPORT_EXC` and convert to ``transport-error``. """ data = json.dumps(body).encode() if body is not None else None headers = {} if data is not None: headers["Content-Type"] = "application/json" req = urllib.request.Request(url, data=data, headers=headers, method=method) with urllib.request.urlopen(req, timeout=timeout) as resp: payload = resp.read() if not payload: return None ct = resp.headers.get("Content-Type", "") if "application/json" in ct: try: return json.loads(payload) except (ValueError, TypeError): return None return payload # Read-only retry budget: a single transient flap (DNS hiccup, OpenCode # 503 during a graceful restart, brief network partition) on a polling # GET should not trash a 10-minute worker session. The retry covers # READ-ONLY operations only — every write (POST /session, POST # /prompt_async, POST /abort, DELETE /session) keeps one-shot # semantics because OpenCode writes are not idempotent (the comment on # :data:`_TRANSPORT_EXC` explains the leak risk). See the audit note # in the auto-agents Tier 2/3 plan, item 4. _READ_RETRY_ATTEMPTS = 3 _READ_RETRY_BACKOFF_SECONDS = 0.5 def _request_read( url: str, *, timeout: int = 30, attempts: int = _READ_RETRY_ATTEMPTS, backoff_seconds: float = _READ_RETRY_BACKOFF_SECONDS, ) -> Any: """Issue one GET request with a small retry budget on transient transport errors. Use this only for idempotent reads where a partial success is indistinguishable from a fresh request — i.e. ``GET /session/status`` and ``GET /session/{id}/message``. POST / DELETE callers must keep using :func:`_request` directly so a transient failure does not silently double-issue the side-effecting call. The retry only fires on members of :data:`_TRANSPORT_EXC` (DNS, connection refused, ``HTTPError``, timeout). Non-transport failures (e.g. malformed JSON the server actually returned) come back as ``None`` from :func:`_request` and are NOT retried — those are bugs in the server's response, not transient flaps, and retrying would mask the regression. Backoff is linear (``backoff_seconds * attempt``) so three attempts at the default 0.5s cap total added latency at ~1.5s against a server that is fully down — the watchdog will fire in its own time. Each retry attempt logs at INFO so an operator can see flap density. """ last_err: BaseException | None = None for attempt in range(1, attempts + 1): try: return _request("GET", url, timeout=timeout) except _TRANSPORT_EXC as e: last_err = e if attempt < attempts: logger.info( "OpenCode read flap on %s (attempt %d/%d): %s — retrying in %.1fs", url, attempt, attempts, type(e).__name__, backoff_seconds * attempt, ) time.sleep(backoff_seconds * attempt) continue raise # Unreachable: the loop either returns a value or raises through # the final attempt's ``raise``. raise RuntimeError( # pragma: no cover f"_request_read exhausted {attempts} attempts: {last_err}" ) # ─── Tolerant JSON extraction ────────────────────────────────────────────── # # The worker's prompt asks for "exactly one JSON object somewhere in # your final assistant message". Real LLMs surround it with prose, code # fences, trailing whitespace, and frequently emit stray ``{`` inside # explanatory prose ("set REPO_DIR={your repo path}"). We scan # left-to-right looking at every ``{`` independently — taking only the # LAST successfully-parsed object — so a single unmatched opener # earlier in the message cannot shadow the structured exit at the end. def _extract_last_json_object(text: str) -> dict[str, Any] | None: """Return the last ``{...}`` substring of ``text`` that parses as a JSON object, or ``None`` if no such substring exists. Strategy: walk left-to-right; at every ``{`` start, run a string-aware bracket matcher to see whether THIS opener closes at depth 0 within the buffer. If it does, attempt ``json.loads`` on the span and record success; in any case we then advance to the next character (NOT to ``end + 1``) so an unmatched outer ``{`` cannot hide a valid object that nests inside it or that follows it later. The LAST successful parse wins so the worker's structured exit (always at the end of the message by prompt convention) is what we return. Cost: best-case O(N) when no stray openers appear; worst case O(N²) when the buffer contains many unterminated ``{`` runs (the pathological run-of-N-opens-no-closes string). LLM output is bounded by ``max_tokens`` so the worst case is irrelevant in practice; the pre-P1-6 implementation was the same complexity for a much weaker reason (quadratic shrink). Replacing the unmatched- opener ``break`` with ``i += 1`` is what closes the regression flagged in the architect's review. Tolerates code fences (``\\`\\`\\`json`` …), trailing prose, partial fragments, multiple JSON objects, and stray ``{`` inside prose. Does NOT tolerate JSON arrays at the outer level — by contract the worker emits an object. """ if not text: return None last_match: dict[str, Any] | None = None n = len(text) i = 0 while i < n: if text[i] != "{": i += 1 continue # Scan forward from here, respecting JSON string syntax, until # we either close at depth 0 (complete object) or run off the # end (truncated). Inside a string, ``{`` and ``}`` are literal # data and must not move the depth counter. depth = 0 in_string = False escape = False end = -1 for j in range(i, n): ch = text[j] if in_string: if escape: escape = False elif ch == "\\": escape = True elif ch == '"': in_string = False continue if ch == '"': in_string = True continue if ch == "{": depth += 1 elif ch == "}": depth -= 1 if depth == 0: end = j + 1 break if end < 0: # Unmatched opener — every attempt that includes this ``{`` # in its span will fail. Advance ONE position so subsequent # ``{``s (which may close fine) still get a chance. The # earlier ``break`` here was a regression: it caused # ``"prose with { stray. Then {valid: true}"`` to return # None instead of finding the trailing object. i += 1 continue slice_ = text[i:end] try: parsed = json.loads(slice_) except (ValueError, TypeError): parsed = None if isinstance(parsed, dict): last_match = parsed # The outer span parsed cleanly — skip past it. Without this # the scan would re-enter the body and overwrite ``last_match`` # with a nested sub-object, returning ``{"inner": 1}`` for # input ``{"outer": {"inner": 1}}``. The contract is "last # complete object", not "last opener". i = end else: # Span closed at depth 0 but did NOT parse (e.g. invalid # syntax inside a stray-bracketed prose block). Advance one # position so a valid object that nests inside this # unparseable span is still discoverable. i += 1 return last_match def _last_assistant_text(messages: list[dict[str, Any]]) -> str: """Return the concatenated text of the most recent assistant message in ``messages`` (oldest-first ordering, as returned by ``GET /session/{id}/message``). Empty string when no assistant message exists yet. """ for m in reversed(messages): info = m.get("info") or {} if info.get("role") != "assistant": continue parts = m.get("parts") or [] # Concatenate visible text parts only; ignore reasoning / tool / # patch / step-* entries — none of them are the structured exit # the worker is asked to emit. text_parts = [p.get("text", "") for p in parts if p.get("type") == "text"] return "\n".join(t for t in text_parts if t) return "" # ─── Idle detection ──────────────────────────────────────────────────────── def _session_busy_state(server_url: str, session_id: str) -> str: """Return ``"busy"``, ``"idle"``, or ``"unknown"`` for ``session_id``. ``GET /session/status`` returns a map of ``session_id -> {"type": "busy"|"idle"}``. The map is updated asynchronously by OpenCode after a session is created, so a missing entry is **ambiguous**: it can mean either "session has not been registered yet" (immediately after ``POST /session/{id}/prompt_async``) or "session has run to completion and been removed" (a real terminal state). Distinguishing these two cases is the caller's responsibility. This helper returns ``"unknown"`` for a missing entry and lets :func:`run_session_blocking` apply the ``seen_busy_at_least_once`` heuristic that disambiguates the pre-launch race from a genuine terminal idle. Mirrors the data shape of ``.opencode/skills/auto-agents-system/scripts/session_wait_till_idle.ts``; that script tolerates the race window because its 5 s poll interval + ``Promise.all`` round-trip lets OpenCode register the session before the first observation. """ statuses = _request_read(f"{server_url}/session/status") or {} if not isinstance(statuses, dict): return "unknown" entry = statuses.get(session_id) if not isinstance(entry, dict): return "unknown" return "busy" if entry.get("type") == "busy" else "idle" def _session_assistant_in_flight(server_url: str, session_id: str) -> bool | None: """Authoritative "is the worker still doing work?" check, used to disambiguate the two flavours of ``state == "unknown"``: - **Transient blip between assistant turns.** OpenCode briefly drops the session entry from ``/session/status`` while spinning up the next assistant message; the in-flight assistant message has ``time.completed == None``. - **Terminal completion.** All assistant turns are done and OpenCode has permanently removed the session from the status map; every assistant message has ``time.completed`` set. Returns ``True`` if the latest assistant message is in-flight, ``False`` if all assistant messages are complete (or there are no assistant messages yet), and ``None`` on transport / parse failure so the caller can fall back to the conservative behaviour (assume in-flight, keep polling) without crashing the polling loop. This helper is called from :func:`run_session_blocking` ONLY when the cheap status-map probe returned ``unknown && seen_busy``; it is intentionally kept off the happy-path so a normal busy session does NOT incur an extra HTTP call per poll. """ try: messages = _request_read(f"{server_url}/session/{session_id}/message") except _TRANSPORT_EXC: return None if not isinstance(messages, list): return None last_assistant: dict[str, Any] | None = None for m in messages: if not isinstance(m, dict): continue info = m.get("info") if isinstance(m.get("info"), dict) else m if isinstance(info, dict) and info.get("role") == "assistant": last_assistant = info if last_assistant is None: return False time_info = last_assistant.get("time") if not isinstance(time_info, dict): return False return time_info.get("completed") is None # Wait at most this many seconds for the session to first appear as # ``busy`` after ``POST /session/{id}/prompt_async``. If the session # never becomes busy within this window we assume the prompt never # reached the agent (OpenCode internal error, malformed agent, etc.) # and short-circuit with ``transport-error`` rather than silently # reporting ``completed`` with an empty raw_response. _STARTUP_GRACE_SECONDS = 30.0 # ─── Doom-loop / hang-pattern early-abort knobs (G4 harvest) ─────────────── # # Default-OFF feature flag, per the dmpipeline safety contract for # new consequential behaviour. When ON, the polling loop runs # :mod:`_session_pattern_checks` over the session's recent messages # once the session has burned past # ``_DOOM_PROBE_SOFT_THRESHOLD_FRACTION`` of its wallclock budget, and # re-probes every ``_DOOM_PROBE_INTERVAL_SECONDS`` thereafter. A # detector hit short-circuits the loop with ``status=timeout`` and a # named ``error_kind`` (e.g. ``doom-loop``, ``retry-cascade``) # instead of letting the full timeout elapse and returning the opaque # ``watchdog-timeout``. # # Tuning: the soft threshold is set high enough that a healthy long # session (e.g. an implementer at Tier 2 churning legitimately on a # hard PR) does not have its messages inspected at all, which keeps # the per-cycle network surface unchanged for the common case. _DOOM_LOOP_ABORT_ENV = "WORKER_DOOM_LOOP_ABORT_ENABLED" _DOOM_PROBE_SOFT_THRESHOLD_FRACTION = 0.30 _DOOM_PROBE_INTERVAL_SECONDS = 60.0 def _is_doom_loop_abort_enabled() -> bool: """Return ``True`` iff the operator has opted in to mid-flight doom-loop / hang-pattern early-abort via the :data:`_DOOM_LOOP_ABORT_ENV` env var. Default OFF. When OFF, the probe is skipped entirely — the polling loop's per-iteration cost is byte-identical to the pre-G4 build. When ON, the abort fires on any of the four patterns detected by :func:`_session_pattern_checks.check_message_stream`. """ raw = os.environ.get(_DOOM_LOOP_ABORT_ENV, "").strip().lower() return raw in ("1", "true", "yes", "on") def _maybe_doom_pattern_finding( *, server_url: str, session_id: str, ) -> Any: """Best-effort: fetch the session's message stream and run the pattern detectors. Returns a :class:`_session_pattern_checks.PatternFinding` on a hit; ``None`` on healthy, on transport failure, or on any unexpected payload shape. Failure to fetch is silently treated as "no signal" — we would rather miss one probe than abort a healthy session because of a transient OpenCode 503 on the message endpoint. """ try: messages = _request_read(f"{server_url}/session/{session_id}/message") except _TRANSPORT_EXC as e: logger.debug( "doom-loop probe: message fetch failed for session %s: %s " "(treating as no-signal)", session_id[:32], e, ) return None if not isinstance(messages, list): return None return _patterns.check_message_stream(messages) # ─── Session archival + per-turn metrics ─────────────────────────────────── # # The OpenCode session and all its messages are deleted in the finally # block of run_session_blocking, which means the worker's reasoning, # tool calls, and per-turn token counts are lost the moment the # dispatcher cycle finishes. That makes post-mortem analysis of a bad # run almost impossible — the dispatcher's terminal_state row in # dispatch_review_cycles tells you "transport-error" but not what the # worker actually said in the 6 minutes leading up to it. # # These helpers snapshot the full /session/{id}/message payload to a # JSON file under ``.dispatcher-logs/sessions/.json`` (or an # operator-overridable directory) before the session is deleted, and # emit per-assistant-turn summary lines so the dispatcher log shows # tool calls + token counts at-a-glance. # # The archive is best-effort: any error (disk full, permission denied, # transport failure during the message fetch) is logged + swallowed. # The dispatcher's correctness contract lives in the SessionResult # return value, never in the archive file. # Default archive location: ``/.dispatcher-logs/sessions``. # ``.dispatcher-logs/`` is already in .gitignore. _DEFAULT_ARCHIVE_SUBPATH = Path(".dispatcher-logs") / "sessions" # Override the archive directory entirely. Useful for tests (point at # tmp_path) or operators who want archives on a separate volume. _ARCHIVE_DIR_ENV = "OPENCODE_WORKER_ARCHIVE_DIR" # Disable archiving altogether. The default is enabled; tests opt out # via a conftest autouse fixture so they don't pollute the repo. _ARCHIVE_DISABLED_ENV = "OPENCODE_WORKER_ARCHIVE_DISABLED" # Schema version for the archive payload. Bump when the on-disk JSON # shape changes; the telemetry console reads this to decide which # fields to render. # # v1 (initial) — top-level wrapper sessions only; carries # session_id, agent, tag, status, started_at, # archived_at, wallclock_seconds, per_turn, # state_history, messages. # v2 (2026-05-10) — adds optional ``parent_session_id``, # ``subagent_title``, and ``subagent_depth`` fields so # the dispatcher can archive the entire ``task`` tool # subagent tree before deleting the root session. The # first three fields are absent / null on top-level # (wrapper) archives so a v1 reader treating them as # missing is forward-compatible. _ARCHIVE_SCHEMA_VERSION = 2 # Pattern that matches the trailing ``(@ subagent)`` suffix # OpenCode appends to a subagent session's title (e.g. # ``"Implement PR fix #30 (@task-implementor subagent)"``). The capture # group is the agent name and is what we use as the ``agent`` field in # the subagent archive payload. If the title does not match — which # would mean OpenCode changed its title convention or this is a # manually-titled top-level session — we fall back to a sentinel string # so the archive still lands on disk. _SUBAGENT_TITLE_RE = re.compile(r"\(@([\w.-]+)\s+subagent\)\s*$") def _resolve_archive_dir() -> Path | None: """Return the directory archives should be written into, or ``None`` if archiving is disabled / unavailable. Resolution order: 1. ``OPENCODE_WORKER_ARCHIVE_DISABLED=1`` set → ``None`` (off). 2. ``OPENCODE_WORKER_ARCHIVE_DIR=`` set → that path (created on demand if missing). 3. Auto-detect: ``/.dispatcher-logs/sessions`` derived from ``__file__`` (``tools/_opencode_worker.py`` → repo root is two parents up). A failure to ``mkdir(parents=True, exist_ok=True)`` is logged at WARNING and returns ``None`` — the rest of the worker keeps running, just without archives for this run. """ if os.environ.get(_ARCHIVE_DISABLED_ENV): return None explicit = os.environ.get(_ARCHIVE_DIR_ENV) if explicit: candidate = Path(explicit) else: repo_root = Path(__file__).resolve().parent.parent candidate = repo_root / _DEFAULT_ARCHIVE_SUBPATH try: candidate.mkdir(parents=True, exist_ok=True) except (OSError, PermissionError) as e: logger.warning("session archive disabled: cannot create %s (%s)", candidate, e) return None return candidate def _safe_filename_component(s: str) -> str: """Return ``s`` with everything outside ``[A-Za-z0-9_.-]`` replaced by ``_``. Keeps archive filenames readable + portable across filesystems (no colons, no slashes, no shell-meaningful chars).""" return re.sub(r"[^A-Za-z0-9_.-]", "_", s) or "_" def _build_archive_payload( *, session_id: str, agent: str, tag: str, status: str, started_at_iso: str, wallclock_seconds: float, messages: list[dict[str, Any]], state_history: list[dict[str, Any]] | None = None, parent_session_id: str | None = None, subagent_title: str | None = None, subagent_depth: int | None = None, ) -> dict[str, Any]: """Build the JSON-serialisable archive payload for a session. Pulled out into a helper so :func:`_archive_session` and the test suite share the same shape; the on-disk schema is therefore derivable from this function alone. ``parent_session_id`` / ``subagent_title`` / ``subagent_depth`` are populated only for subagent archives (sessions spawned by a parent's ``task`` tool call). For the top-level wrapper session they remain ``None`` and serialise as ``null`` so a downstream reader can tell the two apart without checking the schema version. """ return { "schema_version": _ARCHIVE_SCHEMA_VERSION, "session_id": session_id, "parent_session_id": parent_session_id, "subagent_title": subagent_title, "subagent_depth": subagent_depth, "agent": agent, "tag": tag, "status": status, "started_at": started_at_iso, "archived_at": _dt.datetime.now(_dt.timezone.utc).isoformat(), "wallclock_seconds": wallclock_seconds, "per_turn": _summarize_per_turn(messages), "state_history": state_history or [], "messages": messages, } # Anything shorter than this is rejected by :func:`_redact_secret_values` # to keep accidental false positives out of the archive. PATs at the # Forgejo / GitHub end are 40+ chars of high entropy; short ``${VAR}`` # names or single-letter literals would otherwise mangle every # occurrence of that substring in the payload (turning ``-h`` into # ``-``). _REDACT_MIN_LENGTH = 12 _REDACT_PLACEHOLDER = "" def _redact_secret_values(serialised: str, secrets: Iterable[str] | None) -> str: """Replace every occurrence of each ``secrets`` value in ``serialised`` with ````. Operates on the JSON-serialised archive string rather than on the Python payload tree so it catches the secret no matter where the worker echoed it (prompt text, tool ``input.command``, environment capture, error message). High-entropy PATs do not collide with legitimate substrings in practice, but we still enforce a 12-char minimum length to avoid pathological mangling if a caller passes a short or empty value by accident. Empty / ``None`` ``secrets`` is a no-op so callers can pass the parameter unconditionally. """ if not secrets: return serialised for value in secrets: if not value or not isinstance(value, str): continue if len(value) < _REDACT_MIN_LENGTH: logger.warning( "skipping redaction of secret shorter than %d chars " "(would mangle archive); caller must pass full PAT/credential", _REDACT_MIN_LENGTH, ) continue serialised = serialised.replace(value, _REDACT_PLACEHOLDER) return serialised def _archive_session( *, archive_dir: Path, session_id: str, agent: str, tag: str, status: str, started_at_iso: str, wallclock_seconds: float, messages: list[dict[str, Any]], state_history: list[dict[str, Any]] | None = None, redact_values: Iterable[str] | None = None, parent_session_id: str | None = None, subagent_title: str | None = None, subagent_depth: int | None = None, ) -> Path | None: """Write the session archive JSON file. Returns the resolved path on success, or ``None`` on any I/O / serialisation failure. ``redact_values`` is a list of secret strings (typically the ``FORGEJO_PAT`` the dispatcher embedded in the worker prompt) that must be replaced with ```` before the archive lands on disk. The redaction runs on the serialised JSON so it catches the value in any nested location (prompt text, tool input, error string). See :func:`_redact_secret_values` for the safety floor. ``parent_session_id`` / ``subagent_title`` / ``subagent_depth`` are forwarded into the payload to support subagent archiving (see :func:`_archive_subagent_tree`). For top-level (wrapper) archives they should remain ``None``. """ fname_parts = [ _safe_filename_component(started_at_iso), _safe_filename_component(tag), _safe_filename_component(agent), _safe_filename_component(session_id), ] # Subagent archives get a ``sub`` infix so a directory listing # sorts the wrapper first, then its subagents in BFS order # (sub1 before sub2 before sub3...). This makes ``ls -1`` a usable # post-mortem trace tool without needing the telemetry console. if subagent_depth is not None: fname_parts.insert(1, f"sub{subagent_depth:02d}") fname = "__".join(fname_parts) + ".json" path = archive_dir / fname payload = _build_archive_payload( session_id=session_id, agent=agent, tag=tag, status=status, started_at_iso=started_at_iso, wallclock_seconds=wallclock_seconds, messages=messages, state_history=state_history, parent_session_id=parent_session_id, subagent_title=subagent_title, subagent_depth=subagent_depth, ) try: # Render with a default=str so anything urllib hands back that # is not directly JSON-serialisable (rare, but defensive) # falls through to its repr instead of raising. serialised = json.dumps(payload, default=str, indent=2, sort_keys=False) serialised = _redact_secret_values(serialised, redact_values) path.write_text(serialised, encoding="utf-8") except (OSError, PermissionError, TypeError, ValueError) as e: logger.warning("session archive write failed for %s: %s", path, e) return None return path def _walk_subagent_descendants( server_url: str, root_session_id: str, ) -> list[dict[str, Any]] | None: """Walk OpenCode's ``GET /session`` response BFS starting from ``root_session_id`` and return every descendant session record. The returned list excludes ``root_session_id`` itself and is sorted by depth-then-creation-time, which matches the natural order a human reads a trace ("the wrapper called tier-dispatcher first, which called the estimator, which called…"). Each record is the raw session dict from OpenCode with one added field: ``_subagent_depth`` (1-based distance from the root). Returns ``None`` on any transport error or unexpected response shape — distinct from ``[]`` which means "walk succeeded, no subagents existed". The caller's archive flow must continue regardless (we never raise), but Phase 4 telemetry treats the two cases differently: ``None`` -> ``subagent_max_depth = None`` (unknown), ``[]`` -> ``subagent_max_depth = 0`` (measured, flat). A short warning is logged so a missed walk is visible in the dispatcher log. """ try: all_sessions = _request("GET", f"{server_url}/session") except _TRANSPORT_EXC as e: logger.warning("subagent walk failed (transport): %s; archiving root only", e) return None if not isinstance(all_sessions, list): logger.warning( "subagent walk: /session returned unexpected shape " "(%s); archiving root only", type(all_sessions).__name__, ) return None # parent_id -> [child_session_dict, ...] children_by_parent: dict[str, list[dict[str, Any]]] = {} for s in all_sessions: if not isinstance(s, dict): continue pid = s.get("parentID") or s.get("parent_id") or "" if not pid: continue children_by_parent.setdefault(pid, []).append(s) # Sort children by creation time so the trace order is stable. for kids in children_by_parent.values(): kids.sort(key=lambda k: (k.get("time") or {}).get("created", 0) or 0) descendants: list[dict[str, Any]] = [] seen: set[str] = {root_session_id} # BFS frontier: list of (session_id, depth). frontier: list[tuple[str, int]] = [(root_session_id, 0)] while frontier: parent_id, parent_depth = frontier.pop(0) for child in children_by_parent.get(parent_id, []): cid = child.get("id") if not isinstance(cid, str) or cid in seen: continue seen.add(cid) depth = parent_depth + 1 child_with_depth = dict(child) child_with_depth["_subagent_depth"] = depth descendants.append(child_with_depth) frontier.append((cid, depth)) return descendants def _extract_subagent_agent_name(title: str | None) -> str: """Return the ``agent`` name from a subagent session title. Subagent sessions in OpenCode are titled ``" (@ subagent)"`` — see :data:`_SUBAGENT_TITLE_RE`. Returns ``"unknown-subagent"`` if the title is missing or does not match the convention, so the archive still lands on disk with a useful (if generic) filename component. """ if not isinstance(title, str) or not title: return "unknown-subagent" m = _SUBAGENT_TITLE_RE.search(title) if not m: return "unknown-subagent" return m.group(1) def _ms_to_iso(ms: int | float | None) -> str | None: """Convert a Unix-epoch millisecond timestamp to an ISO-8601 UTC string, or return ``None`` if ``ms`` is not a usable number. OpenCode's ``/session`` payload reports ``time.created`` / ``time.updated`` in milliseconds since the epoch. We carry that forward into the archive as an ISO string so the on-disk format matches the wrapper archive's ``started_at`` field. """ if ms is None: return None try: seconds = float(ms) / 1000.0 except (TypeError, ValueError): return None return _dt.datetime.fromtimestamp(seconds, tz=_dt.timezone.utc).isoformat() def _archive_subagent_tree( *, archive_dir: Path, server_url: str, root_session_id: str, root_tag: str, root_started_at_iso: str, redact_values: Iterable[str] | None = None, ) -> tuple[list[Path], int | None]: """Walk every subagent descendant of ``root_session_id``, fetch its message stream, and write an archive file per session. Returns ``(paths, max_depth)``: - ``paths`` — list of archive files written (in BFS order), one per subagent for which we successfully fetched messages. - ``max_depth`` — the deepest BFS depth observed in the walk. ``0`` when the walk succeeded but the wrapper had no subagents. ``None`` when the walk itself failed (transport error / unexpected response shape); the Phase 4 ``subagent_max_depth`` field treats that as "unknown", distinct from "measured flat". Always runs best-effort: a transport error on any single subagent's message fetch logs a warning and continues to the next descendant. The caller MUST invoke this BEFORE deleting the root session — OpenCode may garbage-collect subagent sessions when the root is removed. ``root_tag`` and ``root_started_at_iso`` are inherited so a directory listing groups every session from one dispatcher cycle together. The subagent's own per-session ``time.created`` is captured under the new ``subagent_started_at`` slot for the timeline view in the telemetry console. """ descendants = _walk_subagent_descendants(server_url, root_session_id) if descendants is None: # Walk failed (transport / shape) — we don't know the depth. return [], None if not descendants: # Walk succeeded, no subagents — flat session at depth 0. return [], 0 # Depth from the BFS walk is 1-based; the deepest descendant's # ``_subagent_depth`` is the max-depth we report. ``.get(key, 0)`` # (NOT ``.get(key) or 0``) — the latter would silently coerce a # legitimate depth of 0 to 0 and mask a real bug, even though # _walk_subagent_descendants never emits depth-0 today. Defensive # against future refactors that may use depth-0 for the root. max_depth = max(int(d.get("_subagent_depth", 0)) for d in descendants) paths: list[Path] = [] for desc in descendants: sid = desc.get("id") if not isinstance(sid, str) or not sid: continue title = desc.get("title") or "" agent_name = _extract_subagent_agent_name(title) depth = int(desc.get("_subagent_depth") or 1) time_info = desc.get("time") or {} created_ms = time_info.get("created") updated_ms = time_info.get("updated") sub_started_iso = _ms_to_iso(created_ms) or root_started_at_iso sub_wallclock = 0.0 if isinstance(created_ms, (int, float)) and isinstance( updated_ms, (int, float) ): sub_wallclock = max(0.0, (float(updated_ms) - float(created_ms)) / 1000.0) try: messages = _request("GET", f"{server_url}/session/{sid}/message") except _TRANSPORT_EXC as e: logger.warning( "subagent archive: skipping %s (%s) — transport: %s", sid, agent_name, e, ) continue if not isinstance(messages, list): logger.warning( "subagent archive: skipping %s (%s) — unexpected messages shape %s", sid, agent_name, type(messages).__name__, ) continue # Inherit the parent's tag so a ``ls -1 .dispatcher-logs/sessions`` # listing groups every session from this dispatcher cycle # together. The subagent's own creation time goes into # ``subagent_started_at`` inside the payload. try: written = _archive_session( archive_dir=archive_dir, session_id=sid, agent=agent_name, tag=root_tag, status="subagent", started_at_iso=sub_started_iso, wallclock_seconds=sub_wallclock, messages=messages, state_history=None, redact_values=redact_values, parent_session_id=desc.get("parentID") or root_session_id, subagent_title=title, subagent_depth=depth, ) except Exception as e: # noqa: BLE001 — best-effort, never raise logger.warning( "subagent archive raised for %s (%s): %s", sid, agent_name, e, ) continue if written is not None: paths.append(written) logger.info( "subagent %s (depth=%d agent=%s) archived to %s", sid, depth, agent_name, written, ) return paths, max_depth def _summarize_per_turn(messages: list[dict[str, Any]]) -> list[dict[str, Any]]: """Aggregate per-assistant-turn metrics out of a /session/{id}/message payload. Returns a list of dicts (one per assistant message in chronological order) with:: { "tools": ["skill", ...], "input_tokens": int | None, "output_tokens": int | None, "reasoning_tokens": int | None, "wallclock_seconds": float | None, "completed": bool, } All fields are ``None``-tolerant — the OpenCode payload omits ``time.completed`` for in-flight turns, and may emit ``tokens.input`` of 0 for cached prompts. """ turns: list[dict[str, Any]] = [] for m in messages: if not isinstance(m, dict): continue info = m.get("info") if isinstance(m.get("info"), dict) else m if not isinstance(info, dict): continue if info.get("role") != "assistant": continue parts = m.get("parts") or [] tools = [ p.get("tool", "?") for p in parts if isinstance(p, dict) and p.get("type") == "tool" ] token_info = info.get("tokens") if isinstance(info.get("tokens"), dict) else {} time_info = info.get("time") if isinstance(info.get("time"), dict) else {} created = time_info.get("created") completed = time_info.get("completed") wallclock: float | None if isinstance(created, (int, float)) and isinstance(completed, (int, float)): wallclock = (completed - created) / 1000.0 else: wallclock = None turns.append( { "tools": tools, "input_tokens": token_info.get("input") if isinstance(token_info, dict) else None, "output_tokens": token_info.get("output") if isinstance(token_info, dict) else None, "reasoning_tokens": token_info.get("reasoning") if isinstance(token_info, dict) else None, "wallclock_seconds": wallclock, "completed": completed is not None, } ) return turns def _log_per_turn_metrics( turns: list[dict[str, Any]], *, session_id: str, tag: str ) -> None: """Emit one INFO line per assistant turn so an operator tailing the dispatcher log can see what the worker did per turn at-a-glance. Marks in-flight turns (C3, 2026-05-13): when a session is archived mid-turn (timeout, transport-error), the assistant message hasn't been finalised — token counts may be 0 or stale. The ``in-flight`` marker on the wallclock + ``*`` suffix on token counts make this obvious to an operator without parsing the archive JSON. Without this, the live test 2026-05-13 showed ``tools=[task] input=0tok output=0tok wallclock=in-flight`` and it looked like a real zero-token turn rather than an interrupted one. """ for i, t in enumerate(turns, start=1): wc = t.get("wallclock_seconds") completed = t.get("completed", True) wc_str = f"{wc:.1f}s" if isinstance(wc, (int, float)) else "in-flight" tools = t.get("tools") or [] tools_str = "[" + ",".join(tools) + "]" if tools else "[]" in_tok = t.get("input_tokens") out_tok = t.get("output_tokens") in_str = str(in_tok) if isinstance(in_tok, int) else "?" out_str = str(out_tok) if isinstance(out_tok, int) else "?" # Mark token counts as provisional when the turn was archived # mid-flight — token counts at that point may be 0 (model # never delivered) or partial (some tokens streamed before # the abort). if not completed: in_str = f"{in_str}*" out_str = f"{out_str}*" logger.info( "session %s [%s] turn %d: tools=%s input=%stok output=%stok wallclock=%s", session_id[:32], tag, i, tools_str, in_str, out_str, wc_str, ) # ─── Model resolution ────────────────────────────────────────────────────── # # Each agent is wired to a model via a single-line text file under # ``.opencode/models/.txt``. The dispatcher reads that file at # session-create time and passes the resolved model to ``POST /session`` # so operators can swap models between dispatches without restarting # OpenCode. The same files are referenced from ``opencode.json``'s # ``agent..model`` block via ``{file:...}`` interpolation for the # static (non-dispatcher) resolution path. # # See ``.opencode/models/README.md`` for the full contract. # Module-level constant resolved from the worker module's location, so # tests can override via ``models_dir=`` without monkey-patching. _MODELS_DIR_DEFAULT = Path(__file__).resolve().parent.parent / ".opencode" / "models" @dataclass(frozen=True) class ResolvedModel: """A parsed model assignment ready to send to OpenCode's API. Attributes ---------- provider_id: The provider key as declared in ``opencode.json``'s ``provider.`` block (e.g. ``"anthropic"``, ``"CleverThis-15"``). full_id: The full ``providerID/modelID`` string OpenCode expects in ``POST /session``'s ``model.id`` field (e.g. ``"anthropic/claude-haiku-4-5"``). role_file: The on-disk path the value was read from. Captured for diagnostic logging; not used by the API call. """ provider_id: str full_id: str role_file: Path def _resolve_role_model( agent_name: str, *, models_dir: Path | None = None, ) -> ResolvedModel | None: """Resolve the model assignment for ``agent_name`` via the ``.opencode/models/`` registry. Lookup order: 1. ``/.txt`` (per-agent override). 2. ``/default.txt`` (fallback for the bulk of agents that share a single worker model). Returns ``None`` when neither file exists, or when the file is empty / contains an unparseable line. Callers must treat ``None`` as "skip the runtime override and let OpenCode resolve from its agent registry"; the worker keeps functioning even with the registry absent. File format: exactly one non-empty line whose content matches ``providerID/modelID``. Leading/trailing whitespace and a single trailing newline are stripped. Comments and blank lines are NOT supported — a single ``.txt`` file maps to exactly one model. Parameters ---------- agent_name: OpenCode agent name (e.g. ``"pr-review-worker"``, ``"tier-0"``). Used directly as the filename stem before the ``default.txt`` fallback. models_dir: Override for the registry directory. Defaults to ``/.opencode/models``. Used by tests; production callers leave it unset. """ base = models_dir if models_dir is not None else _MODELS_DIR_DEFAULT # We deliberately do NOT log when both files are missing — that # is the documented "no override; use OpenCode's own resolution" # path and is expected for ad-hoc / interactive callers that # invoke run_session_blocking with an agent whose registry entry # has not been provisioned yet. candidates = [base / f"{agent_name}.txt", base / "default.txt"] for candidate in candidates: if not candidate.is_file(): continue try: raw = candidate.read_text(encoding="utf-8") except OSError as e: logger.warning("model registry: read failed for %s: %s", candidate, e) return None line = raw.strip() if not line: logger.warning( "model registry: %s is empty; cannot resolve model for %s", candidate, agent_name, ) return None # Reject obvious corruption: multi-line files are NOT # supported in Stage 1 (the per-role file holds exactly one # model). A fallback chain — multiple candidates per role — # is reserved for Stage 2's ``.fallbacks.txt`` companion # file. Treat extra lines as a config bug to flag. # # Note: this check runs AFTER ``strip()``, so a single trailing # newline (the normal case for an editor-saved file) has # already been stripped. Any newline remaining means the file # body has at least two non-empty lines, which is the # condition we want to reject. if "\n" in line: logger.warning( "model registry: %s has multiple lines; expected exactly one. " "Treating as malformed and skipping override.", candidate, ) return None # Split on the FIRST slash only — provider names never contain # slashes (validated by OpenCode's provider config), but # model ids occasionally do (e.g. ``DavidAU/Qwen3.6-...`` # variants under CleverThis-7). A naive ``str.split("/")`` # would corrupt those. if "/" not in line: logger.warning( "model registry: %s does not contain '/' separator; " "got %r. Expected 'providerID/modelID'.", candidate, line, ) return None provider_id, _ = line.split("/", 1) if not provider_id: logger.warning( "model registry: %s has empty provider prefix; got %r.", candidate, line, ) return None return ResolvedModel(provider_id=provider_id, full_id=line, role_file=candidate) return None # ─── Public entry point ──────────────────────────────────────────────────── def list_sessions(server_url: str) -> list[dict[str, Any]]: """Return ``GET /session`` from an OpenCode server, or ``[]`` on transport / parse failure. Public wrapper around the internal :func:`_request` helper so other deterministic drivers (e.g. the dispatcher coexistence guard) can enumerate live sessions without reaching into private API. Failures are swallowed and returned as an empty list because the sole consumer today is best-effort startup detection; if your use case needs the distinction between "no sessions" and "server unreachable", call :func:`_request` directly and handle :data:`_TRANSPORT_EXC`. """ try: sessions = _request("GET", f"{server_url}/session") except _TRANSPORT_EXC as e: logger.warning("OpenCode list_sessions failed (%s): %s", server_url, e) return [] if not isinstance(sessions, list): return [] return [s for s in sessions if isinstance(s, dict)] def run_session_blocking( *, server_url: str, agent: str, tag: str, prompt: str, timeout_seconds: int = 900, poll_interval_seconds: float = 2.0, on_poll: Callable[[], None] | None = None, redact_values: Iterable[str] | None = None, ) -> SessionResult: """Spin up a one-shot OpenCode agent session, wait for it to finish, and return a typed :class:`SessionResult`. This is the outcome-agnostic entry point used by drivers whose workers do not emit a fixed ``{"outcome": ...}`` JSON schema (review, implementation). For the conflict driver, see :func:`run_worker_blocking`, which thin-wraps this function and adds the conflict-specific JSON-outcome classification. Lifecycle (matches plan § 6.2): 1. ``POST /session`` with title ``"[{tag}] {agent}"`` → ``session_id``. The ``[TAG]`` prefix is the same convention used by the TypeScript ``session_start.ts`` helper so existing operator tooling (``session_find_by_tag``) finds these sessions. 2. ``POST /session/{id}/prompt_async`` with the agent name + prompt. 3. Poll ``GET /session/status`` every ``poll_interval_seconds`` until the session is idle. The watchdog wallclock at ``timeout_seconds`` short-circuits with ``status="timeout"``; the session is asked to abort and DELETEd in the finally block. 4. ``GET /session/{id}/message`` to read the final assistant turn. 5. Tolerant JSON extraction (last parseable ``{...}`` substring of the last assistant text) → ``parsed_json``. May be ``None`` for workers that do not emit a JSON object — that is not a failure. 6. ``DELETE /session/{id}`` in a finally block. On DELETE failure log a warning and continue — the session leaks until OpenCode server restart. Acceptable for this driver's expected dispatch rate (single-digit per day per plan § 6.2). Any HTTP / network error short-circuits to ``status="transport-error"``; the finally block still attempts session teardown. ``on_poll`` is invoked once per polling iteration (i.e., approximately every ``poll_interval_seconds``) before the session-status GET is issued. The dispatcher uses this to refresh its liveness heartbeat while a long-running worker is in flight; without it, the launcher or systemd unit would treat a 30-minute review session as a hung process and restart the dispatcher mid-cycle, orphaning the OpenCode session and its ``auto/claimed-*`` lock. Exceptions raised from ``on_poll`` are logged and swallowed so a transient heartbeat-write failure (disk full, EROFS, EACCES) cannot mask a successful worker completion. The callback fires: - On every iteration of the polling loop, including the first one after the post-``prompt_async`` warmup sleep. - Before the deadline check, so even a callback that takes longer than the remaining wallclock budget still gets at least one attempt to update liveness state before timeout aborts the session. The callback should be idempotent and side-effect-only (write a file, touch a TCP socket, etc.) — its return value is ignored. ``redact_values`` is an optional collection of secret strings (e.g. the ``FORGEJO_PAT`` the dispatcher embedded in the prompt). Every occurrence of each value is replaced with ```` in the on-disk session archive. The redaction is best-effort — the live OpenCode session message store keeps the value verbatim until the session is DELETEd in this function's finally block, so the redaction protects the persistent record, not the in-flight session. """ started_at = time.monotonic() started_at_iso = _dt.datetime.now(_dt.timezone.utc).isoformat() session_id = "" # Snapshot the redact list once so a generator passed by the caller # is not exhausted on the first invocation. ``None`` stays ``None`` # so the redaction helper short-circuits without iterating. redact_snapshot: tuple[str, ...] | None = ( tuple(redact_values) if redact_values is not None else None ) # Track the terminal status the function will return so the finally # block can include it in the archive payload. The variable is read # in finally even though the return statements assign it just # before — Python keeps locals alive through the unwind. terminal_status: SessionStatus = "transport-error" # Per-poll state-transition history. A bounded list (1000 entries) # so a 30-minute session at 2s polls cannot OOM the dispatcher # process via a runaway log; once full, only transitions are kept, # never repeats. The same data drives both the archive payload and # the at-INFO-level state-transition log lines. state_history: list[dict[str, Any]] = [] _STATE_HISTORY_CAP = 1000 # Closure holder for the SessionResult that THIS invocation will # return. Python's try/finally semantics evaluate the return # expression first, then run finally, then deliver the value to the # caller — so the finally block can still mutate fields on the # already-evaluated dataclass instance, provided we keep a # reference to it. ``_emit_result`` records every return-path # SessionResult into this list; the finally block reads # ``_session_result_holder[-1]`` and sets ``subagent_max_depth`` # from the BFS walk's max-depth signal. Without this plumbing, # ``subagent_max_depth`` lands as ``None`` in Phase 4 telemetry # (the previous behaviour) even though the wrapper has the data. _session_result_holder: list[SessionResult] = [] def _emit_result(sr: SessionResult) -> SessionResult: _session_result_holder.append(sr) return sr def _elapsed() -> float: return time.monotonic() - started_at def _record_state(new_state: str, prev_state: str, seen_busy_flag: bool) -> None: if len(state_history) >= _STATE_HISTORY_CAP: return state_history.append( { "elapsed_s": round(_elapsed(), 3), "state": new_state, "seen_busy": seen_busy_flag, } ) if new_state != prev_state: logger.info( "session %s [%s] state %s -> %s at +%.1fs (seen_busy=%s)", (session_id or "?")[:32], tag, prev_state, new_state, _elapsed(), seen_busy_flag, ) try: title = f"[{tag}] {agent}" # Resolve the model from the .opencode/models/ registry. A # ``None`` return is the documented "no override; let OpenCode # use its own static resolution" path and is NOT a failure; # only true I/O or parse errors get logged as warnings inside # the helper. See ``.opencode/models/README.md`` for the # contract. resolved_model = _resolve_role_model(agent) session_body: dict[str, Any] = {"title": title} if resolved_model is not None: # OpenCode's POST /session accepts ``model = {providerID, # id}`` where ``id`` is the FULL ``providerID/modelID`` # string. We pass it for two reasons: # # 1. OBSERVABILITY: every dispatched session is tagged in # OpenCode's session metadata with the resolved model, # so operators reading the session list see the # intended model regardless of any later cache # interactions. # 2. CONSISTENCY CHECK: if opencode.json's # ``agent..model`` (loaded from # ``{file:./.opencode/models/.txt}`` at OpenCode # startup) has DRIFTED from what's in the registry now # (operator edited a .txt file but did not restart # OpenCode), the session record will at least carry # the registry's intent — even though OpenCode's actual # generation will use its cached value. # # IMPORTANT BEHAVIORAL NOTE: the session-level model does # NOT propagate to ``POST /session/{id}/prompt_async``. # OpenCode re-resolves the agent's model on every prompt # against its cached ``agent..model`` from # opencode.json (loaded at server startup; stale w.r.t. # the live ``.opencode/models/`` registry). Stage 1 # therefore requires an OpenCode restart after editing a # ``.opencode/models/.txt`` file for the change to # actually affect generation. # # We tried adding ``model`` to prompt_async (see git # history of this file) — OpenCode silently dropped those # requests (200 OK, no assistant message produced). # Schema for prompt_async-side runtime override is not # documented; deferred to Stage 2. session_body["model"] = { "providerID": resolved_model.provider_id, "id": resolved_model.full_id, } logger.info( "agent=%s tag=%s: POST /session model hint=%s (from %s) " "— observability only; OpenCode generates from " "opencode.json's agent.%s.model, not this field", agent, tag, resolved_model.full_id, resolved_model.role_file, agent, ) try: session = _request("POST", f"{server_url}/session", body=session_body) except _TRANSPORT_EXC as e: logger.warning("OpenCode session create failed: %s", e) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), raw_response=str(e), error_kind="transport-create-session", ) ) if not isinstance(session, dict) or not session.get("id"): logger.warning( "OpenCode session create returned malformed payload: %r", session ) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), raw_response=json.dumps(session) if session is not None else "", error_kind="transport-create-session-malformed", ) ) session_id = str(session["id"]) logger.info( "OpenCode session %s created (tag=%s, agent=%s)", session_id, tag, agent, ) try: # NOTE: We deliberately do NOT pass a ``model`` override on # prompt_async. Empirical testing against OpenCode 0.x # showed that when the body contains ``model`` here, # OpenCode returns 200 OK but silently never produces an # assistant message (the request is dropped without # diagnostic). Schema docs are unclear; ``id`` and # ``modelID`` were both tried, both broke generation. # # Operator workflow for swapping models is therefore: # 1. Edit ``.opencode/models/.txt`` # 2. Restart OpenCode (so ``{file:./.opencode/models/.txt}`` # in opencode.json's static ``agent..model`` # block re-resolves) # Stage 2 may revisit this once the OpenCode prompt_async # schema for runtime model overrides is documented or a # ``/session/{id}/model`` PATCH endpoint emerges. _request( "POST", f"{server_url}/session/{session_id}/prompt_async", body={ "agent": agent, "parts": [{"type": "text", "text": prompt}], }, ) except _TRANSPORT_EXC as e: logger.warning( "OpenCode prompt_async failed for session %s: %s", session_id, e ) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), session_id=session_id, raw_response=str(e), error_kind="transport-prompt-async", ) ) deadline = started_at + timeout_seconds # Pre-sleep so OpenCode has a chance to register the freshly # dispatched session as ``busy``. Without this the first # status poll fires before the agent starts and the # "missing entry == idle" interpretation in older code paths # would short-circuit the loop with no work done. time.sleep(min(poll_interval_seconds, max(0.0, deadline - time.monotonic()))) seen_busy = False prev_state = "init" warmup_deadline = started_at + _STARTUP_GRACE_SECONDS # Doom-loop probe state — see :data:`_DOOM_LOOP_ABORT_ENV` # and the ``_DOOM_PROBE_*`` constants. ``None`` means the # probe has not yet run for this session (the soft threshold # may not have been crossed); once set, it carries the # wallclock at which the most recent probe ran so subsequent # probes throttle via :data:`_DOOM_PROBE_INTERVAL_SECONDS`. # The flag is read ONCE per session — operators changing it # mid-flight must restart the dispatcher to pick up the # change, same as every other env-var knob in this module. last_doom_probe_at: float | None = None doom_abort_enabled = _is_doom_loop_abort_enabled() doom_soft_threshold = _DOOM_PROBE_SOFT_THRESHOLD_FRACTION * timeout_seconds while True: if on_poll is not None: try: on_poll() except Exception as e: # noqa: BLE001 - callback errors must not kill the worker logger.warning( "on_poll callback raised %s; continuing — heartbeat " "may be stale but the worker session is unaffected", type(e).__name__, exc_info=True, ) # Mid-flight doom-loop probe (G4 harvest 2026-05-15). # Flag-gated, default-OFF. Runs only past # ``_DOOM_PROBE_SOFT_THRESHOLD_FRACTION * timeout_seconds`` # of wallclock and at most every # ``_DOOM_PROBE_INTERVAL_SECONDS`` so the per-cycle # network cost stays bounded. A detector hit aborts the # session with a named ``error_kind`` (e.g. ``doom-loop``) # instead of the opaque ``watchdog-timeout`` the hard # timeout below would produce. When the flag is OFF this # block compiles to a single dict-membership test # (``doom_abort_enabled`` is ``False``) — byte-equivalent # cost to the pre-G4 build. if doom_abort_enabled and _elapsed() >= doom_soft_threshold: _now_probe = time.monotonic() if ( last_doom_probe_at is None or (_now_probe - last_doom_probe_at) >= _DOOM_PROBE_INTERVAL_SECONDS ): last_doom_probe_at = _now_probe finding = _maybe_doom_pattern_finding( server_url=server_url, session_id=session_id, ) if finding is not None: logger.warning( "OpenCode worker doom-loop probe tripped " "after %.0f s (session=%s, kind=%s): %s", _elapsed(), session_id, finding.error_kind, finding.reason, ) try: _request( "POST", f"{server_url}/session/{session_id}/abort", ) except _TRANSPORT_EXC as e: logger.warning( "OpenCode abort failed for session %s: %s", session_id, e, ) terminal_status = "timeout" return _emit_result( SessionResult( status="timeout", wallclock_seconds=_elapsed(), session_id=session_id, error_kind=finding.error_kind, raw_response=finding.reason, ) ) if time.monotonic() >= deadline: logger.warning( "OpenCode worker timed out after %.0f s (session=%s)", timeout_seconds, session_id, ) try: _request( "POST", f"{server_url}/session/{session_id}/abort", ) except _TRANSPORT_EXC as e: logger.warning( "OpenCode abort failed for session %s: %s", session_id, e, ) terminal_status = "timeout" return _emit_result( SessionResult( status="timeout", wallclock_seconds=_elapsed(), session_id=session_id, error_kind="watchdog-timeout", ) ) try: state = _session_busy_state(server_url, session_id) except _TRANSPORT_EXC as e: logger.warning( "OpenCode status poll failed for session %s after %d retries: %s", session_id, _READ_RETRY_ATTEMPTS, e, ) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), session_id=session_id, raw_response=str(e), error_kind="transport-poll-status", ) ) _record_state(state, prev_state, seen_busy) prev_state = state if state == "busy": seen_busy = True elif state == "idle" and seen_busy: # Genuine terminal idle: the session was observed busy # at some point and is now idle. Done. break elif state == "unknown" and seen_busy: # The session entry has disappeared from the # /session/status map even though we previously # observed ``busy``. This is ambiguous: it can mean # either the session is between assistant turns # (transient blip — OpenCode drops the entry while it # spins up the next message and re-adds it as ``busy``) # OR the session has truly completed and OpenCode has # permanently removed it from the status map. # # The first version of this branch (2026-05-07 morning # commit) treated every ``unknown`` after busy as a # transient blip and kept polling forever, which made # the dispatcher hang for the full watchdog timeout # whenever a session terminated normally — exactly the # opposite failure mode of the warmup-deadline bug it # was meant to fix. The disambiguator below uses # the per-message ``time.completed`` field as the # authoritative signal: if the last assistant message # is still in-flight (completed=None) we keep polling; # if every assistant message is finished we break. in_flight = _session_assistant_in_flight(server_url, session_id) if in_flight is False: # Authoritative terminal: all assistant turns done. logger.info( "session %s [%s] terminated cleanly (all assistant " "turns complete; status map cleared) at +%.1fs", session_id[:32], tag, _elapsed(), ) break # ``in_flight is True`` (transient between turns) or # ``in_flight is None`` (transport error reading # /message — be conservative, assume still in-flight). # Continue polling. else: # ``state == "idle"`` and we've never seen busy # (pre-launch race: ``prompt_async`` accepted but the # status map briefly lists the session as idle before # the agent boots), or ``state == "unknown"`` and we # have never seen busy (the prompt never reached the # agent at all). Both states are acceptable transiently # but must resolve to busy before the warmup deadline # or we treat this as a transport error rather than # silently reporting an empty completion. The empty- # completion footgun is exactly what bit # dispatch_review's first cycle on 2026-05-07 — a # reviewer "completed" in 0.025 s of wallclock with # raw_response="". if time.monotonic() >= warmup_deadline: logger.warning( "OpenCode session %s never transitioned to busy " "within %.0fs of prompt_async — treating as " "transport-error (agent dispatch failed)", session_id, _STARTUP_GRACE_SECONDS, ) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), session_id=session_id, raw_response="session never became busy", error_kind="startup-grace-elapsed", ) ) time.sleep(poll_interval_seconds) try: messages = _request_read(f"{server_url}/session/{session_id}/message") except _TRANSPORT_EXC as e: logger.warning( "OpenCode message fetch failed for session %s after %d retries: %s", session_id, _READ_RETRY_ATTEMPTS, e, ) terminal_status = "transport-error" return _emit_result( SessionResult( status="transport-error", wallclock_seconds=_elapsed(), session_id=session_id, raw_response=str(e), error_kind="transport-message-fetch", ) ) if not isinstance(messages, list): messages = [] raw_text = _last_assistant_text(messages) parsed = _extract_last_json_object(raw_text) terminal_status = "completed" return _emit_result( SessionResult( status="completed", wallclock_seconds=_elapsed(), session_id=session_id, raw_response=raw_text, parsed_json=parsed, ) ) finally: # Best-effort archive + per-turn-metrics emit BEFORE the # session is deleted. We swallow every error class here: an # archive failure must NEVER mask a successful worker # completion or the dispatcher's understanding of why a run # failed. Disable archiving entirely via # OPENCODE_WORKER_ARCHIVE_DISABLED=1. if session_id: archive_dir = _resolve_archive_dir() archive_messages: list[dict[str, Any]] | None = None if archive_dir is not None: try: fetched = _request( "GET", f"{server_url}/session/{session_id}/message" ) except _TRANSPORT_EXC: fetched = None if isinstance(fetched, list): archive_messages = fetched if archive_messages is not None: try: turns = _summarize_per_turn(archive_messages) _log_per_turn_metrics(turns, session_id=session_id, tag=tag) except Exception as e: # noqa: BLE001 — never mask the worker outcome logger.warning( "per-turn metrics emit failed for %s: %s", session_id, e, ) if archive_dir is not None: try: archive_path = _archive_session( archive_dir=archive_dir, session_id=session_id, agent=agent, tag=tag, status=terminal_status, started_at_iso=started_at_iso, wallclock_seconds=_elapsed(), messages=archive_messages, state_history=state_history, redact_values=redact_snapshot, ) if archive_path is not None: logger.info( "session %s archived to %s", session_id, archive_path, ) except Exception as e: # noqa: BLE001 — defensive belt-and-braces logger.warning( "session archive raised for %s: %s", session_id, e, ) # Subagent tree archive — runs AFTER the root archive # but BEFORE the DELETE below, so the dispatcher # captures the complete ``task`` tool chain even if # OpenCode garbage-collects subagent sessions on # root deletion. Best-effort: any failure logs a # warning and lets the DELETE proceed. # # Phase 4 telemetry: the walk also returns the # deepest BFS distance observed in the tree. We # mutate it onto the SessionResult that the return # expression evaluated just before finally so the # dispatcher's post-session action can read it. # See the ``_session_result_holder`` block at the # top of this function for the mechanism. try: _, subagent_max_depth = _archive_subagent_tree( archive_dir=archive_dir, server_url=server_url, root_session_id=session_id, root_tag=tag, root_started_at_iso=started_at_iso, redact_values=redact_snapshot, ) if _session_result_holder: _session_result_holder[ -1 ].subagent_max_depth = subagent_max_depth except Exception as e: # noqa: BLE001 — defensive belt-and-braces logger.warning( "subagent tree archive raised for root %s: %s", session_id, e, ) try: _request("DELETE", f"{server_url}/session/{session_id}") except _TRANSPORT_EXC as e: logger.warning( "OpenCode session DELETE failed for %s: %s — leaking " "session, will be reclaimed at server restart", session_id, e, ) def run_worker_blocking( *, server_url: str, agent: str, tag: str, prompt: str, timeout_seconds: int = 900, poll_interval_seconds: float = 2.0, on_poll: Callable[[], None] | None = None, ) -> WorkerResult: """Run a one-shot OpenCode agent and classify the result against the conflict-driver outcome schema. Thin wrapper around :func:`run_session_blocking` that interprets the last parseable JSON object (per plan § 6.2 step 5) into one of: - ``resolved`` / ``unresolvable`` / ``rebase-failed`` from the worker's structured exit, or - ``rebase-failed`` if the session completed but the JSON contract was violated (no JSON, missing ``outcome`` key, or unrecognized value), or - ``timeout`` / ``transport-error`` propagated from the session lifecycle. Drivers whose workers do not emit this schema (review, implementer) should call :func:`run_session_blocking` directly and treat ``status == "completed"`` as success. The ``on_poll`` callback is forwarded verbatim to :func:`run_session_blocking`; see that function's docstring for the contract. Used by ``conflict_drive.py`` to refresh its heartbeat while a long-running rebase + conflict-resolve session is in flight. """ session = run_session_blocking( server_url=server_url, agent=agent, tag=tag, prompt=prompt, timeout_seconds=timeout_seconds, poll_interval_seconds=poll_interval_seconds, on_poll=on_poll, ) if session.status == "timeout": return WorkerResult( outcome="timeout", wallclock_seconds=session.wallclock_seconds, session_id=session.session_id, raw_response=session.raw_response, error_kind=session.error_kind, ) if session.status == "transport-error": return WorkerResult( outcome="transport-error", wallclock_seconds=session.wallclock_seconds, session_id=session.session_id, raw_response=session.raw_response, error_kind=session.error_kind, ) parsed = session.parsed_json or {} outcome_str = parsed.get("outcome") if outcome_str not in ("resolved", "unresolvable", "rebase-failed"): logger.warning( "OpenCode worker returned unexpected outcome %r — classifying " "as rebase-failed (session=%s)", outcome_str, session.session_id, ) outcome_str = "rebase-failed" files_touched = parsed.get("files_touched") or [] if not isinstance(files_touched, list): files_touched = [] return WorkerResult( outcome=outcome_str, # type: ignore[arg-type] files_touched=[str(f) for f in files_touched], wallclock_seconds=session.wallclock_seconds, session_id=session.session_id, raw_response=session.raw_response, )