Files
cleveragents-core/tools/_pipeline_cache.py
T
drew 016b348117 feat(controller): grooming gate (Phase 0 + Phase 1 worker-shape dispatch)
Phase 0 (foundation):
- Cause enum (controller_events.cause) for action attribution
- Schema: grooming_decisions audit table; workflows gains
  grooming_evaluated_at + deferred_reason + deferred_at +
  deferred_target_workflow_id; pulls gains touched_files
- audit_comments: CLOSE / DEFER templates + render_comment_template
- forgejo_writes: close_issue + defer_issue 5-step crash-safe protocol
  (fingerprint dedup, error matrix, dry-run)
- patch_pr_state callback in forgejo_http
- grooming_config: 22-env-var frozen-dataclass config + log_effective
- pulls.touched_files cache extension (_pipeline_cache.py schema v8)
- reaper.reap_grooming_decisions audit-retention sweep
- reconciliation RESUME guard (deferred_reason)

Phase 1 (worker-queue shape, 2026-05-25):
- New state: GROOMING. New events: grooming_started, groom_verdict_
  {proceed,defer,close}. 5 new transitions; all invariants still clean
- GroomingInputV1 + GroomingOutputV1 Pydantic contracts
- outcomes._map_grooming_outcome routes verdicts to state-machine events
- prefetch.build_grooming_stage_b_input + list_open_prs callback
- scheduler GROOMING -> grooming_stage_b role
- promote: cfg-gated DISCOVERED -> GROOMING when CONTROLLER_GROOMING_
  ENABLED=true; issues skip grooming
- forgejo_writes decomposed: close_act/defer_act (Forgejo writes only;
  state-machine already transitioned) + close_decide_and_act/
  defer_decide_and_act (Phase 0 callers); _apply_workflow_transition
  is underscore-private
- grooming.py library: tokenization, suspicion scoring (Jaccard +
  weighted overlap), deterministic checks, action -> verdict mapping
- mcp/grooming_builder.py: 14-tool FastMCP server emits GroomingOutputV1
- .opencode/agents/grooming-stage-b.md: duplicate-detection agent
  prompt (claude-haiku-4-5)
- grooming_side_effects.run_grooming_side_effects_tick: per-state tick
  performs Forgejo writes after groom_verdict_{defer,close} fires.
  Filters on event_type='transition' + payload.event (centralizes the
  convention pending Phase 2's latest_transition_event helper)
- GroomingCallbacks frozen dataclass; loop.py + __main__.py wired

Worker role registry (single source of truth):
- worker/roles.py: WORKER_ROLES + WorkerRoleSpec + default_roles_csv
  + output_filename_for. agent_runner.ROLE_TO_MCP_MODULE / ROLE_TO_
  OUTPUT_MODEL derive from it; opencode_session.agent_name_for reads
  it for flat cases; all 6 prompt builders use output_filename_for;
  worker --roles default = default_roles_csv(); launcher script
  derives --roles via shell substitution. Cross-site invariant test
  enforces alignment across 5 sites + opencode.json MCP registry.

Phase 0 silent-bug fix:
- reconciliation.py RESUME guard SELECT now includes deferred_reason
  (was missing since Phase 0; guard was a silent no-op). Tightened
  from getattr to attribute access to fail fast on future omissions.

Tests (1456 total, +91 grooming-specific):
- test_grooming_phase0.py: 34 tests (orchestrator matrix, crash
  recovery, idempotency, dry-run)
- test_grooming_phase1.py: 60 tests (library, contracts, state
  machine, outcomes, scheduler, promote, prefetch, act-variants
  with signature parity, side-effect tick incl. natural-idempotency
  + executed-flag-skip + verdict-mismatch + reconciliation RESUME)
- test_mcp_builders.py TestGroomingBuilder: 29 tests (happy paths
  + 22 validation rules + Pydantic round-trip + master-tick-read-
  path companion)
- test_worker_agent_runner.py TestRoleMaps: cross-role wiring
  alignment + agent-prompt-vs-worker-fallback filename contract +
  inspect.signature equality (close_act/defer_act vs
  close_issue/defer_issue)
- test_state_machine.py: transition count 51 -> 56 +
  events_from_grooming

Live-validated end-to-end on 4 staged sentinel PRs (#55-#58) in
dry_run: agent emits verdicts via MCP, state-machine transitions
fire, side-effect tick writes audit row, deferred_reason gates
reconciliation RESUME correctly.

Deferred refinements + Phase 2 prerequisite (latest_transition_event
helper) tracked in .drew/regressions-plan.md "Phase 1 follow-up
backlog".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:05:29 -04:00

1826 lines
78 KiB
Python

"""Pipeline SQLite cache for the auto-agents stack.
Started life as a Forgejo-only cache (commits / pulls / reachability)
— hence the on-disk filename ``forgejo.sqlite`` and the
``FORGEJO_OWNER`` / ``FORGEJO_REPO`` / ``FORGEJO_PAT`` env vars that
gate sync targets. Grew tables for pipeline telemetry (merge_cycle,
ci_gate_events, llm_activity, conflict_drive_cycles,
dispatch_*_cycles) without splitting them out. The module and class
were renamed to reflect the broader scope (``_forgejo_cache`` →
``_pipeline_cache``, ``ForgejoCache`` → ``PipelineCache``); the
on-disk filename and FORGEJO_* env vars are kept for compatibility
with existing cache files and operator muscle memory.
Used by ``count-master-merges.py``, ``pr-stats.py``, the dispatchers,
the warmer, the OpenCode archive scraper, and the telemetry console
— anything that needs fast queries over Forgejo or pipeline state.
Design
------
- **Single SQLite file** at ``tools/.cache/forgejo.sqlite`` (gitignored).
- **Forgejo data tables**:
``commits`` - master branch, keyed by full SHA, append-only
``pulls`` - every PR we've seen, keyed by PR number
``reachability`` - cached compare/master...sha results
- **Pipeline telemetry tables** (1B / 1.5 / 2):
``merge_cycle``, ``ci_gate_events``, ``llm_activity``,
``conflict_drive_cycles``, ``dispatch_review_cycles``,
``dispatch_implementer_cycles``
Plus ``sync_meta`` for cursors and ``schema_version`` for migrations.
- **Every row stores the raw API JSON** in a ``raw`` column so new fields can
be surfaced without a cache rebuild; parsed columns are indexed
projections for efficient queries.
- **Delta sync** walks the commits and PR endpoints from page 1 (newest
first) and stops as soon as the current page is entirely already-cached
(commits) or its max ``updated_at`` is older than the last sync
(PRs). Initial seeding is just "delta sync from an empty cache" and
continues until the API pages are exhausted.
- **PR enrichment** (``merged_by``) is lazy: on read, any terminal-state PR
that lacks ``merged_by_login`` but needs it gets one PR-detail fetch,
cached permanently.
- **Concurrency**: ``fcntl.flock`` on ``forgejo.sqlite.lock`` during syncs.
SQLite handles concurrent reads natively.
- **Force-push detection**: sync_meta tracks ``master_head_sha``; if the
previous HEAD is not seen anywhere in the first N pages of current master,
we emit a warning. The user can pass ``--full`` to rebuild.
Public API (used by the scripts)
--------------------------------
::
cache = PipelineCache.open()
cache.sync(token) # delta-sync forward
commits = cache.commits_on_master(start, end) # list of raw commit dicts
closed = cache.pulls_closed_in_window(start, end)
merged = cache.pulls_merged_in_window(start, end)
opened = cache.pulls_opened_in_window(start, end)
sha_map = cache.sha_to_pr_map(start, end)
reach = cache.is_sha_reachable_from_master(sha, token)
detail = cache.enrich_pr(number, token)
All list/map queries return the raw PR/commit dicts exactly as the Forgejo
API returned them (deserialised from the ``raw`` column), so callers that
already expect raw API shape need no changes.
Tables NOT owned here
---------------------
``_mode_marker`` (one-row table: ``mode TEXT PRIMARY KEY, stamped_at TEXT``)
is created and maintained outside this module by
``tools/controller/deploy/validate_db_mode.py`` — the controller launcher
stamps it on first use and asserts it on subsequent launches so a fork-mode
cache and a prod-mode cache can never be confused. Ignored by every query
in this module; the schema-rebuild path leaves it alone.
Cache file path: see ``tools/_cache_path.py`` for the single source of
truth that the telemetry server, the scrapers, and the launcher all agree
on (``forgejo.sqlite`` for the canonical org pair, partitioned by
``forgejo.<owner>.<repo>.sqlite`` otherwise).
"""
from __future__ import annotations
import fcntl
import json
import os
import re
import sqlite3
import sys
import time
import urllib.error
import urllib.request
from datetime import datetime, timedelta, timezone
from pathlib import Path
from typing import Any, Iterable, Sequence
# ─── Config ─────────────────────────────────────────────────────────────────
# Env-driven so fork-mode reporting (`FORGEJO_OWNER=drew` etc.) reads
# from the fork's data instead of the canonical repo. The defaults
# preserve the original behaviour: when run without env vars in a
# canonical-mode shell, the cache, every API call, and every cache file
# resolve exactly as before.
REPO_OWNER = os.environ.get("FORGEJO_OWNER", "cleveragents")
REPO_NAME = os.environ.get("FORGEJO_REPO", "cleveragents-core")
API_BASE = os.environ.get(
"FORGEJO_API_BASE", "https://git.cleverthis.com/api/v1"
).rstrip("/")
REQUEST_TIMEOUT_SEC = 30
REQUEST_RETRIES = 3
RETRY_BACKOFF_SEC = 2.0
DEFAULT_CACHE_DIR = Path(__file__).resolve().parent / ".cache"
# Cache path is partitioned per-(owner, repo) so a fork-mode run never
# clobbers the canonical cache (and vice versa). The default
# ``cleveragents/cleveragents-core`` keeps the historical filename
# ``forgejo.sqlite`` for backward compatibility with existing checkouts.
if REPO_OWNER == "cleveragents" and REPO_NAME == "cleveragents-core":
DEFAULT_CACHE_PATH = DEFAULT_CACHE_DIR / "forgejo.sqlite"
else:
_safe_owner = re.sub(r"[^a-zA-Z0-9._-]+", "-", REPO_OWNER)
_safe_repo = re.sub(r"[^a-zA-Z0-9._-]+", "-", REPO_NAME)
DEFAULT_CACHE_PATH = (
DEFAULT_CACHE_DIR / f"forgejo.{_safe_owner}.{_safe_repo}.sqlite"
)
# Schema versions:
# 1 — original (commits, pulls, reachability, sync_meta)
# 2 — Tier 1B telemetry: merge_cycle, ci_gate_events, llm_activity
# 3 — Tier 1.5 telemetry: conflict_drive_cycles
# 4 — Tier 2 telemetry: dispatch_review_cycles, dispatch_implementer_cycles
# 5 — In-flight row visibility (2026-05-07): drop the NOT NULL
# constraint on ``ended_at`` for ``dispatch_review_cycles`` and
# ``dispatch_implementer_cycles`` so the dispatchers can insert a
# row at cycle start (``ended_at IS NULL``) and update it on
# completion. The migration recreates each affected table only
# when its ``ended_at`` column still carries the old NOT NULL
# constraint; pre-existing rows are preserved verbatim.
# 6 — llm_activity ingest keys (2026-05-17): add ``session_id``,
# ``message_id``, ``provider``, ``parent_session_id``,
# ``subagent_depth`` so the OpenCode archive scraper
# (``tools/llm_activity_scraper.py``) can emit one row per
# assistant turn with stable dedup. The partial unique index on
# ``message_id`` makes re-scrapes idempotent — running the scraper
# N times produces the same row count as running it once.
# 7 — llm_activity prompt_hash (2026-05-17): add ``prompt_hash`` so
# we can measure how often the SAME prompt drives multiple LLM
# calls (i.e. whether a response cache would have leverage).
# Populated by the scraper as SHA-256 of the session's first
# user-message text, applied to every assistant turn of that
# session. No UNIQUE constraint — duplicates are exactly the
# signal we want to count.
SCHEMA_VERSION = 8
# ─── Tier 2 dispatcher cycle-table schema (shared, single source of truth) ──
#
# The two dispatcher cycle tables (``dispatch_review_cycles``,
# ``dispatch_implementer_cycles``) share the same shape and the same
# v4→v5 migration. Both ``PipelineCache._migrate_to_v5_in_flight_rows``
# and ``_dispatch_runtime.ensure_cycle_table`` need this DDL — the
# former runs when the cache file is opened through ``PipelineCache``,
# the latter when the dispatcher writes a cycle row via raw
# ``sqlite3.connect`` and may run *before* anyone opens the cache via
# the high-level wrapper. Defining the DDL twice was a real drift risk
# (a future column add to the cycle table would have to be made in two
# places, in lockstep). Instead, both sites import from here.
DISPATCH_CYCLE_TABLES: tuple[str, ...] = (
"dispatch_review_cycles",
"dispatch_implementer_cycles",
)
def _dispatch_cycle_create_sql(table_name: str) -> str:
"""``CREATE TABLE IF NOT EXISTS`` for a v5 dispatch cycle table.
``ended_at`` is nullable so the dispatcher can insert a cycle row
at start (``ended_at IS NULL`` ⇒ in flight) and update it on
completion. The trailing UNIQUE INDEX on ``cycle_id`` is created
separately by :func:`_dispatch_cycle_index_sqls`.
"""
if table_name not in DISPATCH_CYCLE_TABLES:
raise ValueError(f"unexpected dispatch table: {table_name}")
return f"""
CREATE TABLE IF NOT EXISTS {table_name} (
id INTEGER PRIMARY KEY AUTOINCREMENT,
cycle_id TEXT NOT NULL,
started_at TEXT NOT NULL,
ended_at TEXT,
driver TEXT NOT NULL,
candidates_count INTEGER NOT NULL,
claims_acquired INTEGER NOT NULL,
swept_count INTEGER NOT NULL,
processed_count INTEGER NOT NULL,
terminal_state TEXT,
worker_outcome TEXT,
session_id TEXT,
worker_wallclock_seconds REAL,
raw TEXT NOT NULL
)
"""
def _dispatch_cycle_index_sqls(table_name: str) -> tuple[str, ...]:
"""Index DDL for a v5 dispatch cycle table.
Three indexes:
- ``idx_<t>_started_at`` — common ORDER BY for the telemetry tables.
- ``idx_<t>_terminal`` — used when filtering by terminal_state.
- ``uniq_<t>_cycle_id`` — UNIQUE: prevents accidental double-INSERT
from a re-entrant ``begin_cycle``; required for the in-flight
visibility contract.
"""
if table_name not in DISPATCH_CYCLE_TABLES:
raise ValueError(f"unexpected dispatch table: {table_name}")
return (
f"CREATE INDEX IF NOT EXISTS idx_{table_name}_started_at "
f"ON {table_name}(started_at)",
f"CREATE INDEX IF NOT EXISTS idx_{table_name}_terminal "
f"ON {table_name}(terminal_state)",
f"CREATE UNIQUE INDEX IF NOT EXISTS uniq_{table_name}_cycle_id "
f"ON {table_name}(cycle_id)",
)
def migrate_dispatch_cycle_table_to_v5(
conn: sqlite3.Connection, table_name: str
) -> bool:
"""Rebuild ``table_name`` with the v5 schema (``ended_at`` nullable).
Returns ``True`` if the rebuild ran (the table existed with a v4
``ended_at NOT NULL`` constraint) or ``False`` if it was a no-op
(table missing or already v5). Pre-existing rows are preserved
verbatim and indexes are recreated as part of the rebuild.
Caller is responsible for transaction boundaries (the function
uses ``executescript`` which auto-commits, but does not start an
explicit transaction). Designed to be safely runnable from both
a ``PipelineCache`` migration step and a raw
``sqlite3.connect`` ``ensure_cycle_table`` invocation.
"""
if table_name not in DISPATCH_CYCLE_TABLES:
raise ValueError(f"unexpected dispatch table: {table_name}")
cols = list(conn.execute(f"PRAGMA table_info({table_name})"))
if not cols:
return False
ended_at_col = next((c for c in cols if c["name"] == "ended_at"), None)
if ended_at_col is None or ended_at_col["notnull"] == 0:
return False
# Inside the rebuild we use plain ``CREATE TABLE`` (the IF NOT
# EXISTS form would also work because the table was just renamed,
# but a plain ``CREATE TABLE`` is more honest about intent).
create_inner = _dispatch_cycle_create_sql(table_name).replace(
"CREATE TABLE IF NOT EXISTS", "CREATE TABLE", 1
)
indexes = ";\n ".join(_dispatch_cycle_index_sqls(table_name))
conn.executescript(
f"""
ALTER TABLE {table_name} RENAME TO {table_name}_v4;
{create_inner};
INSERT INTO {table_name} SELECT * FROM {table_name}_v4;
DROP TABLE {table_name}_v4;
{indexes};
"""
)
return True
def ensure_dispatch_cycle_schema(conn: sqlite3.Connection, table_name: str) -> None:
"""Create the v5 dispatch cycle table + indexes, migrating from v4
if needed. Idempotent and self-healing.
Used by both ``_dispatch_runtime.ensure_cycle_table`` (raw
sqlite3.connect path) and ``PipelineCache._migrate_to_v5_in_flight_rows``
(high-level wrapper path) so the schema can never drift between
the two write paths.
"""
if table_name not in DISPATCH_CYCLE_TABLES:
raise ValueError(f"unexpected dispatch table: {table_name}")
conn.execute(_dispatch_cycle_create_sql(table_name))
migrate_dispatch_cycle_table_to_v5(conn, table_name)
for idx_sql in _dispatch_cycle_index_sqls(table_name):
conn.execute(idx_sql)
# Safety ceiling on pagination during a single sync pass.
SYNC_PAGE_HARD_LIMIT = 4000 # 4000 * 50 = 200,000 rows; far above anything real.
# How many consecutive pages with no new rows before we declare the sync done.
SYNC_EMPTY_PAGE_TOLERANCE = 2
# ─── HTTP helper (mirrors count-master-merges.py) ──────────────────────────
def api_get(url: str, token: str) -> Any:
"""GET a Forgejo API URL with retries on transient network errors."""
req = urllib.request.Request(url, headers={"Authorization": f"token {token}"})
last_err: Exception | None = None
for attempt in range(1, REQUEST_RETRIES + 1):
try:
with urllib.request.urlopen(req, timeout=REQUEST_TIMEOUT_SEC) as resp:
return json.loads(resp.read())
except urllib.error.HTTPError:
# HTTP errors are authoritative: propagate so caller can decide
# (e.g. 404 on compare endpoint means "not reachable").
raise
except (urllib.error.URLError, TimeoutError, OSError) as e:
last_err = e
if attempt < REQUEST_RETRIES:
time.sleep(RETRY_BACKOFF_SEC * attempt)
continue
raise RuntimeError(
f"cannot reach {url} after {REQUEST_RETRIES} attempts: {last_err}"
)
# ─── Cache ─────────────────────────────────────────────────────────────────
SCHEMA_SQL = """
CREATE TABLE IF NOT EXISTS schema_version (
version INTEGER PRIMARY KEY
);
CREATE TABLE IF NOT EXISTS commits (
sha TEXT PRIMARY KEY,
committer_date TEXT NOT NULL,
author_name TEXT,
committer_name TEXT,
message TEXT,
parents TEXT,
raw TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_commits_committer_date ON commits(committer_date);
CREATE TABLE IF NOT EXISTS pulls (
number INTEGER PRIMARY KEY,
state TEXT,
merged INTEGER,
created_at TEXT,
updated_at TEXT,
closed_at TEXT,
merged_at TEXT,
base_ref TEXT,
head_sha TEXT,
merge_commit_sha TEXT,
user_login TEXT,
merged_by_login TEXT,
additions INTEGER,
deletions INTEGER,
changed_files INTEGER,
labels TEXT,
has_detail INTEGER NOT NULL DEFAULT 0,
-- Phase 1 grooming plan (v8): JSON array of file paths touched by the PR.
-- Populated via GET /pulls/{n}/files alongside the detail-fetch flow.
-- Used by Stage A pre-filter's suspicion-score file-overlap signal.
touched_files TEXT,
raw TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_pulls_closed_at ON pulls(closed_at);
CREATE INDEX IF NOT EXISTS idx_pulls_merged_at ON pulls(merged_at);
CREATE INDEX IF NOT EXISTS idx_pulls_created_at ON pulls(created_at);
CREATE INDEX IF NOT EXISTS idx_pulls_updated_at ON pulls(updated_at);
CREATE INDEX IF NOT EXISTS idx_pulls_head_sha ON pulls(head_sha);
CREATE INDEX IF NOT EXISTS idx_pulls_merge_commit_sha ON pulls(merge_commit_sha);
CREATE INDEX IF NOT EXISTS idx_pulls_state_merged ON pulls(state, merged);
-- ─── PR classification cache (2026-05-16) ─────────────────────────────────
--
-- Used by ``_pr_classification_cache.refresh_then_filter`` to skip per-PR
-- review/CI/commit fetches on PRs whose ``updated_at`` hasn't advanced
-- since the last cycle. Replaces the per-cycle ``list_prs_*.ts`` subprocess
-- calls in ``dispatch_review.run_outer_loop`` — see
-- ``.drew/planning/fix list_prs_by_filter.md`` for the full plan.
--
-- The 8 classification axes are exactly the inputs the 5 reviewer-filter
-- predicates need:
-- ci_status — 'passing' | 'failing' | 'pending' | 'unknown'
-- approvals_count — count of non-dismissed APPROVE reviews
-- has_active_request_changes — 0/1 (any non-dismissed REQUEST_CHANGES)
-- has_unaddressed_request_changes — 0/1 (any RC not followed by a new commit)
-- is_claimed — 0/1 (any auto/claimed-* label present)
-- is_mergeable — 0/1/NULL (NULL = Forgejo still computing)
-- stale_state — 'not_stale' | 'stale_no_conflicts' | 'stale_with_conflicts' | 'stale_unknown' | 'compute_error'
-- Plus bookkeeping: head_sha + updated_at (for delta-fresh check),
-- last_checked_at (for TTL check), labels_json (for any future filter
-- re-classification without re-fetch), classification_schema_version (so a
-- future axis addition can invalidate stale rows en masse).
--
-- Cache hit predicate:
-- cached.head_sha == pr.head.sha
-- AND cached.updated_at >= pr.updated_at
-- AND (now - cached.last_checked_at) < ttl_seconds
-- AND cached.classification_schema_version == current_version
CREATE TABLE IF NOT EXISTS pr_classifications (
pr_number INTEGER PRIMARY KEY,
head_sha TEXT NOT NULL,
updated_at TEXT NOT NULL,
last_checked_at TEXT NOT NULL,
ci_status TEXT NOT NULL,
approvals_count INTEGER NOT NULL,
has_active_request_changes INTEGER NOT NULL,
has_unaddressed_request_changes INTEGER NOT NULL,
is_claimed INTEGER NOT NULL,
is_mergeable INTEGER,
stale_state TEXT NOT NULL,
labels_json TEXT NOT NULL,
classification_schema_version INTEGER NOT NULL DEFAULT 1
);
CREATE INDEX IF NOT EXISTS idx_pr_classifications_updated_at
ON pr_classifications(updated_at);
CREATE INDEX IF NOT EXISTS idx_pr_classifications_last_checked_at
ON pr_classifications(last_checked_at);
CREATE TABLE IF NOT EXISTS reachability (
merge_sha TEXT PRIMARY KEY,
reachable INTEGER NOT NULL,
checked_at TEXT NOT NULL,
master_head_sha TEXT
);
CREATE TABLE IF NOT EXISTS sync_meta (
key TEXT PRIMARY KEY,
value TEXT
);
-- ─── Tier 1B telemetry tables ────────────────────────────────────────────
--
-- merge_cycle: one row per merge driver attempt (single-PR or train-merge).
--
-- Phase columns are wall-clock seconds; NULL means the phase did not run
-- for this terminal_state. Action columns let the canvas surface "what
-- the driver actually did" without having to grep logs.
CREATE TABLE IF NOT EXISTS merge_cycle (
id INTEGER PRIMARY KEY AUTOINCREMENT,
started_at TEXT NOT NULL,
ended_at TEXT,
pr_numbers TEXT NOT NULL, -- JSON array of PR numbers
train_id TEXT, -- e.g. "auto/train/<ts>" or NULL for single-PR
bisect_depth INTEGER NOT NULL DEFAULT 0,
rebase_seconds REAL,
ci_seconds REAL,
merge_seconds REAL,
total_seconds REAL,
terminal_state TEXT NOT NULL, -- merged | merged_in_train | rebase-conflict-vs-master
-- | ci-fail-on-rebased-sha | ci-timeout | bisect-budget-exhausted
-- | restart-budget-exhausted | umbrella-pr-creation-failed
-- | merge-409-race | crashed
action_taken TEXT, -- short slug e.g. "merged"|"released"|"escalated"
action_detail TEXT, -- free-form JSON or string with details
last_released_at TEXT, -- when the PR(s) was released back to the pool
flake_rate REAL, -- forward-fill from ci_gate_events for this cycle, if computed
raw TEXT -- optional JSON blob with extra fields for debug
);
CREATE INDEX IF NOT EXISTS idx_merge_cycle_started_at ON merge_cycle(started_at);
CREATE INDEX IF NOT EXISTS idx_merge_cycle_ended_at ON merge_cycle(ended_at);
CREATE INDEX IF NOT EXISTS idx_merge_cycle_terminal ON merge_cycle(terminal_state);
CREATE INDEX IF NOT EXISTS idx_merge_cycle_pr_numbers ON merge_cycle(pr_numbers);
-- ci_gate_events: one row per individual CI gate observation (per-PR per-run).
--
-- gate is the status_check_context name (e.g. "CI / lint*"), outcome is
-- "passed" | "failed" | "timeout" | "skipped". Used for flake-rate, gate
-- fall-through analysis, and feeding the future T5A duplicate-failure
-- clusterer.
CREATE TABLE IF NOT EXISTS ci_gate_events (
id INTEGER PRIMARY KEY AUTOINCREMENT,
observed_at TEXT NOT NULL,
pr_number INTEGER,
head_sha TEXT NOT NULL,
gate TEXT NOT NULL,
outcome TEXT NOT NULL, -- passed | failed | timeout | skipped
duration_s REAL,
cycle_id INTEGER, -- FK to merge_cycle.id (nullable)
raw TEXT -- optional log/error excerpt
);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_observed ON ci_gate_events(observed_at);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_gate ON ci_gate_events(gate);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_outcome ON ci_gate_events(outcome);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_head_sha ON ci_gate_events(head_sha);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_pr ON ci_gate_events(pr_number);
CREATE INDEX IF NOT EXISTS idx_ci_gate_events_cycle ON ci_gate_events(cycle_id);
-- llm_activity: one row per LLM call by any agent in the auto-agents pipeline.
--
-- Both for cost accounting (T1A/T1C/T2C in the token-reduction plan) and
-- for activity-stream rendering. tokens_in/tokens_out/cached_tokens are
-- denormalised here for fast aggregation.
CREATE TABLE IF NOT EXISTS llm_activity (
id INTEGER PRIMARY KEY AUTOINCREMENT,
started_at TEXT NOT NULL,
ended_at TEXT,
agent TEXT NOT NULL, -- e.g. "pr-review-worker", "task-implementor"
pr_number INTEGER,
session_tag TEXT, -- AUTO-* session tag, when applicable
model TEXT NOT NULL,
tokens_in INTEGER,
tokens_out INTEGER,
cached_tokens INTEGER,
decision TEXT, -- e.g. "approved", "request_changes", "ci_fix_pushed"
outcome TEXT, -- e.g. "success", "escalated", "timeout"
cycle_id INTEGER, -- FK to merge_cycle.id (nullable)
-- v6 ingest-key columns (populated by llm_activity_scraper.py):
session_id TEXT, -- OpenCode session id (e.g. ses_…)
message_id TEXT, -- OpenCode assistant-message id (dedup key)
provider TEXT, -- OpenCode providerID (e.g. "openai", "anthropic")
parent_session_id TEXT, -- set on subagent turns; NULL on top-level
subagent_depth INTEGER, -- 0 at top level, N for depth-N subagent (scraper-normalised)
prompt_hash TEXT, -- SHA-256 of session's first user-message text (v7); duplicates = same prompt
raw TEXT -- optional JSON blob (per-turn raw metrics)
);
CREATE INDEX IF NOT EXISTS idx_llm_activity_started ON llm_activity(started_at);
CREATE INDEX IF NOT EXISTS idx_llm_activity_agent ON llm_activity(agent);
CREATE INDEX IF NOT EXISTS idx_llm_activity_pr ON llm_activity(pr_number);
CREATE INDEX IF NOT EXISTS idx_llm_activity_model ON llm_activity(model);
CREATE INDEX IF NOT EXISTS idx_llm_activity_outcome ON llm_activity(outcome);
CREATE INDEX IF NOT EXISTS idx_llm_activity_cycle ON llm_activity(cycle_id);
-- v6 indexes for ``session_id`` / ``message_id`` are created by
-- ``_migrate_to_v6_llm_activity_ingest_keys``. They reference v6
-- columns, so emitting them here would fail on legacy v5 caches before
-- the migration has a chance to add the columns. Fresh databases still
-- get the indexes because the migration runs unconditionally when
-- ``current < SCHEMA_VERSION`` (and ``current == 0`` for new DBs).
-- ─── conflict_drive_cycles (Tier 1.5 — conflict driver telemetry) ────────
--
-- One row per (cycle, PR). The cycle-level totals (candidates_count,
-- resolved_count, escalated_count, timeout_count, push_rejected_count)
-- are duplicated across all PR rows from the same cycle for cheap
-- per-cycle aggregation; the per-PR specifics live in details_json.
--
-- ``outcome`` and ``failure_kind`` are promoted to top-level columns so
-- the retry-budget query at § 5.1 of conflict-drive-plan.md can use the
-- ``(pr_number, started_at)`` index directly without parsing JSON.
CREATE TABLE IF NOT EXISTS conflict_drive_cycles (
cycle_id TEXT NOT NULL, -- shared across rows from one cycle
pr_number INTEGER NOT NULL, -- top-level for indexed lookups
started_at TEXT NOT NULL,
ended_at TEXT NOT NULL,
outcome TEXT NOT NULL, -- 'resolved' | 'resolved-no-conflict'
-- | 'unresolvable' | 'verification-fail'
-- | 'rebase-failed' | 'timeout'
-- | 'transport-error' | 'lease-violation'
failure_kind TEXT, -- 'definite' | 'transient' | NULL on success
candidates_count INTEGER NOT NULL, -- cycle-level totals duplicated across PR rows
resolved_count INTEGER NOT NULL,
escalated_count INTEGER NOT NULL,
timeout_count INTEGER NOT NULL,
push_rejected_count INTEGER NOT NULL,
details_json TEXT NOT NULL, -- per-PR breakdown
PRIMARY KEY (cycle_id, pr_number)
);
CREATE INDEX IF NOT EXISTS idx_cdc_pr_number_started ON conflict_drive_cycles(pr_number, started_at);
CREATE INDEX IF NOT EXISTS idx_cdc_started_at ON conflict_drive_cycles(started_at);
-- ─── Tier 2 dispatcher telemetry ─────────────────────────────────────────
--
-- One row per dispatcher cycle. ``raw`` contains the per-group counts, swept
-- claims, and per-item worker details; the promoted columns cover dashboard
-- and operator queries without JSON parsing.
CREATE TABLE IF NOT EXISTS dispatch_review_cycles (
id INTEGER PRIMARY KEY AUTOINCREMENT,
cycle_id TEXT NOT NULL,
started_at TEXT NOT NULL,
ended_at TEXT NOT NULL,
driver TEXT NOT NULL,
candidates_count INTEGER NOT NULL,
claims_acquired INTEGER NOT NULL,
swept_count INTEGER NOT NULL,
processed_count INTEGER NOT NULL,
terminal_state TEXT,
worker_outcome TEXT,
session_id TEXT,
worker_wallclock_seconds REAL,
raw TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dispatch_review_cycles_started_at
ON dispatch_review_cycles(started_at);
CREATE INDEX IF NOT EXISTS idx_dispatch_review_cycles_terminal
ON dispatch_review_cycles(terminal_state);
CREATE TABLE IF NOT EXISTS dispatch_implementer_cycles (
id INTEGER PRIMARY KEY AUTOINCREMENT,
cycle_id TEXT NOT NULL,
started_at TEXT NOT NULL,
ended_at TEXT NOT NULL,
driver TEXT NOT NULL,
candidates_count INTEGER NOT NULL,
claims_acquired INTEGER NOT NULL,
swept_count INTEGER NOT NULL,
processed_count INTEGER NOT NULL,
terminal_state TEXT,
worker_outcome TEXT,
session_id TEXT,
worker_wallclock_seconds REAL,
raw TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_dispatch_implementer_cycles_started_at
ON dispatch_implementer_cycles(started_at);
CREATE INDEX IF NOT EXISTS idx_dispatch_implementer_cycles_terminal
ON dispatch_implementer_cycles(terminal_state);
"""
class PipelineCache:
"""SQLite-backed cache with delta sync."""
def __init__(self, db_path: Path):
self.db_path = Path(db_path)
self.db_path.parent.mkdir(parents=True, exist_ok=True)
# timeout=30s raises SQLite's busy-timeout so writers wait/retry when
# a sibling process (e.g. a subprocess opening the same DB) momentarily
# holds a lock, instead of failing immediately. Python's default is 5s,
# which isn't enough when a parent renderer spawns count-master-merges.
self._conn = sqlite3.connect(str(self.db_path), timeout=30.0)
self._conn.row_factory = sqlite3.Row
self._conn.execute("PRAGMA journal_mode = WAL")
self._conn.execute("PRAGMA synchronous = NORMAL")
# Belt-and-suspenders: the PRAGMA is redundant with timeout=30.0 above,
# but explicit documents intent and is useful if someone opens the
# connection differently in the future.
self._conn.execute("PRAGMA busy_timeout = 30000")
self._conn.executescript(SCHEMA_SQL)
self._migrate()
@classmethod
def open(cls, db_path: Path | None = None) -> "PipelineCache":
return cls(db_path or DEFAULT_CACHE_PATH)
def close(self) -> None:
"""Close the underlying SQLite connection.
Long-running sidecars (warmer, dispatchers) that open a cache
per cycle should call this rather than reaching into
``_conn``. Idempotent — re-closing is a no-op so callers can
use it from ``finally`` blocks without guards.
"""
conn = getattr(self, "_conn", None)
if conn is None:
return
try:
conn.close()
except Exception: # noqa: BLE001 — best-effort cleanup
pass
self._conn = None
def __enter__(self) -> "PipelineCache":
return self
def __exit__(self, exc_type, exc, tb) -> None:
self.close()
# ─── Schema management ──────────────────────────────────────────────
def _migrate(self) -> None:
row = self._conn.execute(
"SELECT MAX(version) AS v FROM schema_version"
).fetchone()
current = row["v"] if row and row["v"] is not None else 0
# The new tables in SCHEMA_SQL are all CREATE TABLE IF NOT EXISTS, so
# they are added on first ``__init__`` after upgrade. Any future
# migration that mutates an existing table goes here, gated on
# ``current < N``.
if current < 5:
self._migrate_to_v5_in_flight_rows()
if current < 6:
self._migrate_to_v6_llm_activity_ingest_keys()
if current < 7:
self._migrate_to_v7_prompt_hash()
if current < 8:
self._migrate_to_v8_pulls_touched_files()
if current < SCHEMA_VERSION:
self._conn.execute(
"INSERT INTO schema_version(version) VALUES (?)", (SCHEMA_VERSION,)
)
self._conn.commit()
def _migrate_to_v6_llm_activity_ingest_keys(self) -> None:
"""Add ingest-key columns to ``llm_activity`` so the OpenCode
archive scraper can dedup on ``message_id``.
Idempotent: each ALTER is gated on the column not already existing
so re-running the migration on a partially-upgraded DB is a no-op.
Fresh databases get the columns via the SCHEMA_SQL DDL directly;
this helper exists only for existing v5 caches in the wild.
"""
existing = {
row["name"] for row in self._conn.execute("PRAGMA table_info(llm_activity)")
}
new_cols = [
("session_id", "TEXT"),
("message_id", "TEXT"),
("provider", "TEXT"),
("parent_session_id", "TEXT"),
("subagent_depth", "INTEGER"),
]
for name, decl in new_cols:
if name not in existing:
self._conn.execute(f"ALTER TABLE llm_activity ADD COLUMN {name} {decl}")
# Indexes are CREATE … IF NOT EXISTS so safe to re-run.
self._conn.execute(
"CREATE INDEX IF NOT EXISTS idx_llm_activity_session "
"ON llm_activity(session_id)"
)
self._conn.execute(
"CREATE UNIQUE INDEX IF NOT EXISTS uniq_llm_activity_message_id "
"ON llm_activity(message_id) WHERE message_id IS NOT NULL"
)
self._conn.commit()
def _migrate_to_v8_pulls_touched_files(self) -> None:
"""Add ``touched_files`` (JSON array) column to ``pulls``.
Phase 1 grooming plan: powers the file-overlap signal in Stage A's
suspicion-score pre-filter. Idempotent: gated on column existence
so re-running the migration on a partially-upgraded DB is a no-op.
"""
existing = {
row["name"] for row in self._conn.execute("PRAGMA table_info(pulls)")
}
if "touched_files" not in existing:
self._conn.execute("ALTER TABLE pulls ADD COLUMN touched_files TEXT")
self._conn.commit()
def _migrate_to_v7_prompt_hash(self) -> None:
"""Add ``prompt_hash`` column + index to ``llm_activity``.
Idempotent: gated on column existence; CREATE INDEX is
IF NOT EXISTS. The column is nullable — existing rows
(pre-instrumentation) get NULL; new rows get a hash from the
scraper. No UNIQUE constraint: duplicates are the measurement.
"""
existing = {
row["name"] for row in self._conn.execute("PRAGMA table_info(llm_activity)")
}
if "prompt_hash" not in existing:
self._conn.execute("ALTER TABLE llm_activity ADD COLUMN prompt_hash TEXT")
self._conn.execute(
"CREATE INDEX IF NOT EXISTS idx_llm_activity_prompt_hash "
"ON llm_activity(prompt_hash)"
)
self._conn.commit()
def _migrate_to_v5_in_flight_rows(self) -> None:
"""Drop the NOT NULL constraint on ``ended_at`` for the
``dispatch_*_cycles`` tables so the Tier 2 dispatchers can record
a cycle row at start (``ended_at IS NULL``) and update it on
completion.
Delegates to the shared
:func:`migrate_dispatch_cycle_table_to_v5` helper at module
scope, which is also imported by
:func:`_dispatch_runtime.ensure_cycle_table`. Single source of
truth for the v5 cycle-table schema.
Skips tables that don't exist yet (created lazily by
``ensure_cycle_table`` on the first dispatcher invocation) and
tables whose ``ended_at`` column is already nullable. All
pre-existing rows are preserved verbatim.
"""
for table in DISPATCH_CYCLE_TABLES:
migrate_dispatch_cycle_table_to_v5(self._conn, table)
# ─── Tier 1B telemetry helpers ─────────────────────────────────────
def record_merge_cycle(self, cycle: dict[str, Any]) -> int:
"""Insert a ``merge_cycle`` row. Returns the new row id.
``cycle`` is a dict matching the table columns (extra keys are
captured into the ``raw`` JSON blob). Used by the merge driver to
record one full attempt (single-PR or train-merge).
"""
cols = (
"started_at",
"ended_at",
"pr_numbers",
"train_id",
"bisect_depth",
"rebase_seconds",
"ci_seconds",
"merge_seconds",
"total_seconds",
"terminal_state",
"action_taken",
"action_detail",
"last_released_at",
"flake_rate",
)
# Anything not in cols goes into raw for forensics.
raw_blob = {k: v for k, v in cycle.items() if k not in cols}
values = [cycle.get(c) for c in cols]
# pr_numbers MUST be a JSON-encoded array of ints — coerce here so
# callers can pass either a list or a string.
prs = values[cols.index("pr_numbers")]
if isinstance(prs, (list, tuple)):
values[cols.index("pr_numbers")] = json.dumps(list(prs))
elif prs is None:
values[cols.index("pr_numbers")] = "[]"
# bisect_depth is NOT NULL DEFAULT 0 in the schema; supply 0 when the
# caller doesn't pass it so the explicit-NULL doesn't violate the
# constraint.
if values[cols.index("bisect_depth")] is None:
values[cols.index("bisect_depth")] = 0
cur = self._conn.execute(
f"INSERT INTO merge_cycle ({','.join(cols)}, raw) "
f"VALUES ({','.join('?' for _ in cols)}, ?)",
(*values, json.dumps(raw_blob) if raw_blob else None),
)
self._conn.commit()
rid = cur.lastrowid
assert rid is not None
return rid
def record_ci_gate_event(self, event: dict[str, Any]) -> int:
"""Insert one ``ci_gate_events`` row. Commits immediately — use
:meth:`record_ci_gate_events_batch` for high-volume callers (the
merge driver collects all 6 gate outcomes per PR ramp-up; an
eager fsync per row is wasteful).
"""
return self.record_ci_gate_events_batch([event])[-1]
def record_ci_gate_events_batch(self, events: list[dict[str, Any]]) -> list[int]:
"""Insert N ``ci_gate_events`` rows in a single transaction (one
fsync rather than N). Returns the list of new row ids in input
order."""
if not events:
return []
cols = (
"observed_at",
"pr_number",
"head_sha",
"gate",
"outcome",
"duration_s",
"cycle_id",
)
rows: list[tuple] = []
for event in events:
raw_blob = {k: v for k, v in event.items() if k not in cols}
rows.append(
(
*[event.get(c) for c in cols],
json.dumps(raw_blob) if raw_blob else None,
)
)
sql = (
f"INSERT INTO ci_gate_events ({','.join(cols)}, raw) "
f"VALUES ({','.join('?' for _ in cols)}, ?)"
)
ids: list[int] = []
# executemany commits implicitly only inside a ``with self._conn``
# block; we want a single explicit commit at the end so partial
# batches don't half-write on error.
try:
for row in rows:
cur = self._conn.execute(sql, row)
rid = cur.lastrowid
assert rid is not None
ids.append(rid)
self._conn.commit()
except Exception:
self._conn.rollback()
raise
return ids
# Column tuple shared by the single-row and batch ``llm_activity``
# writers. The trailing ``raw`` column is always appended by the
# SQL builder so it is not listed here.
_LLM_ACTIVITY_COLS: tuple[str, ...] = (
"started_at",
"ended_at",
"agent",
"pr_number",
"session_tag",
"model",
"tokens_in",
"tokens_out",
"cached_tokens",
"decision",
"outcome",
"cycle_id",
"session_id",
"message_id",
"provider",
"parent_session_id",
"subagent_depth",
"prompt_hash",
)
def upsert_llm_activity(self, activity: dict[str, Any]) -> int:
"""Insert one ``llm_activity`` row, or skip it if a row with
the same ``message_id`` already exists (dedup). Returns the
new row id on insert, or ``0`` if dedup'd.
Used by the OpenCode archive scraper (one row per assistant
turn) and by direct callers that record cost outside the
scraper path. The v6 schema added ingest-key columns
(``session_id``, ``message_id``, ``provider``,
``parent_session_id``, ``subagent_depth``); when ``message_id``
is supplied the partial unique index makes the insert
idempotent via ``INSERT OR IGNORE`` so re-scraping the same
archive does not double-count tokens.
High-volume callers (the scraper) should prefer
:meth:`upsert_llm_activity_batch` to amortise the per-row
fsync.
"""
cols = self._LLM_ACTIVITY_COLS
raw_blob = {k: v for k, v in activity.items() if k not in cols}
cur = self._conn.execute(
f"INSERT OR IGNORE INTO llm_activity ({','.join(cols)}, raw) "
f"VALUES ({','.join('?' for _ in cols)}, ?)",
(
*[activity.get(c) for c in cols],
json.dumps(raw_blob) if raw_blob else None,
),
)
self._conn.commit()
# On INSERT OR IGNORE, SQLite leaves ``lastrowid`` at the
# connection's previous successful insert (not the just-ignored
# one), so ``rowcount == 0`` is the reliable dedup signal.
if cur.rowcount == 0:
return 0
rid = cur.lastrowid
assert rid is not None
return rid
def upsert_llm_activity_batch(
self,
activities: list[dict[str, Any]],
) -> dict[str, int]:
"""Insert a batch of ``llm_activity`` rows in a single
transaction. Returns ``{"inserted": N, "duplicate": M}``.
Mirrors :meth:`record_ci_gate_events_batch`: one fsync per
batch rather than per row, which matters on the scraper's
first backfill (thousands of turns) and on busy steady-state
cycles. The partial unique index on ``message_id`` still does
the dedup work — duplicates are counted and skipped without
raising.
An empty list is a no-op.
"""
if not activities:
return {"inserted": 0, "duplicate": 0}
cols = self._LLM_ACTIVITY_COLS
sql = (
f"INSERT OR IGNORE INTO llm_activity ({','.join(cols)}, raw) "
f"VALUES ({','.join('?' for _ in cols)}, ?)"
)
inserted = duplicate = 0
try:
for activity in activities:
raw_blob = {k: v for k, v in activity.items() if k not in cols}
cur = self._conn.execute(
sql,
(
*[activity.get(c) for c in cols],
json.dumps(raw_blob) if raw_blob else None,
),
)
if cur.rowcount == 0:
duplicate += 1
else:
inserted += 1
self._conn.commit()
except Exception:
self._conn.rollback()
raise
return {"inserted": inserted, "duplicate": duplicate}
def merge_cycles_in_window(
self, start: datetime, end: datetime
) -> list[dict[str, Any]]:
"""Return ``merge_cycle`` rows whose ``started_at`` is within
[start, end). Used by the pr-velocity canvas renderer."""
rows = self._conn.execute(
"SELECT * FROM merge_cycle "
"WHERE started_at >= ? AND started_at < ? "
"ORDER BY started_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [dict(r) for r in rows]
def ci_gate_events_in_window(
self, start: datetime, end: datetime
) -> list[dict[str, Any]]:
rows = self._conn.execute(
"SELECT * FROM ci_gate_events "
"WHERE observed_at >= ? AND observed_at < ? "
"ORDER BY observed_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [dict(r) for r in rows]
def llm_activity_in_window(
self, start: datetime, end: datetime
) -> list[dict[str, Any]]:
rows = self._conn.execute(
"SELECT * FROM llm_activity "
"WHERE started_at >= ? AND started_at < ? "
"ORDER BY started_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [dict(r) for r in rows]
# ─── conflict_drive_cycles helpers ─────────────────────────────────
def record_conflict_drive_cycle(self, row: dict[str, Any]) -> None:
"""Insert one ``conflict_drive_cycles`` row.
``row`` must carry every NOT NULL column. ``details_json`` is
encoded if a dict is passed; if the caller already JSON-encoded
it, the string is stored as-is. ``outcome`` and ``failure_kind``
live both as top-level indexed columns and inside ``details_json``;
callers MUST keep them in sync (a single
``record_telemetry`` helper in ``conflict_drive.py`` is the
intended source of truth — see plan § 9).
"""
details = row.get("details_json")
if isinstance(details, (dict, list)):
details_str = json.dumps(details)
else:
details_str = details if details is not None else "{}"
cols = (
"cycle_id",
"pr_number",
"started_at",
"ended_at",
"outcome",
"failure_kind",
"candidates_count",
"resolved_count",
"escalated_count",
"timeout_count",
"push_rejected_count",
"details_json",
)
values = list(row.get(c) for c in cols)
values[cols.index("details_json")] = details_str
# Use INSERT OR REPLACE so a same-cycle re-record (e.g. a retry
# mid-cycle) doesn't trip the PRIMARY KEY (cycle_id, pr_number).
# Cycle ids are intended to be unique-per-cycle; replace is the
# safe default if a caller emits twice by accident.
self._conn.execute(
f"INSERT OR REPLACE INTO conflict_drive_cycles "
f"({','.join(cols)}) VALUES ({','.join('?' for _ in cols)})",
values,
)
self._conn.commit()
def mark_conflict_drive_cycle_escalated(self, cycle_id: str, pr_number: int) -> int:
"""Set ``escalated_count = 1`` on the row keyed by
``(cycle_id, pr_number)``.
Telemetry semantics (plan § 9): ``escalated_count`` is a
cycle-level event — it fires exactly once across the rolling
24-hour retry budget when ``conflict_drive.run_one_cycle``
decides the budget is exhausted and applies
``auto/needs-implementer``. ``record_conflict_drive_cycle``
always writes ``escalated_count = 0`` because at row-write
time the budget verdict has not yet been computed (the
verdict needs the row that's being written). This helper is
the second-phase update; calling it twice for the same row
is idempotent.
Returns the number of rows updated. Zero rows means the
telemetry write earlier in the cycle either failed or used a
different ``cycle_id`` — the caller may use this to decide
whether to log a diagnostic. The caller is responsible for
any logging; this helper stays silent so the cache module
does not depend on a specific logger.
"""
cur = self._conn.execute(
"UPDATE conflict_drive_cycles SET escalated_count = 1 "
"WHERE cycle_id = ? AND pr_number = ?",
(cycle_id, pr_number),
)
self._conn.commit()
return cur.rowcount or 0
def count_conflict_drive_definite_failures_24h(
self, pr_number: int, *, now: datetime | None = None
) -> int:
"""Return the number of ``failure_kind='definite'`` rows for
``pr_number`` whose ``started_at`` is within the trailing 24 h
window. This is the retry-budget query at plan § 5.1 — used by
``conflict_drive.resolve_one`` to decide whether the next failure
triggers escalation to ``auto/needs-implementer``.
Uses the ``idx_cdc_pr_number_started`` index for an O(log N) lookup.
``now`` is injectable so unit tests can pin the window without
time-traveling the wall clock.
"""
cutoff = (now or datetime.now(timezone.utc)) - timedelta(hours=24)
row = self._conn.execute(
"SELECT COUNT(*) AS n FROM conflict_drive_cycles "
"WHERE pr_number = ? AND failure_kind = 'definite' "
"AND started_at > ?",
(pr_number, cutoff.isoformat()),
).fetchone()
return int(row["n"]) if row else 0
def conflict_drive_cycles_in_window(
self, start: datetime, end: datetime
) -> list[dict[str, Any]]:
"""Return ``conflict_drive_cycles`` rows whose ``started_at`` is
within [start, end). Used by ``tools/render-pr-velocity.py`` for
the "Conflict-resolution activity" section.
Uses the ``idx_cdc_started_at`` index for an O(log N) seek.
"""
rows = self._conn.execute(
"SELECT * FROM conflict_drive_cycles "
"WHERE started_at >= ? AND started_at < ? "
"ORDER BY started_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [dict(r) for r in rows]
def last_conflict_drive_cycle_for_pr(self, pr_number: int) -> dict[str, Any] | None:
"""Return the most recent ``conflict_drive_cycles`` row for
``pr_number`` (highest ``started_at``), or ``None`` if no row
exists.
Used by ``tools/conflict_drive.py`` for the per-cycle cooldown
check. Uses the ``idx_cdc_pr_number_started`` index, so the
lookup is O(log N) — important because this fires once per
candidate per cycle.
"""
row = self._conn.execute(
"SELECT * FROM conflict_drive_cycles "
"WHERE pr_number = ? "
"ORDER BY started_at DESC LIMIT 1",
(pr_number,),
).fetchone()
return dict(row) if row else None
def last_released_at_for_pr(self, pr_number: int) -> str | None:
"""Most recent ``last_released_at`` from merge_cycle for one PR.
Used by the merge driver's candidate filter to enforce the
cooldown window without depending on labels (S4 simplification:
labels persist for visibility, but eligibility is computed from
telemetry).
"""
# Pull cycles that mention this PR and sort by ended_at desc.
# SQLite has no native JSON containment, so we use a LIKE filter
# and then verify in Python (cheap; merge_cycle is small).
rows = self._conn.execute(
"SELECT pr_numbers, last_released_at FROM merge_cycle "
"WHERE last_released_at IS NOT NULL "
" AND pr_numbers LIKE ? "
"ORDER BY last_released_at DESC LIMIT 50",
(f"%{pr_number}%",),
).fetchall()
for r in rows:
try:
prs = json.loads(r["pr_numbers"])
except (TypeError, ValueError):
continue
if pr_number in prs:
return r["last_released_at"]
return None
# ─── sync_meta helpers ──────────────────────────────────────────────
def _meta_get(self, key: str) -> str | None:
row = self._conn.execute(
"SELECT value FROM sync_meta WHERE key = ?", (key,)
).fetchone()
return row["value"] if row else None
def _meta_set(self, key: str, value: str) -> None:
self._conn.execute(
"INSERT INTO sync_meta(key, value) VALUES(?, ?) "
"ON CONFLICT(key) DO UPDATE SET value = excluded.value",
(key, value),
)
# ─── Locking ────────────────────────────────────────────────────────
def _lock(self):
"""fcntl.flock on ``<db_path>.lock`` for exclusive sync."""
lock_path = Path(str(self.db_path) + ".lock")
lock_path.touch()
fd = os.open(str(lock_path), os.O_RDWR)
fcntl.flock(fd, fcntl.LOCK_EX)
return fd
@staticmethod
def _unlock(fd: int) -> None:
try:
fcntl.flock(fd, fcntl.LOCK_UN)
finally:
os.close(fd)
# ─── Commits sync ───────────────────────────────────────────────────
def _upsert_commit(self, c: dict) -> bool:
"""Insert a commit if new. Returns True if a new row was inserted."""
sha = c.get("sha")
if not sha:
return False
row = self._conn.execute(
"INSERT OR IGNORE INTO commits "
"(sha, committer_date, author_name, committer_name, message, parents, raw) "
"VALUES (?, ?, ?, ?, ?, ?, ?)",
(
sha,
(c.get("commit", {}).get("committer", {}) or {}).get("date"),
(c.get("commit", {}).get("author", {}) or {}).get("name"),
(c.get("commit", {}).get("committer", {}) or {}).get("name"),
(c.get("commit", {}) or {}).get("message"),
json.dumps([p.get("sha") for p in (c.get("parents") or [])]),
json.dumps(c),
),
)
return row.rowcount > 0
def _sync_commits(self, token: str, full: bool, progress: bool) -> dict:
"""Pull master commits forward from HEAD until we reach cached content.
When ``full=True`` the cache is truncated first (use after a force-push
or to rebuild from scratch).
"""
if full:
self._conn.execute("DELETE FROM commits")
self._conn.commit()
prev_head = self._meta_get("master_head_sha")
known_shas = {
row["sha"] for row in self._conn.execute("SELECT sha FROM commits")
}
new_count = 0
consecutive_all_cached = 0
first_sha: str | None = None
saw_prev_head = False
for page in range(1, SYNC_PAGE_HARD_LIMIT + 1):
commits = api_get(
f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/commits"
f"?sha=master&limit=50&page={page}",
token,
)
if not commits:
break
if first_sha is None and commits:
first_sha = commits[0].get("sha")
page_new = 0
for c in commits:
sha = c.get("sha")
if sha == prev_head:
saw_prev_head = True
if sha in known_shas:
continue
if self._upsert_commit(c):
known_shas.add(sha)
new_count += 1
page_new += 1
if progress and (page == 1 or page % 10 == 0):
print(
f"# commits: page {page}, +{page_new} new (running total {new_count})",
file=sys.stderr,
)
# Stop when we've hit a page that's entirely old content (meaning
# we've caught up to the cache's existing HEAD).
if page_new == 0:
consecutive_all_cached += 1
if consecutive_all_cached >= SYNC_EMPTY_PAGE_TOLERANCE:
break
else:
consecutive_all_cached = 0
if len(commits) < 50:
break
self._conn.commit()
if first_sha:
self._meta_set("master_head_sha", first_sha)
self._meta_set(
"last_commit_sync_at", datetime.now(timezone.utc).isoformat()
)
return {
"new_commits": new_count,
"master_head_sha": first_sha,
"force_push_suspected": (
prev_head is not None and not saw_prev_head and new_count > 50
),
}
# ─── PR sync ────────────────────────────────────────────────────────
# ─── PR touched-files (Phase 1 grooming plan) ───────────────────────
def set_touched_files(self, number: int, files: list[str]) -> None:
"""Persist the list of file paths a PR touched.
Phase 1 grooming plan: powers the file-overlap signal in Stage A's
pre-filter. Populated via ``GET /pulls/{n}/files`` alongside the
detail-fetch flow; safe to call standalone for one-off backfills.
"""
self._conn.execute(
"UPDATE pulls SET touched_files = ? WHERE number = ?",
(json.dumps(files), number),
)
self._conn.commit()
def get_touched_files(self, number: int) -> list[str] | None:
"""Return the cached list of file paths the PR touched, or None
if not yet populated. Returns [] if the PR exists but the column
is empty/non-list (defensive — treat malformed cache as 'no data')."""
row = self._conn.execute(
"SELECT touched_files FROM pulls WHERE number = ?",
(number,),
).fetchone()
if row is None or row["touched_files"] is None:
return None
try:
files = json.loads(row["touched_files"])
except (ValueError, TypeError):
return []
return files if isinstance(files, list) else []
def _upsert_pr(self, pr: dict, has_detail: bool | None = None) -> bool:
"""Insert or update a PR. Returns True if row was new or updated_at advanced."""
num = pr.get("number")
if num is None:
return False
existing = self._conn.execute(
"SELECT updated_at, has_detail FROM pulls WHERE number = ?",
(num,),
).fetchone()
new_updated = pr.get("updated_at")
if existing is not None and existing["updated_at"] == new_updated:
# No change; but we may have just pulled detail — update has_detail if so.
if has_detail and not existing["has_detail"]:
self._conn.execute(
"UPDATE pulls SET has_detail = 1, merged_by_login = ?, "
"additions = ?, deletions = ?, changed_files = ?, raw = ? "
"WHERE number = ?",
(
(pr.get("merged_by") or {}).get("login"),
pr.get("additions"),
pr.get("deletions"),
pr.get("changed_files"),
json.dumps(pr),
num,
),
)
return True
return False
self._conn.execute(
"INSERT INTO pulls "
"(number, state, merged, created_at, updated_at, closed_at, merged_at, "
" base_ref, head_sha, merge_commit_sha, user_login, merged_by_login, "
" additions, deletions, changed_files, labels, has_detail, raw) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) "
"ON CONFLICT(number) DO UPDATE SET "
" state = excluded.state, merged = excluded.merged, "
" created_at = excluded.created_at, updated_at = excluded.updated_at, "
" closed_at = excluded.closed_at, merged_at = excluded.merged_at, "
" base_ref = excluded.base_ref, head_sha = excluded.head_sha, "
" merge_commit_sha = excluded.merge_commit_sha, "
" user_login = excluded.user_login, "
" merged_by_login = COALESCE(excluded.merged_by_login, pulls.merged_by_login), "
" additions = COALESCE(excluded.additions, pulls.additions), "
" deletions = COALESCE(excluded.deletions, pulls.deletions), "
" changed_files = COALESCE(excluded.changed_files, pulls.changed_files), "
" labels = excluded.labels, "
" has_detail = MAX(excluded.has_detail, pulls.has_detail), "
" raw = excluded.raw",
(
num,
pr.get("state"),
1 if pr.get("merged") else 0,
pr.get("created_at"),
pr.get("updated_at"),
pr.get("closed_at"),
pr.get("merged_at"),
(pr.get("base") or {}).get("ref"),
(pr.get("head") or {}).get("sha"),
pr.get("merge_commit_sha"),
(pr.get("user") or {}).get("login"),
(pr.get("merged_by") or {}).get("login"),
pr.get("additions"),
pr.get("deletions"),
pr.get("changed_files"),
json.dumps([lbl.get("name") for lbl in (pr.get("labels") or [])]),
1 if has_detail else 0,
json.dumps(pr),
),
)
return True
def _sync_pulls(self, token: str, full: bool, progress: bool) -> dict:
"""Sync all PRs by walking ``pulls?state=all&sort=recentupdate``.
Stops when we see a page whose most-recent updated_at is older than
our last-sync cursor (everything older is already cached and stable).
"""
if full:
self._conn.execute("DELETE FROM pulls")
self._conn.commit()
last_sync = self._meta_get("last_pull_sync_updated_at")
if full or last_sync is None:
last_sync_dt = None
else:
last_sync_dt = last_sync # string comparison is safe on ISO 8601
upserts = 0
max_updated: str | None = None
consecutive_stale_pages = 0
for page in range(1, SYNC_PAGE_HARD_LIMIT + 1):
prs = api_get(
f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/pulls"
f"?state=all&sort=recentupdate&limit=50&page={page}",
token,
)
if not prs:
break
page_new_or_updated = 0
page_min_updated: str | None = None
for pr in prs:
upd = pr.get("updated_at")
if upd is not None:
if page_min_updated is None or upd < page_min_updated:
page_min_updated = upd
if max_updated is None or upd > max_updated:
max_updated = upd
if last_sync_dt is not None and upd is not None and upd <= last_sync_dt:
continue
if self._upsert_pr(pr):
upserts += 1
page_new_or_updated += 1
if progress and (page == 1 or page % 20 == 0):
print(
f"# pulls: page {page}, +{page_new_or_updated} upserts "
f"(running total {upserts})",
file=sys.stderr,
)
# Once the oldest updated_at on this page is older than our last
# sync cursor, any subsequent page is entirely older still.
if (
last_sync_dt is not None
and page_min_updated is not None
and page_min_updated <= last_sync_dt
):
break
if page_new_or_updated == 0:
consecutive_stale_pages += 1
if consecutive_stale_pages >= SYNC_EMPTY_PAGE_TOLERANCE:
break
else:
consecutive_stale_pages = 0
if len(prs) < 50:
break
self._conn.commit()
if max_updated:
self._meta_set("last_pull_sync_updated_at", max_updated)
self._meta_set("last_pull_sync_at", datetime.now(timezone.utc).isoformat())
return {"pulls_upserted": upserts, "max_updated_at": max_updated}
# ─── PR enrichment (lazy, on read) ──────────────────────────────────
def enrich_pr(self, number: int, token: str) -> dict:
"""Return the full PR dict, fetching detail on demand.
The list endpoint omits ``merged_by``, ``additions``, ``deletions``,
and ``changed_files``. We fetch the detail once per PR and persist it.
"""
row = self._conn.execute(
"SELECT raw, has_detail FROM pulls WHERE number = ?",
(number,),
).fetchone()
if row is None:
# Not in cache at all — fetch and store.
pr = api_get(
f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/pulls/{number}",
token,
)
self._upsert_pr(pr, has_detail=True)
self._fetch_and_cache_touched_files(number, token)
self._conn.commit()
return pr
if row["has_detail"]:
return json.loads(row["raw"])
# Need to fetch detail.
pr = api_get(
f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/pulls/{number}",
token,
)
self._upsert_pr(pr, has_detail=True)
self._fetch_and_cache_touched_files(number, token)
self._conn.commit()
return pr
def _fetch_and_cache_touched_files(self, number: int, token: str) -> None:
"""Phase 1 grooming plan: alongside the PR detail fetch, ask
Forgejo for the list of files this PR touches and persist them.
Non-fatal on error: a Forgejo flake here shouldn't block the
detail-fetch flow that called us. The grooming Stage A pre-filter
treats a missing ``touched_files`` cache entry as "no file-overlap
signal" rather than refusing to evaluate.
"""
try:
files = api_get(
f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/pulls/{number}/files",
token,
)
except Exception as exc: # noqa: BLE001
logger.warning(
"_fetch_and_cache_touched_files: GET /pulls/%d/files failed: %s",
number,
exc,
)
return
# Forgejo returns a list of file objects; ``filename`` is the path.
# Defensive: skip malformed entries (non-dict, missing filename).
paths: list[str] = []
if isinstance(files, list):
for f in files:
if isinstance(f, dict):
name = f.get("filename")
if isinstance(name, str) and name:
paths.append(name)
self.set_touched_files(number, paths)
def backfill_merged_by(self, token: str, progress: bool = True) -> int:
"""One-shot: fetch detail for every in-cache closed+merged PR that
lacks ``merged_by_login``. Called optionally after a fresh sync to
warm the merged-by cache. Returns number of PRs enriched."""
rows = self._conn.execute(
"SELECT number FROM pulls "
"WHERE state = 'closed' AND merged = 1 AND merged_by_login IS NULL "
"AND has_detail = 0"
).fetchall()
n = 0
for i, row in enumerate(rows):
self.enrich_pr(row["number"], token)
n += 1
if progress and (i + 1) % 50 == 0:
print(f"# merged_by backfill: {i + 1}/{len(rows)}", file=sys.stderr)
return n
# ─── Reachability (compare endpoint) ────────────────────────────────
def is_sha_reachable_from_master(
self,
sha: str,
token: str,
*,
master_head_sha: str | None = None,
) -> bool:
"""Cached reachability check.
Cache key is ``(merge_sha, master_head_sha)``; a positive result is
permanent (a commit that was ever an ancestor of master stays so,
absent a force-push), but negative results are re-checked if the
master HEAD has advanced.
"""
if not sha:
return False
row = self._conn.execute(
"SELECT reachable, master_head_sha FROM reachability WHERE merge_sha = ?",
(sha,),
).fetchone()
if row is not None:
if row["reachable"]:
return True
# Negative result: only trust it if master HEAD hasn't changed.
cur_head = master_head_sha or self._meta_get("master_head_sha")
if cur_head and row["master_head_sha"] == cur_head:
return False
# Miss or stale negative — do the API call.
url = f"{API_BASE}/repos/{REPO_OWNER}/{REPO_NAME}/compare/master...{sha}"
try:
data = api_get(url, token)
except urllib.error.HTTPError as e:
if e.code == 404:
reachable = False
else:
raise
else:
total = data.get("total_commits")
if total is None:
total = len(data.get("commits") or [])
reachable = total == 0
cur_head = master_head_sha or self._meta_get("master_head_sha")
self._conn.execute(
"INSERT INTO reachability(merge_sha, reachable, checked_at, master_head_sha) "
"VALUES(?, ?, ?, ?) "
"ON CONFLICT(merge_sha) DO UPDATE SET "
" reachable = excluded.reachable, "
" checked_at = excluded.checked_at, "
" master_head_sha = excluded.master_head_sha",
(
sha,
1 if reachable else 0,
datetime.now(timezone.utc).isoformat(),
cur_head,
),
)
self._conn.commit()
return reachable
# ─── Top-level sync ─────────────────────────────────────────────────
def sync(
self,
token: str,
*,
full: bool = False,
backfill_details: bool = False,
progress: bool = True,
) -> dict:
"""Delta-sync the cache. Returns a summary dict."""
lock_fd = self._lock()
try:
if progress:
print("# syncing commits ...", file=sys.stderr)
commits_summary = self._sync_commits(token, full, progress)
if commits_summary.get("force_push_suspected"):
print(
"# WARN: previous master HEAD not seen in recent commits — "
"possible force-push. Re-run with --full to rebuild from scratch.",
file=sys.stderr,
)
if progress:
print("# syncing pulls ...", file=sys.stderr)
pulls_summary = self._sync_pulls(token, full, progress)
enriched = 0
if backfill_details:
if progress:
print(
"# back-filling merged_by for merged PRs ...", file=sys.stderr
)
enriched = self.backfill_merged_by(token, progress=progress)
return {
"commits": commits_summary,
"pulls": pulls_summary,
"enriched_prs": enriched,
}
finally:
self._unlock(lock_fd)
# ─── Query API (all return raw API-shaped dicts) ────────────────────
def commits_on_master(
self,
start: datetime,
end: datetime,
) -> list[dict]:
rows = self._conn.execute(
"SELECT raw FROM commits "
"WHERE committer_date >= ? AND committer_date < ? "
"ORDER BY committer_date DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [json.loads(r["raw"]) for r in rows]
def pulls_closed_in_window(self, start: datetime, end: datetime) -> list[dict]:
rows = self._conn.execute(
"SELECT raw FROM pulls "
"WHERE state = 'closed' AND closed_at IS NOT NULL "
" AND closed_at >= ? AND closed_at < ? "
"ORDER BY closed_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [json.loads(r["raw"]) for r in rows]
def pulls_merged_in_window(self, start: datetime, end: datetime) -> list[dict]:
rows = self._conn.execute(
"SELECT raw FROM pulls "
"WHERE merged = 1 AND merged_at IS NOT NULL "
" AND merged_at >= ? AND merged_at < ? "
"ORDER BY merged_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [json.loads(r["raw"]) for r in rows]
def pulls_opened_in_window(self, start: datetime, end: datetime) -> list[dict]:
rows = self._conn.execute(
"SELECT raw FROM pulls "
"WHERE created_at >= ? AND created_at < ? "
"ORDER BY created_at DESC",
(start.isoformat(), end.isoformat()),
).fetchall()
return [json.loads(r["raw"]) for r in rows]
def sha_to_pr_map(self, start: datetime, end: datetime) -> dict[str, dict]:
"""SHA -> PR dict for PRs merged in window (indexes both head.sha
and merge_commit_sha), for count-master-merges Phase 1 cross-reference.
"""
rows = self._conn.execute(
"SELECT raw, head_sha, merge_commit_sha FROM pulls "
"WHERE merged = 1 AND merged_at IS NOT NULL "
" AND merged_at >= ? AND merged_at < ?",
(start.isoformat(), end.isoformat()),
).fetchall()
out: dict[str, dict] = {}
for r in rows:
pr = json.loads(r["raw"])
if r["head_sha"]:
out[r["head_sha"]] = pr
if r["merge_commit_sha"]:
out[r["merge_commit_sha"]] = pr
return out
# ─── Introspection helpers ──────────────────────────────────────────
def stats(self) -> dict:
c = self._conn
row = c.execute(
"SELECT COUNT(*) AS n, MIN(committer_date) AS oldest, "
"MAX(committer_date) AS newest FROM commits"
).fetchone()
pull_row = c.execute(
"SELECT COUNT(*) AS n, MIN(created_at) AS oldest_created, "
"MAX(updated_at) AS latest_update FROM pulls"
).fetchone()
reach_row = c.execute("SELECT COUNT(*) AS n FROM reachability").fetchone()
return {
"cache_path": str(self.db_path),
"cache_size_bytes": self.db_path.stat().st_size
if self.db_path.exists()
else 0,
"commits": {
"count": row["n"],
"oldest": row["oldest"],
"newest": row["newest"],
},
"pulls": {
"count": pull_row["n"],
"oldest_created": pull_row["oldest_created"],
"latest_update": pull_row["latest_update"],
},
"reachability_cached": reach_row["n"],
"last_commit_sync_at": self._meta_get("last_commit_sync_at"),
"last_pull_sync_at": self._meta_get("last_pull_sync_at"),
"master_head_sha": self._meta_get("master_head_sha"),
}
# ─── PR classification cache (2026-05-16) ───────────────────────────
#
# Backing store for ``_pr_classification_cache.refresh_then_filter``.
# Read/write API kept tight (one upsert, one read, one bulk-purge) so
# the consumer module owns the freshness logic.
PR_CLASSIFICATION_SCHEMA_VERSION = 1
def upsert_pr_classification(self, row: dict[str, Any]) -> None:
"""Insert or replace a ``pr_classifications`` row. All fields
listed in the schema must be present in ``row`` (the consumer
builds them in ``_classify_pr``)."""
self._conn.execute(
"""
INSERT OR REPLACE INTO pr_classifications (
pr_number, head_sha, updated_at, last_checked_at,
ci_status, approvals_count,
has_active_request_changes, has_unaddressed_request_changes,
is_claimed, is_mergeable, stale_state, labels_json,
classification_schema_version
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
int(row["pr_number"]),
str(row["head_sha"]),
str(row["updated_at"]),
str(row["last_checked_at"]),
str(row["ci_status"]),
int(row["approvals_count"]),
int(bool(row["has_active_request_changes"])),
int(bool(row["has_unaddressed_request_changes"])),
int(bool(row["is_claimed"])),
None
if row.get("is_mergeable") is None
else int(bool(row["is_mergeable"])),
str(row["stale_state"]),
str(row["labels_json"]),
int(
row.get(
"classification_schema_version",
self.PR_CLASSIFICATION_SCHEMA_VERSION,
)
),
),
)
self._conn.commit()
def get_pr_classification(self, pr_number: int) -> dict[str, Any] | None:
"""Return the cached classification row for ``pr_number``, or
``None`` if absent. Returned as a plain dict (not sqlite3.Row)
so the consumer doesn't take a row-class dependency."""
cur = self._conn.execute(
"SELECT * FROM pr_classifications WHERE pr_number = ?",
(int(pr_number),),
)
row = cur.fetchone()
if row is None:
return None
return dict(row)
def purge_pr_classifications_older_than(self, cutoff_iso: str) -> int:
"""Delete cache rows whose ``last_checked_at`` is older than
``cutoff_iso``. Returns the number of rows deleted.
Intended for periodic GC: PRs that haven't been seen in days
are usually closed/merged and don't need cache slots. Not
called automatically by the consumer — the dispatcher can
invoke this on a slow cadence (e.g. once per N cycles)."""
cur = self._conn.execute(
"DELETE FROM pr_classifications WHERE last_checked_at < ?",
(str(cutoff_iso),),
)
self._conn.commit()
return cur.rowcount
def close(self) -> None:
self._conn.close()