docs(development): add system watchdog architecture and audit reference #2449
@@ -0,0 +1,350 @@
|
||||
# System Watchdog
|
||||
|
||||
The **System Watchdog** (`ca-system-watchdog`) is the 16th supervisor agent in the
|
||||
CleverAgents autonomous build system. It runs continuously with a 5-minute polling
|
||||
cycle and monitors the entire agent ecosystem for correctness, health, and policy
|
||||
compliance.
|
||||
|
||||
> **Audience:** Contributors and operators who need to understand how the autonomous
|
||||
> agent system is monitored, what violations are detected, and how the watchdog
|
||||
> responds to findings.
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
The watchdog operates exclusively through the **Forgejo API** and the
|
||||
**OpenCode Server HTTP API** — it does not clone the repository or access the
|
||||
filesystem. It performs 12 audits per cycle, ranging from quality gate compliance
|
||||
to deep session introspection of individual supervisor conversations.
|
||||
|
||||
```
|
||||
cycle = 0
|
||||
LOOP FOREVER (5-minute intervals):
|
||||
cycle += 1
|
||||
findings = []
|
||||
|
||||
findings += audit_quality_gates() # Audit 1 — MOST CRITICAL
|
||||
findings += audit_branch_protection() # Audit 2
|
||||
findings += audit_ticket_states() # Audit 3
|
||||
findings += audit_priority_ordering() # Audit 4
|
||||
findings += audit_pr_pipeline() # Audit 5
|
||||
findings += audit_supervisor_health() # Audit 6 (session introspection)
|
||||
findings += audit_labels_and_dependencies() # Audit 7
|
||||
findings += audit_ticket_hierarchy() # Audit 8
|
||||
findings += audit_test_health() # Audit 9
|
||||
findings += audit_improvement_generation() # Audit 10
|
||||
findings += audit_session_spot_check() # Audit 11 (every cycle)
|
||||
|
||||
if cycle % 6 == 0:
|
||||
findings += audit_deep_session_introspection() # Audit 12 (every 30 min)
|
||||
|
||||
for finding in findings:
|
||||
take_action(finding)
|
||||
|
||||
if cycle % 6 == 0 and findings:
|
||||
post_summary(cycle, findings)
|
||||
|
||||
sleep(300) # 5 minutes
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Audit Descriptions
|
||||
|
||||
### Audit 1 — Quality Gate Compliance (CRITICAL)
|
||||
|
||||
Ensures **no code reaches `master` without passing all CI checks**.
|
||||
|
||||
- Checks the last 10 commits on `master` for passing `status-check` CI status.
|
||||
- Verifies recently merged PRs had passing CI at merge time.
|
||||
- Detects direct pushes to `master` (all changes must go through PRs).
|
||||
|
||||
**Severity:** CRITICAL for any violation.
|
||||
|
||||
---
|
||||
|
||||
### Audit 2 — Branch Protection Verification
|
||||
|
||||
Verifies that Forgejo branch protection is active and correctly configured for
|
||||
`master`.
|
||||
|
||||
Checks:
|
||||
- Branch protection rules exist for `master`.
|
||||
- `enable_status_check` is `true`.
|
||||
- The `status-check` context is required.
|
||||
- At least 2 approving reviews are required (per `CONTRIBUTING.md`).
|
||||
|
||||
---
|
||||
|
||||
### Audit 3 — Ticket State Integrity
|
||||
|
||||
Ensures all issues have correct `State/` labels matching their actual state.
|
||||
|
||||
Checks:
|
||||
- Closed issues have `State/Completed` (not `State/Unverified`).
|
||||
- Issues with `State/In Review` have an open PR referencing them.
|
||||
- No issue has multiple `State/` labels simultaneously.
|
||||
|
||||
---
|
||||
|
||||
### Audit 4 — Priority and Milestone Ordering
|
||||
|
||||
Ensures Critical bugs on lower milestones are addressed before feature work on
|
||||
later milestones.
|
||||
|
||||
- Finds the lowest milestone with `Priority/Critical` or `MoSCoW/Must Have` bugs.
|
||||
- Flags any non-bug issue in `State/In Progress` on a later milestone.
|
||||
|
||||
---
|
||||
|
||||
### Audit 5 — PR Pipeline Health
|
||||
|
||||
Tracks PR aging, review coverage, and merge throughput.
|
||||
|
||||
| Condition | Severity |
|
||||
|-----------|----------|
|
||||
| PR open >24h with no reviews | MEDIUM |
|
||||
| PR approved but not merged >6h | HIGH |
|
||||
| PR CI failing >2h | HIGH |
|
||||
|
||||
---
|
||||
|
||||
### Audit 6 — Supervisor Health (Zombie Detection via Session Introspection)
|
||||
|
||||
Uses the **OpenCode Server API** to read actual session messages and detect
|
||||
zombie or stuck supervisors.
|
||||
|
||||
**API endpoints used:**
|
||||
- `GET /session` — list all sessions
|
||||
- `GET /session/status` — get status of all sessions
|
||||
- `GET /session/{id}/message?limit=5` — read last 5 messages
|
||||
- `GET /session/{id}/todo` — read the agent's todo list
|
||||
|
||||
**Detection patterns:**
|
||||
|
||||
| Pattern | Signal | Severity |
|
||||
|---------|--------|----------|
|
||||
| Last 3+ tool calls are all `sleep` with zero productive calls | Zombie (likely context exhaustion) | HIGH |
|
||||
| Last 3+ tool calls all returned errors | Stuck in error loop | HIGH |
|
||||
| Same tool+args repeated 3+ times in last 5 messages | Circular loop | HIGH |
|
||||
| Expected supervisor session not running | Missing supervisor | HIGH |
|
||||
|
||||
**Expected supervisors (16 total):**
|
||||
`implementor-pool`, `reviewer-pool`, `tester-pool`, `hunter-pool`,
|
||||
`test-infra-pool`, `architect`, `epic-planner`, `human-liaison`,
|
||||
`agent-evolver`, `arch-guard`, `spec-updater`, `backlog-groomer`,
|
||||
`docs-writer`, `timeline-updater`, `project-owner`, `system-watchdog`
|
||||
|
||||
---
|
||||
|
||||
### Audit 7 — Label and Dependency Compliance
|
||||
|
||||
Checks all open issues for required labels and dependency links.
|
||||
|
||||
| Check | Severity |
|
||||
|-------|----------|
|
||||
| Missing `State/` label | MEDIUM |
|
||||
| Missing `Type/` label | MEDIUM |
|
||||
| Missing `Priority/` label (non-Unverified) | LOW |
|
||||
| Non-Epic/Legendary issue beyond Unverified with no milestone | MEDIUM |
|
||||
| Issue with no parent Epic dependency link | LOW |
|
||||
|
||||
---
|
||||
|
||||
### Audit 8 — Ticket Hierarchy Integrity
|
||||
|
||||
Verifies the Issue → Epic → Legendary hierarchy is intact.
|
||||
|
||||
- Epics must have a parent Legendary dependency link.
|
||||
- Non-completed Epics should have at least 2 child issues.
|
||||
|
||||
---
|
||||
|
||||
### Audit 9 — Test Infrastructure Health
|
||||
|
||||
Monitors CI execution times and failure rates for signs of test suite bloat or
|
||||
flaky tests. (Currently a lightweight check; expanded in future milestones.)
|
||||
|
||||
---
|
||||
|
||||
### Audit 10 — Improvement Generation
|
||||
|
||||
Verifies the system is generating `needs feedback` improvement tickets.
|
||||
|
||||
- Checks for `needs feedback` issues created in the last 24 hours.
|
||||
- Flags if none exist (the `ca-spec-updater` and `ca-agent-evolver` should be
|
||||
generating proposals regularly).
|
||||
|
||||
---
|
||||
|
||||
### Audit 11 — Quick Session Spot-Check (Every Cycle)
|
||||
|
||||
Fast scan of the **3 most recently active** supervisor sessions for obvious
|
||||
policy violations. Runs every cycle (~5 min) because catching `force_merge` or
|
||||
direct master pushes quickly is critical.
|
||||
|
||||
**API endpoints used:**
|
||||
- `GET /session/status` — find active sessions
|
||||
- `GET /session/{id}/message?limit=3` — read last 3 messages only
|
||||
|
||||
**Checks:**
|
||||
|
||||
| Violation | Severity |
|
||||
|-----------|----------|
|
||||
| `force_merge: true` in any tool call arguments | CRITICAL |
|
||||
| `git push` directly to `master`/`main` without a feature branch | CRITICAL |
|
||||
| `type: ignore` suppression written to a file | HIGH |
|
||||
|
||||
---
|
||||
|
||||
### Audit 12 — Deep Session Introspection (Every 6th Cycle, ~30 min)
|
||||
|
||||
Comprehensive analysis of **all 16 supervisor sessions**. Reads recent messages,
|
||||
tool calls, and todo lists to understand what each agent is actually doing.
|
||||
|
||||
**API endpoints used:**
|
||||
- `GET /session/{id}/message?limit=10` — read last 10 messages
|
||||
- `GET /session/{id}/todo` — read the agent's todo list
|
||||
- `GET /session/{id}/children` — list child worker sessions
|
||||
|
||||
**Analysis passes:**
|
||||
|
||||
#### Analysis 1 — Misbehavior Detection
|
||||
|
||||
Scans all tool calls in the last 10 messages for policy violations:
|
||||
- `force_merge: true` usage (CRITICAL)
|
||||
- Closing PRs as duplicates of their tracking issues (HIGH)
|
||||
- Creating issues without `State/` or `Type/` labels (MEDIUM)
|
||||
|
||||
#### Analysis 2 — Progress Assessment via Todo List
|
||||
|
||||
Reads the agent's todo list to assess forward progress:
|
||||
- Flags agents with 0 completed todos and 1+ in-progress todos out of 3+ total
|
||||
(possible stall or spin).
|
||||
|
||||
#### Analysis 3 — Conversation Health Metrics
|
||||
|
||||
Computes per-session metrics from the last 10 messages:
|
||||
|
||||
| Metric | Threshold | Severity |
|
||||
|--------|-----------|----------|
|
||||
| Error rate (errors / total tool calls) | >50% with ≥4 calls | HIGH |
|
||||
| Sleep ratio (sleep calls / total tool calls) | >80% with ≥5 calls | MEDIUM |
|
||||
|
||||
#### Analysis 4 — Context Exhaustion Signals
|
||||
|
||||
Scans text outputs for phrases indicating the agent is hitting its context window:
|
||||
`"context limit"`, `"context window"`, `"running out of"`, `"token limit"`,
|
||||
`"maximum length"`, etc.
|
||||
|
||||
#### Analysis 5 — Cross-Agent Conflict Detection
|
||||
|
||||
Looks for multiple agents interacting with the same PR or issue:
|
||||
- 3+ different agents touching the same PR → `cross_agent_pr_conflict` (MEDIUM)
|
||||
- 3+ different agents modifying the same issue labels → `cross_agent_issue_conflict` (MEDIUM)
|
||||
|
||||
---
|
||||
|
||||
## Finding Severity Levels
|
||||
|
||||
| Severity | Response |
|
||||
|----------|----------|
|
||||
| **CRITICAL** | Immediate one-off agent dispatch + bug issue creation |
|
||||
| **HIGH** | Issue creation or zombie alert posted to session state issue |
|
||||
| **MEDIUM** | Tracked; issue created if finding persists for 3+ cycles |
|
||||
| **LOW** | Tracked; no immediate action |
|
||||
|
||||
---
|
||||
|
||||
## Action Dispatch
|
||||
|
||||
When a CRITICAL finding is detected, the watchdog dispatches a one-off fix agent
|
||||
via the OpenCode Server API:
|
||||
|
||||
```bash
|
||||
# Create a one-off session
|
||||
SESSION_ID=$(curl -s -X POST "${SERVER}/session" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"title": "[CA-AUTO] one-off: ca-quality-enforcer — missing_ci"}' \
|
||||
| jq -r '.id')
|
||||
|
||||
# Dispatch the agent
|
||||
curl -s -X POST "${SERVER}/session/${SESSION_ID}/prompt_async" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"agent": "ca-quality-enforcer",
|
||||
"parts": [{"type": "text", "text": "Fix this finding: ..."}]}'
|
||||
```
|
||||
|
||||
**Dispatch table:**
|
||||
|
||||
| Finding Type | Action |
|
||||
|-------------|--------|
|
||||
| `no_branch_protection`, `status_check_disabled` | Dispatch `ca-quality-enforcer` |
|
||||
| `merged_without_ci`, `failing_ci_on_master` | Dispatch `ca-quality-enforcer` + create bug issue |
|
||||
| `force_merge_detected` | Create bug issue + post CRITICAL alert |
|
||||
| `direct_push_to_master` | Create bug issue + post CRITICAL alert |
|
||||
| `closed_wrong_state`, `in_review_but_merged` | Dispatch `ca-state-reconciler` |
|
||||
| `zombie_supervisor`, `stuck_supervisor`, `looping_supervisor` | Post zombie alert |
|
||||
| `context_exhaustion` | Post alert (product-builder re-launches the supervisor) |
|
||||
| `high_error_rate` | Post diagnostic comment |
|
||||
| `type_ignore_suppression` | Create `Priority/High` bug issue |
|
||||
|
||||
---
|
||||
|
||||
## Health Reporting
|
||||
|
||||
Every 6 cycles (~30 min), the watchdog posts a health report to the session
|
||||
state issue:
|
||||
|
||||
```
|
||||
[WATCHDOG] Health report — cycle N:
|
||||
- Quality gate violations: 0
|
||||
- State label mismatches: 2
|
||||
- Priority ordering issues: 0
|
||||
- PR pipeline issues: 1
|
||||
- Zombie/stuck/looping supervisors: 0
|
||||
- Missing labels/links: 5
|
||||
- Session introspection findings: 3
|
||||
- Misbehavior (force_merge, direct push): 0
|
||||
- Stuck/looping agents: 1
|
||||
- Context exhaustion signals: 0
|
||||
- Cross-agent conflicts: 2
|
||||
- One-off agents dispatched this period: 1
|
||||
- Issues created this period: 2
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
The watchdog is configured in `.opencode/agents/ca-system-watchdog.md`.
|
||||
|
||||
Key parameters:
|
||||
- **Model:** `anthropic/claude-opus-4-6` (highest capability for complex analysis)
|
||||
- **Temperature:** `0.1` (deterministic auditing)
|
||||
- **Cycle interval:** 5 minutes
|
||||
- **Deep introspection interval:** Every 6th cycle (~30 min)
|
||||
- **Context self-management:** Every 20 cycles, discard accumulated tool outputs
|
||||
|
||||
---
|
||||
|
||||
## Constraints
|
||||
|
||||
The watchdog operates under strict constraints to avoid interfering with the
|
||||
build system:
|
||||
|
||||
- **No filesystem access** — works exclusively through APIs
|
||||
- **No merging PRs** — never calls `forgejo_merge_pull_request`
|
||||
- **No modifying agent definitions** — creates `needs feedback` issues instead
|
||||
- **No direct master pushes** — all corrections go through one-off agents
|
||||
- **Corrections limited to:** state label fixes, dependency link fixes, and
|
||||
issue creation
|
||||
|
||||
---
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [Quality Automation](quality-automation.md) — CI/CD quality gates
|
||||
- [CI/CD Pipeline](ci-cd.md) — pipeline architecture
|
||||
- [Review Playbook](review_playbook.md) — PR review process
|
||||
@@ -29,6 +29,7 @@ nav:
|
||||
- Review Playbook: development/review_playbook.md
|
||||
- Scale Testing: development/scale_testing.md
|
||||
- Ops Runbook: development/ops-runbook.md
|
||||
- System Watchdog: development/system-watchdog.md
|
||||
- Implementation Timeline: timeline.md
|
||||
- FAQ: faq.md
|
||||
- Reference: reference/
|
||||
|
||||
Reference in New Issue
Block a user