docs(development): add system watchdog architecture and audit reference #2449

Merged
freemo merged 1 commits from docs/system-watchdog-documentation into master 2026-04-05 21:24:06 +00:00
2 changed files with 351 additions and 0 deletions
+350
View File
@@ -0,0 +1,350 @@
# System Watchdog
The **System Watchdog** (`ca-system-watchdog`) is the 16th supervisor agent in the
CleverAgents autonomous build system. It runs continuously with a 5-minute polling
cycle and monitors the entire agent ecosystem for correctness, health, and policy
compliance.
> **Audience:** Contributors and operators who need to understand how the autonomous
> agent system is monitored, what violations are detected, and how the watchdog
> responds to findings.
---
## Overview
The watchdog operates exclusively through the **Forgejo API** and the
**OpenCode Server HTTP API** — it does not clone the repository or access the
filesystem. It performs 12 audits per cycle, ranging from quality gate compliance
to deep session introspection of individual supervisor conversations.
```
cycle = 0
LOOP FOREVER (5-minute intervals):
cycle += 1
findings = []
findings += audit_quality_gates() # Audit 1 — MOST CRITICAL
findings += audit_branch_protection() # Audit 2
findings += audit_ticket_states() # Audit 3
findings += audit_priority_ordering() # Audit 4
findings += audit_pr_pipeline() # Audit 5
findings += audit_supervisor_health() # Audit 6 (session introspection)
findings += audit_labels_and_dependencies() # Audit 7
findings += audit_ticket_hierarchy() # Audit 8
findings += audit_test_health() # Audit 9
findings += audit_improvement_generation() # Audit 10
findings += audit_session_spot_check() # Audit 11 (every cycle)
if cycle % 6 == 0:
findings += audit_deep_session_introspection() # Audit 12 (every 30 min)
for finding in findings:
take_action(finding)
if cycle % 6 == 0 and findings:
post_summary(cycle, findings)
sleep(300) # 5 minutes
```
---
## Audit Descriptions
### Audit 1 — Quality Gate Compliance (CRITICAL)
Ensures **no code reaches `master` without passing all CI checks**.
- Checks the last 10 commits on `master` for passing `status-check` CI status.
- Verifies recently merged PRs had passing CI at merge time.
- Detects direct pushes to `master` (all changes must go through PRs).
**Severity:** CRITICAL for any violation.
---
### Audit 2 — Branch Protection Verification
Verifies that Forgejo branch protection is active and correctly configured for
`master`.
Checks:
- Branch protection rules exist for `master`.
- `enable_status_check` is `true`.
- The `status-check` context is required.
- At least 2 approving reviews are required (per `CONTRIBUTING.md`).
---
### Audit 3 — Ticket State Integrity
Ensures all issues have correct `State/` labels matching their actual state.
Checks:
- Closed issues have `State/Completed` (not `State/Unverified`).
- Issues with `State/In Review` have an open PR referencing them.
- No issue has multiple `State/` labels simultaneously.
---
### Audit 4 — Priority and Milestone Ordering
Ensures Critical bugs on lower milestones are addressed before feature work on
later milestones.
- Finds the lowest milestone with `Priority/Critical` or `MoSCoW/Must Have` bugs.
- Flags any non-bug issue in `State/In Progress` on a later milestone.
---
### Audit 5 — PR Pipeline Health
Tracks PR aging, review coverage, and merge throughput.
| Condition | Severity |
|-----------|----------|
| PR open >24h with no reviews | MEDIUM |
| PR approved but not merged >6h | HIGH |
| PR CI failing >2h | HIGH |
---
### Audit 6 — Supervisor Health (Zombie Detection via Session Introspection)
Uses the **OpenCode Server API** to read actual session messages and detect
zombie or stuck supervisors.
**API endpoints used:**
- `GET /session` — list all sessions
- `GET /session/status` — get status of all sessions
- `GET /session/{id}/message?limit=5` — read last 5 messages
- `GET /session/{id}/todo` — read the agent's todo list
**Detection patterns:**
| Pattern | Signal | Severity |
|---------|--------|----------|
| Last 3+ tool calls are all `sleep` with zero productive calls | Zombie (likely context exhaustion) | HIGH |
| Last 3+ tool calls all returned errors | Stuck in error loop | HIGH |
| Same tool+args repeated 3+ times in last 5 messages | Circular loop | HIGH |
| Expected supervisor session not running | Missing supervisor | HIGH |
**Expected supervisors (16 total):**
`implementor-pool`, `reviewer-pool`, `tester-pool`, `hunter-pool`,
`test-infra-pool`, `architect`, `epic-planner`, `human-liaison`,
`agent-evolver`, `arch-guard`, `spec-updater`, `backlog-groomer`,
`docs-writer`, `timeline-updater`, `project-owner`, `system-watchdog`
---
### Audit 7 — Label and Dependency Compliance
Checks all open issues for required labels and dependency links.
| Check | Severity |
|-------|----------|
| Missing `State/` label | MEDIUM |
| Missing `Type/` label | MEDIUM |
| Missing `Priority/` label (non-Unverified) | LOW |
| Non-Epic/Legendary issue beyond Unverified with no milestone | MEDIUM |
| Issue with no parent Epic dependency link | LOW |
---
### Audit 8 — Ticket Hierarchy Integrity
Verifies the Issue → Epic → Legendary hierarchy is intact.
- Epics must have a parent Legendary dependency link.
- Non-completed Epics should have at least 2 child issues.
---
### Audit 9 — Test Infrastructure Health
Monitors CI execution times and failure rates for signs of test suite bloat or
flaky tests. (Currently a lightweight check; expanded in future milestones.)
---
### Audit 10 — Improvement Generation
Verifies the system is generating `needs feedback` improvement tickets.
- Checks for `needs feedback` issues created in the last 24 hours.
- Flags if none exist (the `ca-spec-updater` and `ca-agent-evolver` should be
generating proposals regularly).
---
### Audit 11 — Quick Session Spot-Check (Every Cycle)
Fast scan of the **3 most recently active** supervisor sessions for obvious
policy violations. Runs every cycle (~5 min) because catching `force_merge` or
direct master pushes quickly is critical.
**API endpoints used:**
- `GET /session/status` — find active sessions
- `GET /session/{id}/message?limit=3` — read last 3 messages only
**Checks:**
| Violation | Severity |
|-----------|----------|
| `force_merge: true` in any tool call arguments | CRITICAL |
| `git push` directly to `master`/`main` without a feature branch | CRITICAL |
| `type: ignore` suppression written to a file | HIGH |
---
### Audit 12 — Deep Session Introspection (Every 6th Cycle, ~30 min)
Comprehensive analysis of **all 16 supervisor sessions**. Reads recent messages,
tool calls, and todo lists to understand what each agent is actually doing.
**API endpoints used:**
- `GET /session/{id}/message?limit=10` — read last 10 messages
- `GET /session/{id}/todo` — read the agent's todo list
- `GET /session/{id}/children` — list child worker sessions
**Analysis passes:**
#### Analysis 1 — Misbehavior Detection
Scans all tool calls in the last 10 messages for policy violations:
- `force_merge: true` usage (CRITICAL)
- Closing PRs as duplicates of their tracking issues (HIGH)
- Creating issues without `State/` or `Type/` labels (MEDIUM)
#### Analysis 2 — Progress Assessment via Todo List
Reads the agent's todo list to assess forward progress:
- Flags agents with 0 completed todos and 1+ in-progress todos out of 3+ total
(possible stall or spin).
#### Analysis 3 — Conversation Health Metrics
Computes per-session metrics from the last 10 messages:
| Metric | Threshold | Severity |
|--------|-----------|----------|
| Error rate (errors / total tool calls) | >50% with ≥4 calls | HIGH |
| Sleep ratio (sleep calls / total tool calls) | >80% with ≥5 calls | MEDIUM |
#### Analysis 4 — Context Exhaustion Signals
Scans text outputs for phrases indicating the agent is hitting its context window:
`"context limit"`, `"context window"`, `"running out of"`, `"token limit"`,
`"maximum length"`, etc.
#### Analysis 5 — Cross-Agent Conflict Detection
Looks for multiple agents interacting with the same PR or issue:
- 3+ different agents touching the same PR → `cross_agent_pr_conflict` (MEDIUM)
- 3+ different agents modifying the same issue labels → `cross_agent_issue_conflict` (MEDIUM)
---
## Finding Severity Levels
| Severity | Response |
|----------|----------|
| **CRITICAL** | Immediate one-off agent dispatch + bug issue creation |
| **HIGH** | Issue creation or zombie alert posted to session state issue |
| **MEDIUM** | Tracked; issue created if finding persists for 3+ cycles |
| **LOW** | Tracked; no immediate action |
---
## Action Dispatch
When a CRITICAL finding is detected, the watchdog dispatches a one-off fix agent
via the OpenCode Server API:
```bash
# Create a one-off session
SESSION_ID=$(curl -s -X POST "${SERVER}/session" \
-H "Content-Type: application/json" \
-d '{"title": "[CA-AUTO] one-off: ca-quality-enforcer — missing_ci"}' \
| jq -r '.id')
# Dispatch the agent
curl -s -X POST "${SERVER}/session/${SESSION_ID}/prompt_async" \
-H "Content-Type: application/json" \
-d '{"agent": "ca-quality-enforcer",
"parts": [{"type": "text", "text": "Fix this finding: ..."}]}'
```
**Dispatch table:**
| Finding Type | Action |
|-------------|--------|
| `no_branch_protection`, `status_check_disabled` | Dispatch `ca-quality-enforcer` |
| `merged_without_ci`, `failing_ci_on_master` | Dispatch `ca-quality-enforcer` + create bug issue |
| `force_merge_detected` | Create bug issue + post CRITICAL alert |
| `direct_push_to_master` | Create bug issue + post CRITICAL alert |
| `closed_wrong_state`, `in_review_but_merged` | Dispatch `ca-state-reconciler` |
| `zombie_supervisor`, `stuck_supervisor`, `looping_supervisor` | Post zombie alert |
| `context_exhaustion` | Post alert (product-builder re-launches the supervisor) |
| `high_error_rate` | Post diagnostic comment |
| `type_ignore_suppression` | Create `Priority/High` bug issue |
---
## Health Reporting
Every 6 cycles (~30 min), the watchdog posts a health report to the session
state issue:
```
[WATCHDOG] Health report — cycle N:
- Quality gate violations: 0
- State label mismatches: 2
- Priority ordering issues: 0
- PR pipeline issues: 1
- Zombie/stuck/looping supervisors: 0
- Missing labels/links: 5
- Session introspection findings: 3
- Misbehavior (force_merge, direct push): 0
- Stuck/looping agents: 1
- Context exhaustion signals: 0
- Cross-agent conflicts: 2
- One-off agents dispatched this period: 1
- Issues created this period: 2
```
---
## Configuration
The watchdog is configured in `.opencode/agents/ca-system-watchdog.md`.
Key parameters:
- **Model:** `anthropic/claude-opus-4-6` (highest capability for complex analysis)
- **Temperature:** `0.1` (deterministic auditing)
- **Cycle interval:** 5 minutes
- **Deep introspection interval:** Every 6th cycle (~30 min)
- **Context self-management:** Every 20 cycles, discard accumulated tool outputs
---
## Constraints
The watchdog operates under strict constraints to avoid interfering with the
build system:
- **No filesystem access** — works exclusively through APIs
- **No merging PRs** — never calls `forgejo_merge_pull_request`
- **No modifying agent definitions** — creates `needs feedback` issues instead
- **No direct master pushes** — all corrections go through one-off agents
- **Corrections limited to:** state label fixes, dependency link fixes, and
issue creation
---
## Related Documentation
- [Quality Automation](quality-automation.md) — CI/CD quality gates
- [CI/CD Pipeline](ci-cd.md) — pipeline architecture
- [Review Playbook](review_playbook.md) — PR review process
+1
View File
@@ -29,6 +29,7 @@ nav:
- Review Playbook: development/review_playbook.md
- Scale Testing: development/scale_testing.md
- Ops Runbook: development/ops-runbook.md
- System Watchdog: development/system-watchdog.md
- Implementation Timeline: timeline.md
- FAQ: faq.md
- Reference: reference/