Commit Graph

14 Commits

Author SHA1 Message Date
HAL9000 c87fc3bb2a fix: Scale implementation orchestrator to 32 parallel workers
- Reduce main dispatch loop sleep from 10s to 2s (5x faster cycles)
- Simplify worker verification from 5 retries to 1 quick check
- Remove unnecessary delays between dispatch operations
- Reduce retry delays from 15s to 2s for faster recovery
- Reduce idle sleep from 60s to 10s for quicker response
- Add optimistic verification to trust dispatch success

These changes enable the orchestrator to scale from 1-4 workers to the
full 32 workers within seconds instead of minutes, dramatically increasing
system throughput and allowing autonomous unblocking of CI failures.
2026-04-09 01:23:21 -04:00
freemo b72b827525 fix: centralize automation tracking to prevent cycle reuse issues
- Create automation-tracking-manager subagent as single source of truth
- Migrate 7 key agents to use centralized tracking manager
- Fix AUTO-WATCHDOG skipping cycles 22-23 (was commenting on old issues)
- Fix AUTO-IMP-POOL creating duplicate tracking issues for same cycle
- Fix AUTO-TIME and AUTO-PROJ-OWN potential issue reuse patterns
- Ensure cycle numbers persist across agent restarts
- Delete shared/automation_tracking.md in favor of subagent pattern

The new system ensures:
- One tracking issue per cycle (never reuse old issues)
- Sequential cycle numbers that persist across restarts
- Proper cleanup of previous cycles before creating new ones
- Consistent tracking patterns across all agents
- Impossible for agents to comment on old tracking issues

Migrated agents:
- system-watchdog (most problematic - missing cycles)
- implementation-orchestrator (duplicate issues)
- timeline-updater (potential reuse)
- project-owner (potential reuse)
- product-builder (critical orchestrator)
- backlog-groomer (for consistency)

Fixes the issue where agents incorrectly report future cycles as comments
on older status update tickets instead of creating new tracking issues.
2026-04-09 01:08:08 -04:00
freemo 50096391b5 feat: complete comprehensive worker tracking system implementation
- Enhanced product-builder with detailed session monitoring via OpenCode API
- Added comprehensive worker status reporting with session details, targets, and activity
- Enhanced implementation-orchestrator with detailed worker tracking and health monitoring
- Enhanced continuous-pr-reviewer with detailed worker status and progress reporting
- Enhanced uat-tester with detailed worker monitoring and testing progress
- All supervisors now provide detailed visibility into worker activities and health
- Added actual cycle time calculation using timestamps throughout all agents
- Implemented proper tracking issue lifecycle (delete previous, create new each cycle)
- Added stale worker detection and restart functionality across all pool supervisors
- Ensured redundant monitoring between product-builder and individual supervisors

This completes the comprehensive worker tracking system providing full visibility
into all 16 supervisors and their workers with detailed status reporting, automatic
restarts, and actual timing data.
2026-04-09 02:44:00 +00:00
HAL9000 30b457b090 Fix implementation-orchestrator worker dispatch verification
The orchestrator was failing to dispatch workers due to incorrect parsing
of the OpenCode API /session/status response format.

Changes:
- Fixed verify_worker_started() to handle dict response format instead of array
- Check for session_id key and type='busy' instead of status='active'
- Increased verification retries from 3 to 5 with progressive delays
- Enhanced error messages to show actual verification results
- Improved session state handling with longer initialization wait times

This fix allows the orchestrator to correctly verify that workers have
started, preventing it from incorrectly deleting valid worker sessions.
Workers should now dispatch successfully and PRs will be processed.
2026-04-08 22:12:23 -04:00
freemo 1b83d15920 fix: comprehensive tracking issue system improvements
- Fix tracking issue lifecycle: each cycle closes old issue and creates new one
- Add tracking functionality to 4 missing supervisors (architect, timeline-updater, docs-writer, architecture-guard)
- Enhance product-builder to report all 16 supervisors with worker counts
- Add actual cycle time calculation based on elapsed timestamps
- Standardize tracking issue format across all agents
- Implement automatic supervisor re-launch when missing
- Add comprehensive supervisor and worker count monitoring

Fixes tracking issue problems where agents were appending to old issues
instead of creating fresh ones each cycle, and ensures all 16 supervisors
are properly monitored and tracked.
2026-04-09 01:46:44 +00:00
CleverAgents Build Agent 0edc1bf13d refactor!: migrate agents from session state to individual tracking issues
BREAKING CHANGE: Migrate all CleverAgents from shared session state issue
system to individual tracking issues with 'Automation Tracking' labels

Changes:
- Replace SESSION_STATE_ISSUE_NUMBER with individual tracking issues
- Add automation tracking systems to 10 core agents
- Implement standardized agent prefixes (AUTO-UAT-POOL, AUTO-PROJ-OWN, etc.)
- Add cleanup protocols for one-issue-per-cycle management
- Remove session state dependencies from supervisor launch prompts
- Update health signaling to create individual tracking issues
- Preserve announcement issues while cleaning up cycle reports

Affected agents:
- agent-evolver.md: Added AUTO-EVLV tracking system
- bug-hunter.md: Updated tracking documentation
- epic-planner.md: Fixed remaining session state reference
- implementation-orchestrator.md: Updated health signaling
- product-builder.md: Major refactor of supervisor coordination
- project-owner.md: Added AUTO-PROJ-OWN tracking system
- spec-updater.md: Added AUTO-SPEC-UPD tracking system
- test-infra-improver.md: Added AUTO-TEST-INFRA tracking system
- uat-tester.md: Added AUTO-UAT-POOL tracking system

Benefits:
- Better isolation: no shared state conflicts between agents
- Cleaner tracking: one issue per agent per cycle
- Full traceability: each agent's work is independently tracked
- Systematic discovery: standardized labels enable monitoring

This migration follows the automation tracking specification in
.opencode/agents/shared/automation_tracking.md and maintains
compatibility with existing CleverAgents infrastructure.
2026-04-08 19:57:38 -04:00
freemo 014033eed9 feat: enhance automation tracking with health monitoring and recovery
Add comprehensive automated health monitoring and recovery capabilities
to the automation tracking system for proactive agent management.

**Major Enhancements:**

1. **Standardized Interval Reporting**
   - Mandatory interval declaration in all tracking issues
   - Format: 'Reporting Interval: <interval> (Next report expected: <timestamp>)'
   - Enables precise staleness detection and recovery triggering

2. **Automated Health Monitoring (system-watchdog)**
   - New audit_automation_tracking_health() function runs every 5 minutes
   - Monitors all issues with 'Automation Tracking' label
   - Detects stalled agents when >20% overdue from expected interval
   - Calculates staleness ratios and time overdue metrics

3. **Automated Recovery System**
   - Kills stalled agent sessions via OpenCode Server API (port 4096)
   - Performs root cause analysis of session messages and agent definitions
   - Creates high-priority diagnostic issues with detailed findings
   - Automatically closes stale tracking issues with recovery notes
   - Provides human-readable remediation recommendations

**Agent Updates with Standardized Format:**

- **implementation-orchestrator**: Status updates (5 cycles) + health reports (10 cycles)
- **backlog-groomer**: Grooming reports (5 min) + health reports (50 min)
- **human-liaison**: Status updates (20 min monitoring cycles)
- **session-persister**: Event-driven checkpoints with standardized format
- **system-watchdog**: Enhanced with comprehensive recovery capabilities

**Template Standardization:**
- Unified header format across all tracking issues
- Health indicators and next actions sections
- Consistent metadata and automation signatures
- Support for active/warning/error status indicators

**Documentation Updates:**
- Comprehensive automated recovery process documentation
- Agent interval reference table with all timing details
- Recovery issue format and diagnostic workflow
- Health check algorithm and staleness threshold explanation

**Benefits:**
- Proactive detection of crashed or stuck agents (20% staleness threshold)
- Automated recovery reduces manual intervention requirements
- Root cause analysis provides actionable diagnostic information
- Standardized format improves searchability and monitoring
- Comprehensive health metrics enable system-wide visibility

This enhancement transforms the automation tracking system from passive
logging to active health monitoring with automated recovery capabilities.
2026-04-08 22:34:23 +00:00
freemo a323f07783 feat: implement new automation tracking system for agent supervision
Replace shared session state issue tracking with individual tracking issues
per agent to reduce noise and improve searchability.

**Agent Updates:**
- session-persister: [AUTO-SESSION] prefix with cycle management
- implementation-orchestrator: [AUTO-IMP-POOL] prefix for health reports
- system-watchdog: [AUTO-WATCHDOG] prefix for system health
- backlog-groomer: [AUTO-GROOMER] prefix + backup cleanup functionality
- human-liaison: [AUTO-LIAISON] prefix for status updates

**New Features:**
- Standardized issue title format: [AUTO-<PREFIX>] <TYPE> (Cycle <N>)
- Announcement format: [AUTO-<PREFIX>] Announce: <message>
- Automatic cleanup to prevent issue accumulation
- Required 'Automation Tracking' label for filtering
- Validation script for format compliance

**Documentation:**
- Complete system documentation at docs/development/automation-tracking.md
- Added to mkdocs.yml navigation
- Validation script at scripts/validate_automation_tracking.py

**Benefits:**
- Reduced noise from shared tracking issue
- Better searchability with agent-specific prefixes
- Cleaner history per agent type
- Easier debugging with focused issue threads
- Automatic cleanup prevents accumulation

Closes automation tracking system implementation requirements.
2026-04-08 21:28:46 +00:00
freemo 3b1d6d1931 fix(agents): standardize label handling and prevent label creation
- Quote all specific label references ("State/Verified", "Priority/High", etc.)
- Add explicit 'NEVER create new labels' warnings to all agents
- Ensure agents assume labels exist on Forgejo server
- Fix unquoted label patterns across 12+ agent files
- Standardize label reference format for consistency

Key changes:
* issue-state-updater.md: Fixed state transition label references
* human-liaison.md: Quoted all triage and verification labels
* project-owner.md: Fixed MoSCoW and priority label handling
* backlog-groomer.md: Updated auto-fix label compliance
* pr-api-creator.md: Fixed PR metadata label references
* quality-enforcer.md: Fixed CI-Blocker label handling
* state-reconciler.md: Fixed reconciliation label patterns
* new-issue-creator.md: Added comprehensive label usage rules
* issue-finder.md: Fixed priority sorting label references
* spec-updater.md: Fixed proposal label handling
* implementation-orchestrator.md: Fixed CI-Blocker prioritization
* milestone-reviewer.md: Fixed issue creation label references

Resolves label capitalization, spelling, spacing, and creation issues
across the entire agent system to ensure exact Forgejo server matching.
2026-04-08 21:04:34 +00:00
freemo 7ddd6a7e2d feat(agents): add Priority/CI-Blocker label to break PR-first deadlock
**Problem**: 
- Broken CI blocks all PR merges
- PR-first rule blocks CI-fixing issues  
- Creates deadlock where system can't fix itself

**Solution**:
- Created Priority/CI-Blocker label (ID: 1396)
- Added ONE exception to absolute PR-first rule
- Priority/CI-Blocker issues can be worked immediately

**Changes**:
- quality-enforcer: Use Priority/CI-Blocker for CI violations
- implementation-orchestrator: Exception for Priority/CI-Blocker 
- issue-finder: Priority/CI-Blocker as absolute highest priority
- system-watchdog: Create Priority/CI-Blocker for CI failures
- +4 supporting agents updated with new label

**Impact**: 
Prevents CI deadlock while preserving PR-first priority for all other work.
2026-04-08 18:15:32 +00:00
freemo e5f75c5c83 refactor: remove parallelism cap and backpressure throttling
- Remove maximum cap (16) on CA_MAX_PARALLEL_WORKERS in resources.yaml
  - Can now be set to any positive value (32, 64, etc.)
  - Only minimum validation remains (must be > 0)

- Remove dynamic backpressure/throttling from implementation-orchestrator
  - Dispatch always runs at full configured speed
  - Resource monitoring remains for visibility only
  - No automatic reduction of slots_available based on failures

- Convert system-watchdog from auto-degradation to monitoring + suggestions
  - Renamed DEGRADATION_THRESHOLDS to HEALTH_THRESHOLDS
  - Removed apply_system_degradation() and check_degradation_recovery()
  - Changed findings to include suggestions instead of actions
  - Watchdog now reports issues with fix recommendations
  - No automatic throttling or pausing of agents

The system now operates at maximum configured speed at all times,
with the watchdog providing diagnostic insights when issues arise.
2026-04-07 01:13:27 -04:00
freemo 1dd38020d4 feat: strengthen CONTRIBUTING.md compliance in key agents
- Add CRITICAL compliance sections to implementation-orchestrator
- Add explicit instructions to pass CONTRIBUTING.md to all workers
- Strengthen behave-tester compliance requirements
- Improve architect compliance section prominence
- Enhance human-liaison CODE_OF_CONDUCT emphasis

These changes ensure agents don't spend time rediscovering project rules.
2026-04-06 23:40:24 +00:00
freemo 4591ae053d feat(agents): remove ca- prefix to make agents generic
- Rename 72 agent files: ca-{name}.md → {name}.md
- Update all agent references across 76 files:
  - Permission blocks: "ca-agent": allow → "agent": allow
  - Invocations: invoke ca-agent → invoke agent
  - Bot signatures: Agent: ca-agent → Agent: agent
  - Temporary paths: /tmp/ca-* → /tmp/*
  - Clone directories: /tmp/ca-{id} → /tmp/{id}
- Preserve CleverAgents references (190 legitimate uses)
- All agents now have generic names suitable for any project
- Zero broken references remaining
2026-04-06 16:43:49 -04:00
freemo 8b44b265d1 refactor(agents): restructure implementation system for PR-first priority with cost-optimized escalation
- Rename ca-issue-worker → ca-implementation-worker for dual-mode operation (PR fixing + issue implementation)
- Rename issue-implementor → implementation-orchestrator with PR-first priority
- Create new pr-fix-orchestrator for aggressive parallel PR fixing with common cause analysis
- Update escalation order from sonnet→codex→opus to codex→sonnet→opus for better cost optimization
- Update all implementer and tester agents to reflect new escalation tiers
- Add web-based CI log access support since Forgejo Actions API returns 404s
- Update all cross-references and bot signatures throughout codebase
- Add CI_LOG_ACCESS_GUIDE.md with web authentication functions

The system now prioritizes fixing failing PRs over implementing new issues,
uses cost-effective escalation starting with codex, and supports aggressive
parallel execution with intelligent root cause analysis.
2026-04-06 18:46:00 +00:00