- Add 170+ lines of test determinism requirements to behave-tester with forbidden/required patterns
- Add 180+ lines of integration test stability rules to robot-tester
- Enhance pr-self-reviewer with 150+ lines of flaky test detection during code review
- Add emergency master CI monitoring to system-watchdog with auto-skip failing tests
- Implement automatic test skipping system with framework-specific instructions
- Add cross-PR analysis to detect master branch CI issues vs PR-specific failures
- Prohibit label creation in epic-planner and new-issue-creator to prevent duplicates
- Add test stability awareness to implementation-worker for all implementers
This comprehensive system prevents flaky tests from reaching master, automatically
handles CI failures through emergency test skipping, and eliminates label duplication
issues. Includes detailed detection patterns, emergency response workflows, and
framework-specific guidance for Behave, Robot Framework, and generic test systems.
PROBLEM: Primary agents refused to use ci-log-fetcher because documentation incorrectly
suggested they needed to provide forgejo_username/forgejo_password parameters.
SOLUTION: Updated all agents to clarify that ci-log-fetcher handles credentials automatically.
Changes made:
- ci-log-fetcher.md: Updated description and added prominent warning that NO CREDENTIALS are needed
- implementation-worker.md: Removed forgejo_username/forgejo_password from 3 usage examples
- pr-fix-orchestrator.md: Removed credential parameters from 2 usage examples, clarified env var usage
- pr-checker.md: Removed credential parameters from 2 usage examples
Now all agents clearly understand that ci-log-fetcher automatically uses FORGEJO_USERNAME
and FORGEJO_PASSWORD environment variables without any credential parameters needed.
Updated multiple agents to understand and properly handle TDD (Test-Driven
Development) tags as documented in CONTRIBUTING.md. This prevents confusion
when agents encounter tests with @tdd_expected_fail that invert their behavior.
Key changes:
- Test writers (behave-tester, robot-tester) now understand when to use TDD tags
- Implementers know to remove @tdd_expected_fail tags when fixing bugs
- Test-fixer won't try to "fix" correctly passing TDD tests
- PR reviewers check for proper TDD tag removal in bug fix PRs
- Human liaison can explain TDD tags to confused developers
- Coverage improver avoids modifying TDD tests
- Reference reader includes TDD tag info in summaries
This ensures all agents work correctly with the TDD workflow where tests are
written before bug fixes and use special tags to prove bugs exist.
- Remove maximum cap (16) on CA_MAX_PARALLEL_WORKERS in resources.yaml
- Can now be set to any positive value (32, 64, etc.)
- Only minimum validation remains (must be > 0)
- Remove dynamic backpressure/throttling from implementation-orchestrator
- Dispatch always runs at full configured speed
- Resource monitoring remains for visibility only
- No automatic reduction of slots_available based on failures
- Convert system-watchdog from auto-degradation to monitoring + suggestions
- Renamed DEGRADATION_THRESHOLDS to HEALTH_THRESHOLDS
- Removed apply_system_degradation() and check_degradation_recovery()
- Changed findings to include suggestions instead of actions
- Watchdog now reports issues with fix recommendations
- No automatic throttling or pausing of agents
The system now operates at maximum configured speed at all times,
with the watchdog providing diagnostic insights when issues arise.
- Add system-watchdog audit for PRs with 3+ failed attempts
- Implement automatic human assistance requests with detailed analysis
- Add deep context gathering to implementation-worker before fixes
- Enhance all agents with enriched context propagation
- Add loop detection to prevent repetitive failed attempts
- Improve PR reviewer with anti-pattern detection
- Update human-liaison to provide targeted help for struggling PRs
- Add historical awareness to PR fix orchestrator
- Enhance epic-planner with context-aware issue creation
- Create documentation for improvements and future agent ideas
These changes enable the system to:
- Recognize when it's stuck and needs human help
- Learn from previous failures to avoid repetition
- Understand full context including comments and history
- Provide detailed debugging information to humans
BREAKING CHANGE: All PR-related agents must now use ci-log-fetcher for CI logs
Issues fixed:
- pr-checker: Now invokes ci-log-fetcher instead of manual web scraping
- implementation-worker: Uses ci-log-fetcher for pr-fix mode CI analysis
- pr-self-reviewer: Checks CI status and fetches logs before reviewing
- human-liaison: Fetches CI logs when responding to PR comments/reviews
- pr-fix-orchestrator: Uses ci-log-fetcher instead of manual implementation
- pr-status-checker: Uses ci-log-fetcher when include_logs=true
Key changes:
1. Added ci-log-fetcher permission to all PR agents
2. Replaced manual CI log fetching implementations with ci-log-fetcher calls
3. Updated human-liaison with new PR response behavior that checks CI status
4. Fixed pr-fix-orchestrator visibility (now properly hidden)
5. Ensured read-only agents understand full PR context including CI failures
This ensures:
- Consistent CI log access across all agents
- No duplicate web scraping implementations
- Better context for PR reviews and human interactions
- Agents never run tests locally to understand failures
- All agents have complete picture of PR status before acting