docs/v360/align-depth-reduction-devcontainer
22 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
435e409df9
|
build: moved all sonnet agents to haiku
CI / benchmark-regression (push) Failing after 0s
CI / benchmark-publish (push) Failing after 0s
CI / push-validation (push) Successful in 32s
CI / helm (push) Failing after 42s
CI / build (push) Successful in 3m59s
CI / lint (push) Successful in 4m10s
CI / quality (push) Successful in 4m37s
CI / typecheck (push) Successful in 4m48s
CI / security (push) Successful in 4m57s
CI / e2e_tests (push) Successful in 7m13s
CI / integration_tests (push) Successful in 10m40s
CI / unit_tests (push) Successful in 11m47s
CI / docker (push) Failing after 46s
CI / coverage (push) Successful in 14m54s
CI / status-check (push) Failing after 3s
CI / helm (pull_request) Successful in 37s
CI / push-validation (pull_request) Successful in 22s
CI / build (pull_request) Successful in 4m0s
CI / lint (pull_request) Successful in 4m37s
CI / quality (pull_request) Successful in 4m37s
CI / typecheck (pull_request) Successful in 4m55s
CI / security (pull_request) Successful in 5m23s
CI / integration_tests (pull_request) Successful in 8m16s
CI / e2e_tests (pull_request) Successful in 8m20s
CI / unit_tests (pull_request) Successful in 9m27s
CI / docker (pull_request) Successful in 1m48s
CI / coverage (pull_request) Successful in 15m1s
CI / status-check (pull_request) Successful in 3s
|
||
|
|
0257841825
|
Revert "refactor(agents): migrate all agent definitions to use skills for universal rules"
CI / build (push) Successful in 18s
CI / helm (push) Successful in 30s
CI / typecheck (push) Successful in 50s
CI / push-validation (push) Successful in 21s
CI / lint (push) Successful in 3m19s
CI / e2e_tests (push) Failing after 3m21s
CI / quality (push) Successful in 3m50s
CI / security (push) Successful in 4m13s
CI / integration_tests (push) Successful in 9m9s
CI / unit_tests (push) Successful in 9m9s
CI / docker (push) Successful in 8s
CI / coverage (push) Successful in 8m26s
CI / status-check (push) Failing after 1s
CI / benchmark-publish (push) Successful in 1h14m43s
CI / benchmark-regression (push) Has been skipped
This reverts commit
|
||
|
|
bb97f1450e
|
refactor(agents): migrate all agent definitions to use skills for universal rules
CI / push-validation (push) Successful in 16s
CI / lint (push) Successful in 18s
CI / typecheck (push) Successful in 31s
CI / helm (push) Successful in 31s
CI / build (push) Successful in 32s
CI / e2e_tests (push) Successful in 3m27s
CI / quality (push) Successful in 3m43s
CI / integration_tests (push) Successful in 4m0s
CI / security (push) Successful in 4m11s
CI / unit_tests (push) Successful in 8m38s
CI / coverage (push) Successful in 5m38s
CI / docker (push) Successful in 1m30s
CI / status-check (push) Successful in 1s
CI / benchmark-publish (push) Successful in 1h13m4s
CI / benchmark-regression (push) Has been skipped
Replace ~600 chars of verbatim per-agent boilerplate with skill references. All 91 agents now load cleveragents-agent-rules for exhaustive pagination, label management, bot signatures, and credential flow rules. Adds explicit skill: "*": deny + targeted allows to every agent permission block, matching the existing bash: and task: deny-first convention. Tier selectors carry no skill permissions since they are pure pass-through with no skill references in their bodies. forgejo-label-manager also grants forgejo-api for its curl pattern reference. |
||
|
|
93c349d531 |
Build: Stopped using codex for most of our agents
CI / helm (push) Successful in 29s
CI / push-validation (push) Successful in 34s
CI / lint (push) Successful in 3m22s
CI / build (push) Successful in 3m48s
CI / integration_tests (push) Successful in 3m57s
CI / quality (push) Successful in 4m13s
CI / typecheck (push) Successful in 4m30s
CI / security (push) Successful in 4m50s
CI / e2e_tests (push) Successful in 6m28s
CI / unit_tests (push) Successful in 9m8s
CI / coverage (push) Has been cancelled
CI / docker (push) Has been cancelled
CI / benchmark-publish (push) Has been cancelled
CI / status-check (push) Has been cancelled
CI / benchmark-regression (push) Has been cancelled
|
||
|
|
5a57eb9a07
|
Build: doom loops are detected and killed
CI / push-validation (push) Successful in 17s
CI / build (push) Successful in 36s
CI / lint (push) Successful in 38s
CI / helm (push) Successful in 43s
CI / typecheck (push) Successful in 50s
CI / quality (push) Successful in 3m42s
CI / integration_tests (push) Successful in 3m59s
CI / security (push) Successful in 4m8s
CI / e2e_tests (push) Successful in 4m17s
CI / unit_tests (push) Successful in 8m29s
CI / docker (push) Successful in 10s
CI / coverage (push) Successful in 13m52s
CI / status-check (push) Successful in 1s
|
||
|
|
acb901abf1
|
Build: Refined some of the wording in the supervisors to get more reliable performance out of them. Made permissions stricter so we will get less circumvention of intended permissions
CI / push-validation (push) Successful in 16s
CI / helm (push) Successful in 17s
CI / typecheck (push) Successful in 1m4s
CI / build (push) Successful in 3m19s
CI / lint (push) Successful in 3m19s
CI / quality (push) Successful in 3m49s
CI / integration_tests (push) Successful in 4m0s
CI / security (push) Successful in 4m5s
CI / e2e_tests (push) Successful in 6m12s
CI / unit_tests (push) Successful in 9m35s
CI / docker (push) Successful in 1m31s
CI / coverage (push) Successful in 13m52s
CI / status-check (push) Successful in 1s
|
||
|
|
78cfdc9b1b
|
Build: Improved merge, review, and implementor logic to have better and more clear priorities
CI / push-validation (push) Successful in 10s
CI / helm (push) Successful in 24s
CI / lint (push) Successful in 35s
CI / build (push) Successful in 40s
CI / typecheck (push) Successful in 48s
CI / e2e_tests (push) Successful in 3m26s
CI / quality (push) Successful in 3m53s
CI / security (push) Successful in 4m5s
CI / integration_tests (push) Successful in 6m26s
CI / unit_tests (push) Successful in 7m28s
CI / docker (push) Successful in 1m36s
CI / coverage (push) Successful in 17m32s
CI / status-check (push) Successful in 1s
|
||
|
|
38bcd41338
|
Build: Better protection against agents editing the main working directory
CI / lint (push) Successful in 24s
CI / typecheck (push) Successful in 54s
CI / quality (push) Successful in 45s
CI / security (push) Successful in 1m15s
CI / build (push) Successful in 29s
CI / push-validation (push) Successful in 30s
CI / helm (push) Successful in 37s
CI / e2e_tests (push) Successful in 3m39s
CI / integration_tests (push) Successful in 4m28s
CI / unit_tests (push) Successful in 5m22s
CI / docker (push) Successful in 21s
CI / coverage (push) Successful in 11m39s
CI / status-check (push) Successful in 1s
|
||
|
|
a0664ad662
|
Build: enforce pagination with agents
CI / status-check (push) Blocked by required conditions
CI / push-validation (push) Successful in 17s
CI / helm (push) Successful in 31s
CI / quality (push) Successful in 43s
CI / typecheck (push) Successful in 55s
CI / lint (push) Successful in 3m20s
CI / build (push) Successful in 3m23s
CI / security (push) Successful in 4m5s
CI / integration_tests (push) Successful in 4m14s
CI / e2e_tests (push) Successful in 7m21s
CI / unit_tests (push) Successful in 8m22s
CI / docker (push) Successful in 10s
CI / coverage (push) Failing after 21m53s
|
||
|
|
ad069c2012
|
Build: Fixed forgejo tool permissions in the agents and fixed the label manager so it reliably works with org level labels
CI / push-validation (push) Successful in 20s
CI / helm (push) Successful in 22s
CI / build (push) Successful in 3m18s
CI / lint (push) Successful in 3m21s
CI / quality (push) Successful in 3m40s
CI / typecheck (push) Successful in 3m59s
CI / security (push) Successful in 4m5s
CI / e2e_tests (push) Successful in 6m10s
CI / integration_tests (push) Successful in 6m12s
CI / unit_tests (push) Successful in 7m31s
CI / docker (push) Successful in 1m32s
CI / coverage (push) Successful in 13m52s
CI / status-check (push) Successful in 1s
|
||
|
|
5a9aaa79ed |
Build: Reinforced label enforcement, and ensure implementation workers dont continue work on a mergable PR.
CI / lint (push) Successful in 39s
CI / quality (push) Successful in 41s
CI / typecheck (push) Successful in 52s
CI / build (push) Successful in 36s
CI / helm (push) Successful in 27s
CI / push-validation (push) Successful in 18s
CI / security (push) Successful in 4m5s
CI / e2e_tests (push) Successful in 3m43s
CI / integration_tests (push) Successful in 4m2s
CI / unit_tests (push) Successful in 5m37s
CI / docker (push) Successful in 22s
CI / coverage (push) Successful in 11m0s
CI / status-check (push) Successful in 1s
CI / benchmark-publish (push) Has been cancelled
CI / benchmark-regression (push) Has been cancelled
|
||
|
|
8692bb46e5 |
build: Refactored agent definitions to be simpler and less contention
CI / benchmark-publish (push) Waiting to run
CI / push-validation (push) Successful in 18s
CI / helm (push) Successful in 25s
CI / lint (push) Successful in 28s
CI / quality (push) Successful in 55s
CI / e2e_tests (push) Successful in 3m4s
CI / build (push) Successful in 3m20s
CI / typecheck (push) Successful in 3m59s
CI / security (push) Successful in 4m5s
CI / benchmark-regression (push) Waiting to run
CI / unit_tests (push) Successful in 7m44s
CI / docker (push) Successful in 1m19s
CI / integration_tests (push) Successful in 9m56s
CI / coverage (push) Successful in 11m47s
CI / status-check (push) Successful in 1s
|
||
|
|
0eca98103e |
fix: replace async-agent-starter with comprehensive async-agent-manager
CI / lint (push) Successful in 20s
CI / quality (push) Successful in 32s
CI / push-validation (push) Successful in 21s
CI / helm (push) Successful in 24s
CI / typecheck (push) Successful in 54s
CI / security (push) Successful in 59s
CI / benchmark-regression (push) Has been skipped
CI / build (push) Successful in 47s
CI / e2e_tests (push) Successful in 3m8s
CI / integration_tests (push) Successful in 4m1s
CI / unit_tests (push) Successful in 4m58s
CI / docker (push) Successful in 10s
CI / coverage (push) Successful in 10m16s
CI / status-check (push) Successful in 2s
CI / benchmark-publish (push) Has been cancelled
- Created new async-agent-manager to handle all async operations centrally - Fixed permission issues where agents couldn't execute curl commands - Updated all agents to use async-agent-manager instead of direct curl - Only async-agent-manager has curl permissions to localhost:4096 - All other agents use it via Task tool with proper permissions - Tested and verified all curl commands work correctly - Added comprehensive operations: start, status, messages, search, cleanup, health monitoring - Improved error handling with structured JSON responses - Enhanced security with proper input escaping This fixes the blocking issue where supervisors couldn't launch workers due to environment restrictions on curl commands. Now all async operations go through a single, well-tested agent with proper permissions. |
||
|
|
5c584c1cab |
feat(agents): Harden label creation restrictions
CI / build (push) Successful in 24s
CI / quality (push) Successful in 31s
CI / push-validation (push) Successful in 16s
CI / helm (push) Successful in 23s
CI / security (push) Successful in 59s
CI / e2e_tests (push) Successful in 3m8s
CI / lint (push) Successful in 3m20s
CI / integration_tests (push) Successful in 4m0s
CI / typecheck (push) Successful in 4m4s
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Successful in 5m9s
CI / docker (push) Successful in 1m20s
CI / coverage (push) Successful in 10m34s
CI / status-check (push) Successful in 1s
CI / benchmark-publish (push) Has been cancelled
- Block REST API endpoints for label creation at the bash level for all agents. - Restrict `forgejo_create_label` and related MCP tools for all agents. - Restrict `forgejo_add_issue_labels` to only the `forgejo-label-manager`. - Ensure all label operations are centralized through the `forgejo-label-manager`. - Update agent definitions to use the label manager instead of direct API calls or MCP tools for adding labels. This prevents agents from creating new project-level labels and enforces the use of organization-level labels, resolving the issue of duplicate labels being created. |
||
|
|
96a70c170e |
feat(agents): add struggling PR detection and deep context understanding
ci.yml / feat(agents): add struggling PR detection and deep context understanding (push) Failing after 0s
- Add system-watchdog audit for PRs with 3+ failed attempts - Implement automatic human assistance requests with detailed analysis - Add deep context gathering to implementation-worker before fixes - Enhance all agents with enriched context propagation - Add loop detection to prevent repetitive failed attempts - Improve PR reviewer with anti-pattern detection - Update human-liaison to provide targeted help for struggling PRs - Add historical awareness to PR fix orchestrator - Enhance epic-planner with context-aware issue creation - Create documentation for improvements and future agent ideas These changes enable the system to: - Recognize when it's stuck and needs human help - Learn from previous failures to avoid repetition - Understand full context including comments and history - Provide detailed debugging information to humans |
||
|
|
2db0646369 |
feat: add async parallel execution to subtask-loop using established abstractions
ci.yml / feat: add async parallel execution to subtask-loop using established abstractions (push) Failing after 0s
- Add permissions for async-agent-starter, async-agent-monitor, async-agent-cleanup
- Convert test writers (behave-tester, robot-tester, asv-benchmarker) to async
- Launch all test writers in parallel via async-agent-starter
- Monitor completion with async-agent-monitor
- Clean up sessions with async-agent-cleanup
- Convert quality gates to async parallel execution
- Launch all 5 quality gates (coverage, lint, typecheck, unit tests, integration tests) in parallel
- Support both first-pass (all gates) and subsequent passes (selective re-runs)
- Monitor and handle failures/restarts
- Implement unique tagging system for session recovery:
- Test writers: AUTO-SUBTASK-{TYPE}-{attempt}
- Quality gates: AUTO-SUBTASK-{GATE}-A{attempt}-P{pass}
- Add monitoring loops with health checks and automatic restart capability
- Maintain existing escalation logic while gaining 60-80% speedup from parallelization
This change eliminates the synchronous bottleneck in subtask-loop where quality
gates and test writers were running sequentially despite the "IN PARALLEL"
pseudo-code. Now they truly run in parallel using the established async
infrastructure, significantly reducing subtask completion time.
|
||
|
|
97aa29d68e |
refactor: restore original high-level behavior with tier selector implementation
ci.yml / refactor: restore original high-level behavior with tier selector implementation (push) Failing after 0s
- Remove quality-gate-escalator.md (unnecessary orchestration layer) - Update subtask-loop to directly manage quality gates with escalation - Quality gates now start at haiku tier for cost efficiency - Inner stabilization loop remains (5 passes) - Failed gates escalate individually after inner loop exhaustion - Maintains original parallel execution behavior - Update test-fixer invocation to use tier selectors - Fix issue-comment-formatter to use new agent naming convention - Change "implementer-opus" to "implementer (tier: opus)" etc. The system now maintains the exact same high-level behavior as before: - subtask-loop directly invokes all quality gates - No new primary agents or orchestration layers - Tier selector architecture is purely an implementation detail - All agents support escalation but primary agent behavior unchanged |
||
|
|
db7e044b18 |
refactor: complete tier selector architecture migration for all agents
ci.yml / refactor: complete tier selector architecture migration for all agents (push) Failing after 0s
BREAKING CHANGE: Removed all tier-specific agents in favor of model-agnostic versions
Key changes:
- Create model-agnostic agents:
- behave-tester.md (replaces 4 tier-specific versions)
- robot-tester.md (replaces 4 tier-specific versions)
- coverage-improver.md (replaces coverage-checker with escalation)
- Convert existing agents to support escalation:
- lint-fixer.md (now model-agnostic)
- test-fixer.md (now model-agnostic)
- integration-test-runner.md (now model-agnostic)
- Update tier selectors to support all new agents
- Update quality-gate-escalator to handle all quality fixers
- Update subtask-loop to use quality-gate-escalator for all quality gates
- Empty redundant tier-specific agents for deletion:
- All implementer-{tier}.md files
- All behave-tester-{tier}.md files
- All robot-tester-{tier}.md files
- coverage-checker.md (replaced by coverage-improver.md)
Benefits:
- Eliminates ~90% code duplication
- All agents now support full 4-tier escalation (haiku→codex→sonnet→opus)
- Consistent escalation behavior across all agent types
- Single source of truth for each agent's logic
- Significant cost savings by defaulting to haiku for all quality gates
The system now uses tier selectors (tier-haiku, tier-codex, tier-sonnet,
tier-opus) that set the model and invoke model-agnostic worker agents,
eliminating the need for separate implementations per model tier.
|
||
|
|
5270987624 |
refactor: remove version suffixes from subtask-loop agent
ci.yml / refactor: remove version suffixes from subtask-loop agent (push) Failing after 0s
- Update subtask-loop.md to use the new tier selector architecture - Remove v2 references - we maintain only the latest version - Effectively remove subtask-loop-v2.md (emptied for deletion) - Keep single source of truth without version suffixes The subtask-loop agent now uses the tier selector architecture without any version references. |
||
|
|
4aa236b7df |
feat: implement tier selector architecture for escalation without redundancy
BREAKING CHANGE: New escalation architecture eliminates duplicate agent code Key changes: - Add tier selector agents (tier-haiku, tier-codex, tier-sonnet, tier-opus) that set the model and invoke worker agents - Create model-agnostic implementer.md that inherits model from caller - Update unit-test-runner and typecheck-fixer to support escalation - Add quality-gate-escalator to manage escalation for quality fixers - Create subtask-loop-v2.md demonstrating the new architecture - Document the new architecture in escalation-architecture-proposal.md Benefits: - Eliminates 90% code duplication across tier-specific agents - Single source of truth for implementation logic - Easy to add new tiers or modify escalation paths - Consistent behavior across all model tiers - Quality gate fixers now support full escalation starting from haiku The architecture leverages the fact that agents without a model specification inherit the model from their caller, allowing thin tier selectors to control which model executes the actual work. |
||
|
|
58b1e50410 |
feat: add 4-tier escalation system with Haiku bottom tier
- Add implementer-haiku.md as ultra-fast, cost-effective first tier - Add behave-tester-haiku.md and robot-tester-haiku.md for fast testing - Update difficulty-evaluator.md to assess 4 tiers (haiku/codex/sonnet/opus) - Update subtask-loop.md with 4-tier escalation logic and state persistence - Add PR comment state tracking for escalation recovery after restarts - Conservative evaluator defaults to Haiku when in doubt for cost efficiency - New escalation path: haiku → haiku → codex → sonnet → opus (forever) - Includes state persistence to resume from correct tier after agent restarts |
||
|
|
4591ae053d |
feat(agents): remove ca- prefix to make agents generic
ci.yml / feat(agents): remove ca- prefix to make agents generic (push) Failing after 0s
- Rename 72 agent files: ca-{name}.md → {name}.md
- Update all agent references across 76 files:
- Permission blocks: "ca-agent": allow → "agent": allow
- Invocations: invoke ca-agent → invoke agent
- Bot signatures: Agent: ca-agent → Agent: agent
- Temporary paths: /tmp/ca-* → /tmp/*
- Clone directories: /tmp/ca-{id} → /tmp/{id}
- Preserve CleverAgents references (190 legitimate uses)
- All agents now have generic names suitable for any project
- Zero broken references remaining
|