Files
cleveragents-core/robot/cli.robot
T
brent.edwards 01b6eb1804
CI / build (push) Successful in 17s
CI / helm (push) Successful in 22s
CI / lint (push) Successful in 28s
CI / typecheck (push) Successful in 47s
CI / benchmark-regression (push) Has been skipped
CI / quality (push) Successful in 3m49s
CI / security (push) Successful in 4m11s
CI / unit_tests (push) Successful in 9m19s
CI / docker (push) Successful in 1m22s
CI / coverage (push) Successful in 12m35s
CI / e2e_tests (push) Successful in 16m17s
CI / integration_tests (push) Successful in 25m5s
CI / status-check (push) Successful in 2s
CI / benchmark-publish (push) Successful in 28m31s
feat(autonomy): parallel execution scales to 10+ concurrent subplans (#1201)
## Summary

Add M6 parallel-scaling coverage for 10+ concurrent subplans:

- **15-subplan parallel scenario** with explicit peak-concurrency bound checks (`max_parallel=10`) and thread-safe concurrency tracking via `_build_executor()`.
- **Deep hierarchical decomposition** coverage (4+ levels) with adjusted leaf condition that only stops early when hitting `max_depth` or when the workset is trivially small (`min_files_per_subplan`).
- **Non-progress guard** in `_build_hierarchy` to prevent pathological recursion when clustering cannot meaningfully split the file set.
- **Small-project regression test** (< 50 files) verifying decomposition depth does not increase unexpectedly with the relaxed leaf condition.
- **ASV benchmark** for 15-subplan parallel execution with `max_parallel=10` to track scaling behavior.

### Removed from this PR

The `_build_hierarchy` child-linkage correctness fix (returning `node_id` from recursive calls instead of using `nodes[-1].node_id`) has been **removed** per review feedback — it is a separate bug fix and will be submitted as an independent issue/PR per CONTRIBUTING.md §Atomic Commits.

## Approach

- **Concurrency tracking:** The `_build_executor()` closure in step definitions detects `context.concurrency_counter` / `context.concurrency_lock` and performs thread-safe peak tracking in a try/finally block.
- **Leaf condition:** Replaced the `max_files_per_subplan` / `max_tokens_per_subplan` leaf check with a `min_files_per_subplan` check to allow deeper decomposition for large projects. Added a non-progress guard so clustering that cannot split the file set terminates immediately rather than recursing to `max_depth`.
- **Deterministic IDs:** `_ids_for_count()` preserves legacy fixed IDs for the first 5 subplans and generates additional deterministic IDs for scale scenarios.

## Validation

### Passing
- `nox -s lint` — all checks passed
- `nox -s typecheck` — 0 errors, 0 warnings
- `nox -s unit_tests` — 12,988 scenarios passed, 0 failed
- `nox -s coverage_report` — 97% (passes `--fail-under=97`)

Closes #855

Reviewed-on: #1201
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-31 23:57:39 +00:00

89 lines
4.3 KiB
Plaintext

*** Settings ***
Documentation CLI Benchmark and Integration Tests for CleverAgents
Resource ${CURDIR}/common.resource
Library Process
Library OperatingSystem
Library String
Library Collections
Suite Setup Setup Test Environment
Suite Teardown Cleanup Test Environment
*** Variables ***
# PYTHON variable is passed from nox via --variable PYTHON:/path/to/python
# Default to 'python' if not passed (for standalone runs)
${PYTHON} python
*** Test Cases ***
CLI Help Command Performance
[Documentation] Benchmark help command execution time
${start_time}= Get Time epoch
${result}= Run Process ${PYTHON} -m cleveragents --help timeout=120s on_timeout=kill
${end_time}= Get Time epoch
${duration}= Evaluate ${end_time} - ${start_time}
Should Be True ${duration} < 90 Help command took too long: ${duration}s
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} AI-powered development assistant
Should Contain ${result.stdout} Usage:
CLI Version Command Performance
[Documentation] Benchmark version command execution time
${start_time}= Get Time epoch
${result}= Run Process ${PYTHON} -m cleveragents --version timeout=120s on_timeout=kill
${end_time}= Get Time epoch
${duration}= Evaluate ${end_time} - ${start_time}
Should Be True ${duration} < 90 Version command took too long: ${duration}s
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} 1.0.0
Python Module Entry Works
[Documentation] Verify Python module entry point works correctly
${result}= Run Process ${PYTHON} -m cleveragents --help timeout=120s on_timeout=kill
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} AI-powered development assistant
Should Contain ${result.stdout} Usage:
CLI Diagnostics Command
[Documentation] Test diagnostics command functionality
${result}= Run Process ${PYTHON} -m cleveragents diagnostics timeout=120s on_timeout=kill
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} Checks
Should Contain ${result.stdout} Summary
CLI Info Command
[Documentation] Test info command functionality
${result}= Run Process ${PYTHON} -m cleveragents info timeout=120s on_timeout=kill
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} Environment
Should Contain ${result.stdout} Runtime
Invalid Command Error Handling
[Documentation] Verify proper error handling for invalid commands
${result}= Run Process ${PYTHON} -m cleveragents invalid-command-xyz timeout=120s on_timeout=kill
Should Not Be Equal As Integers ${result.rc} 0
Should Contain ${result.stderr} Error: Invalid command
Multiple Commands Benchmark
[Documentation] Run multiple commands and measure overall performance
${commands}= Create List --help --version diagnostics info
${total_start}= Get Time epoch
FOR ${cmd} IN @{commands}
${result}= Run Process ${PYTHON} -m cleveragents ${cmd} timeout=120s on_timeout=kill
Should Be Equal As Integers ${result.rc} 0 Command failed: ${cmd}
END
${total_end}= Get Time epoch
${total_duration}= Evaluate ${total_end} - ${total_start}
Should Be True ${total_duration} < 360 Multiple commands took too long: ${total_duration}s
CLI Response Time Consistency
[Documentation] Verify consistent response times across multiple runs
@{durations}= Create List
FOR ${i} IN RANGE 5
${start}= Get Time epoch
${result}= Run Process ${PYTHON} -m cleveragents --help timeout=120s on_timeout=kill
${end}= Get Time epoch
${duration}= Evaluate ${end} - ${start}
Append To List ${durations} ${duration}
END
${avg_duration}= Evaluate sum(${durations}) / len(${durations})
Should Be True ${avg_duration} < 90 Average response time too high: ${avg_duration}s