Renames `plan lifecycle-list` to `plan list` and `plan lifecycle-apply` to `plan apply` to align with the specification's canonical command names. Removes legacy V2 plan commands that occupied those names. - Renamed CLI command registrations from lifecycle-list/lifecycle-apply to list/apply - Removed legacy V2 apply and list commands (~200 lines) - Updated apply shortcut in main.py to delegate to v3 lifecycle - Added defensive null check for plan existence in apply command - Updated 63+ test, doc, and benchmark files for consistency Closes #881 Co-authored-by: Jeffrey Phillips Freeman <the@jeffreyfreeman.me> Co-committed-by: Jeffrey Phillips Freeman <the@jeffreyfreeman.me>
76 KiB
Testing Guide
Overview
CleverAgents uses a comprehensive testing strategy with three complementary frameworks:
- Behave (BDD/Gherkin) for unit-level and scenario tests under
features/ - Robot Framework for integration and end-to-end tests under
robot/ - ASV (airspeed velocity) for performance benchmarks under
benchmarks/ - Coverage:
coverage.pywith branch coverage, enforced at >=97%
All tests are executed exclusively through nox sessions. Never invoke behave, robot, or other test runners directly.
Running Tests
# Run all default sessions (lint, typecheck, unit tests, integration tests, coverage)
nox
# Unit tests only (Behave)
nox -s unit_tests
# Integration tests only (Robot Framework)
nox -s integration_tests
# Coverage report with 97% enforcement
nox -s coverage_report
# Run a specific feature file
nox -s unit_tests -- features/plan_model.feature
# Benchmarks (ASV)
nox -s benchmark
# Type checking (pyright)
nox -s typecheck
# Linting (ruff)
nox -s lint
# Formatting check
nox -s format -- --check
Coverage Requirements
Threshold: 97%
The project enforces a minimum 97% code coverage at all times. This is a hard gate:
- The
nox -s coverage_reportsession fails if coverage drops below 97%. - The Forgejo CI
coveragejob runsnox -s coverage_reportand blocks merges on failure. - Branch coverage is enabled (
branch = trueinpyproject.toml).
How Coverage Is Measured
Coverage is measured by running Behave tests under slipcover (bytecode-based instrumentation) in parallel via behave-parallel. Each worker produces a per-feature JSON coverage file, then slipcover merges them:
# Each behave-parallel worker runs:
python -m slipcover --json --out build/.slipcover.<uuid>.json --source src -m behave <feature>
# After all workers finish:
python -m slipcover --merge build/.slipcover.*.json --json --out build/coverage.json
python -m slipcover --merge build/coverage.json --fail-under=97
Coverage Output
On success, the nox session emits:
COVERAGE OK: 97.2% (threshold: 97%)
On failure, the nox session emits and exits non-zero:
COVERAGE FAILED: 95.3% < 97% threshold
Both messages are single-line and CI-parseable. The CI pipeline greps for these lines to surface them in job summaries.
Coverage Reports
Three report formats are generated under build/:
| Report | Path | Description |
|---|---|---|
| Terminal | stdout | Per-file summary with --show-missing |
| HTML | build/htmlcov/index.html |
Interactive browser report |
| XML | build/coverage.xml |
Cobertura-format for CI tools |
| JSON | build/coverage.json |
Machine-readable totals and per-file data |
Coverage Configuration
Coverage is collected by slipcover (bytecode-based, faster than coverage.py's sys.settrace).
Source and omit patterns are passed directly to slipcover via the coverage_report nox session.
The pyproject.toml [tool.coverage.run] section is retained for any tools that still read it:
[tool.coverage.run]
source = ["src", "scripts"]
branch = true
parallel = false
omit = [
"*/tests/*",
"*/test_*",
"features/*",
"*/features/*",
"*/__pycache__/*",
"*/site-packages/*",
"*/dependency_injector/*",
"*/venv/*",
"*/.venv/*",
"*/.nox/*",
"src/cleveragents/discovery/*",
]
data_file = "build/.coverage"
[tool.coverage.html]
directory = "build/htmlcov"
[tool.coverage.xml]
output = "build/coverage.xml"
The --fail-under=97 threshold is enforced in the coverage_report nox session (noxfile.py) via slipcover --fail-under=97.
How to Improve Coverage
- Run
nox -s coverage_reportto generate the HTML report. - Open
build/htmlcov/index.htmlin a browser to identify uncovered files. - Sort by "Missing" column to find files with the most uncovered lines.
- Write Behave scenarios targeting the uncovered code paths.
- Re-run
nox -s coverage_reportto verify improvement.
Do not lower the threshold, add # pragma: no cover comments, or expand the omit list without explicit team agreement.
All test files go under features/ (Behave) or robot/ (Robot Framework). Never create a tests/ directory.
Unit Tests (Behave)
Unit tests use the Behave BDD framework with Gherkin-syntax .feature files.
Structure
features/
*.feature # Feature files (Gherkin scenarios)
steps/ # Step definitions (Python)
mocks/ # Mock implementations (test doubles)
fixtures/ # Test fixture data
environment.py # Behave hooks (before/after scenario, etc.)
Guidelines
- No pytest: All unit tests must be Behave-based. There is intentionally no
tests/directory. - Step naming: Use unique, descriptive step names. Behave loads all step files globally; ambiguous steps cause
AmbiguousSteperrors. - Feature naming: Name feature-specific step files after their feature (e.g.,
foo.feature->foo_steps.py). - Mocks in features/ only: All mock implementations live in
features/mocks/. Production code (src/) must never contain test doubles.
Running Specific Features
# Run a single feature file
nox -s unit_tests -- features/plan_model.feature
# Run scenarios with a specific tag
nox -s unit_tests -- --tags=@smoke
Integration Tests (Robot Framework)
Integration tests use Robot Framework for end-to-end and system-level testing.
Structure
robot/
*.robot # Test suites
*.resource # Shared keywords and variables
common.resource # Common setup/teardown keywords
v2_paths.resource # Path configuration
Guidelines
- Resource paths: Always use
${CURDIR}/prefix forResourceimports (e.g.,Resource ${CURDIR}/common.resource). - Python injection: Use
${PYTHON}variable (injected by nox) instead of barepythoninRun Processcalls. - Timeouts: Always add
timeout=toRun Processcalls that invoke long-running commands. - Slow tests: Tag tests that require external services (API keys, running servers) with
slow. These are excluded in CI via--exclude slow.
Parallel Execution Isolation (pabot)
Robot integration tests run in parallel via pabot (CPU-count workers by default).
To prevent race conditions, common.resource provides per-suite isolation:
-
Per-suite CLEVERAGENTS_HOME — Each suite gets a unique temp directory (
${TEMPDIR}/.cleveragents_<SuiteName>) so configuration and data files never collide. -
Per-suite database URL —
CLEVERAGENTS_DATABASE_URLandCLEVERAGENTS_TEST_DATABASE_URLpoint to SQLite files inside the per-suite home directory. Without this, all pabot workers would fall back to the defaultsqlite:///cleveragents.db(relative to CWD), causing write contention. -
Singleton reset in helpers — Robot helpers that chain multiple CLI invocations within a single process (e.g.
full_flow()) must callreset_global_state()between invocations. The reset logic is centralised inrobot/helpers_common.py; each helper imports and delegates tohelpers_common.reset_global_state(). This function resets:Settings._instance(singleton config)reset_container()(DI container)reset_provider_registry()(provider registry)MEMORY_ENGINES(SQLAlchemy engine cache)
-
Environment variable cleanup —
Cleanup Test Environmentremoves all environment variables set during setup (CLEVERAGENTS_HOME,CLEVERAGENTS_AUTO_APPLY_MIGRATIONS,CLEVERAGENTS_TESTING_USE_MOCK_AI,CLEVERAGENTS_DATABASE_URL, andCLEVERAGENTS_TEST_DATABASE_URL) after each suite to prevent leakage into subsequent suites.
Adding a new Robot helper
When writing a new Robot helper that invokes CliRunner or otherwise
exercises the application layer, import the shared reset function from
helpers_common.py:
from helpers_common import reset_global_state
# In the __main__ dispatcher:
if __name__ == "__main__":
reset_global_state()
fn = _COMMANDS[sys.argv[1]]
fn()
Call reset_global_state() once at the start of the __main__ block and
again between chained CLI invocations within the same process. Do not
redefine the reset logic locally — always delegate to helpers_common.
Running Specific Suites
# Run a single Robot suite
nox -s integration_tests -- --suite robot/plan_generation_graph.robot
# Include slow tests
nox -s slow_integration_tests
Performance Benchmarks (ASV)
ASV (airspeed velocity) tracks performance over time.
Structure
benchmarks/
*.py # Benchmark suites (Time*/Mem*/Track* classes)
asv.conf.json # ASV configuration
Running Benchmarks
nox -s benchmark
Writing Effective Tests
Behave Test Guidelines
Feature: Plan lifecycle transitions
Scenario: Execute plan transitions phase from strategize to execute
Given an action "local/build-app" exists
And a plan is created using "local/build-app" on project "local/my-project"
And strategize is completed
When I execute the plan
Then the plan phase should be "execute"
And the processing state should be "queued"
Robot Framework Guidelines
*** Settings ***
Documentation Plan lifecycle integration tests
Library Process
Library OperatingSystem
Resource ${CURDIR}/common.resource
*** Test Cases ***
Plan Execute Transitions Correctly
[Documentation] Verify execute phase transition via CLI
${result}= Run Process ${PYTHON} -m cleveragents plan execute ${PLAN_ID}
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} execute
CLI Lifecycle Test Suites
The CLI lifecycle suites provide comprehensive coverage for action and plan lifecycle commands, including success paths, error paths, multi-project arguments, automation profile overrides, invariants, actor overrides, phase visibility, and terminal outcomes.
Behave Suite: features/cli_lifecycle_coverage.feature
Covers 68 scenarios across these areas:
| Area | Scenarios | Description |
|---|---|---|
| Action CRUD | 13 | create, list, show, archive with success + error paths |
| Plan Use | 10 | basic use, multi-project, automation profile, invariants, actor overrides |
| Plan Status | 9 | phase visibility (strategize, execute, apply) and terminal outcomes (applied, constrained, errored, cancelled) |
| Plan Execute | 5 | execute success, phase transition, error paths |
| Plan Apply | 5 | apply success, terminal outcomes, error paths |
| Plan Cancel | 5 | cancel with reason, already-terminal guard, error paths |
| Plan List | 4 | list with and without filters |
| Negative cases | 17 | missing config, invalid args, invalid project names, unknown actions/resources |
Step definitions: features/steps/cli_lifecycle_coverage_steps.py
All step names are prefixed with lifecycle coverage to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/cli_lifecycle_e2e.robot
End-to-end integration tests exercising the full action-to-apply lifecycle through
the CLI entry points. Uses robot/helper_cli_lifecycle_e2e.py as a Python helper
that mocks the service layer while exercising real CLI argument parsing and output
formatting.
| Test Case | Description |
|---|---|
| Action Create From Config Via CLI | Creates an action from YAML config |
| Plan Use Creates Plan In Strategize Phase | Uses an action to create a plan |
| Plan Execute Transitions To Execute Phase | Executes a plan, checks phase |
| Plan Lifecycle Apply Transitions To Apply Phase | Applies a plan, checks phase |
| Plan Status Shows Plan Details | Verifies status output rendering |
| Plan Cancel Cancels Non-Terminal Plan | Cancels with a reason string |
| Full Lifecycle Action To Apply | End-to-end: create -> use -> execute -> apply |
| Plan List Shows Plans | Verifies list output |
ASV Benchmark: benchmarks/plan_cli_smoke_bench.py
Three benchmark suites measuring CLI argument parsing overhead (service layer is mocked so only parsing + routing cost is measured):
ActionCLIParsingSuite--time_action_create,time_action_list,time_action_list_with_filters,time_action_show,time_action_archivePlanUseArgParsingSuite--time_use_minimal,time_use_multi_project,time_use_with_all_flags,time_use_with_args_onlyPlanLifecycleCmdParsingSuite--time_execute,time_lifecycle_apply,time_status_by_id,time_status_list_all,time_cancel_with_reason,time_lifecycle_list,time_lifecycle_list_filtered
Running the CLI Lifecycle Suites
# Behave only (our feature)
nox -s unit_tests -- features/cli_lifecycle_coverage.feature
# Robot only (our suite)
nox -s integration_tests -- --suite robot/cli_lifecycle_e2e.robot
# Benchmarks
nox -s benchmark
Robot Suite: robot/cli_lifecycle.robot
End-to-end Robot suite for the CLI lifecycle with git_worktree sandbox simulation
and ChangeSet capture verification. Uses robot/helper_cli_lifecycle.py as a
Python helper that exercises the full action → plan → execute → apply flow with
mocked service layer, simulated sandbox workspace, and InMemoryChangeSetStore
for change tracking.
Positive tests:
| Test Case | Description |
|---|---|
| Action Create Via Config CLI | Creates an action from YAML config |
| Plan Use Creates Strategize Plan With Sandbox | Uses action with sandbox workspace mock |
| Plan Execute With ChangeSet Capture | Executes plan, verifies changeset entry recorded |
| Plan Apply With ChangeSet Verification | Applies plan, verifies changeset summary |
| Full Lifecycle Action To Apply With Sandbox | End-to-end with sandbox + changeset |
Negative tests:
| Test Case | Description |
|---|---|
| Invalid Project Name Rejected | Invalid project name format handled gracefully |
| Missing Resource Returns Error | Non-existent action returns error |
| Invalid Arg Format Rejected | --arg without = separator rejected |
Fixtures:
helper_cli_lifecycle.pycreates temporary YAML config files and sandbox directories- Uses
InMemoryChangeSetStorefor ChangeSet capture without database - Each subcommand is self-contained; cleanup runs in
finallyblocks
Aligned Behave feature: features/cli_lifecycle_robot_alignment.feature
mirrors the Robot E2E flow to keep unit/integration expectations aligned.
ASV benchmark: benchmarks/cli_robot_flow_bench.py measures fixture setup
overhead (CLIRobotFixtureSetupSuite) and full lifecycle CLI throughput
(CLIRobotLifecycleSuite).
Running the Robot CLI Lifecycle Suite
# Robot suite
nox -s integration_tests -- --suite robot/cli_lifecycle.robot
# Aligned Behave feature
nox -s unit_tests -- features/cli_lifecycle_robot_alignment.feature
# Benchmarks
nox -s benchmark
Persistence Test Suites
The persistence layer (SQLAlchemy repositories for plans and actions) has dedicated test suites at both unit and integration levels.
Behave: Plan Persistence (features/plan_persistence.feature)
19 scenarios covering LifecyclePlanRepository operations:
- Create: Persist plans via repository, retrieve by ULID, verify all fields round-trip.
- Phase/state transitions: Update
phaseandprocessing_state, verify persistence after each transition. - List filters: Filter plans by phase, processing state, and action name.
- Plan tree links: Parent/child/root hierarchy with
parent_plan_idandroot_plan_id. - Cross-restart: Close the SQLite session, reopen from disk, verify plans survive reconnection (including project links, arguments, and invariants).
Step definitions: features/steps/plan_persistence_steps.py
Fixtures: Each scenario gets a fresh temp SQLite database via _setup_db() / _teardown_db() helper functions. The database is file-backed (not :memory:) so cross-restart scenarios can close and reopen the connection.
Behave: Action Persistence (features/action_persistence.feature)
14 scenarios covering ActionRepository operations:
- CRUD: Create, retrieve by
namespaced_name, list available, filter by namespace, archive, delete. - Argument ordering: Persist 3
ActionArgumententries and verify positional order is preserved. - Invariant ordering: Persist 3 invariant strings and verify order is preserved.
- Terminal state storage: Create plans from actions in Apply phase with each terminal
ProcessingState(applied,constrained,errored,cancelled) and verify persistence.
Step definitions: features/steps/action_persistence_steps.py
Robot: Plan Persistence E2E (robot/plan_persistence_e2e.robot)
5 end-to-end tests exercising the persistence layer through Python helper scripts:
| Test Case | Description |
|---|---|
| Plan Full Lifecycle Persistence | Create action + plan, transition through all phases to applied terminal state |
| Plan Restart Persistence | Create plan, close DB, reopen, verify all fields survive |
| Plan Concurrent Session Access | Two independent sessions against the same DB file |
| Action CRUD Persistence E2E | Create, read, update, delete an action with arguments/invariants |
| Plan Tree Hierarchy Persistence E2E | Parent/child plan hierarchy with root_plan_id verification |
Helper script: robot/helper_plan_persistence_e2e.py
Robot: Persistence Lifecycle (robot/persistence_lifecycle.robot)
7 end-to-end tests exercising persistence lifecycle patterns through a Python helper:
| Test Case | Description |
|---|---|
| Plan Full Lifecycle Phase Transitions | Create action + plan, transition through all phases to applied terminal state |
| Process Restart Simulation | Write plan, close DB, reopen from disk, verify all fields survive |
| Reopen Plan Status After Restart | Close/reopen, verify phase, state, automation level, and tags unchanged |
| Concurrent CLI Access Safeguards | Two independent sessions see each other's committed data |
| Stored Arguments Ordering Persistence | 4 ordered ActionArgument entries survive roundtrip |
| Stored Invariants Ordering Persistence | 3 ordered invariant strings survive roundtrip |
| Project Links Persistence Through Restart | 2 ProjectLink entries survive close/reopen cycle |
Helper script: robot/helper_persistence_lifecycle.py
Behave: Persistence Robot Alignment (features/persistence_robot_alignment.feature)
1 scenario mirroring the Robot restart-simulation flow to keep unit and integration persistence expectations aligned. Step definitions in features/steps/persistence_robot_alignment_steps.py.
ASV Benchmarks: Robot Fixture Overhead (benchmarks/persistence_robot_bench.py)
Performance baselines for Robot fixture setup operations:
TimeRobotFixtureSetupSuite.time_file_db_init-- File-backed SQLite init + schema creationTimeRobotFixtureSetupSuite.time_seed_action-- Action seeding into a fresh file DBTimeRobotFixtureSetupSuite.time_seed_plan-- Plan creation including FK action seedingTimeRobotRestartCycleSuite.time_close_reopen_verify-- Close + reopen + get cycle latency
ASV Benchmarks (benchmarks/persistence_suites_bench.py)
Performance baselines for persistence test operations:
TimePlanPersistenceSuites.time_create_and_retrieve_plan-- Plan round-trip latencyTimePlanPersistenceSuites.time_phase_state_transition-- Phase/state update cycleTimePlanPersistenceSuites.time_list_plans_filtered-- List with filtersTimePlanPersistenceSuites.time_plan_tree_operations-- Parent/child hierarchyTimeActionPersistenceSuites.time_create_and_retrieve_action-- Action round-tripTimeActionPersistenceSuites.time_action_arguments_ordering-- Ordered arguments persistenceTimeActionPersistenceSuites.time_cross_restart_persistence-- Close/reopen cycle
Run with: nox -s benchmark
ASV Benchmarks (benchmarks/decision_persistence_serialization_bench.py)
Performance baselines for decision persistence (serialization round-trip) operations:
DecisionDumpRoundTripSuite.time_root_dump-- Root decision model_dumpDecisionDumpRoundTripSuite.time_root_validate-- Root decision model_validateDecisionDumpRoundTripSuite.time_root_roundtrip-- Root decision full round-tripDecisionDumpRoundTripSuite.time_child_roundtrip-- Child decision full round-tripDecisionDumpRoundTripSuite.time_full_roundtrip-- Fully-populated decision round-tripDecisionJsonRoundTripSuite.time_root_dump_json-- Root decision model_dump_jsonDecisionJsonRoundTripSuite.time_root_json_roundtrip-- Root decision JSON round-tripDecisionJsonRoundTripSuite.time_full_json_roundtrip-- Full decision JSON round-tripDecisionJsonRoundTripSuite.time_root_json_validate-- Root decision model_validate from JSONDecisionJsonRoundTripSuite.time_full_json_validate-- Full decision model_validate from JSONDecisionSnapshotPersistenceSuite.time_empty_snapshot_roundtrip-- Empty snapshot round-tripDecisionSnapshotPersistenceSuite.time_small_snapshot_roundtrip-- Small snapshot round-tripDecisionSnapshotPersistenceSuite.time_large_snapshot_roundtrip-- Large snapshot (50 resources) round-tripDecisionTreeSerializationSuite.time_small_tree_roundtrip-- 3-node tree round-tripDecisionTreeSerializationSuite.time_medium_tree_roundtrip-- 7-node tree round-tripDecisionTreeSerializationSuite.time_large_tree_roundtrip-- 15-node tree round-tripDecisionTreeSerializationSuite.time_large_tree_json_roundtrip-- 15-node tree JSON round-trip
Run with: nox -s benchmark
Behave: Decision Persistence Serialization (features/decision_persistence_serialization.feature)
21 scenarios covering Decision serialization round-trips, context-snapshot persistence, correction-chain reconstruction, and decision-tree reconstruction from serialized data:
- CRUD round-trips: Root and child decisions survive
model_dump/model_validate. - JSON round-trips: Decisions with artifacts survive
model_dump_json/model_validate. - Snapshot persistence: Empty snapshots, snapshots with resources, and actor state refs.
- Correction chains: Correction metadata, superseded_by, and 3-decision chain reconstruction.
- Tree reconstruction: 3-level tree (7 nodes) with parent-link integrity verification.
- Edge cases: Empty alternatives, confidence boundaries (0.0, 1.0, None), all types, long rationale.
- Downstream IDs:
downstream_decision_idsanddownstream_plan_idssurvive round-trip.
Step definitions: features/steps/decision_persistence_serialization_steps.py
Robot Framework: Decision Persistence Serialization (robot/decision_persistence_serialization.robot)
7 integration smoke tests exercising decision persistence round-trips via
a helper script (robot/helper_decision_persistence_serialization.py):
- Root decision round-trip
- Child decision round-trip
- Context snapshot persistence
- JSON serialization round-trip
- Correction chain persistence
- Decision tree reconstruction
- All decision types round-trip
Behave: Skill Registry Persistence (features/skill_registry.feature)
23 scenarios covering SkillRepository and SkillRegistryService operations:
- Register: Persist skills with each item type (
tool_ref,include,inline_tool,mcp_source,agent_source), verify round-trip. - Duplicate rejection: Second registration with the same name raises
DuplicateSkillError. - Retrieve / list: Get by name, list all, list with namespace filter.
- Update / remove: Change description, remove with cascade, non-existent skill error paths.
- Overrides round-trip: Override metadata survives persistence.
- Ordering stability: Mixed item types maintain stable
item_orderacross round-trips. - Name validation: Empty, no-slash, and special-character names rejected at persistence layer.
- Large payload: 50 tool refs persist correctly.
- Domain objects: Listed results are
Skilldomain instances.
Step definitions: features/steps/skill_registry_steps.py
Fixtures: Each scenario gets a fresh in-memory SQLite database. A single
shared Session is reused across all repository calls within each scenario via a
lambda session factory (lambda: shared_session). This mirrors the production
UnitOfWork pattern (ADR-019) and avoids transaction-scoping issues that arise
when multiple ephemeral sessions share one in-memory SQLite connection
(StaticPool): abandoned sessions can be garbage-collected at unpredictable
times, rolling back previously flushed but uncommitted data.
Robot: Skill Registry (robot/skill_registry.robot)
5 smoke tests exercising the registry through a Python helper script:
| Test Case | Description |
|---|---|
| Register And Retrieve A Skill | Round-trip: register then get by name |
| List Skills With Namespace Filter | Register multiple skills, filter by namespace |
| Update A Skill | Change description after registration |
| Reject Duplicate Skill Name | Duplicate name produces error |
| Remove A Skill | Remove and verify absence |
Helper script: robot/helper_skill_registry.py
ASV Benchmarks: Skill Registry (benchmarks/skill_registry_bench.py)
Performance baselines for skill registry persistence operations:
SkillModelConstruction.time_from_domain-- Domain-to-ORM mappingSkillModelReconstruction.time_to_domain-- ORM-to-domain reconstructionSkillRepositoryCRUD.time_create_skill-- Single skill insert + commitSkillRepositoryCRUD.time_get_skill-- Lookup by nameSkillRepositoryCRUD.time_list_all-- List all skillsSkillRepositoryCRUD.time_list_namespace-- List with namespace filterSkillRegistryListPerformance.time_list_100_skills-- List with 100 seeded skillsSkillRegistryListPerformance.time_list_100_skills_filtered-- Filtered list with 100 skills
Running Persistence Tests
# Unit tests (Behave) -- runs all features including persistence
nox -s unit_tests
# Run only plan persistence scenarios
nox -s unit_tests -- features/plan_persistence.feature
# Run only action persistence scenarios
nox -s unit_tests -- features/action_persistence.feature
# Run only skill registry persistence scenarios
nox -s unit_tests -- features/skill_registry.feature
# Integration tests (Robot) -- runs all robot suites including persistence E2E
nox -s integration_tests
# Run only plan persistence E2E
python -m robot robot/plan_persistence_e2e.robot
# Run only persistence lifecycle E2E
python -m robot robot/persistence_lifecycle.robot
# Run only skill registry smoke tests
nox -s integration_tests -- robot/skill_registry.robot
# Run only persistence robot alignment (Behave)
nox -s unit_tests -- features/persistence_robot_alignment.feature
Subplan Model Test Suites
The subplan model (hierarchy, configuration, failure handling) has dedicated test suites at unit, integration, and benchmark levels.
Behave: Subplan Model (features/subplan_model.feature)
33 scenarios covering Plan hierarchy properties, SubplanConfig defaults, SubplanStatus tracking, SubplanFailureHandler decisions, and CLI dict rendering:
| Area | Scenarios | Description |
|---|---|---|
| Root plan identity | 2 | is_root_plan, is_subplan, depth for root plans |
| Child plan identity | 3 | Parent reference, attempt counter, root propagation |
| Default values | 4 | Phase, state, automation level, subplan_config defaults |
| SubplanConfig storage | 3 | Sequential/parallel config, standalone defaults |
| Hierarchy helpers | 3 | has_subplans, depth for non-root plans |
| Lifecycle transitions | 2 | Child plans follow same phase transition rules |
| Dependency guardrails | 5 | FailureHandler stop/retry decisions, phase independence |
| CLI dict rendering | 3 | parent_plan_id, subplan_count in output |
| Validation guardrails | 4 | Invalid ULIDs, out-of-range config values |
Step definitions: features/steps/subplan_model_steps.py
All step names are prefixed with subplan-test to avoid AmbiguousStep conflicts
with existing plan_model_steps.py steps.
Robot: Subplan Model (robot/subplan_model.robot)
11 smoke tests exercising subplan model properties via Python helper scripts:
| Test Case | Description |
|---|---|
| Root Plan Identity Flags | is_root_plan=True, is_subplan=False, depth=0 |
| Child Plan Identity Flags | is_subplan=True, is_root_plan=False, depth=-1 |
| Three Level Hierarchy Root Propagation | root_plan_id propagates through 3 levels |
| SubplanConfig Defaults | All default values are sensible |
| SubplanConfig Custom Settings | Parallel mode with custom retry settings |
| Failure Handler Stop Others On Fail Fast | should_stop_others with fail_fast=True |
| Failure Handler Retry Retriable Error | should_retry for TimeoutError |
| Failure Handler Skip Non Retriable Error | should_retry skips ConfigurationError |
| CLI Dict Renders Parent Plan Id For Child | parent_plan_id in child plan CLI dict |
| CLI Dict Renders Subplan Count | subplan_count in parent plan CLI dict |
| Plan With Subplan Config And Statuses | Full plan model with config + statuses |
Helper script: robot/helper_subplan_model.py
ASV Benchmarks (benchmarks/subplan_model_validation_bench.py)
Performance baselines for subplan model operations:
SubplanConfigValidationSuite--time_default_config_creation,time_custom_config_creation,time_config_model_dump,time_config_model_dump_jsonSubplanStatusValidationSuite--time_minimal_status_creation,time_full_status_creation,time_status_model_dumpSubplanFailureHandlerSuite--time_should_stop_others_fail_fast,time_should_stop_others_no_fail_fast,time_should_retry_retriable,time_should_retry_non_retriableSubplanHierarchySuite--time_is_root_plan,time_is_subplan,time_depth_root,time_depth_child,time_has_subplans,time_cli_dict_with_subplans,time_child_plan_creation,time_plan_with_subplans_creation
Running Subplan Model Tests
# Behave only (subplan model feature)
nox -s unit_tests -- features/subplan_model.feature
# Robot only (subplan model suite)
nox -s integration_tests -- --suite robot/subplan_model.robot
# Benchmarks
nox -s benchmark
CI Pipeline
The Forgejo CI pipeline runs these jobs in order:
- lint —
nox -s lint(ruff check) - typecheck —
nox -s typecheck(pyright) - security —
nox -s security_scan(bandit + vulture) - quality —
nox -s complexity(radon) - unit_tests —
nox -s unit_tests(behave) - integration_tests —
nox -s integration_tests(robot) - coverage —
nox -s coverage_report(fail-under 97%) - build —
nox -s build(wheel)
All jobs use Python 3.13 and route commands through nox. See docs/development/ci-cd.md for full CI/CD documentation.
Troubleshooting
Coverage report shows 0% or very low coverage
Ensure you're running in serial mode (not parallel). The coverage_report nox session handles this correctly. Do not run behave directly outside of coverage run.
Also ensure:
parallel = falseis set in[tool.coverage.run]COVERAGE_FILEandCOVERAGE_RCFILEenv vars are set correctly by the nox session
AmbiguousStep errors in Behave
Two step files define the same @given/@when/@then pattern. Behave loads all step files globally. Make step names unique or use more specific patterns.
Robot Resource file does not exist
Use Resource ${CURDIR}/filename.resource instead of bare Resource filename.resource. Robot resolves bare paths relative to CWD, not the test file directory.
Robot No module named cleveragents
Ensure Robot tests use ${PYTHON} (injected by nox from the venv) instead of bare python in Run Process calls.
In-memory SQLite tests lose data between repository calls
When writing Behave tests that use sqlite:///:memory: with a sessionmaker,
do not create a new Session per repository call. SQLite in-memory databases
use StaticPool (a single shared connection). If each repository method creates
its own session, the abandoned sessions may be garbage-collected at unpredictable
times, triggering a ROLLBACK that discards previously flushed data.
Fix: Create a single shared Session at scenario setup and pass a factory
that always returns it:
real_factory = sessionmaker(bind=engine)
shared_session = real_factory()
repo = MyRepository(session_factory=lambda: shared_session)
This mirrors the production UnitOfWork pattern (ADR-019) where all
repositories share one session within a transaction. See
features/steps/skill_registry_steps.py for a working example.
Tests hang indefinitely
Add timeout parameters to Run Process calls in Robot tests:
${result}= Run Process ${PYTHON} -m cleveragents ... timeout=30s
Validation Edge Case Test Catalog
The validation edge case suites exercise boundary conditions, error paths, and conflict detection across the validation subsystem.
Shared Fixtures (features/fixtures/validation/)
| Fixture File | Description |
|---|---|
malformed_tool_output.json |
Missing required fields, wrong types, extra nulls, non-object outputs |
missing_resources.json |
Dangling references, unresolvable paths, circular dependencies |
validation_timeouts.json |
Slow validators, zero/negative timeouts, partial results |
invalid_schema_transforms.json |
Transforms returning non-dict, missing type, null, circular refs |
mixed_ordering.json |
Required vs informational ordering, duplicate attachment IDs |
Each fixture file is a JSON document with a "fixtures" array. Each entry has
"name", "description", "input", and "expected_error" fields.
Behave Suite: features/validation_edge_cases.feature
25 scenarios covering:
| Area | Scenarios | Description |
|---|---|---|
| Malformed tool output | 6 | Missing fields, wrong types, nulls, empty/string outputs |
| Missing resources | 5 | Dangling actions, unresolvable paths, missing actors, circular deps |
| Validation timeouts | 4 | Slow validators, zero timeout, negative timeout, partial results |
| Invalid schema transforms | 3 | Non-dict returns, missing type field, null schema |
| Mixed ordering | 4 | Required-before-informational, duplicate attachment IDs, unknown levels |
| Concurrent conflicts | 1 | Two validators on same resource detect conflict |
| Rollback on failure | 2 | Required failure triggers rollback of pending changes |
Step definitions: features/steps/validation_edge_case_steps.py
All step names are prefixed with validation edge to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/validation_edge.robot
6 integration tests exercising the validation edge case helper:
| Test Case | Description |
|---|---|
| Validation Edge Load All Fixtures | Loads all fixture files and verifies structure |
| Validation Edge Malformed Output Detection | Detects malformed tool output patterns |
| Validation Edge Missing Resource Error Paths | Validates missing resource reference detection |
| Validation Edge Timeout Simulation | Validates timeout simulation fixture data |
| Validation Edge Schema Validation Errors | Validates schema transform error detection |
| Validation Edge Unknown Command Returns Error | Verifies unknown command exits with code 1 |
Helper script: robot/helper_validation_edge.py
ASV Benchmarks: benchmarks/validation_edge_bench.py
Two benchmark suites:
ValidationEdgeFixtureLoadSuite--time_load_malformed_tool_output,time_load_missing_resources,time_load_validation_timeouts,time_load_invalid_schema_transforms,time_load_mixed_ordering,time_load_all_fixturesValidationEdgeDetectionSuite--time_malformed_output_detection,time_missing_resource_detection,time_validation_ordering,time_combined_edge_detection
Running the Validation Edge Case Suites
# Behave only (validation edge cases)
nox -s unit_tests -- features/validation_edge_cases.feature
# Robot only (validation edge suite)
nox -s integration_tests -- --suite robot/validation_edge.robot
# Benchmarks
nox -s benchmark
Scale Testing
Scale tests validate performance baselines for repository indexing and decomposition across different repository sizes (1K, 5K, and 10K files).
Fixtures
Scale test fixtures live in features/fixtures/scale/:
scale_metadata.json— Scale profile definitions (file count, language mix, expected timings)baseline_thresholds.json— Performance threshold matrices (p50/p95/p99 per operation per file count)generator_instructions.md— Manual instructions for generating synthetic repos
Running Scale Tests
# Behave scale test scenarios
nox -s unit_tests -- features/scale_test.feature
# Robot scale test integration
nox -s integration_tests -- --suite robot/scale_test.robot
# ASV scale fixture benchmarks
nox -s benchmark
Scale Test Suites
| Suite | Framework | Scenarios | Description |
|---|---|---|---|
features/scale_test.feature |
Behave | 20 | Fixture loading, metadata validation, threshold checks, distribution simulation |
robot/scale_test.robot |
Robot | 6 | Integration-level fixture validation via helper script |
benchmarks/scale_fixture_bench.py |
ASV | 6 | Performance benchmarks for fixture processing |
For full scale testing documentation, see docs/development/scale_testing.md.
CLI Extension Test Fixtures
The CLI extension tests cover automation profile, invariant, and actor override functionality added in #325. Deeper coverage added in #326 validates resolution logic, ordering semantics, error cases, and output format snapshots.
Test Fixture Patterns
The CLI extension test suites use a mocked lifecycle service pattern. The
_get_lifecycle_service function in both plan.py and action.py is patched
so that tests exercise real CLI argument parsing, validation, and output
rendering without requiring a database or running server.
Common Fixtures
| Fixture | Module | Description |
|---|---|---|
_make_plan() |
cli_extensions_steps.py |
Creates a Plan instance with configurable automation profile, invariants, and actor overrides |
_make_action() |
cli_extensions_steps.py |
Creates an Action with configurable optional fields (estimation_actor, invariant_actor, inputs_schema, automation_profile) |
CliRunner |
typer.testing |
Invokes CLI commands in-process without subprocess overhead |
MagicMock service |
unittest.mock |
Mocks PlanLifecycleService for plan/action operations |
Automation Profile Resolution Fixtures
Tests validate that each builtin profile name (manual, review, supervised,
cautious, trusted, auto) resolves correctly through the CLI and that
invalid profile names are rejected with a helpful error listing available
profiles. The service mock returns a Plan with an AutomationProfileRef to
verify end-to-end profile assignment.
Invariant Ordering Fixtures
Tests pass multiple --invariant flags and verify the service receives them in
exact insertion order. This tests the list-accumulation behaviour of Typer's
repeated option handling.
Actor Override Error Fixtures
Tests exercise the validate_namespaced_actor regex
(^[a-z][a-z0-9-]*/[a-z][a-z0-9._-]*$) with various invalid inputs:
- Empty string
- Missing slash (
no-slash) - Uppercase characters (
UPPER/case) - Leading slash (
/leading-slash) - Trailing slash (
trail/ing/) - Double slashes (
open//ai) - Special characters (
open@ai/gpt-4) - Numeric-leading namespace (
123/model) - Whitespace (
open ai/gpt)
Output Snapshot Assertions
Tests render plan status, list, and action show output in JSON, YAML, and table formats, then assert that:
- JSON output is valid (parses without error)
- Required keys are present (
namespaced_name,phase,automation_profile, etc.) - Invariant text values survive serialization round-trips
- YAML output contains expected field names
- Table output contains human-readable field values
Behave Suite: features/cli_extensions.feature
Covers scenarios across these areas:
| Area | Scenarios | Description |
|---|---|---|
| Automation profile flags | 2 | Valid and invalid profile names |
| Profile resolution (deep) | 9 | All builtin profiles + error cases with special chars, spaces, empty |
| Invariant flags | 2 | Single and dual invariant flags |
| Invariant ordering (deep) | 2 | Three and five invariant insertion-order preservation |
| Actor overrides (valid) | 4 | Strategy, execution, estimation, invariant actors |
| Actor overrides (invalid) | 4 | Malformed namespace/name formats |
| Actor error cases (deep) | 11 | Empty, double-slash, special chars, numeric-leading, whitespace |
| Plan status display | 4 | Table, JSON, single plan with invariants/profile |
| Plan list | 2 | Table and JSON with invariant count |
| Action show | 8 | Optional actors, invariants, inputs_schema, automation profile |
| Combined flags | 1 | Profile + invariant together |
| validate_namespaced_actor | 2 | Unit-level accept/reject assertions |
| Output snapshots (deep) | 9 | JSON/YAML/table format validation for plan and action |
Step definitions: features/steps/cli_extensions_steps.py
Robot Suite: robot/cli_extensions.robot
Integration tests exercising CLI extensions through Python helper scripts:
| Test Case | Description |
|---|---|
| Plan Use With Invariants | Verifies --invariant flags pass through correctly |
| Plan Use With Automation Profile | Verifies --automation-profile works |
| Plan Use Actor Validation Valid | Valid namespaced actor format accepted |
| Plan Use Actor Validation Invalid | Invalid actor format rejected |
| Plan Use Combined Profile And Invariants | Both profile and invariants together |
| Action Show With Optional Actors And Invariants | action show renders estimation_actor, invariant_actor, invariants, inputs_schema |
| Action Show JSON With Optional Fields | JSON output includes all optional fields |
| Invariant Ordering Preserved | Multiple invariant flags preserve insertion order |
| Profile Resolution All Builtins | All builtin profile names resolve correctly |
Helper script: robot/helper_cli_extensions.py
ASV Benchmarks
Two benchmark files cover CLI extension performance:
benchmarks/cli_extensions_bench.py (original):
PlanUseWithProfileSuite— plan use with automation profilePlanUseWithInvariantsSuite— plan use with invariant flagsPlanUseActorValidationSuite— plan use with actor overridesPlanStatusWithExtendedFieldsSuite— status/list renderingActionShowExtendedSuite— action show in rich and JSON formats
benchmarks/cli_extension_tests_bench.py (extended #326):
ProfileResolutionSuite— resolve all builtin profiles, single profile, invalid profileInvariantOrderingSuite— three and five invariant flag parsingActorValidationErrorSuite— valid/invalid actor regex, CLI rejectionOutputFormatRenderingSuite— JSON/YAML/table rendering for action show and plan status
Running CLI Extension Tests
# Behave only (CLI extensions feature)
nox -s unit_tests -- features/cli_extensions.feature
# Robot only (CLI extensions suite)
nox -s integration_tests -- --suite robot/cli_extensions.robot
# Benchmarks
nox -s benchmark
M2 Actor + Tool Source Smoke Suite
The M2 smoke suite validates the foundation for the Actor Graphs + Tool Sources epic (#356). It covers actor YAML loading, skill registry, tool lifecycle, and MCP stub tool discovery.
Fixtures
M2 fixtures live in features/fixtures/m2/:
| Fixture File | Description |
|---|---|
m2_hierarchical_actor.yaml |
Graph-type actor with planner → executor → reviewer nodes |
m2_skill_pack.yaml |
Skill pack with tool references and an inline tool |
MCP Stub Server
The MCP stub server (features/mocks/mcp_stub_server.py) simulates MCP tool discovery and invocation in-process without requiring the real MCP adapter (#159). It exposes three deterministic tools: mcp/search, mcp/fetch, and mcp/transform.
Usage in tests:
from features.mocks.mcp_stub_server import McpStubServer
server = McpStubServer()
server.start()
tools = server.discover() # Returns 3 stub tools
result = server.invoke("mcp/search", {"query": "hello"})
server.stop()
Test Suites
| Suite | Framework | Scenarios | Description |
|---|---|---|---|
features/m2_actor_tool_smoke.feature |
Behave | 10 | Actor loading, skill registry, tool lifecycle, MCP stub |
robot/m2_actor_tool_smoke.robot |
Robot | 6 | CLI-level smoke for actor/skill/tool/MCP operations |
benchmarks/m2_actor_tool_smoke_bench.py |
ASV | 12 | Baseline runtime for actor, skill, tool, and MCP operations |
Running the M2 Smoke Suite
# Behave only (M2 feature)
nox -s unit_tests -- features/m2_actor_tool_smoke.feature
# Robot only (M2 suite)
nox -s integration_tests -- --suite robot/m2_actor_tool_smoke.robot
# Benchmarks
nox -s benchmark
What Is NOT Tested (deferred to dependent issues)
- Real MCP adapter integration (#159, assigned to Aditya)
- Agent skills loader (#160, assigned to Aditya)
- Hierarchical actor compiler (#158, assigned to Jeff)
The MCP stub server provides a test seam that can be replaced with the real adapter once #159 lands.
M1 Source-Code Plan Lifecycle Smoke Tests
The M1 source-code smoke suites verify the minimal end-to-end source-code
workflow: a git repo fixture, a git-checkout resource config, and a minimal
action YAML with strategy/execution actors. Helper steps create temp projects,
link resources, and capture plan IDs for subsequent CLI verification.
Fixtures (features/fixtures/m1/)
| Fixture File | Description |
|---|---|
git_repo.json |
Minimal git repo definitions with file listings and branch info |
git_checkout_resource.json |
Git-checkout resource configs with path and branch properties |
action_sourcecode.yaml |
Minimal action YAML with strategy/execution actors and arguments |
Behave Suite: features/m1_sourcecode_smoke.feature
16 scenarios covering:
| Area | Scenarios | Description |
|---|---|---|
| Fixture loading | 3 | Load and validate git repo, checkout resource, action YAML |
| Action create | 1 | Create action from YAML config via CLI |
| Project/resource | 2 | Create temp project, link resource |
| Plan use | 3 | Use action to create plan, with project and args |
| Plan execute | 1 | Execute transitions to execute phase |
| Plan diff | 1 | Show changeset via diff |
| Plan apply | 1 | Apply transitions to terminal state |
| Negative cases | 3 | Unknown action, wrong phase, invalid actor |
Step definitions: features/steps/m1_sourcecode_smoke_steps.py
All step names are prefixed with m1 smoke to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/m1_sourcecode_smoke.robot
8 integration tests exercising the M1 lifecycle through CLI:
| Test Case | Description |
|---|---|
| M1 Action Create From Config | Creates action from fixture YAML |
| M1 Plan Use Creates Strategize Plan | Uses action to create plan |
| M1 Plan Use With Project Link | Plan use with project argument |
| M1 Plan Execute Transitions Phase | Execute phase transition |
| M1 Plan Diff Shows Changeset | Diff command for changeset |
| M1 Plan Apply Reaches Terminal | Apply to terminal state |
| M1 Full Lifecycle Action To Apply | End-to-end flow |
| M1 Plan Use With Plain Format | --format plain for stable assertions |
Helper script: robot/helper_m1_sourcecode_smoke.py
ASV Benchmarks: benchmarks/m1_sourcecode_smoke_bench.py
Six benchmark suites measuring M1 CLI overhead:
M1ActionCreateSuite--time_action_create_from_yamlM1PlanUseSuite--time_plan_use_minimal,time_plan_use_with_project,time_plan_use_with_args,time_plan_use_plain_formatM1PlanExecuteSuite--time_plan_executeM1PlanApplySuite--time_plan_lifecycle_applyM1FixtureLoadSuite--time_load_git_repo_fixture,time_load_git_checkout_fixture
Running the M1 Source-Code Smoke Suites
# Behave only (M1 smoke feature)
nox -s unit_tests -- features/m1_sourcecode_smoke.feature
# Robot only (M1 smoke suite)
nox -s integration_tests -- --suite robot/m1_sourcecode_smoke.robot
# Benchmarks
nox -s benchmark
M1 Smoke Run Instructions
To perform a quick M1 source-code smoke run:
- Run the Behave feature:
nox -s unit_tests -- features/m1_sourcecode_smoke.feature - Run the Robot suite:
nox -s integration_tests -- --suite robot/m1_sourcecode_smoke.robot - Run benchmarks:
nox -s benchmark
Failure Triage Tips
AmbiguousSteperrors: All M1 smoke steps are prefixed withm1 smoke. If ambiguous, check that no other step file defines a conflicting pattern.- Fixture file not found: Verify
features/fixtures/m1/contains all three fixture files (git_repo.json,git_checkout_resource.json,action_sourcecode.yaml). - Robot
FAILsentinel: Each Robot helper subcommand prints a detailedFAIL:line with exit code and output when something goes wrong. Check the Robot log for these messages. InvalidPhaseTransitionError: Verify the plan is in the correct phase before attempting a transition. The M1 smoke tests mock phase state carefully.- Coverage drops: The M1 smoke tests cover fixture loading, CLI argument
parsing, and mock service integration. If coverage drops, look at
build/htmlcov/index.htmlfor uncovered lines in plan/action CLI commands.
M3 Decision Tree, Validation Gating, and Invariant Enforcement Smoke Tests
Overview
The M3 smoke suite validates the decision tree, correction workflow, validation gating (add / attach / detach), and invariant constraint (add / list / remove) subsystems via CLI integration tests.
Issue: #179
Files
| File | Purpose |
|---|---|
features/m3_decision_validation_smoke.feature |
Behave BDD scenarios (25 scenarios) |
features/steps/m3_decision_validation_smoke_steps.py |
Step implementations |
features/fixtures/m3/decision_tree_outputs.json |
Decision tree fixture data |
features/fixtures/m3/validation_attachments.json |
Validation attachment fixture data |
features/fixtures/m3/invariant_configs.json |
Invariant config fixture data |
robot/m3_decision_validation_smoke.robot |
Robot Framework integration tests (8 cases) |
robot/helper_m3_decision_validation_smoke.py |
Robot helper with 8 subcommands |
benchmarks/m3_smoke_bench.py |
ASV benchmark suites (5 classes) |
Scenario Categories
- Decision tree fixture loading -- Loads JSON fixtures and verifies decision tree structure, invariant nodes, and correction metadata.
- Validation fixture loading -- Verifies validation attachment data including required and informational modes.
- Invariant fixture loading -- Verifies invariant config data including merge set precedence.
- Invariant CLI -- Tests
agents invariant add,list, andremovecommands with global and project scopes. - Validation CLI -- Tests
agents validation add,attach, anddetachcommands. - Plan correct CLI -- Tests
agents plan correctwith--dry-run,--mode revert, and--mode append. - Negative cases -- Empty guidance, invalid correction mode.
Mock Strategy
- Invariant CLI patches
cleveragents.cli.commands.invariant._get_service - Validation CLI patches
cleveragents.cli.commands.validation._get_tool_registry_service - Plan CLI patches
cleveragents.cli.commands.plan._get_lifecycle_service - Correction patches
CorrectionServiceconstructor at module level
Failure Triage Tips
AmbiguousSteperrors: All M3 smoke steps are prefixed withm3 smoke. If ambiguous, check that no other step file defines a conflicting pattern.- Fixture file not found: Verify
features/fixtures/m3/contains all three JSON fixture files. - Robot
FAILsentinel: Each Robot helper subcommand prints a detailedFAIL:line with exit code and output. Check the Robot log. - Coverage drops: The M3 smoke tests cover invariant/validation CLI
commands, plan correct workflow, and fixture loading. Check
build/htmlcov/index.htmlfor uncovered lines.
M4 Correction + Subplan Smoke Suite
The M4 smoke suite validates correction flows, subplan execution, and conflict simulation for the M4 milestone. It covers decision tree corrections (revert and append modes, dry-run impact analysis), subplan lifecycle (sequential, parallel, dependency-ordered execution), failure handling (stop-others, retry logic), and merge conflict simulation (git-three-way, fail-on-conflict, last-wins, sequential-apply strategies).
Fixtures (features/fixtures/m4/)
| Fixture File | Description |
|---|---|
correction_flows.json |
Revert, append, dry-run, and high-risk correction scenarios |
subplan_execution.json |
Sequential, parallel, dependency-ordered, and retry subplan configs |
conflict_simulations.json |
Merge conflict scenarios for all SubplanMergeStrategy variants |
Behave Suite: features/m4_correction_subplan_smoke.feature
20 scenarios covering:
| Area | Scenarios | Description |
|---|---|---|
| Fixture loading | 3 | Load and validate correction, subplan, and conflict fixtures |
| Correction flows | 4 | Revert, append, dry-run, and invalid mode via CLI |
| Subplan execution | 2 | Sequential and parallel subplan status rendering |
| Failure handling | 3 | Stop-others, retry retriable, skip non-retriable |
| Conflict simulation | 3 | No-conflict, fail-on-conflict, last-wins merge strategies |
| Negative cases | 2 | Non-existent plan, empty decision ID |
Step definitions: features/steps/m4_correction_subplan_smoke_steps.py
All step names are prefixed with m4 smoke to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/m4_correction_subplan_smoke.robot
8 integration tests exercising M4 correction and subplan flows through CLI:
| Test Case | Description |
|---|---|
| M4 Correction Revert Via CLI | Invoke plan correct --mode revert |
| M4 Correction Append Via CLI | Invoke plan correct --mode append |
| M4 Correction Dry Run Via CLI | Invoke plan correct --dry-run |
| M4 Subplan Status Sequential | Plan status with sequential subplan config |
| M4 Subplan Status Parallel | Plan status with parallel subplan config |
| M4 Failure Handler Evaluation | SubplanFailureHandler decision logic |
| M4 Fixture Loading | Load all M4 fixture files |
| M4 Full Correction And Subplan Flow | End-to-end correction + subplan + failure handler |
Helper script: robot/helper_m4_correction_subplan_smoke.py
ASV Benchmarks: benchmarks/m4_smoke_bench.py
Four benchmark suites measuring M4 suite runtime:
M4CorrectionCLISuite--time_correction_revert,time_correction_append,time_correction_dry_runM4SubplanStatusSuite--time_status_with_subplans_json,time_status_with_subplans_richM4FailureHandlerSuite--time_should_stop_others_fail_fast,time_should_retry_retriable,time_should_retry_non_retriableM4FixtureLoadSuite--time_load_correction_flows,time_load_subplan_execution,time_load_conflict_simulations,time_load_all_fixtures
Running the M4 Correction + Subplan Smoke Suites
# Behave only (M4 smoke feature)
nox -s unit_tests -- features/m4_correction_subplan_smoke.feature
# Robot only (M4 smoke suite)
nox -s integration_tests -- --suite robot/m4_correction_subplan_smoke.robot
# Benchmarks
nox -s benchmark
Failure Triage Tips
AmbiguousSteperrors: All M4 smoke steps are prefixed withm4 smoke. If ambiguous, check that no other step file defines a conflicting pattern.- Fixture file not found: Verify
features/fixtures/m4/contains all three fixture files (correction_flows.json,subplan_execution.json,conflict_simulations.json). - Robot
FAILsentinel: Each Robot helper subcommand prints a detailedFAIL:line with exit code and output when something goes wrong. - Correction service mock: The M4 steps patch
CorrectionServicedirectly in the plan CLI module. If the import path changes, update the patch target. - Coverage drops: The M4 smoke tests cover correction CLI commands,
SubplanFailureHandler logic, and fixture loading. Check
build/htmlcov/index.htmlfor uncovered lines in correction/plan CLI code.
M5 ACMS Pipeline + Large-Project Context Smoke Tests
The M5 smoke suites verify the ACMS context pipeline foundation: context assembly, budget enforcement, multi-project context, context analysis, and project context policy resolution.
Behave: features/m5_acms_smoke.feature
27 scenarios covering:
| Area | Scenarios |
|---|---|
| Fixture loading | Load ACMS context policy, large project context, analysis results |
| Context policy resolution | Default view, inheritance chain, strategize/execute/apply overrides |
| Budget enforcement | max_file_size, max_total_size, zero/negative validation, unlimited |
| Context assembly (CLI) | list, add, show, clear via mocked ContextService |
| Project context policy CLI | show policy, inspect/simulate stubs (NotImplementedError) |
| Context analysis agent | Summary, dependencies, relevance scores (mocked) |
| Multi-project context | Independent context per project |
| Exclusion patterns | Glob-based path filtering |
Step definitions: features/steps/m5_acms_smoke_steps.py
Fixtures: features/fixtures/m5/
acms_context_policy.json-- Phase-specific context policy with size limitslarge_project_context.json-- Simulated 500-file project with tier metadatacontext_analysis_results.json-- Pre-computed analysis with dependencies and relevance scores
Robot: robot/m5_acms_smoke.robot
| Test Case | Description |
|---|---|
| Load ACMS Context Policy Fixture | Validates fixture structure |
| Load Large Project Context Fixture | Validates file count and entries |
| Resolve Default View From Empty Policy | Empty policy includes all paths |
| Resolve Strategize Inherits From Default | Phase inheritance works |
| Resolve Strategize With Override | Override takes precedence |
| Budget Max File Size Enforcement | File-level budget check |
| Budget Max Total Size Enforcement | Aggregate budget check |
| Invalid Phase Raises Error | Error handling for invalid phases |
| Context Analysis Fixture Has Required Fields | Analysis result structure |
| Multi Project Independent Context | Independent context per project |
Helper script: robot/helper_m5_acms_smoke.py
ASV Benchmarks: benchmarks/m5_smoke_bench.py
Four benchmark suites measuring M5 suite runtime:
PolicyResolutionSuite--time_resolve_viewacross all 4 phasesBudgetEnforcementSuite--time_check_file_budget,time_check_total_budgetwith 100/1000/10000 filesFixtureLoadSuite--time_load_policy_fixture,time_load_large_project_fixture,time_load_analysis_fixture,time_load_all_fixturesContextViewConstructionSuite--time_construct_view,time_construct_full_policywith 1/10/50 path patterns
Running the M5 ACMS Smoke Suites
# Behave only (M5 smoke feature)
nox -s unit_tests -- features/m5_acms_smoke.feature
# Robot only (M5 smoke suite)
nox -s integration_tests -- --suite robot/m5_acms_smoke.robot
# Benchmarks
nox -s benchmark
Failure Triage Tips
AmbiguousSteperrors: All M5 smoke steps are prefixed withm5 smoke. If ambiguous, check that no other step file defines a conflicting pattern.- Fixture file not found: Verify
features/fixtures/m5/contains all three fixture files (acms_context_policy.json,large_project_context.json,context_analysis_results.json). context_inspect/context_simulatestubs: These commands raiseNotImplementedErrorby design until ACMS wiring is complete.- Coverage drops: The M5 smoke tests cover context policy resolution, budget
enforcement, and CLI context commands. Check
build/htmlcov/index.htmlfor uncovered lines in context-related modules.
M6 Autonomy Acceptance Suite
E2E fixtures covering A2A facade flows and autonomy guardrails. The suite exercises guardrail enforcement (max steps, tool budget, confirmations) and audit trail persistence in integrated scenarios.
Overview
The M6 acceptance suite validates the autonomy hardening layer:
- A2A Local Facade — local-mode dispatch of all 11 supported operations (session, plan, registry, context, event).
- Autonomy Guardrails — denylist/allowlist enforcement, cost budget caps, tool call limits, write approval gates, apply approval gates.
- Automation Profiles — 8 built-in profiles, custom namespaced profiles, YAML loading, threshold validation, and four-level resolution precedence (plan > action > project > global).
- A2A Event Queue — publish/subscribe, local callbacks, unsubscribe, close semantics, remote stub rejection.
- A2A HTTP Transport Stub — all mutating methods raise
A2aNotAvailableErrorin local mode. - A2A Version Negotiation — accept supported versions, reject unsupported.
- A2A Model Validation — non-empty operation, valid status values, non-empty event types, non-empty error codes.
Fixtures (features/fixtures/m6/)
| File | Purpose |
|---|---|
a2a_facade_flows.json |
A2A operation flows (session, plan, registry, context, event) |
autonomy_guardrails.json |
Guard configurations (denylist, allowlist, budget, write/apply approval) |
automation_profiles.json |
Built-in profile assertions, custom profile config, resolution precedence cases |
Behave Suite: features/m6_autonomy_acceptance.feature
- Fixture loading — 3 scenarios verifying JSON structure
- A2A facade dispatch — 10 scenarios covering all 42 operations
- A2A error handling — 2 scenarios (unknown operation, invalid request type)
- Service registration — 2 scenarios (register + list operations)
- Event queue — 5 scenarios (publish, subscribe, unsubscribe, close, remote reject)
- HTTP transport stub — 4 scenarios (send, connect, disconnect, is_connected)
- Version negotiation — 3 scenarios (accept, reject, is_supported)
- Built-in profiles — 3 scenarios (list, manual thresholds, full-auto thresholds)
- Profile creation — 4 scenarios (custom, invalid name, bad threshold, YAML load)
- Guard enforcement — 8 scenarios (denylist, allowlist, call limit, budget, write, apply, no guards)
- Profile resolution — 3 scenarios (plan precedence, action precedence, global default)
- Service guard evaluation — 1 scenario
- Model validation — 4 scenarios (A2aRequest, A2aResponse, A2aEvent, A2aErrorDetail)
Step implementations: features/steps/m6_autonomy_acceptance_steps.py
Robot Suite: robot/m6_autonomy_acceptance.robot
| Test Case | What it checks |
|---|---|
| M6 A2A Facade Session Lifecycle | session.create + session.close dispatch |
| M6 A2A Facade Plan Lifecycle | plan create/execute/status/diff/apply |
| M6 A2A Facade Unknown Operation Error | A2aOperationNotFoundError on bad op |
| M6 A2A Event Queue Publish Subscribe | publish, subscribe, unsubscribe, close |
| M6 A2A Transport Stub Rejects All | send/connect/disconnect raise A2aNotAvailableError |
| M6 A2A Version Negotiation | accept 1.0, reject 2.0 |
| M6 Guard Denylist Enforcement | denylist blocks denied tools |
| M6 Guard Budget Enforcement | cost budget and call limit guards |
| M6 Profile Resolution Precedence | plan > action > project > global |
| M6 Fixture Loading | all 3 fixture files load and validate |
| M6 Full Autonomy Acceptance Flow | end-to-end facade + guard + resolution |
Helper: robot/helper_m6_autonomy_acceptance.py
ASV Benchmarks: benchmarks/m6_acceptance_bench.py
| Suite | Methods |
|---|---|
M6FacadeDispatchSuite |
time_session_create, time_plan_create, time_plan_execute, time_list_operations |
M6GuardEvaluationSuite |
time_denylist_check_allowed, time_denylist_check_denied, time_allowlist_check_allowed, time_allowlist_check_denied, time_budget_check_under, time_no_guards_check |
M6ProfileResolutionSuite |
time_resolve_plan_level, time_resolve_action_level, time_resolve_global_default, time_get_builtin_profile, time_list_profiles, time_evaluate_guard, time_version_negotiation |
M6EventQueueSuite |
time_publish_event, time_subscribe_and_publish, time_get_events |
M6FixtureLoadSuite |
time_load_a2a_facade_flows, time_load_autonomy_guardrails, time_load_automation_profiles, time_load_all_fixtures |
Running the M6 Autonomy Acceptance Suites
# Behave only
nox -s behave -- features/m6_autonomy_acceptance.feature
# Robot only
nox -s robot -- robot/m6_autonomy_acceptance.robot
# ASV benchmarks
nox -s benchmark
# Everything (default nox sessions)
nox
Failure Triage Tips
AmbiguousSteperrors: All M6 smoke steps are prefixed withm6 smoke. If ambiguous, check that no other step file defines a conflicting pattern.- Fixture file not found: Verify
features/fixtures/m6/contains all three fixture files (a2a_facade_flows.json,autonomy_guardrails.json,automation_profiles.json). - Robot
FAILsentinel: Each Robot helper subcommand prints a detailedFAIL:line with exit code and output when something goes wrong. - A2aNotAvailableError expected: The transport stub and remote event subscribe are supposed to raise this error. If a test fails here, it likely means the stub was changed to not raise.
- Guard evaluation failures: Check that the
AutomationGuardmodel fields match the fixture data. Denylist/allowlist/budget fields are all optional. - Coverage drops: The M6 acceptance tests cover the
a2a/package facade, events, transport, versioning, errors, and models, plusautomation_profile.pyandautomation_profile_service.py. Checkbuild/htmlcov/index.htmlfor uncovered lines.