Add comprehensive E2E test suite for M2 (Actor Graphs + Tool Sources) epic: - Behave BDD: 10 scenarios covering actor YAML loading, skill registry, tool lifecycle (discover/activate/execute/deactivate), and MCP stub - Robot Framework: 6 integration tests via CLI helper script - ASV benchmarks: 12 benchmarks for actor loading, skill registry, tool lifecycle, and MCP stub performance baselines - MCP stub server mock: in-process fake with 3 tools (search/fetch/transform) - Fixtures: hierarchical graph actor YAML + skill pack with tool refs and inline tools - Docs: updated testing.md with M2 smoke suite section Closes #169
30 KiB
Testing Guide
Overview
CleverAgents uses a comprehensive testing strategy with three complementary frameworks:
- Behave (BDD/Gherkin) for unit-level and scenario tests under
features/ - Robot Framework for integration and end-to-end tests under
robot/ - ASV (airspeed velocity) for performance benchmarks under
benchmarks/ - Coverage:
coverage.pywith branch coverage, enforced at >=97%
All tests are executed exclusively through nox sessions. Never invoke behave, robot, or other test runners directly.
Running Tests
# Run all default sessions (lint, typecheck, unit tests, integration tests, coverage)
nox
# Unit tests only (Behave)
nox -s unit_tests
# Integration tests only (Robot Framework)
nox -s integration_tests
# Coverage report with 97% enforcement
nox -s coverage_report
# Run a specific feature file
nox -s unit_tests -- features/plan_model.feature
# Benchmarks (ASV)
nox -s benchmark
# Type checking (pyright)
nox -s typecheck
# Linting (ruff)
nox -s lint
# Formatting check
nox -s format -- --check
Coverage Requirements
Threshold: 97%
The project enforces a minimum 97% code coverage at all times. This is a hard gate:
- The
nox -s coverage_reportsession fails if coverage drops below 97%. - The Forgejo CI
coveragejob runsnox -s coverage_reportand blocks merges on failure. - Branch coverage is enabled (
branch = trueinpyproject.toml).
How Coverage Is Measured
Coverage is measured by running Behave tests under coverage.py in serial mode:
coverage run --source=src -m behave -q --tags=-discovery --no-capture features/
coverage report --show-missing --fail-under=97
Coverage Output
On success, the nox session emits:
COVERAGE OK: 97.2% (threshold: 97%)
On failure, the nox session emits and exits non-zero:
COVERAGE FAILED: 95.3% < 97% threshold
Both messages are single-line and CI-parseable. The CI pipeline greps for these lines to surface them in job summaries.
Coverage Reports
Three report formats are generated under build/:
| Report | Path | Description |
|---|---|---|
| Terminal | stdout | Per-file summary with --show-missing |
| HTML | build/htmlcov/index.html |
Interactive browser report |
| XML | build/coverage.xml |
Cobertura-format for CI tools |
| JSON | build/coverage.json |
Machine-readable totals and per-file data |
Coverage Configuration
Coverage settings live in pyproject.toml:
[tool.coverage.run]
source = ["src", "scripts"]
branch = true
parallel = false
omit = [
"*/tests/*",
"*/test_*",
"features/*",
"*/features/*",
"*/__pycache__/*",
"*/site-packages/*",
"*/dependency_injector/*",
"*/venv/*",
"*/.venv/*",
"*/.nox/*",
"src/cleveragents/discovery/*",
]
data_file = "build/.coverage"
[tool.coverage.html]
directory = "build/htmlcov"
[tool.coverage.xml]
output = "build/coverage.xml"
The --fail-under=97 threshold is enforced in the coverage_report nox session (noxfile.py).
How to Improve Coverage
- Run
nox -s coverage_reportto generate the HTML report. - Open
build/htmlcov/index.htmlin a browser to identify uncovered files. - Sort by "Missing" column to find files with the most uncovered lines.
- Write Behave scenarios targeting the uncovered code paths.
- Re-run
nox -s coverage_reportto verify improvement.
Do not lower the threshold, add # pragma: no cover comments, or expand the omit list without explicit team agreement.
All test files go under features/ (Behave) or robot/ (Robot Framework). Never create a tests/ directory.
Unit Tests (Behave)
Unit tests use the Behave BDD framework with Gherkin-syntax .feature files.
Structure
features/
*.feature # Feature files (Gherkin scenarios)
steps/ # Step definitions (Python)
mocks/ # Mock implementations (test doubles)
fixtures/ # Test fixture data
environment.py # Behave hooks (before/after scenario, etc.)
Guidelines
- No pytest: All unit tests must be Behave-based. There is intentionally no
tests/directory. - Step naming: Use unique, descriptive step names. Behave loads all step files globally; ambiguous steps cause
AmbiguousSteperrors. - Feature naming: Name feature-specific step files after their feature (e.g.,
foo.feature->foo_steps.py). - Mocks in features/ only: All mock implementations live in
features/mocks/. Production code (src/) must never contain test doubles.
Running Specific Features
# Run a single feature file
nox -s unit_tests -- features/plan_model.feature
# Run scenarios with a specific tag
nox -s unit_tests -- --tags=@smoke
Integration Tests (Robot Framework)
Integration tests use Robot Framework for end-to-end and system-level testing.
Structure
robot/
*.robot # Test suites
*.resource # Shared keywords and variables
common.resource # Common setup/teardown keywords
v2_paths.resource # Path configuration
Guidelines
- Resource paths: Always use
${CURDIR}/prefix forResourceimports (e.g.,Resource ${CURDIR}/common.resource). - Python injection: Use
${PYTHON}variable (injected by nox) instead of barepythoninRun Processcalls. - Timeouts: Always add
timeout=toRun Processcalls that invoke long-running commands. - Slow tests: Tag tests that require external services (API keys, running servers) with
slow. These are excluded in CI via--exclude slow.
Running Specific Suites
# Run a single Robot suite
nox -s integration_tests -- --suite robot/plan_generation_graph.robot
# Include slow tests
nox -s slow_integration_tests
Performance Benchmarks (ASV)
ASV (airspeed velocity) tracks performance over time.
Structure
benchmarks/
*.py # Benchmark suites (Time*/Mem*/Track* classes)
asv.conf.json # ASV configuration
Running Benchmarks
nox -s benchmark
Writing Effective Tests
Behave Test Guidelines
Feature: Plan lifecycle transitions
Scenario: Execute plan transitions phase from strategize to execute
Given an action "local/build-app" exists
And a plan is created using "local/build-app" on project "local/my-project"
And strategize is completed
When I execute the plan
Then the plan phase should be "execute"
And the processing state should be "queued"
Robot Framework Guidelines
*** Settings ***
Documentation Plan lifecycle integration tests
Library Process
Library OperatingSystem
Resource ${CURDIR}/common.resource
*** Test Cases ***
Plan Execute Transitions Correctly
[Documentation] Verify execute phase transition via CLI
${result}= Run Process ${PYTHON} -m cleveragents plan execute ${PLAN_ID}
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} execute
CLI Lifecycle Test Suites
The CLI lifecycle suites provide comprehensive coverage for action and plan lifecycle commands, including success paths, error paths, multi-project arguments, automation profile overrides, invariants, actor overrides, phase visibility, and terminal outcomes.
Behave Suite: features/cli_lifecycle_coverage.feature
Covers 68 scenarios across these areas:
| Area | Scenarios | Description |
|---|---|---|
| Action CRUD | 13 | create, list, show, archive with success + error paths |
| Plan Use | 10 | basic use, multi-project, automation profile, invariants, actor overrides |
| Plan Status | 9 | phase visibility (strategize, execute, apply) and terminal outcomes (applied, constrained, errored, cancelled) |
| Plan Execute | 5 | execute success, phase transition, error paths |
| Plan Apply | 5 | lifecycle-apply success, terminal outcomes, error paths |
| Plan Cancel | 5 | cancel with reason, already-terminal guard, error paths |
| Plan List | 4 | lifecycle-list with and without filters |
| Negative cases | 17 | missing config, invalid args, invalid project names, unknown actions/resources |
Step definitions: features/steps/cli_lifecycle_coverage_steps.py
All step names are prefixed with lifecycle coverage to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/cli_lifecycle_e2e.robot
End-to-end integration tests exercising the full action-to-apply lifecycle through
the CLI entry points. Uses robot/helper_cli_lifecycle_e2e.py as a Python helper
that mocks the service layer while exercising real CLI argument parsing and output
formatting.
| Test Case | Description |
|---|---|
| Action Create From Config Via CLI | Creates an action from YAML config |
| Plan Use Creates Plan In Strategize Phase | Uses an action to create a plan |
| Plan Execute Transitions To Execute Phase | Executes a plan, checks phase |
| Plan Lifecycle Apply Transitions To Apply Phase | Applies a plan, checks phase |
| Plan Status Shows Plan Details | Verifies status output rendering |
| Plan Cancel Cancels Non-Terminal Plan | Cancels with a reason string |
| Full Lifecycle Action To Apply | End-to-end: create -> use -> execute -> apply |
| Plan Lifecycle List Shows Plans | Verifies lifecycle-list output |
ASV Benchmark: benchmarks/plan_cli_smoke_bench.py
Three benchmark suites measuring CLI argument parsing overhead (service layer is mocked so only parsing + routing cost is measured):
ActionCLIParsingSuite--time_action_create,time_action_list,time_action_list_with_filters,time_action_show,time_action_archivePlanUseArgParsingSuite--time_use_minimal,time_use_multi_project,time_use_with_all_flags,time_use_with_args_onlyPlanLifecycleCmdParsingSuite--time_execute,time_lifecycle_apply,time_status_by_id,time_status_list_all,time_cancel_with_reason,time_lifecycle_list,time_lifecycle_list_filtered
Running the CLI Lifecycle Suites
# Behave only (our feature)
nox -s unit_tests -- features/cli_lifecycle_coverage.feature
# Robot only (our suite)
nox -s integration_tests -- --suite robot/cli_lifecycle_e2e.robot
# Benchmarks
nox -s benchmark
Robot Suite: robot/cli_lifecycle.robot
End-to-end Robot suite for the CLI lifecycle with git_worktree sandbox simulation
and ChangeSet capture verification. Uses robot/helper_cli_lifecycle.py as a
Python helper that exercises the full action → plan → execute → apply flow with
mocked service layer, simulated sandbox workspace, and InMemoryChangeSetStore
for change tracking.
Positive tests:
| Test Case | Description |
|---|---|
| Action Create Via Config CLI | Creates an action from YAML config |
| Plan Use Creates Strategize Plan With Sandbox | Uses action with sandbox workspace mock |
| Plan Execute With ChangeSet Capture | Executes plan, verifies changeset entry recorded |
| Plan Apply With ChangeSet Verification | Applies plan, verifies changeset summary |
| Full Lifecycle Action To Apply With Sandbox | End-to-end with sandbox + changeset |
Negative tests:
| Test Case | Description |
|---|---|
| Invalid Project Name Rejected | Invalid project name format handled gracefully |
| Missing Resource Returns Error | Non-existent action returns error |
| Invalid Arg Format Rejected | --arg without = separator rejected |
Fixtures:
helper_cli_lifecycle.pycreates temporary YAML config files and sandbox directories- Uses
InMemoryChangeSetStorefor ChangeSet capture without database - Each subcommand is self-contained; cleanup runs in
finallyblocks
Aligned Behave feature: features/cli_lifecycle_robot_alignment.feature
mirrors the Robot E2E flow to keep unit/integration expectations aligned.
ASV benchmark: benchmarks/cli_robot_flow_bench.py measures fixture setup
overhead (CLIRobotFixtureSetupSuite) and full lifecycle CLI throughput
(CLIRobotLifecycleSuite).
Running the Robot CLI Lifecycle Suite
# Robot suite
nox -s integration_tests -- --suite robot/cli_lifecycle.robot
# Aligned Behave feature
nox -s unit_tests -- features/cli_lifecycle_robot_alignment.feature
# Benchmarks
nox -s benchmark
Persistence Test Suites
The persistence layer (SQLAlchemy repositories for plans and actions) has dedicated test suites at both unit and integration levels.
Behave: Plan Persistence (features/plan_persistence.feature)
19 scenarios covering LifecyclePlanRepository operations:
- Create: Persist plans via repository, retrieve by ULID, verify all fields round-trip.
- Phase/state transitions: Update
phaseandprocessing_state, verify persistence after each transition. - List filters: Filter plans by phase, processing state, and action name.
- Plan tree links: Parent/child/root hierarchy with
parent_plan_idandroot_plan_id. - Cross-restart: Close the SQLite session, reopen from disk, verify plans survive reconnection (including project links, arguments, and invariants).
Step definitions: features/steps/plan_persistence_steps.py
Fixtures: Each scenario gets a fresh temp SQLite database via _setup_db() / _teardown_db() helper functions. The database is file-backed (not :memory:) so cross-restart scenarios can close and reopen the connection.
Behave: Action Persistence (features/action_persistence.feature)
14 scenarios covering ActionRepository operations:
- CRUD: Create, retrieve by
namespaced_name, list available, filter by namespace, archive, delete. - Argument ordering: Persist 3
ActionArgumententries and verify positional order is preserved. - Invariant ordering: Persist 3 invariant strings and verify order is preserved.
- Terminal state storage: Create plans from actions in Apply phase with each terminal
ProcessingState(applied,constrained,errored,cancelled) and verify persistence.
Step definitions: features/steps/action_persistence_steps.py
Robot: Plan Persistence E2E (robot/plan_persistence_e2e.robot)
5 end-to-end tests exercising the persistence layer through Python helper scripts:
| Test Case | Description |
|---|---|
| Plan Full Lifecycle Persistence | Create action + plan, transition through all phases to applied terminal state |
| Plan Restart Persistence | Create plan, close DB, reopen, verify all fields survive |
| Plan Concurrent Session Access | Two independent sessions against the same DB file |
| Action CRUD Persistence E2E | Create, read, update, delete an action with arguments/invariants |
| Plan Tree Hierarchy Persistence E2E | Parent/child plan hierarchy with root_plan_id verification |
Helper script: robot/helper_plan_persistence_e2e.py
Robot: Persistence Lifecycle (robot/persistence_lifecycle.robot)
7 end-to-end tests exercising persistence lifecycle patterns through a Python helper:
| Test Case | Description |
|---|---|
| Plan Full Lifecycle Phase Transitions | Create action + plan, transition through all phases to applied terminal state |
| Process Restart Simulation | Write plan, close DB, reopen from disk, verify all fields survive |
| Reopen Plan Status After Restart | Close/reopen, verify phase, state, automation level, and tags unchanged |
| Concurrent CLI Access Safeguards | Two independent sessions see each other's committed data |
| Stored Arguments Ordering Persistence | 4 ordered ActionArgument entries survive roundtrip |
| Stored Invariants Ordering Persistence | 3 ordered invariant strings survive roundtrip |
| Project Links Persistence Through Restart | 2 ProjectLink entries survive close/reopen cycle |
Helper script: robot/helper_persistence_lifecycle.py
Behave: Persistence Robot Alignment (features/persistence_robot_alignment.feature)
1 scenario mirroring the Robot restart-simulation flow to keep unit and integration persistence expectations aligned. Step definitions in features/steps/persistence_robot_alignment_steps.py.
ASV Benchmarks: Robot Fixture Overhead (benchmarks/persistence_robot_bench.py)
Performance baselines for Robot fixture setup operations:
TimeRobotFixtureSetupSuite.time_file_db_init-- File-backed SQLite init + schema creationTimeRobotFixtureSetupSuite.time_seed_action-- Action seeding into a fresh file DBTimeRobotFixtureSetupSuite.time_seed_plan-- Plan creation including FK action seedingTimeRobotRestartCycleSuite.time_close_reopen_verify-- Close + reopen + get cycle latency
ASV Benchmarks (benchmarks/persistence_suites_bench.py)
Performance baselines for persistence test operations:
TimePlanPersistenceSuites.time_create_and_retrieve_plan-- Plan round-trip latencyTimePlanPersistenceSuites.time_phase_state_transition-- Phase/state update cycleTimePlanPersistenceSuites.time_list_plans_filtered-- List with filtersTimePlanPersistenceSuites.time_plan_tree_operations-- Parent/child hierarchyTimeActionPersistenceSuites.time_create_and_retrieve_action-- Action round-tripTimeActionPersistenceSuites.time_action_arguments_ordering-- Ordered arguments persistenceTimeActionPersistenceSuites.time_cross_restart_persistence-- Close/reopen cycle
Run with: nox -s benchmark
Running Persistence Tests
# Unit tests (Behave) -- runs all features including persistence
nox -s unit_tests
# Run only plan persistence scenarios
nox -s unit_tests -- features/plan_persistence.feature
# Run only action persistence scenarios
nox -s unit_tests -- features/action_persistence.feature
# Integration tests (Robot) -- runs all robot suites including persistence E2E
nox -s integration_tests
# Run only plan persistence E2E
python -m robot robot/plan_persistence_e2e.robot
# Run only persistence lifecycle E2E
python -m robot robot/persistence_lifecycle.robot
# Run only persistence robot alignment (Behave)
nox -s unit_tests -- features/persistence_robot_alignment.feature
Subplan Model Test Suites
The subplan model (hierarchy, configuration, failure handling) has dedicated test suites at unit, integration, and benchmark levels.
Behave: Subplan Model (features/subplan_model.feature)
33 scenarios covering Plan hierarchy properties, SubplanConfig defaults, SubplanStatus tracking, SubplanFailureHandler decisions, and CLI dict rendering:
| Area | Scenarios | Description |
|---|---|---|
| Root plan identity | 2 | is_root_plan, is_subplan, depth for root plans |
| Child plan identity | 3 | Parent reference, attempt counter, root propagation |
| Default values | 4 | Phase, state, automation level, subplan_config defaults |
| SubplanConfig storage | 3 | Sequential/parallel config, standalone defaults |
| Hierarchy helpers | 3 | has_subplans, depth for non-root plans |
| Lifecycle transitions | 2 | Child plans follow same phase transition rules |
| Dependency guardrails | 5 | FailureHandler stop/retry decisions, phase independence |
| CLI dict rendering | 3 | parent_plan_id, subplan_count in output |
| Validation guardrails | 4 | Invalid ULIDs, out-of-range config values |
Step definitions: features/steps/subplan_model_steps.py
All step names are prefixed with subplan-test to avoid AmbiguousStep conflicts
with existing plan_model_steps.py steps.
Robot: Subplan Model (robot/subplan_model.robot)
11 smoke tests exercising subplan model properties via Python helper scripts:
| Test Case | Description |
|---|---|
| Root Plan Identity Flags | is_root_plan=True, is_subplan=False, depth=0 |
| Child Plan Identity Flags | is_subplan=True, is_root_plan=False, depth=-1 |
| Three Level Hierarchy Root Propagation | root_plan_id propagates through 3 levels |
| SubplanConfig Defaults | All default values are sensible |
| SubplanConfig Custom Settings | Parallel mode with custom retry settings |
| Failure Handler Stop Others On Fail Fast | should_stop_others with fail_fast=True |
| Failure Handler Retry Retriable Error | should_retry for TimeoutError |
| Failure Handler Skip Non Retriable Error | should_retry skips ConfigurationError |
| CLI Dict Renders Parent Plan Id For Child | parent_plan_id in child plan CLI dict |
| CLI Dict Renders Subplan Count | subplan_count in parent plan CLI dict |
| Plan With Subplan Config And Statuses | Full plan model with config + statuses |
Helper script: robot/helper_subplan_model.py
ASV Benchmarks (benchmarks/subplan_model_validation_bench.py)
Performance baselines for subplan model operations:
SubplanConfigValidationSuite--time_default_config_creation,time_custom_config_creation,time_config_model_dump,time_config_model_dump_jsonSubplanStatusValidationSuite--time_minimal_status_creation,time_full_status_creation,time_status_model_dumpSubplanFailureHandlerSuite--time_should_stop_others_fail_fast,time_should_stop_others_no_fail_fast,time_should_retry_retriable,time_should_retry_non_retriableSubplanHierarchySuite--time_is_root_plan,time_is_subplan,time_depth_root,time_depth_child,time_has_subplans,time_cli_dict_with_subplans,time_child_plan_creation,time_plan_with_subplans_creation
Running Subplan Model Tests
# Behave only (subplan model feature)
nox -s unit_tests -- features/subplan_model.feature
# Robot only (subplan model suite)
nox -s integration_tests -- --suite robot/subplan_model.robot
# Benchmarks
nox -s benchmark
CI Pipeline
The Forgejo CI pipeline runs these jobs in order:
- lint —
nox -s lint(ruff check) - typecheck —
nox -s typecheck(pyright) - security —
nox -s security_scan(bandit + vulture) - quality —
nox -s complexity(radon) - unit_tests —
nox -s unit_tests(behave) - integration_tests —
nox -s integration_tests(robot) - coverage —
nox -s coverage_report(fail-under 97%) - build —
nox -s build(wheel)
All jobs use Python 3.13 and route commands through nox. See docs/development/ci-cd.md for full CI/CD documentation.
Troubleshooting
Coverage report shows 0% or very low coverage
Ensure you're running in serial mode (not parallel). The coverage_report nox session handles this correctly. Do not run behave directly outside of coverage run.
Also ensure:
parallel = falseis set in[tool.coverage.run]COVERAGE_FILEandCOVERAGE_RCFILEenv vars are set correctly by the nox session
AmbiguousStep errors in Behave
Two step files define the same @given/@when/@then pattern. Behave loads all step files globally. Make step names unique or use more specific patterns.
Robot Resource file does not exist
Use Resource ${CURDIR}/filename.resource instead of bare Resource filename.resource. Robot resolves bare paths relative to CWD, not the test file directory.
Robot No module named cleveragents
Ensure Robot tests use ${PYTHON} (injected by nox from the venv) instead of bare python in Run Process calls.
Tests hang indefinitely
Add timeout parameters to Run Process calls in Robot tests:
${result}= Run Process ${PYTHON} -m cleveragents ... timeout=30s
Validation Edge Case Test Catalog
The validation edge case suites exercise boundary conditions, error paths, and conflict detection across the validation subsystem.
Shared Fixtures (features/fixtures/validation/)
| Fixture File | Description |
|---|---|
malformed_tool_output.json |
Missing required fields, wrong types, extra nulls, non-object outputs |
missing_resources.json |
Dangling references, unresolvable paths, circular dependencies |
validation_timeouts.json |
Slow validators, zero/negative timeouts, partial results |
invalid_schema_transforms.json |
Transforms returning non-dict, missing type, null, circular refs |
mixed_ordering.json |
Required vs informational ordering, duplicate attachment IDs |
Each fixture file is a JSON document with a "fixtures" array. Each entry has
"name", "description", "input", and "expected_error" fields.
Behave Suite: features/validation_edge_cases.feature
25 scenarios covering:
| Area | Scenarios | Description |
|---|---|---|
| Malformed tool output | 6 | Missing fields, wrong types, nulls, empty/string outputs |
| Missing resources | 5 | Dangling actions, unresolvable paths, missing actors, circular deps |
| Validation timeouts | 4 | Slow validators, zero timeout, negative timeout, partial results |
| Invalid schema transforms | 3 | Non-dict returns, missing type field, null schema |
| Mixed ordering | 4 | Required-before-informational, duplicate attachment IDs, unknown levels |
| Concurrent conflicts | 1 | Two validators on same resource detect conflict |
| Rollback on failure | 2 | Required failure triggers rollback of pending changes |
Step definitions: features/steps/validation_edge_case_steps.py
All step names are prefixed with validation edge to avoid AmbiguousStep
conflicts with existing steps.
Robot Suite: robot/validation_edge.robot
6 integration tests exercising the validation edge case helper:
| Test Case | Description |
|---|---|
| Validation Edge Load All Fixtures | Loads all fixture files and verifies structure |
| Validation Edge Malformed Output Detection | Detects malformed tool output patterns |
| Validation Edge Missing Resource Error Paths | Validates missing resource reference detection |
| Validation Edge Timeout Simulation | Validates timeout simulation fixture data |
| Validation Edge Schema Validation Errors | Validates schema transform error detection |
| Validation Edge Unknown Command Returns Error | Verifies unknown command exits with code 1 |
Helper script: robot/helper_validation_edge.py
ASV Benchmarks: benchmarks/validation_edge_bench.py
Two benchmark suites:
ValidationEdgeFixtureLoadSuite--time_load_malformed_tool_output,time_load_missing_resources,time_load_validation_timeouts,time_load_invalid_schema_transforms,time_load_mixed_ordering,time_load_all_fixturesValidationEdgeDetectionSuite--time_malformed_output_detection,time_missing_resource_detection,time_validation_ordering,time_combined_edge_detection
Running the Validation Edge Case Suites
# Behave only (validation edge cases)
nox -s unit_tests -- features/validation_edge_cases.feature
# Robot only (validation edge suite)
nox -s integration_tests -- --suite robot/validation_edge.robot
# Benchmarks
nox -s benchmark
Scale Testing
Scale tests validate performance baselines for repository indexing and decomposition across different repository sizes (1K, 5K, and 10K files).
Fixtures
Scale test fixtures live in features/fixtures/scale/:
scale_metadata.json— Scale profile definitions (file count, language mix, expected timings)baseline_thresholds.json— Performance threshold matrices (p50/p95/p99 per operation per file count)generator_instructions.md— Manual instructions for generating synthetic repos
Running Scale Tests
# Behave scale test scenarios
nox -s unit_tests -- features/scale_test.feature
# Robot scale test integration
nox -s integration_tests -- --suite robot/scale_test.robot
# ASV scale fixture benchmarks
nox -s benchmark
Scale Test Suites
| Suite | Framework | Scenarios | Description |
|---|---|---|---|
features/scale_test.feature |
Behave | 20 | Fixture loading, metadata validation, threshold checks, distribution simulation |
robot/scale_test.robot |
Robot | 6 | Integration-level fixture validation via helper script |
benchmarks/scale_fixture_bench.py |
ASV | 6 | Performance benchmarks for fixture processing |
For full scale testing documentation, see docs/development/scale_testing.md.
M2 Actor + Tool Source Smoke Suite
The M2 smoke suite validates the foundation for the Actor Graphs + Tool Sources epic (#356). It covers actor YAML loading, skill registry, tool lifecycle, and MCP stub tool discovery.
Fixtures
M2 fixtures live in features/fixtures/m2/:
| Fixture File | Description |
|---|---|
m2_hierarchical_actor.yaml |
Graph-type actor with planner → executor → reviewer nodes |
m2_skill_pack.yaml |
Skill pack with tool references and an inline tool |
MCP Stub Server
The MCP stub server (features/mocks/mcp_stub_server.py) simulates MCP tool discovery and invocation in-process without requiring the real MCP adapter (#159). It exposes three deterministic tools: mcp/search, mcp/fetch, and mcp/transform.
Usage in tests:
from features.mocks.mcp_stub_server import McpStubServer
server = McpStubServer()
server.start()
tools = server.discover() # Returns 3 stub tools
result = server.invoke("mcp/search", {"query": "hello"})
server.stop()
Test Suites
| Suite | Framework | Scenarios | Description |
|---|---|---|---|
features/m2_actor_tool_smoke.feature |
Behave | 11 | Actor loading, skill registry, tool lifecycle, MCP stub |
robot/m2_actor_tool_smoke.robot |
Robot | 6 | CLI-level smoke for actor/skill/tool/MCP operations |
benchmarks/m2_actor_tool_smoke_bench.py |
ASV | 12 | Baseline runtime for actor, skill, tool, and MCP operations |
Running the M2 Smoke Suite
# Behave only (M2 feature)
nox -s unit_tests -- features/m2_actor_tool_smoke.feature
# Robot only (M2 suite)
nox -s integration_tests -- --suite robot/m2_actor_tool_smoke.robot
# Benchmarks
nox -s benchmark
What Is NOT Tested (deferred to dependent issues)
- Real MCP adapter integration (#159, assigned to Aditya)
- Agent skills loader (#160, assigned to Aditya)
- Hierarchical actor compiler (#158, assigned to Jeff)
The MCP stub server provides a test seam that can be replaced with the real adapter once #159 lands.