Files
cleveragents-core/docs/development/TEST_ARCHITECTURE.md
T
HAL9000 5134e6e719 refactor(tests): add TEST_ARCHITECTURE.md clarifying behave vs robot framework roles (#9049)
Create comprehensive documentation defining Behave (business-facing BDD/Gherkin)
and Robot Framework (technical integration/E2E) roles in the test architecture.
Includes decision criteria for framework selection, known duplication areas,
and consolidation recommendations.

Changes:
- docs/development/TEST_ARCHITECTURE.md: new file with framework guidelines
- CHANGELOG.md: added entry under [Unreleased] / ### Added
- CONTRIBUTORS.md: added contribution entry for this change

ISSUES CLOSED: #9049
2026-06-02 16:17:42 -04:00

7.2 KiB

Test Architecture

Issue: #9049

This document defines the roles and responsibilities of each testing framework in the CleverAgents test suite, establishes clear guidelines for selecting the appropriate framework when writing new tests, and documents known areas where test logic is duplicated across frameworks.

Frameworks Overview

CleverAgents uses two complementary BDD/testing frameworks:

Framework Purpose Typical Test Level Audience Language
Behave Business-facing behavior verification (BDD/Gherkin) Unit and integration Technical + non-technical Gherkin .feature files with Python step definitions
Robot Framework Technical integration and end-to-end testing Integration and E2E Primarily technical Robot .robot files with Python/Robot keywords

Behave: Business-Facing BDD Tests

Use Behave when writing tests that describe the system from a user or customer perspective.

When to Use Behave

  • The test scenario can be expressed as Given / When / Then steps readable by non-developers
  • You're testing business logic rules, validation, or domain constraints
  • Stakeholders need to review or modify test scenarios without understanding implementation details
  • You're testing CLI command behavior through a step-mapping abstration
  • Coverage goals require high-scenario-count verification of edge cases and boundary conditions

Behave Convention

Feature: Example feature name
  As a [role]
  I want [behavior/goal]
  So that [benefit]

  @example_tag
  Scenario: Specific behavior description
    Given the system is in state X
    When the user performs action Y
    Then the result should be Z

Location: features/*.feature Step definitions: features/steps/*_steps.py Run command: nox -s unit_tests -- features/<name>.feature

Common Behave Testing Patterns

  • Fixture loading and validation (JSON/YAML fixtures in features/fixtures/)
  • Repository persistence tests with in-memory SQLite databases
  • Domain model lifecycle and transition tests
  • CLI command argument parsing, validation, and output rendering
  • Mocked service integration scenarios
  • Validation edge cases and boundary conditions

Robot Framework: Technical Integration and E2E Tests

Use Robot Framework when writing tests that verify technical behavior, system interactions, and integrated workflows.

When to Use Robot Framework

  • The test requires process execution, subprocess commands, or shell operations
  • You're testing end-to-end workflows across multiple services or components
  • The test involves actual file system operations, database connections, or network calls
  • You need structured logging for debugging complex multi-step procedures
  • You're testing CLI tools as they run in production-like conditions

Robot Framework Convention

*** Test Cases ***
Example Test Case Name
    [Documentation]    Brief description of what the test verifies
    ${result}=    ${PYTHON}    -m    cleveragents    <command>    <args>
    Should Be Equal As Strings    ${result.returncode}    0
    Log    ${result.stdout}

Location: robot/*.robot Helper scripts: robot/helper_*.py Run command: nox -s integration_tests -- robot/<name>.robot

Common Robot Framework Testing Patterns

  • CLI end-to-end execution with actual subprocess invocation
  • Tool and skill lifecycle smoke tests
  • Integration-level fixture validation via helper Python scripts
  • Performance benchmarks (ASV) often paired with Robot execution
  • Database persistence and migration testing in real environments
  • File system operations, sandboxing, and resource management
  • Concurrency and threading behavior under production-like conditions

Decision Criteria: Behe vs. Robot Framework

When deciding which framework to use for a new test, follow this decision tree:

  1. Can the scenario be described in natural language (Given/When/That) without implementation details?

    • Yes → Use Behave
    • No → Continue to question 2
  2. Does the test require subprocess execution, process spawning, or shell commands?

    • Yes → Use Robot Framework
    • No → Continue to question 3
  3. Are you testing domain logic, validation rules, or business constraints?

    • Yes → Use Behave
    • No → Continue to question 4
  4. Is the test an end-to-end workflow involving multiple system components?

    • Yes → Use Robot Framework
    • No → Use Robot Framework for infrastructure/service tests, Behave for pure unit tests
  5. Are you writing performance benchmarks?

    • Both can be used: ASV (via Python) for baseline measurements, paired with Robot for integration validation

Summary Rule of Thumb

  • Business-facing, behavior-driven, stakeholder-readable → Behave
  • Technical, integration-focused, subprocess-heavy, E2E workflows → Robot Framework

Known Test Duplications and Consolidation Candidates

The following areas have overlapping test coverage across both frameworks and should be considered for consolidation:

Area Behave Location Robot Location Recommendation
Actor Compiler features/actor_cli_yaml.feature / *_steps.py robot/actor_compiler.robot / helper_actor_compiler.py Consolidate into Behave for unit, keep Robot for E2E validation only
Provider Registry features/provider_registry_coverage.feature Not applicable No duplication — Behave only
Session CLI features/session_cli.feature / *_steps.py robot/session_commands.robot Consolidate session commands into one framework per feature area
Actor Subgraph Cycle Detection features/actor_subgraph_cycle_detection.feature / *_steps.py robot/actor_compiler.robot cycle detection tests Deduplicate — Robot should not repeat compiler unit tests
Git Tools Concurrency features/git_tools.feature / git_tools_thread_safety_steps.py None No duplication
Provider Registry Coverage Boost features/provider_registry_coverage_boost.feature None No duplication
Actor Remove CLI Not applicable robot/actor_remove_cli.robot / helper_actor_remove_cli.py Robot only — this is an E2E CLI smoke test

Audit Recommendations

  1. Quarterly Duplication Review: Run both frameworks side by side and compare scenario lists to identify newly overlapping coverage
  2. New Test Checklist: Before creating a new feature file or robot suite, check for existing coverage in the alternative framework
  3. Framework Rationale Documentation: Each new test suite should include a comment at the top explaining why its chosen framework was selected over the alternative
  4. Coverage Monitoring: Ensure that neither framework's deletion causes the total project coverage to fall below the 97% threshold