From cd06fe58f6def2f73ccd93777898799a7cee7749 Mon Sep 17 00:00:00 2001 From: Jeffrey Phillips Freeman Date: Thu, 12 Mar 2026 23:12:07 +0000 Subject: [PATCH] =?UTF-8?q?test(e2e):=20workflow=20example=202=20=E2=80=94?= =?UTF-8?q?=20automated=20test=20generation=20for=20a=20module=20(trusted?= =?UTF-8?q?=20profile)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit E2E Robot Framework test for WF02: Automated Test Generation using the trusted automation profile. Exercises the full plan lifecycle (action create, resource add, project create, validation add/attach, plan use/execute/artifacts/diff/apply) against a temporary repository with a low-coverage auth module. Key design decisions: - Registers coverage validation (pytest --cov=src/auth --cov-fail-under=80) and attaches it to the project, per spec WF02 Step 1. - Verifies plan artifacts listing after execute (spec §36971). - Verifies coverage improvement after apply — checks apply output for validation results, falls back to running pytest --cov directly. - All LLM-dependent commands use expected_rc=None for flexible assertions. - Plan ID extraction uses Safe Parse Json Field from common_e2e.resource (eliminates duplicated Extract Plan Id keyword). - Invariant check uses regex matching any src/ path in diff output (not just src/auth.py), and fails hard on violation (not just warn). - Post-apply verification detects both new test files and modifications to existing files; handles zero-artifact LLM responses gracefully. - Commit Fixture Files checks git return codes consistently with common_e2e.resource patterns. - Traceback/INTERNAL checks on all plan execute and apply outputs. - Skip If No LLM Keys guard for graceful CI degradation. - Actor selection based on available API key (Anthropic/OpenAI). - CHANGELOG.md updated per CONTRIBUTING.md requirements. ISSUES CLOSED: #748 --- CHANGELOG.md | 4 + robot/e2e/wf02_test_generation.robot | 542 +++++++++++++++++++++++++++ 2 files changed, 546 insertions(+) create mode 100644 robot/e2e/wf02_test_generation.robot diff --git a/CHANGELOG.md b/CHANGELOG.md index 5d2eb97cf..f788a3402 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,10 @@ ## Unreleased +- Added E2E Robot Framework test for Workflow Example 2: automated test + generation with trusted automation profile, covering coverage validation + registration and attachment, plan artifacts verification, and coverage + improvement checks after apply. (#748) - Added TDD bug-capture tests for bug #1076 — `use_action()` does not propagate `automation_profile` to Plan. Three Behave BDD scenarios (`@tdd_bug @tdd_bug_1076 @tdd_expected_fail`) verify the full precedence diff --git a/robot/e2e/wf02_test_generation.robot b/robot/e2e/wf02_test_generation.robot new file mode 100644 index 000000000..86a62e3f0 --- /dev/null +++ b/robot/e2e/wf02_test_generation.robot @@ -0,0 +1,542 @@ +*** Settings *** +Documentation E2E test for Workflow Example 2: Automated Test Generation +... for a Module (Trusted Profile). +... +... Scenario: A team increases test coverage for an ``auth`` +... module by having CleverAgents autonomously analyse coverage +... gaps and generate comprehensive test files. The trusted +... profile auto-runs strategize and execute; only apply +... requires human approval. +... +... **Zero mocking** — real CLI, real LLM API keys, real +... subprocess execution. +Resource common_e2e.resource +Suite Setup WF02 Suite Setup +Suite Teardown E2E Suite Teardown +Force Tags E2E + +*** Variables *** +${ACTION_YAML} test-gen-action.yaml +${ACTION_NAME} local/test-gen +${PROJECT_NAME} local/test-project +${RESOURCE_NAME} local/test-repo +${TARGET_MODULE} src/auth.py +${VALIDATION_NAME} local/auth-coverage +${VALIDATION_YAML} auth-coverage-validation.yaml +# Note: ACTION_NAME, PROJECT_NAME, RESOURCE_NAME, and VALIDATION_NAME are +# overridden in WF02 Suite Setup with a unique per-run suffix to avoid +# UNIQUE constraint collisions on repeated E2E runs. + +*** Test Cases *** +WF02 Automated Test Generation With Trusted Profile + [Documentation] End-to-end workflow: create a project with a low-coverage + ... auth module, define a test-generation action with + ... invariants, run the plan under the trusted automation + ... profile, and verify that only test files are generated + ... and coverage improves. + [Teardown] Log WF02 test completed (cleanup via suite teardown) + Skip If No LLM Keys + + # ── Step 1: Create temp repo with auth module and minimal tests ── + ${repo_dir}= Create Temp Git Repo wf02-test-repo-${RUN_SUFFIX} + Create Auth Module Fixture ${repo_dir} + Create Minimal Test Fixture ${repo_dir} + Commit Fixture Files ${repo_dir} + + # Record initial test file count and fixture content length before workflow + @{initial_test_files}= List Files In Directory ${repo_dir}${/}tests + ... pattern=test_*.py absolute=${FALSE} + ${initial_count}= Get Length ${initial_test_files} + ${fixture_content}= Build Minimal Test Content + ${fixture_content_length}= Get Length ${fixture_content} + Log Initial test file count: ${initial_count}, fixture content length: ${fixture_content_length} + + # ── Step 2: Create action YAML for test generation ── + ${yaml_path}= Create Test Gen Action YAML + + # ── Step 3: Register the action ── + ${result}= Run CleverAgents Command + ... action create --config ${yaml_path} + Output Should Contain ${result} test-gen + + # ── Step 4: Register resource and create project ── + ${result}= Run CleverAgents Command + ... resource add git-checkout ${RESOURCE_NAME} + ... --path ${repo_dir} + Output Should Contain ${result} ${RESOURCE_NAME} + + ${result}= Run CleverAgents Command + ... project create ${PROJECT_NAME} + ... --resource ${RESOURCE_NAME} + ... --description Auth module test coverage project + Output Should Contain ${result} ${PROJECT_NAME} + + # ── Step 5: Register coverage validation and attach to project ── + ${val_yaml_path}= Create Coverage Validation YAML + ${result}= Run CleverAgents Command + ... validation add --config ${val_yaml_path} + Output Should Contain ${result} ${VALIDATION_NAME} + + ${result}= Run CleverAgents Command + ... validation attach --project ${PROJECT_NAME} + ... ${RESOURCE_NAME} ${VALIDATION_NAME} + Output Should Contain ${result} ${VALIDATION_NAME} + + # ── Step 6: Plan use with trusted profile ── + # TODO: --arg target_module="src/auth" --arg coverage_target=80 are omitted + # because PlanLifecycleService.use_action triggers a UNIQUE constraint violation + # when --arg values duplicate action argument definitions with defaults. + # Re-add --arg flags once that bug is fixed. + ${result}= Run CleverAgents Command + ... plan use ${ACTION_NAME} ${PROJECT_NAME} + ... --automation-profile trusted + ... --format json + ... expected_rc=None timeout=120s + Log Plan use output: ${result.stdout} + IF ${result.rc} != 0 + Fail plan use failed (rc=${result.rc}): ${result.stderr} + END + ${plan_id}= Safe Parse Json Field ${result.stdout} plan_id + Should Not Be Empty ${plan_id} msg=Failed to extract plan ID from plan use output + # Verify the trusted automation profile was applied to the plan + ${profile}= Safe Parse Json Field ${result.stdout} automation_profile + IF '${profile}' != '' + Should Be Equal As Strings ${profile} trusted + ELSE + Log automation_profile field not found in plan use output — profile verification skipped WARN + END + + # ── Step 7: Execute plan (strategize phase — trusted auto-proceeds) ── + ${result}= Run CleverAgents Command + ... plan execute ${plan_id} + ... --format json + ... expected_rc=None timeout=300s + Log Plan execute (strategize) output: ${result.stdout} + Should Not Contain ${result.stdout}${result.stderr} Traceback + Should Not Contain ${result.stdout}${result.stderr} INTERNAL + IF ${result.rc} != 0 + Fail plan execute (strategize) failed (rc=${result.rc}): ${result.stderr} + END + + # ── Step 8: Check plan status between execution phases ── + ${status_result}= Run CleverAgents Command + ... plan status ${plan_id} + ... --format json + ... expected_rc=None timeout=60s + Log Plan status after strategize: ${status_result.stdout} + IF ${status_result.rc} != 0 + Log plan status returned rc=${status_result.rc} after strategize — status may be unavailable WARN + END + ${status_phase}= Safe Parse Json Field ${status_result.stdout} phase + Log Plan phase after strategize: ${status_phase} + + # ── Step 9: Execute plan (execute phase) ── + ${result}= Run CleverAgents Command + ... plan execute ${plan_id} + ... --format json + ... expected_rc=None timeout=300s + Log Plan execute (execute) output: ${result.stdout} + Should Not Contain ${result.stdout}${result.stderr} Traceback + Should Not Contain ${result.stdout}${result.stderr} INTERNAL + IF ${result.rc} != 0 + Fail plan execute (execute) failed (rc=${result.rc}): ${result.stderr} + END + + # ── Step 10: Verify artifacts listing shows generated test files ── + ${artifacts_result}= Run CleverAgents Command + ... plan artifacts ${plan_id} + ... --format json + ... expected_rc=None timeout=60s + Log Plan artifacts output: ${artifacts_result.stdout} + ${artifacts_total}= Set Variable 0 + IF ${artifacts_result.rc} == 0 + ${combined}= Set Variable ${artifacts_result.stdout}\n${artifacts_result.stderr} + ${has_test_file}= Run Keyword And Return Status + ... Should Match Regexp ${combined} (?i)tests?/test_\\w+\\.py + IF not ${has_test_file} + Log Artifacts output does not list test file paths — LLM may have used different naming WARN + END + # Check whether files_changed is empty (regex handles whitespace variations in JSON) + ${is_empty}= Run Keyword And Return Status + ... Should Match Regexp ${combined} "files_changed"\\s*:\\s*\\[\\s*\\] + IF not ${is_empty} + # Non-empty files_changed — count by matching path entries + ${path_matches}= Get Regexp Matches ${combined} "path":\\s*"[^"]+" + ${artifacts_total}= Get Length ${path_matches} + END + ELSE + Log plan artifacts returned rc=${artifacts_result.rc} — artifacts may not be available yet WARN + END + # Warn if artifacts command succeeded with output but no files were parsed + IF ${artifacts_total} == 0 and ${artifacts_result.rc} == 0 + ${output_length}= Get Length ${artifacts_result.stdout.strip()} + IF ${output_length} > 10 + Log plan artifacts returned rc=0 with non-trivial output but parsed 0 files — JSON schema may have changed WARN + END + END + Log Artifacts total files changed: ${artifacts_total} + + # ── Step 11: Verify diff shows only test file changes ── + ${result}= Run CleverAgents Command + ... plan diff ${plan_id} + ... --format plain + ... expected_rc=None timeout=60s + Log Plan diff output: ${result.stdout} + Should Not Contain ${result.stdout}${result.stderr} Traceback + Should Not Contain ${result.stdout}${result.stderr} INTERNAL + IF ${result.rc} != 0 + Log plan diff returned rc=${result.rc} — invariant check may be unreliable WARN + END + Verify No Production Code Changes ${result} + + # ── Step 12: Apply the plan ── + ${result}= Run CleverAgents Command + ... plan lifecycle-apply ${plan_id} + ... expected_rc=None timeout=120s + Log Plan apply output: ${result.stdout} + Should Not Contain ${result.stdout}${result.stderr} Traceback + Should Not Contain ${result.stdout}${result.stderr} INTERNAL + IF ${result.rc} != 0 + Fail lifecycle-apply failed (rc=${result.rc}): ${result.stderr} + END + # Verify plan transitioned to apply/applied state + ${post_apply_status}= Run CleverAgents Command + ... plan status ${plan_id} + ... --format json + ... expected_rc=None timeout=60s + ${apply_phase}= Safe Parse Json Field ${post_apply_status.stdout} phase + ${apply_state}= Safe Parse Json Field ${post_apply_status.stdout} state + Log Plan state after apply: phase=${apply_phase} state=${apply_state} + IF '${apply_phase}' != '' + Should Contain ${apply_phase.lower()} apply + ... msg=Expected plan to transition to apply phase after lifecycle-apply, got: ${apply_phase} + ELSE + Log phase field not found in post-apply plan status — state transition verification skipped WARN + END + + # ── Step 13: Verify new test files exist after apply ── + Verify Test Files Exist ${repo_dir} ${initial_count} ${artifacts_total} ${fixture_content_length} + + # ── Step 14: Verify coverage improvement ── + Verify Coverage Improvement ${repo_dir} ${result} ${artifacts_total} + +*** Keywords *** +WF02 Suite Setup + [Documentation] E2E Suite Setup plus workspace initialisation for WF02. + ... Generates a unique per-run suffix for entity names to avoid + ... UNIQUE constraint collisions on repeated E2E runs. + E2E Suite Setup + ${init}= Run CleverAgents Command init --force --yes + Should Be Equal As Integers ${init.rc} 0 + # Generate a unique suffix for resource/project/action names to avoid UNIQUE + # constraint collisions on repeated E2E runs against the same database. + ${suffix}= Evaluate __import__('uuid').uuid4().hex[:12] + Set Suite Variable ${RUN_SUFFIX} ${suffix} + Set Suite Variable ${ACTION_NAME} local/test-gen-${RUN_SUFFIX} + Set Suite Variable ${PROJECT_NAME} local/test-project-${RUN_SUFFIX} + Set Suite Variable ${RESOURCE_NAME} local/test-repo-${RUN_SUFFIX} + Set Suite Variable ${VALIDATION_NAME} local/auth-coverage-${RUN_SUFFIX} + +Build Auth Module Content + [Documentation] Build the auth module fixture content using Catenate for + ... readability. Returns the multi-line Python source string. + ${content}= Catenate SEPARATOR=\n + ... """Authentication module with multiple code paths.""" + ... ${EMPTY} + ... import hashlib + ... import hmac + ... import time + ... from typing import Optional + ... ${EMPTY} + ... ${EMPTY} + ... class AuthError(Exception): + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Raised on authentication failure.""" + ... ${EMPTY} + ... ${EMPTY} + ... class TokenExpiredError(AuthError): + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Raised when a token has expired.""" + ... ${EMPTY} + ... ${EMPTY} + ... def hash_password(password: str, salt: str = "default_salt") -> str: + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Hash a password with the given salt.""" + ... ${SPACE}${SPACE}${SPACE}${SPACE}return hashlib.sha256(f"{salt}{password}".encode()).hexdigest() + ... ${EMPTY} + ... ${EMPTY} + ... def verify_password(password: str, hashed: str, salt: str = "default_salt") -> bool: + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Verify a password against its hash.""" + ... ${SPACE}${SPACE}${SPACE}${SPACE}return hmac.compare_digest(hash_password(password, salt), hashed) + ... ${EMPTY} + ... ${EMPTY} + ... def create_token(user_id: str, ttl_seconds: int = 3600) -> dict: + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Create an authentication token.""" + ... ${SPACE}${SPACE}${SPACE}${SPACE}now = int(time.time()) + ... ${SPACE}${SPACE}${SPACE}${SPACE}return { + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}"user_id": user_id, + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}"issued_at": now, + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}"expires_at": now + ttl_seconds, + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}"token": hashlib.sha256(f"{user_id}{now}".encode()).hexdigest(), + ... ${SPACE}${SPACE}${SPACE}${SPACE}} + ... ${EMPTY} + ... ${EMPTY} + ... def validate_token(token: dict) -> bool: + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Validate that a token has not expired.""" + ... ${SPACE}${SPACE}${SPACE}${SPACE}if "expires_at" not in token: + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}raise AuthError("Invalid token: missing expires_at") + ... ${SPACE}${SPACE}${SPACE}${SPACE}if int(time.time()) > token["expires_at"]: + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}raise TokenExpiredError("Token has expired") + ... ${SPACE}${SPACE}${SPACE}${SPACE}return True + ... ${EMPTY} + ... ${EMPTY} + ... def authenticate(username: str, password: str, user_db: Optional[dict] = None) -> dict: + ... ${SPACE}${SPACE}${SPACE}${SPACE}"""Authenticate a user and return a token.""" + ... ${SPACE}${SPACE}${SPACE}${SPACE}if user_db is None: + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}user_db = {} + ... ${SPACE}${SPACE}${SPACE}${SPACE}record = user_db.get(username) + ... ${SPACE}${SPACE}${SPACE}${SPACE}if record is None: + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}raise AuthError(f"User not found: {username}") + ... ${SPACE}${SPACE}${SPACE}${SPACE}if not verify_password(password, record["password_hash"], record.get("salt", "default_salt")): + ... ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}raise AuthError("Invalid password") + ... ${SPACE}${SPACE}${SPACE}${SPACE}return create_token(username) + RETURN ${content} + +Build Minimal Test Content + [Documentation] Build the minimal test fixture content using Catenate for + ... readability. Returns the multi-line Python source string. + ${content}= Catenate SEPARATOR=\n + ... """Minimal tests for auth module — low coverage baseline.""" + ... ${EMPTY} + ... from src.auth import hash_password + ... ${EMPTY} + ... ${EMPTY} + ... def test_hash_password_returns_string(): + ... ${SPACE}${SPACE}${SPACE}${SPACE}result = hash_password("secret") + ... ${SPACE}${SPACE}${SPACE}${SPACE}assert isinstance(result, str) + ... ${SPACE}${SPACE}${SPACE}${SPACE}assert len(result) > 0 + RETURN ${content} + +Create Auth Module Fixture + [Documentation] Create ``src/auth.py`` with a realistic auth module. + [Arguments] ${repo_dir} + ${auth_content}= Build Auth Module Content + Create Directory ${repo_dir}${/}src + Create File ${repo_dir}${/}src${/}__init__.py \n + Create File ${repo_dir}${/}src${/}auth.py ${auth_content} + +Create Minimal Test Fixture + [Documentation] Create ``tests/test_auth.py`` with minimal test coverage. + [Arguments] ${repo_dir} + ${test_content}= Build Minimal Test Content + Create Directory ${repo_dir}${/}tests + Create File ${repo_dir}${/}tests${/}__init__.py \n + Create File ${repo_dir}${/}tests${/}test_auth.py ${test_content} + +Commit Fixture Files + [Documentation] Stage and commit all fixture files in the temp repo. + [Arguments] ${repo_dir} + ${add_result}= Run Process git add . + ... cwd=${repo_dir} timeout=60s on_timeout=kill + Should Be Equal As Integers ${add_result.rc} 0 + ... git add failed: ${add_result.stderr} + ${commit_result}= Run Process git commit -m Add auth module and minimal tests + ... cwd=${repo_dir} timeout=60s on_timeout=kill + Should Be Equal As Integers ${commit_result.rc} 0 + ... git commit failed: ${commit_result.stderr} + +Create Coverage Validation YAML + [Documentation] Write a coverage validation YAML config and return its path. + ... Registers a validation that runs ``pytest --cov=src/auth + ... --cov-fail-under=80`` to verify coverage meets the 80% + ... threshold. The validation code includes a subprocess timeout + ... to prevent indefinite hangs. + ${yaml_path}= Set Variable ${SUITE_HOME}${/}${VALIDATION_YAML} + # Build YAML content line-by-line using Catenate to avoid + # Robot Framework bracket/brace interpolation in Python code. + ${code_line1}= Set Variable ${SPACE}${SPACE}import subprocess + ${code_line2}= Set Variable ${SPACE}${SPACE}def run(input_data): + ${code_line3}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}cmd = "pytest --cov=src/auth --cov-fail-under=80 -q".split() + ${code_line4}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}try: + # Note: cwd is not set because the validation engine sets the working directory + # to the resource root (git-checkout path) before invoking the validation code. + ${code_line5}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}result = subprocess.run(cmd, capture_output=True, text=True, timeout=120) + ${code_line6}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}except subprocess.TimeoutExpired: + ${code_line7}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}${SPACE}return dict(passed=False, message="timeout") + ${code_line8}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}passed = result.returncode == 0 + ${code_line9}= Set Variable ${SPACE}${SPACE}${SPACE}${SPACE}return dict(passed=passed, message="ok" if passed else "fail") + ${content}= Catenate SEPARATOR=\n + ... name: ${VALIDATION_NAME} + ... description: "Coverage must be at least 80% for auth module" + ... source: custom + ... code: | + ... ${code_line1} + ... ${code_line2} + ... ${code_line3} + ... ${code_line4} + ... ${code_line5} + ... ${code_line6} + ... ${code_line7} + ... ${code_line8} + ... ${code_line9} + ... validation: + ... ${SPACE}${SPACE}mode: required + ... input_schema: + ... ${SPACE}${SPACE}type: object + ... ${SPACE}${SPACE}properties: \{} + ... read_only: true + ... idempotent: true + ... timeout: 240 + ... resource_slots: + ... ${SPACE}${SPACE}- name: repo + ... ${SPACE}${SPACE}${SPACE}${SPACE}resource_type: git-checkout + ... ${SPACE}${SPACE}${SPACE}${SPACE}access: read_only + ... ${SPACE}${SPACE}${SPACE}${SPACE}binding: contextual + Create File ${yaml_path} ${content} + RETURN ${yaml_path} + +Create Test Gen Action YAML + [Documentation] Write a test-generation action YAML config and return + ... its path. Picks the actor based on the available API key, + ... consistent with ``m6_acceptance.robot`` Suite Setup. + ... Includes ``arguments``, ``state``, and three invariants per + ... the spec WF02 action definition. + ${has_anthropic}= Evaluate bool(__import__('os').environ.get('ANTHROPIC_API_KEY', '')) + IF ${has_anthropic} + ${actor}= Set Variable anthropic/claude-sonnet-4-20250514 + ELSE + ${actor}= Set Variable openai/gpt-4o + END + ${yaml_path}= Set Variable ${SUITE_HOME}${/}${ACTION_YAML} + ${content}= Catenate SEPARATOR=\n + ... name: ${ACTION_NAME} + ... description: "Generate comprehensive tests for a target module" + ... long_description: | + ... ${SPACE}${SPACE}Analyse the target module for coverage gaps and generate + ... ${SPACE}${SPACE}comprehensive unit tests to improve coverage. + ... strategy_actor: ${actor} + ... execution_actor: ${actor} + ... definition_of_done: | + ... ${SPACE}${SPACE}New test files are generated that cover ${TARGET_MODULE}. + ... ${SPACE}${SPACE}No production source files are modified. + ... ${SPACE}${SPACE}All new tests follow pytest naming conventions. + ... reusable: true + ... state: available + ... read_only: false + ... arguments: + ... ${SPACE}${SPACE}- name: target_module + ... ${SPACE}${SPACE}${SPACE}${SPACE}type: string + ... ${SPACE}${SPACE}${SPACE}${SPACE}required: false + ... ${SPACE}${SPACE}${SPACE}${SPACE}description: "Module path to generate tests for" + ... ${SPACE}${SPACE}${SPACE}${SPACE}default: "src/auth" + ... ${SPACE}${SPACE}- name: coverage_target + ... ${SPACE}${SPACE}${SPACE}${SPACE}type: integer + ... ${SPACE}${SPACE}${SPACE}${SPACE}required: false + ... ${SPACE}${SPACE}${SPACE}${SPACE}description: "Target coverage percentage" + ... ${SPACE}${SPACE}${SPACE}${SPACE}default: 80 + ... invariants: + ... ${SPACE}${SPACE}- "No production code changes" + ... ${SPACE}${SPACE}- "Follow pytest naming conventions" + ... ${SPACE}${SPACE}- "Use the project's existing test fixtures and conftest.py patterns" + Create File ${yaml_path} ${content} + RETURN ${yaml_path} + +Verify No Production Code Changes + [Documentation] Check that the plan diff does not include changes to + ... production source files under ``src/``. Only test + ... file paths (``tests/`` or ``test_``) should appear. + ... Fails the test if any ``src/`` path appears in the diff. + ... Requires the diff to be in ``plain`` (unified) format. + [Arguments] ${result} + ${combined}= Set Variable ${result.stdout}\n${result.stderr} + # If there is diff output, check it does not touch any src/ files + ${has_diff}= Run Keyword And Return Status + ... Should Not Be Empty ${combined.strip()} + IF ${has_diff} + ${src_match}= Run Keyword And Return Status + ... Should Not Match Regexp ${combined} (?m)^(?:---|\\+\\+\\+)\\s+[ab]/src/ + IF not ${src_match} + Fail Invariant violation: plan diff modifies production files under src/ + END + END + +Verify Test Files Exist + [Documentation] After apply, verify that test files exist in the + ... repository's ``tests/`` directory and that the workflow + ... produced changes. The LLM may generate entirely new + ... files or modify the existing test file; either counts as + ... a successful generation. When the LLM produced zero + ... artifacts, verification is structural only (tests dir + ... still exists, at least the fixture file remains). + [Arguments] ${repo_dir} ${initial_count} ${artifacts_total} ${fixture_content_length}=0 + ${tests_exists}= Run Keyword And Return Status + ... Directory Should Exist ${repo_dir}${/}tests + IF not ${tests_exists} + Fail tests/ directory not found after apply — workflow did not produce test files + END + @{test_files}= List Files In Directory ${repo_dir}${/}tests + ... pattern=test_*.py absolute=${FALSE} + ${count}= Get Length ${test_files} + Log Found ${count} test file(s) in tests/ directory (initial: ${initial_count}, artifacts: ${artifacts_total}) + # When the LLM produced artifacts (total > 0), verify new or modified test files + ${arts_int}= Convert To Integer ${artifacts_total} + IF ${arts_int} > 0 + # The LLM may create new test files (count > initial) or modify existing + # ones (count == initial with more content). Either outcome proves the + # workflow produced test content. + ${min_expected}= Evaluate ${initial_count} + 1 + ${has_new_files}= Evaluate ${count} >= ${min_expected} + IF ${has_new_files} + Log Workflow generated ${count} - ${initial_count} new test file(s) + ELSE + # No new files — check if existing test_auth.py grew in content + ${file_exists}= Run Keyword And Return Status + ... File Should Exist ${repo_dir}${/}tests${/}test_auth.py + IF not ${file_exists} + Fail LLM reported artifacts but test_auth.py not found — LLM may have renamed the file + END + ${test_content}= Get File ${repo_dir}${/}tests${/}test_auth.py + ${content_length}= Get Length ${test_content} + ${minimal_length}= Convert To Integer ${fixture_content_length} + Log test_auth.py content length: ${content_length} (fixture was ${minimal_length}) + Should Be True ${content_length} > ${minimal_length} + ... msg=Workflow did not generate new test files and did not modify existing test_auth.py (${content_length} <= ${minimal_length} chars) + Log Workflow modified existing test_auth.py (content grew from ${minimal_length} to ${content_length} chars) + END + ELSE + # LLM produced zero artifacts — this is a valid outcome for real LLM + # E2E testing. Verify at least the fixture file survives apply. + Should Be True ${count} >= 1 + ... msg=tests/ directory exists but contains no test files after apply + Pass Execution LLM produced zero artifacts — structural verification passed (${count} fixture file(s) intact) + END + +Verify Coverage Improvement + [Documentation] Verify that coverage improved after apply. When the LLM + ... produced artifacts, checks the apply command output for + ... coverage validation results, then falls back to running + ... ``pytest --cov`` directly. If the fallback also fails, + ... the test fails — the LLM produced changes that do not + ... meet the coverage threshold. When the LLM produced zero + ... artifacts, only structural verification is logged. + [Arguments] ${repo_dir} ${apply_result} ${artifacts_total} + ${arts_int}= Convert To Integer ${artifacts_total} + IF ${arts_int} == 0 + Pass Execution LLM produced zero artifacts — coverage improvement check skipped (no test changes to measure) + END + ${combined}= Set Variable ${apply_result.stdout}\n${apply_result.stderr} + # Check if the apply output contains coverage validation results + ${has_coverage_pass}= Run Keyword And Return Status + ... Should Match Regexp ${combined} (?i)coverage.*\\bpass(?:ed)?\\b + IF ${has_coverage_pass} + Log Coverage validation passed (confirmed from apply output) + RETURN + END + # Fallback: run pytest --cov in the repo to verify coverage + ${cov_result}= Run Process ${PYTHON} -m pytest + ... --cov\=src/auth --cov-fail-under\=80 -q + ... cwd=${repo_dir} timeout=120s on_timeout=kill + Log Coverage check stdout: ${cov_result.stdout} + Log Coverage check stderr: ${cov_result.stderr} + IF ${cov_result.rc} == 0 + Log Coverage verification passed (pytest --cov exit code 0) + ELSE + Fail Coverage verification failed (rc=${cov_result.rc}): LLM produced ${arts_int} artifact(s) but coverage did not reach 80% threshold. stdout=${cov_result.stdout} stderr=${cov_result.stderr} + END -- 2.52.0