From 03d2df26ce77719f21290dfd0048988474772517 Mon Sep 17 00:00:00 2001 From: CleverThis Date: Sat, 6 Jun 2026 18:07:09 -0400 Subject: [PATCH] fix(tests): align CI tests with A2A boundary refactor MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The shared `format_data` serializer introduced for the CLI→Application A2A boundary returns raw payloads without the `{"data": ...}` envelope that the legacy CLI `format_output` wraps around. Two test-step definitions (`step_artifacts_json_validation`, `step_artifacts_json_apply_summary`) still unwrapped that envelope and crashed with `KeyError: 'data'`, errrring the Behave scenarios `Plan artifacts shows validation results when available` and `Artifacts include apply summary from metadata`. Also remove the stale `@tdd_expected_fail` tag from the Robot scenario `WF02 Mocked Generation Produces Test Artifacts Only`: the scenario exercises the `_cleveragents/plan/artifacts` A2A dispatch path that this PR added and now passes naturally; the `tdd_expected_fail_listener` inverts the passing result to a failure with "Bug appears to be fixed. Remove the tdd_expected_fail tag". Adds a CHANGELOG entry covering both the boundary refactor and these test alignments. Refs: #9962 Refs: #4253 --- CHANGELOG.md | 1 + features/steps/plan_diff_artifacts_steps.py | 6 ++---- robot/wf02_test_generation_integration.robot | 2 +- 3 files changed, 4 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c78105dfa..f1f295c80 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ Changed `wf10_batch.robot` to be less likely to create files, and ## [Unreleased] - **feat(resources): resource type extension interface** (#9998): New `cleveragents.resources` package providing the stable public API third-party developers use to add custom resource types without modifying core code. Includes `ResourceType` ABC with five abstract lifecycle methods (`provision`, `deprovision`, `status`, `validate_config`, `to_dict`), a `ResourceConfig` Pydantic model (`name`, `resource_type`, `properties`), a `ResourceStatus` StrEnum (`PENDING`, `ACTIVE`, `FAILED`, `DEPROVISIONED`), and registry functions `register_resource_type` / `get_resource_type` / `list_resource_types`. Custom types are registered under namespaced names (e.g. `myorg/database`); registration raises `TypeError` for non-`ResourceType` subclasses and `ValueError` for duplicate names. 25 BDD scenarios in `features/resource_type_extension_interface.feature` cover enum values, config instantiation, ABC enforcement, all lifecycle method return types, and registry CRUD + error paths. +- **refactor(a2a): route CLI→Application communication through A2A boundary** (Refs #9962, #4253): Introduced `cleveragents.shared.output_format` as a layer-neutral serialiser (`format_data` supporting `json`/`yaml`/`plain`/`table`) with no dependency on `cleveragents.cli.*`, eliminating a reverse dependency from `PlanApplyService.artifacts()` on the CLI presentation layer. The shared formatter returns raw payloads with no CLI envelope wrapping (`{"data": ..., "command": ..., "status": ...}`); callers that previously parsed `parsed["data"]` from `apply_service.artifacts(fmt="json")` output now read fields at the top level. Updated `features/steps/plan_diff_artifacts_steps.py` (`step_artifacts_json_validation`, `step_artifacts_json_apply_summary`) to drop the stale envelope unwrap that caused `KeyError: 'data'` under the new boundary. Removed stale `@tdd_expected_fail` tag from `WF02 Mocked Generation Produces Test Artifacts Only` in `robot/wf02_test_generation_integration.robot` — the scenario now passes naturally through the A2A facade dispatch path (`_cleveragents/plan/artifacts`) introduced by this refactor. - **fix(test): move advanced context strategy test doubles to features/mocks** (#7574): Extracted `FakeEmbeddings`, `RelevanceScoringStrategy`, `AdaptiveContextSelector`, `ContextFusionStrategy`, and `_pack_budget` from `features/steps/advanced_context_strategies_steps.py` into a new `features/mocks/advanced_context_strategies_mocks.py` file per CONTRIBUTING.md mock-placement rules. Updated the Robot Framework helper `robot/helper_advanced_context_strategies.py` to import directly from `features.mocks` rather than manipulating `sys.path` to reach the Behave steps file. Added `None` guard in `step_assemble_context_query` before calling `selected.assemble()`, and added explicit `ValueError` for unknown strategy types in both `step_load_yaml_strategy` and `load_strategy_from_yaml_impl`. - **fix(a2a): regression tests for stale cleveragents.acp removal** (#5566): Added two Behave BDD scenarios verifying that `cleveragents.acp` is not importable (raises `ImportError`) and that `src/cleveragents/acp/` does not exist in the source tree. These guard against regression of the `__pycache__`-based import that allowed the removed ACP module to still be loaded from bytecode after the v3.6.0 rename to `a2a`. - **Virtual Resource Type Base Class** (#8610): Implemented `VirtualResource` base class with two example concrete implementations (`MetricResource`, `APIEndpointResource`) for abstract/computed resources that are derived rather than mapped to physical files. Virtual resources are computed on demand via a `compute_fn` callable. Includes Behave BDD scenarios in `features/resource_virtual_types.feature` exercising construction, computation, name validation, kwargs passthrough, exception handling, string representation, and subclassing. Resource names are validated against `^[a-zA-Z][a-zA-Z0-9_-]*$` (must start with a letter; alphanumeric, hyphens, and underscores otherwise). diff --git a/features/steps/plan_diff_artifacts_steps.py b/features/steps/plan_diff_artifacts_steps.py index 049ce5a46..e34bf739b 100644 --- a/features/steps/plan_diff_artifacts_steps.py +++ b/features/steps/plan_diff_artifacts_steps.py @@ -434,8 +434,7 @@ def step_artifacts_has_files(context: Context) -> None: @then("the artifacts JSON should contain validation summary") def step_artifacts_json_validation(context: Context) -> None: - parsed = json.loads(context.artifacts_output) - data = parsed["data"] + data = json.loads(context.artifacts_output) assert "validation_summary" in data assert data["validation_summary"]["total"] == 3 @@ -565,8 +564,7 @@ def step_plan_with_changeset_id_but_no_store_entry(context: Context) -> None: @then("the artifacts JSON should contain apply summary") def step_artifacts_json_apply_summary(context: Context) -> None: - parsed = json.loads(context.artifacts_output) - data = parsed["data"] + data = json.loads(context.artifacts_output) assert "apply_summary" in data assert data["apply_summary"]["files_changed"] == "7" assert data["apply_summary"]["validations_run"] == "4" diff --git a/robot/wf02_test_generation_integration.robot b/robot/wf02_test_generation_integration.robot index eb9bea4d9..4f2eaca08 100644 --- a/robot/wf02_test_generation_integration.robot +++ b/robot/wf02_test_generation_integration.robot @@ -28,7 +28,7 @@ WF02 Mocked Generation Produces Test Artifacts Only [Documentation] Verify mocked LLM responses generate deterministic test ... files, assert user-facing plan artifacts command path, ... and artifacts list only test paths (no prod edits). - [Tags] tdd_issue tdd_issue_4296 tdd_expected_fail mocking artifacts invariants + [Tags] tdd_issue tdd_issue_4296 mocking artifacts invariants ${result}= Run Process ${PYTHON} ${HELPER} mocked-generation-artifacts cwd=${WORKSPACE} timeout=120s on_timeout=kill Log ${result.stdout} Log ${result.stderr}