Files
temp/features/steps/cli_lifecycle_robot_alignment_steps.py
T
hurui200320 414abb1396 fix(cli): add missing --yes flag to plan apply command (#1127)
## Summary

Adds the `--yes`/`-y` flag to the `lifecycle-apply` CLI command as required by the specification (`agents plan apply [--yes|-y] <PLAN_ID>`). Without `--yes`, a confirmation prompt now displays before proceeding with the destructive Apply phase. With `--yes`, the apply proceeds immediately without prompting.

Closes #932

## Changes

### Source Code
- **`src/cleveragents/cli/commands/plan.py`**: Added `yes: Annotated[bool, typer.Option("--yes", "-y", help="Skip confirmation prompt")] = False` parameter to `lifecycle_apply_plan`. Added `typer.confirm()` prompt before the apply operation, consistent with the pattern used by `rollback_plan`, `correct_plan`, and other destructive commands.
  - Confirmation prompt text matches spec exactly: `"Apply changes for plan {plan_id}?"` producing `Apply changes for plan <ID>? [y/N]:`.
  - Fixed redundant plan ID display when `pre_plan` is `None` — now shows `"Apply changes for plan X?"` instead of `"Apply plan X (X)?"`.
  - Added `except ValueError` handler consistent with sibling commands `lifecycle_execute_plan` and `_lifecycle_apply_with_id`.
  - Added `except Exception` catch-all handler with `isinstance(e, (typer.Abort, typer.Exit))` re-raise guard, consistent with `lifecycle_execute_plan`.
  - Moved `PlanPhase` and `ProcessingState` imports to module level per CONTRIBUTING.md §Import Guidelines.

### TDD Tag Removal (Bug Fix Workflow)
- **`features/tdd_plan_apply_yes_flag.feature`**: Removed `@tdd_expected_fail` tag (leaving `@tdd_bug` and `@tdd_bug_932` as permanent regression guards).
- **`robot/tdd_plan_apply_yes_flag.robot`**: Removed `tdd_expected_fail` tag (leaving `tdd_bug` and `tdd_bug_932`).

### Test Updates
Updated all existing `lifecycle-apply` invocations across 17 test/benchmark files to pass `--yes`, since the new confirmation prompt would otherwise abort in non-interactive test environments:
- 9 Behave step definition files
- 3 Robot Framework helper scripts
- 2 Robot Framework e2e acceptance tests
- 3 ASV benchmark files (4 invocations: `cli_robot_flow_bench.py` ×2, `m1_sourcecode_smoke_bench.py` ×1, `plan_cli_smoke_bench.py` ×1)

### Confirmation Prompt Tests (New + Strengthened)
- **`features/tdd_plan_apply_yes_flag.feature`**: 5 scenarios total:
  - `lifecycle-apply recognises the --yes long flag` — verifies flag acceptance, prompt suppression, exit code 0, and `apply_plan` was called
  - `lifecycle-apply recognises the -y short flag` — same as above for short flag
  - `lifecycle-apply without --yes prompts for confirmation and user declines` — verifies `"Apply cancelled."` message, `exit_code == 0`, and `apply_plan` was NOT called
  - `lifecycle-apply without --yes prompts for confirmation and user accepts` — verifies prompt appears, `exit_code == 0`, and `apply_plan` was called
  - `lifecycle-apply catches unexpected exceptions cleanly` — verifies `"Unexpected error"` output, no traceback leak, non-zero exit code (exercises the `except Exception` catch-all)
- **`features/steps/tdd_plan_apply_yes_flag_steps.py`**: Refactored step definitions:
  - `_make_mock_plan` uses `PlanPhase` and `ProcessingState` enum types instead of raw strings
  - `_make_mock_plan` uses `datetime.now(tz=UTC)` instead of timezone-naive `datetime.now()`
  - Unified prompt suppression step handles both `--yes` and `-y` via parameterised step pattern
  - Added `When` step for unexpected error scenario with `RuntimeError` side_effect
  - Added `Then` step for non-zero exit code assertion
- **Feature/Robot documentation**: Updated stale descriptions that said "implementation does not accept --yes" to reflect the flag is now implemented.

### Documentation
- **`docs/reference/plan_cli.md`**: Updated `lifecycle-apply` section with:
  - `### Synopsis` heading with code block
  - `### Options` table listing `--yes/-y` and `--format/-f` flags
  - `### Arguments` table listing `PLAN_ID`
  - Matches the style used by other command sections in the same file

## Review Fixes (Cycle 3 — Luis's review)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| M1 | Medium | `typer.Abort()` on user decline produces exit code 1 and redundant "Aborted." | Changed to `raise typer.Exit(0)` — consistent with `correct_decision` and legacy `apply` |
| M2 | Medium | Missing exit code assertion on decline scenario | Added `And the lifecycle-apply exit code should be 0` to the decline scenario |
| M3 | Medium | Spec compliance: "summary of pending changes" not implemented | Deferred — spec example shows summary *after* confirmation, not before; implementation matches spec. Ticket-vs-spec ambiguity noted. |
| L1 | Low | Missing `except Exception` catch-all handler | Added catch-all matching `lifecycle_execute_plan` pattern; re-raises `typer.Abort`/`typer.Exit` |
| L2 | Low | Documentation description not updated | Expanded description in `plan_cli.md` to explain confirmation prompt and `--yes` |
| L3 | Low | Dead code `is not None` guards | Removed both guards — `get_plan()` raises `NotFoundError`, never returns `None` |
| I1 | Info | Duplicate `PlanPhase` import | Hoisted import to top of `try` block, eliminating duplicate at old line 2087 |
| I2 | Info | `typer.confirm` without explicit `default=False` | Added `default=False` for consistency with sibling commands |
| L4 | Low | No test for `--yes` after positional arg | Not addressed — Typer/Click handles both orderings; low risk |
| L5 | Low | No test for auto-select + interactive prompt | Not addressed — separate concern outside ticket scope |
| I3 | Info | Robot helper only tests flag recognition | By design — noted as informational |

## Review Fixes (Cycle 4 — Self-QA)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| Major-1 | Major | No test for `except Exception` catch-all handler | Added new scenario `"lifecycle-apply catches unexpected exceptions cleanly"` with `RuntimeError` side_effect; asserts `"Unexpected error"` output, no traceback, non-zero exit |
| Minor-2 | Minor | Missing `ValueError` handler inconsistent with siblings | Added `except ValueError as e:` with `"[red]Execution Error:[/red]"` before catch-all, matching `lifecycle_execute_plan` and `_lifecycle_apply_with_id` |
| Minor-3 | Minor | Flag scenarios don't verify `apply_plan` called | Added `And the lifecycle-apply should have called apply` to both `--yes` and `-y` scenarios |
| Minor-4 | Minor | Stale docstring in Robot helper references `tdd_expected_fail` inversion | Updated to reflect bug is fixed and tests serve as regression guards |
| Minor-5 | Minor | `plan_cli.md` lacks Options table for `lifecycle-apply` | Added Synopsis, Options, and Arguments sections matching sibling command style |
| Nit-6 | Nit | Duplicated step defs for `--yes` vs `-y` prompt suppression | Unified into single parameterised step `"the lifecycle-apply {flag} output should not contain the confirmation prompt"` |
| Nit-7 | Nit | `datetime.now()` timezone-naive | Changed to `datetime.now(tz=UTC)` |
| Nit-8 | Nit | `_make_mock_plan` params use `str` instead of enum types | Changed to `PlanPhase` and `ProcessingState` enum types |

## Review Fixes (Cycle 5 — Jeff's approval note)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| Import-1 | Minor | `PlanPhase`/`ProcessingState` imports inside function body instead of module level | Moved to module-level import per CONTRIBUTING.md §Import Guidelines |

## Known Limitations / Deferred Items

- **M3: Ticket AC mentions "summary of pending changes"** but the spec example only shows `"Apply changes for plan <ID>? [y/N]: y"` without a change summary. The implementation shows plan ID only (matching the spec), not a change summary. This is a ticket-vs-spec ambiguity; recommend discussing with ticket author.
- **Legacy `apply` command** accepts `--yes` but does not pass it to `_lifecycle_apply_with_id()`. This is a pre-existing issue outside the scope of this ticket.
- **`pre_plan is None` branch** has no explicit test. Pre-existing architectural issue; no action taken.

## Quality Gates

| Gate | Result |
|------|--------|
| `nox -s lint` |  passed |
| `nox -s typecheck` |  passed (0 errors) |
| `nox -s unit_tests` |  passed (471 features, 12,424 scenarios, 0 failures) |
| `nox -s integration_tests` |  passed (1,727 tests, 0 failures) |
| `nox -s e2e_tests` |  passed (41 tests, 0 failures) |
| `nox -s coverage_report` |  passed (≥97% coverage) |

Reviewed-on: cleveragents/cleveragents-core#1127
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Rui Hu <rui.hu@cleverthis.com>
Co-committed-by: Rui Hu <rui.hu@cleverthis.com>
2026-03-26 07:50:09 +00:00

327 lines
12 KiB
Python

"""Step definitions for CLI lifecycle Robot alignment feature.
Mirrors the Robot E2E lifecycle flow (action -> plan -> execute -> apply)
in Behave to keep unit and integration test expectations aligned.
All step names are prefixed with ``robot alignment`` to avoid
``AmbiguousStep`` conflicts with existing steps.
"""
from __future__ import annotations
import os
import tempfile
from datetime import datetime
from unittest.mock import MagicMock, patch
from behave import given, then, when
from behave.runner import Context
from typer.testing import CliRunner
from cleveragents.cli.commands.action import app as action_app
from cleveragents.cli.commands.plan import app as plan_app
from cleveragents.core.exceptions import NotFoundError
from cleveragents.domain.models.core.action import Action, ActionState
from cleveragents.domain.models.core.change import (
ChangeEntry,
ChangeOperation,
InMemoryChangeSetStore,
)
from cleveragents.domain.models.core.plan import (
AutomationProfileProvenance,
AutomationProfileRef,
NamespacedName,
Plan,
PlanIdentity,
PlanPhase,
PlanTimestamps,
ProcessingState,
ProjectLink,
)
_PLAN_ULID = "01KHDE6WWS2171PWW3GJEBXZ8T"
_VALID_YAML = """\
name: local/lifecycle-action
description: Lifecycle test action
strategy_actor: openai/gpt-4
execution_actor: openai/gpt-4
definition_of_done: All lifecycle tests pass
"""
def _mock_action(name: str = "local/lifecycle-action") -> Action:
"""Create a minimal valid Action for alignment tests."""
return Action(
namespaced_name=NamespacedName.parse(name),
description="Lifecycle test action",
long_description=None,
definition_of_done="All lifecycle tests pass",
strategy_actor="openai/gpt-4",
execution_actor="openai/gpt-4",
state=ActionState.AVAILABLE,
reusable=True,
read_only=False,
created_at=datetime.now(),
updated_at=datetime.now(),
created_by=None,
)
def _mock_plan(
phase: PlanPhase = PlanPhase.STRATEGIZE,
state: ProcessingState = ProcessingState.QUEUED,
) -> Plan:
"""Create a minimal valid Plan for alignment tests."""
now = datetime.now()
return Plan(
identity=PlanIdentity(plan_id=_PLAN_ULID),
namespaced_name=NamespacedName.parse("local/lifecycle-plan"),
description="Lifecycle test plan",
definition_of_done="Tests pass",
action_name="local/lifecycle-action",
phase=phase,
processing_state=state,
project_links=[ProjectLink(project_name="proj-a")],
arguments={"target_coverage": 80},
arguments_order=["target_coverage"],
automation_profile=AutomationProfileRef(
profile_name="trusted",
provenance=AutomationProfileProvenance.PLAN,
),
strategy_actor="openai/gpt-4",
execution_actor="openai/gpt-4",
reusable=True,
read_only=False,
created_by=None,
timestamps=PlanTimestamps(created_at=now, updated_at=now),
)
# ---------------------------------------------------------------------------
# Background
# ---------------------------------------------------------------------------
@given("a robot alignment CLI runner")
def step_robot_alignment_runner(context: Context) -> None:
"""Set up the CLI runner."""
context.ra_runner = CliRunner()
@given("a robot alignment mocked lifecycle service")
def step_robot_alignment_service(context: Context) -> None:
"""Set up a mock PlanLifecycleService for the plan CLI."""
context.ra_mock_service = MagicMock()
context.ra_action_patcher = patch(
"cleveragents.cli.commands.action._get_lifecycle_service",
return_value=context.ra_mock_service,
)
context.ra_plan_patcher = patch(
"cleveragents.cli.commands.plan._get_lifecycle_service",
return_value=context.ra_mock_service,
)
context.ra_executor_patcher = patch(
"cleveragents.cli.commands.plan._get_plan_executor",
return_value=MagicMock(),
)
context.ra_action_patcher.start()
context.ra_plan_patcher.start()
context.ra_executor_patcher.start()
# Set up changeset store for tracking
context.ra_changeset_store = InMemoryChangeSetStore()
context.ra_changeset_id = context.ra_changeset_store.start(_PLAN_ULID)
if not hasattr(context, "_cleanup_handlers"):
context._cleanup_handlers = []
context._cleanup_handlers.append(context.ra_action_patcher.stop)
context._cleanup_handlers.append(context.ra_plan_patcher.stop)
context._cleanup_handlers.append(context.ra_executor_patcher.stop)
# ---------------------------------------------------------------------------
# Given steps
# ---------------------------------------------------------------------------
@given('a robot alignment action "{name}" is created via config')
def step_robot_alignment_action_create(context: Context, name: str) -> None:
"""Create action via CLI config and configure mocks."""
action = _mock_action(name)
context.ra_mock_service.create_action.return_value = action
context.ra_mock_service.get_action_by_name.return_value = action
context.ra_mock_service.use_action.return_value = _mock_plan()
context.ra_mock_service.get_plan.return_value = _mock_plan(
phase=PlanPhase.STRATEGIZE, state=ProcessingState.COMPLETE
)
context.ra_mock_service.execute_plan.return_value = _mock_plan(
phase=PlanPhase.EXECUTE, state=ProcessingState.QUEUED
)
context.ra_mock_service.apply_plan.return_value = _mock_plan(
phase=PlanPhase.APPLY, state=ProcessingState.QUEUED
)
fd, yaml_path = tempfile.mkstemp(suffix=".yaml")
with os.fdopen(fd, "w") as fh:
fh.write(_VALID_YAML)
context.ra_yaml_path = yaml_path
result = context.ra_runner.invoke(action_app, ["create", "--config", yaml_path])
context.ra_action_result = result
os.unlink(yaml_path)
assert result.exit_code == 0, (
f"action create failed ({result.exit_code}): {result.output}"
)
@given('a robot alignment action "{name}" exists')
def step_robot_alignment_action_exists(context: Context, name: str) -> None:
"""Set up an existing action mock."""
context.ra_mock_service.get_action_by_name.return_value = _mock_action(name)
context.ra_mock_service.use_action.return_value = _mock_plan()
# ---------------------------------------------------------------------------
# When steps
# ---------------------------------------------------------------------------
@when('I run robot alignment plan use "{action}" on project "{project}"')
def step_robot_alignment_plan_use(context: Context, action: str, project: str) -> None:
"""Run plan use."""
context.ra_use_result = context.ra_runner.invoke(plan_app, ["use", action, project])
@when("I run robot alignment plan execute for the plan")
def step_robot_alignment_plan_execute(context: Context) -> None:
"""Run plan execute and record a changeset entry."""
# Record a change entry to mirror the Robot test
context.ra_changeset_store.record(
context.ra_changeset_id,
ChangeEntry(
plan_id=_PLAN_ULID,
resource_id="res-a",
tool_name="builtin/file-write",
operation=ChangeOperation.CREATE,
path="src/new_module.py",
),
)
context.ra_execute_result = context.ra_runner.invoke(
plan_app, ["execute", _PLAN_ULID]
)
@when("I run robot alignment plan apply for the plan")
def step_robot_alignment_plan_apply(context: Context) -> None:
"""Run plan apply and record an additional changeset entry."""
# Record a modify entry to bring total to 2
context.ra_changeset_store.record(
context.ra_changeset_id,
ChangeEntry(
plan_id=_PLAN_ULID,
resource_id="res-a",
tool_name="builtin/file-edit",
operation=ChangeOperation.MODIFY,
path="src/existing.py",
),
)
context.ra_apply_result = context.ra_runner.invoke(
plan_app, ["lifecycle-apply", "--yes", _PLAN_ULID]
)
@when('I run robot alignment plan use with invalid arg "{arg_str}"')
def step_robot_alignment_invalid_arg(context: Context, arg_str: str) -> None:
"""Run plan use with an invalid --arg format."""
context.ra_invalid_arg_result = context.ra_runner.invoke(
plan_app,
["use", "local/lifecycle-action", "proj-a", "--arg", arg_str],
)
@when('I run robot alignment plan use for nonexistent action "{action}"')
def step_robot_alignment_missing_action(context: Context, action: str) -> None:
"""Run plan use for an action that doesn't exist."""
context.ra_mock_service.get_action_by_name.side_effect = NotFoundError(
f"Action '{action}' not found"
)
context.ra_missing_result = context.ra_runner.invoke(
plan_app, ["use", action, "proj-a"]
)
# ---------------------------------------------------------------------------
# Then steps
# ---------------------------------------------------------------------------
@then("the robot alignment plan use should succeed")
def step_robot_alignment_plan_use_ok(context: Context) -> None:
"""Verify plan use succeeded."""
assert context.ra_use_result.exit_code == 0, (
f"plan use failed ({context.ra_use_result.exit_code}): "
f"{context.ra_use_result.output}"
)
@then("the robot alignment plan execute should succeed")
def step_robot_alignment_plan_execute_ok(context: Context) -> None:
"""Verify plan execute succeeded."""
assert context.ra_execute_result.exit_code == 0, (
f"plan execute failed ({context.ra_execute_result.exit_code}): "
f"{context.ra_execute_result.output}"
)
@then('the robot alignment changeset should have {count:d} entry with operation "{op}"')
def step_robot_alignment_changeset_entry(context: Context, count: int, op: str) -> None:
"""Verify changeset entries."""
cs = context.ra_changeset_store.get(context.ra_changeset_id)
assert cs is not None, "changeset not found"
assert len(cs.entries) == count, f"expected {count} entries, got {len(cs.entries)}"
assert cs.entries[0].operation == op, (
f"expected operation={op}, got {cs.entries[0].operation}"
)
@then("the robot alignment plan apply should succeed")
def step_robot_alignment_plan_apply_ok(context: Context) -> None:
"""Verify plan apply succeeded."""
assert context.ra_apply_result.exit_code == 0, (
f"plan apply failed ({context.ra_apply_result.exit_code}): "
f"{context.ra_apply_result.output}"
)
@then("the robot alignment changeset summary should have total {total:d}")
def step_robot_alignment_changeset_total(context: Context, total: int) -> None:
"""Verify changeset summary total."""
summary = context.ra_changeset_store.summarize(context.ra_changeset_id)
assert summary.get("total") == total, (
f"expected total={total}, got {summary.get('total')}"
)
@then("the robot alignment CLI should reject the invalid arg format")
def step_robot_alignment_invalid_arg_rejected(context: Context) -> None:
"""Verify the CLI rejected the invalid arg format."""
result = context.ra_invalid_arg_result
output_lower = result.output.lower()
assert (
result.exit_code != 0
or "invalid argument format" in output_lower
or "aborted" in output_lower
), f"expected rejection, got exit_code={result.exit_code}: {result.output}"
@then("the robot alignment CLI should report action not found")
def step_robot_alignment_action_not_found(context: Context) -> None:
"""Verify the CLI reported action not found."""
result = context.ra_missing_result
output_lower = result.output.lower()
assert (
result.exit_code != 0 or "not found" in output_lower or "error" in output_lower
), f"expected not-found error, got exit_code={result.exit_code}: {result.output}"