d3cb534caf
CI / push-validation (pull_request) Successful in 17s
CI / helm (pull_request) Successful in 19s
CI / lint (pull_request) Successful in 27s
CI / security (pull_request) Successful in 32s
CI / quality (pull_request) Successful in 41s
CI / typecheck (pull_request) Successful in 48s
CI / build (pull_request) Successful in 3m18s
CI / e2e_tests (pull_request) Successful in 4m16s
CI / integration_tests (pull_request) Successful in 4m16s
CI / unit_tests (pull_request) Successful in 5m8s
CI / docker (pull_request) Successful in 1m32s
CI / coverage (pull_request) Successful in 10m53s
CI / status-check (pull_request) Successful in 1s
CI / push-validation (push) Successful in 18s
CI / helm (push) Successful in 23s
CI / build (push) Successful in 31s
CI / lint (push) Successful in 43s
CI / typecheck (push) Successful in 52s
CI / security (push) Successful in 52s
CI / e2e_tests (push) Successful in 3m37s
CI / quality (push) Successful in 3m44s
CI / integration_tests (push) Successful in 6m35s
CI / unit_tests (push) Failing after 7m36s
CI / docker (push) Has been skipped
CI / coverage (push) Successful in 13m39s
CI / status-check (push) Failing after 1s
Implement StrategyActor class for the plan strategize phase that uses an LLM to produce hierarchical execution strategies with dependencies, resource requirements, estimated complexity, and risk scores. Key components: - StrategyActor: Core actor with LLM prompt construction, response parsing (JSON and numbered-list fallback), and graceful degradation to StrategizeStubActor when no LLM provider is configured - StrategyAction/StrategyTree: Pydantic models for the hierarchical action tree with dependency links - validate_no_cycles(): Kahns algorithm (deque-based) for dependency graph cycle detection, raising PlanError on circular dependencies - build_strategy_prompt(): Context-aware prompt construction using definition_of_done, resources, project context, and ACMS analysis with XML-delimited user content sections for prompt injection hardening - parse_strategy_response(): Robust LLM output parsing with JSON extraction and numbered-list fallback - resolve_strategy_actor(): Integration point for the existing actor.default.strategy config key (CLEVERAGENTS_DEFAULT_STRATEGY_ACTOR) - Decision conversion producing strategy_choice Decision objects - build_decisions() preserves tree hierarchy via parent_id mapping, populates downstream_decision_ids from dependency edges, and validates plan_id Structural tree hierarchy (B2 review fix): - _build_tree infers parent_id from the dependency graph: each actions first resolved dependency becomes its structural parent. Actions with no dependencies fall back to the root. This produces hierarchical trees for agents plan tree rendering per spec Plan Decision Tree. Downstream decision tracking (B3 review fix): - build_decisions populates downstream_decision_ids from the strategy trees dependency edges using a pre-generated decision_id map so influence relationships between decisions are recorded per the spec Decision Record Structure. Post code-review hardening (PR #1175): - Broadened exception handling in execute() and ACMS retrieval to catch all LLM provider errors (openai, httpx, anthropic, etc.) with graceful fallback to stub mode (H1, H2) - Added warning log for unresolvable dependency references so dropped edges are visible in structured logs (H3) - Added XML-delimited user content sections and explicit data-only instructions in system prompt for prompt injection hardening (H4) - Switched prompt truncation to word-boundary-safe _truncate_at_word() for all prompt input sections (M1) - Fixed _parse_actor_name to preserve user-specified provider or model when only one segment is empty, instead of discarding both (M2) - Annotated _build_invariant_records as placeholder pending the Invariant Reconciliation Actor implementation (M5) - Documented resources/project_context params as future-wired through PlanExecutor.run_strategize() (M8) - Added docstring noting supersession relationship with LLMStrategizeActor in llm_actors.py (M9) - Added __all__ export definition (L2) - Improved validate_no_cycles docstring edge direction semantics (L7) - Cap JSON parse retry loop at _MAX_JSON_PARSE_RETRIES (10) Post second code-review hardening (PR #1175, review cycle 2): - Fixed _truncate_at_word docstring: documented max_chars >= 3 precondition for the result-length guarantee (R-H1) - Added warning log in build_decisions for unresolvable parent_id references, matching the existing _build_tree warning for unresolvable dependency references (R-H2) - Fixed _parse_actor_name to handle whitespace-only input by adding actor_name.strip() check alongside the emptiness check (R-M1) - Tightened ACMS scenario assertions from non-empty to expected count of 5 decisions (R-L3) - Added timeout=60s on_timeout=kill to all Robot test cases for consistency with project patterns (R-M5) Post third code-review hardening (PR #1175, review cycle 3): - Added _sanitize_xml_content() to escape XML special characters (<, >, &) in user content before embedding into XML-delimited prompt sections, preventing prompt injection via forged closing tags (spec Prompt Injection Mitigation) (CR3-M1) - Upgraded _try_parse_json() to multi-anchor retry: collects all [{ positions left-to-right and tries each as a candidate start, fixing false-start anchoring when LLM preamble contains [{ fragments before the real JSON array (CR3-M2) - Added _truncate_at_word() guard for max_chars < 3: returns a hard slice instead of word-boundary truncation when the ellipsis would exceed the limit (CR3-L2) - Changed _build_tree collision fallback key from -(idx+1) to -(1_000_000+idx) to eliminate theoretical collision with LLM-produced negative step numbers (CR3-L3) - Added forward-looking API docstring note to build_decisions() documenting that it is not yet wired into PlanExecutor and will be integrated once Decision persistence lands (CR3-M3) Post fourth code-review hardening (PR #1175, review cycle 4): - Fixed _try_parse_json per-anchor retry counter: reset retries=0 at the start of each anchor iteration so false-start [{ anchors in LLM preamble text no longer exhaust the retry budget for the correct anchor (CR4-B1) - Added known-limitations docstring to module header documenting missing decision types (resource_selection, subplan_spawn, invariant_enforced) as future work (CR4-D1) - Rewrote XML injection assertion in test to use regex extraction instead of fragile chained .split() calls that could IndexError on structural changes (CR4-T5) Post fifth code-review hardening (PR #1175, review cycle 5): - Added warning log in build_decisions for empty-string parent_id (distinct from None) so the silent fallback to root is visible in structured logs for debuggability (CR5-B1) - Added plan_id propagation assertion to build_decisions test scenarios verifying decision.plan_id matches the input (CR5-T1) - Added sequence_number monotonicity assertion verifying decision sequence_numbers are zero-indexed and monotonically increasing (CR5-T2) - Added _truncate_at_word boundary test for max_chars=3 (exactly ellipsis length) verifying correct "..." output (CR5-T3) - Tightened false-start anchor test from permissive len>=1 to specific description match "Sole real action" (CR5-T4) - Added word-boundary truncation test using space-separated input to exercise the rfind(" ") path under oversized DoD (CR5-T5) Post sixth code-review hardening (PR #1175, review cycle 6): - Added _MAX_INVARIANTS cap (100) for invariant list truncation in prompt to prevent token limit overflows, consistent with other prompt section caps (CR6-M4) - Added negative max_chars guard in _truncate_at_word returning empty string instead of slicing from end (CR6-M5) - Added global JSON parse attempt cap _MAX_GLOBAL_JSON_ATTEMPTS (50) across all anchors in _try_parse_json (CR6-L3) - Moved re import to module level in strategy_parsing.py per CONTRIBUTING import guidelines (CR6-L4) - Extracted _DEFAULT_DESCRIPTION constant to eliminate duplication between _default_action() and _build_tree() (CR6-L5) Post seventh code-review hardening (PR #1175, review cycle 7): - Decoupled _execute_stub from StrategizeStubActor._parse_steps private method by delegating to parse_strategy_response, removing cross-class private method dependency (CR7-M1) - Added ULID format validation on plan_id in execute() and build_decisions() for spec-consistent argument validation per §Plan glossary and CONTRIBUTING §Argument Validation (CR7-M2) - Constrained StrategyAction.estimated_complexity to Literal["low", "medium", "high"] at Pydantic model level per CONTRIBUTING §Type Safety (CR7-M5) - Documented XML-tag prompt boundary deviation from spec [USER_CONTENT_START]/[USER_CONTENT_END] markers with rationale for the more structured approach (CR7-M6) - Added _build_tree empty-input guard comment documenting orphaned root_id semantics (CR7-L1) - Added _truncate_at_word > 0 intent comment explaining why position-0 space is intentionally excluded (CR7-L2) - Added build_decisions context_snapshot future-work comment referencing spec §Decision Record Structure (CR7-L5) - Used enumerate() in _build_tree first loop for idiomatic Python (CR7-L7) - Fixed false-start anchor test (CR5-T4) broken by CR6-L3 global cap: reduced preamble fragments from 15 to 3 so total attempts stay within _MAX_GLOBAL_JSON_ATTEMPTS (CR7-T1) - Fixed test plan_ids containing non-Crockford-Base32 characters (L→K) to pass ULID format validation (CR7-T2) Tests: - 105 Behave BDD scenarios in features/strategy_actor_llm.feature adding: global JSON attempt cap exhaustion (CR7-L3), orphaned dependency edge silent drop (CR7-L4), non-ULID plan_id rejection in execute() and build_decisions() (CR7-M2) - 101 Behave BDD scenarios in features/strategy_actor_llm.feature including new scenarios for _truncate_at_word edge cases (L3), create_llm argument verification (L4), non-numeric step field fallback (L5), updated assertions for XML-delimited prompts and _parse_actor_name partial-segment preservation (M2), lifecycle exception fallback (R1), PydanticValidationError re-raise verification (R2), self-loop cycle detection (R3), whitespace-only actor name (R4), XML tag injection sanitisation (CR3-M1), preamble bracket fragment parsing (CR3-M2), _truncate_at_word sub-3 limit (CR3-L2), resolve_strategy_actor with both llm config and registry (CR3-L5), build_decisions unresolvable parent_id fallback (CR3-L7), XML injection in resources/project_context/acms_context fields (CR4-S1), ampersand escaping (CR4-S1d), false-start anchor retry budget (CR4-T3), non-sequential step edge specificity (CR4-T4), plan_id propagation (CR5-T1), sequence_number monotonicity (CR5-T2), max_chars=3 boundary (CR5-T3), false-start anchor specificity (CR5-T4), word-boundary truncation (CR5-T5), invariant prompt constraints (CR6-M2), invariant XML sanitisation (CR6-M3), invariant truncation cap (CR6-M4), negative max_chars (CR6-M5), and no-space truncation (CR6-L8) - 7 Robot Framework integration tests in robot/strategy_actor.robot - Mock LLM provider in features/mocks/mock_strategy_llm.py All nox stages pass: lint, typecheck, unit_tests (13789 scenarios), integration_tests (1863 passed). integration_tests (1863 passed, 2 pre-existing TDD failures unrelated to this change). ISSUES CLOSED: #828
424 lines
20 KiB
Gherkin
424 lines
20 KiB
Gherkin
Feature: PlanExecutor comprehensive coverage
|
|
As a developer
|
|
I want thorough test coverage of plan_executor.py
|
|
So that all branches and edge cases are exercised
|
|
|
|
# ------------------------------------------------------------------
|
|
# StrategizeStubActor._parse_steps parsing branches
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: Parse steps with dash bullet prefix
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "- First item\n- Second item"
|
|
Then the cov2 parsed steps count should be "2"
|
|
And the cov2 parsed step at index "0" should be "First item"
|
|
And the cov2 parsed step at index "1" should be "Second item"
|
|
|
|
Scenario: Parse steps with asterisk bullet prefix
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "* Alpha\n* Beta"
|
|
Then the cov2 parsed steps count should be "2"
|
|
And the cov2 parsed step at index "0" should be "Alpha"
|
|
And the cov2 parsed step at index "1" should be "Beta"
|
|
|
|
Scenario: Parse steps with bullet character prefix
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "\u2022 Uno\n\u2022 Dos"
|
|
Then the cov2 parsed steps count should be "2"
|
|
And the cov2 parsed step at index "0" should be "Uno"
|
|
And the cov2 parsed step at index "1" should be "Dos"
|
|
|
|
Scenario: Parse steps with dot-numbered items
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "1. Do alpha\n2. Do beta\n10. Do gamma"
|
|
Then the cov2 parsed steps count should be "3"
|
|
And the cov2 parsed step at index "0" should be "Do alpha"
|
|
And the cov2 parsed step at index "1" should be "Do beta"
|
|
And the cov2 parsed step at index "2" should be "Do gamma"
|
|
|
|
Scenario: Parse steps with paren-numbered items
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "1) Task one\n2) Task two"
|
|
Then the cov2 parsed steps count should be "2"
|
|
And the cov2 parsed step at index "0" should be "Task one"
|
|
And the cov2 parsed step at index "1" should be "Task two"
|
|
|
|
Scenario: Parse steps skips blank lines in input
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "Line A\n\n\nLine B"
|
|
Then the cov2 parsed steps count should be "2"
|
|
And the cov2 parsed step at index "0" should be "Line A"
|
|
And the cov2 parsed step at index "1" should be "Line B"
|
|
|
|
Scenario: Parse steps with prefix that leaves empty content returns default
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "- "
|
|
Then the cov2 parsed steps should be the default objective
|
|
|
|
Scenario: Parse steps with empty string returns default
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from an empty string
|
|
Then the cov2 parsed steps should be the default objective
|
|
|
|
Scenario: Parse steps with whitespace-only returns default
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from " \n\t "
|
|
Then the cov2 parsed steps should be the default objective
|
|
|
|
Scenario: Parse steps with plain text lines
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 parse steps from "Just a sentence"
|
|
Then the cov2 parsed steps count should be "1"
|
|
And the cov2 parsed step at index "0" should be "Just a sentence"
|
|
|
|
# ------------------------------------------------------------------
|
|
# StrategizeStubActor.execute - validation and edge cases
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: Strategize stub rejects empty plan_id
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 execute strategize with an empty plan_id
|
|
Then a cov2 ValidationError should be raised with message containing "plan_id"
|
|
|
|
Scenario: Strategize stub with None definition of done produces default decision
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 execute strategize with plan_id "PLAN01" and no definition
|
|
Then the cov2 strategize result should have "1" decisions
|
|
And the cov2 first decision text should be "Complete the plan objectives"
|
|
|
|
Scenario: Strategize stub with multi-line definition produces correct decisions
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 execute strategize with plan_id "PLAN02" and definition "Step A\nStep B\nStep C"
|
|
Then the cov2 strategize result should have "3" decisions
|
|
And the cov2 root decision should have no parent
|
|
And the cov2 non-root decisions should reference root as parent
|
|
|
|
Scenario: Strategize stub with invariants produces invariant records
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 execute strategize with invariants
|
|
Then the cov2 strategize result should have "2" invariant records
|
|
And each cov2 invariant record should have enforced set to true
|
|
|
|
Scenario: Strategize stub emits all stream events
|
|
Given a fresh StrategizeStubActor for cov2
|
|
And a cov2 stream event collector
|
|
When I cov2 execute strategize with stream callback
|
|
Then the cov2 stream events should include "strategize_started"
|
|
And the cov2 stream events should include "strategize_decisions"
|
|
And the cov2 stream events should include "strategize_complete"
|
|
|
|
Scenario: Strategize stub without stream callback does not error
|
|
Given a fresh StrategizeStubActor for cov2
|
|
When I cov2 execute strategize with plan_id "PLAN03" and definition "Do thing" without callback
|
|
Then the cov2 strategize result should have "1" decisions
|
|
|
|
# ------------------------------------------------------------------
|
|
# ExecuteStubActor - tool_runner, sandbox_root, stream, read_only
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: Execute stub rejects empty plan_id
|
|
Given a fresh ExecuteStubActor for cov2
|
|
When I cov2 execute execute-stub with empty plan_id
|
|
Then a cov2 ValidationError should be raised with message containing "plan_id"
|
|
|
|
Scenario: Execute stub with tool_runner counts discovered tools
|
|
Given a fresh ExecuteStubActor for cov2
|
|
And a cov2 mock tool runner that discovers 3 tools
|
|
When I cov2 execute execute-stub with tool_runner and 2 decisions
|
|
Then the cov2 execute result tool_calls_count should be "6"
|
|
|
|
Scenario: Execute stub with sandbox_root populates sandbox_refs
|
|
Given a fresh ExecuteStubActor for cov2
|
|
When I cov2 execute execute-stub with sandbox_root "/tmp/sandbox"
|
|
Then the cov2 execute result sandbox_refs should contain "/tmp/sandbox"
|
|
|
|
Scenario: Execute stub without sandbox_root has empty sandbox_refs
|
|
Given a fresh ExecuteStubActor for cov2
|
|
When I cov2 execute execute-stub without sandbox_root
|
|
Then the cov2 execute result sandbox_refs should be empty
|
|
|
|
Scenario: Execute stub emits all stream events for each decision
|
|
Given a fresh ExecuteStubActor for cov2
|
|
And a cov2 stream event collector
|
|
When I cov2 execute execute-stub with 2 decisions and stream callback
|
|
Then the cov2 stream events should include "execute_started"
|
|
And the cov2 stream events should include "execute_step"
|
|
And the cov2 stream events should include "execute_complete"
|
|
|
|
Scenario: Execute stub with read_only passes to changeset capture
|
|
Given a fresh ExecuteStubActor for cov2
|
|
When I cov2 execute execute-stub with read_only true
|
|
Then the cov2 execute result should have a valid changeset_id
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor.__init__ - validation
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: PlanExecutor rejects None lifecycle_service
|
|
When I cov2 construct PlanExecutor with None lifecycle
|
|
Then a cov2 ValidationError should be raised with message containing "lifecycle_service"
|
|
|
|
Scenario: PlanExecutor initialises with valid lifecycle
|
|
Given a cov2 mock lifecycle service
|
|
When I cov2 construct PlanExecutor with that lifecycle
|
|
Then the cov2 PlanExecutor should be initialised without error
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor properties
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: has_runtime is False without execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 PlanExecutor without execution context
|
|
Then the cov2 executor has_runtime should be "False"
|
|
|
|
Scenario: has_runtime is True with execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 PlanExecutor with execution context
|
|
Then the cov2 executor has_runtime should be "True"
|
|
|
|
Scenario: changeset_store is None without execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 PlanExecutor without execution context
|
|
Then the cov2 executor changeset_store should be None
|
|
|
|
Scenario: changeset_store returns store from execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context with changeset store
|
|
And a cov2 PlanExecutor with that execution context
|
|
Then the cov2 executor changeset_store should be the mock store
|
|
|
|
Scenario: execution_context property returns None without context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 PlanExecutor without execution context
|
|
Then the cov2 executor execution_context should be None
|
|
|
|
Scenario: execution_context property returns context when set
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 PlanExecutor with execution context
|
|
Then the cov2 executor execution_context should not be None
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor.run_strategize - guards and happy path
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: run_strategize rejects empty plan_id
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call run_strategize with empty plan_id
|
|
Then a cov2 ValidationError should be raised with message containing "plan_id"
|
|
|
|
Scenario: run_strategize rejects plan not in Strategize phase
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute phase
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call run_strategize
|
|
Then a cov2 PlanError should be raised containing "not in Strategize phase"
|
|
|
|
Scenario: run_strategize happy path succeeds
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Strategize phase with definition "Build feature\nAdd tests"
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call run_strategize
|
|
Then the cov2 strategize run result should be a StrategizeResult
|
|
And the cov2 lifecycle should have called start_strategize
|
|
And the cov2 lifecycle should have called complete_strategize
|
|
And the cov2 lifecycle should have called _commit_plan
|
|
|
|
Scenario: run_strategize sets decision_root_id on execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 plan in Strategize phase with definition "Do work"
|
|
And a cov2 PlanExecutor with execution context
|
|
When I cov2 call run_strategize
|
|
Then the cov2 execution context decision_root_id should be set
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor.run_strategize - exception path
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: run_strategize calls fail_strategize on exception
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Strategize phase with definition "Do work"
|
|
And a cov2 PlanExecutor with failing strategize actor
|
|
When I cov2 call run_strategize expecting exception
|
|
Then a cov2 RuntimeError should have been raised
|
|
And the cov2 lifecycle should have called fail_strategize
|
|
And the cov2 plan error_details should contain exception_type
|
|
|
|
Scenario: run_strategize records error via error_recovery service
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Strategize phase with definition "Do work"
|
|
And a cov2 mock error recovery service
|
|
And a cov2 PlanExecutor with error recovery and failing strategize actor
|
|
When I cov2 call run_strategize expecting exception
|
|
Then the cov2 error recovery should have recorded a strategize error
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor._guard_execute - all branches
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: guard_execute rejects plan not in Execute phase
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Strategize phase with definition "Do work"
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call guard_execute
|
|
Then a cov2 PlanError should be raised containing "not in Execute phase"
|
|
|
|
Scenario: guard_execute rejects plan not in queued state
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute phase but Processing state
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call guard_execute
|
|
Then a cov2 PlanError should be raised containing "not queued"
|
|
|
|
Scenario: guard_execute rejects plan with no decision_root_id
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state without decision root
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call guard_execute
|
|
Then a cov2 PlanError should be raised containing "no decision tree"
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor.run_execute - routing
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: run_execute rejects empty plan_id
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call run_execute with empty plan_id
|
|
Then a cov2 ValidationError should be raised with message containing "plan_id"
|
|
|
|
Scenario: run_execute routes to stub when no execution context
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call run_execute
|
|
Then the cov2 execute run result should be an ExecuteResult
|
|
And the cov2 lifecycle should have called complete_execute
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor._run_execute_with_runtime
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: run_execute routes to runtime when execution context is set
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with execution context and tool runner
|
|
When I cov2 call run_execute
|
|
Then the cov2 execute run result should be a RuntimeExecuteResult
|
|
And the cov2 lifecycle should have called complete_execute
|
|
|
|
Scenario: run_execute runtime mode raises PlanError when no ToolRunner
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with execution context but no tool runner
|
|
When I cov2 call run_execute expecting exception
|
|
Then a cov2 PlanError should be raised containing "ToolRunner"
|
|
|
|
Scenario: run_execute runtime mode calls fail_execute on exception
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 mock execution context
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with execution context and tool runner and failing runtime
|
|
When I cov2 call run_execute expecting exception
|
|
Then a cov2 RuntimeError should have been raised
|
|
And the cov2 lifecycle should have called fail_execute
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor._run_execute_with_stub - error recovery / retry
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: run_execute stub calls fail_execute on exception without recovery
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with failing execute actor and no recovery
|
|
When I cov2 call run_execute expecting exception
|
|
Then a cov2 RuntimeError should have been raised
|
|
And the cov2 lifecycle should have called fail_execute
|
|
|
|
Scenario: run_execute stub retries once with error recovery then succeeds
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with execute actor that fails once then succeeds
|
|
When I cov2 call run_execute
|
|
Then the cov2 execute run result should be an ExecuteResult
|
|
And the cov2 error recovery should have recorded an execute error
|
|
|
|
Scenario: run_execute stub exhausts retries with error recovery and fails
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with always-failing execute actor and exhausted retries
|
|
When I cov2 call run_execute expecting exception
|
|
Then a cov2 RuntimeError should have been raised
|
|
And the cov2 lifecycle should have called fail_execute
|
|
And the cov2 error recovery should have recorded an execute error
|
|
|
|
Scenario: run_execute stub with error recovery records error and no retry
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor with failing execute actor and recovery denying retry
|
|
When I cov2 call run_execute expecting exception
|
|
Then a cov2 RuntimeError should have been raised
|
|
And the cov2 error recovery should have recorded an execute error
|
|
And the cov2 lifecycle should have called fail_execute
|
|
|
|
# ------------------------------------------------------------------
|
|
# PlanExecutor._build_decisions
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: _build_decisions creates decisions from plan definition_of_done
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Execute-Queued state with decision root and definition
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call build_decisions
|
|
Then the cov2 built decisions should have "2" entries
|
|
And the cov2 first built decision should use the plan decision_root_id
|
|
And the cov2 second built decision should have root as parent
|
|
|
|
# ------------------------------------------------------------------
|
|
# Result model edge cases
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: StrategyDecision model validates sequence is non-negative
|
|
When I cov2 create a StrategyDecision with negative sequence
|
|
Then a cov2 validation error should be raised for the model
|
|
|
|
Scenario: ExecuteResult defaults are correct
|
|
When I cov2 create an ExecuteResult with minimal fields
|
|
Then the cov2 ExecuteResult tool_calls_count should default to "0"
|
|
And the cov2 ExecuteResult sandbox_refs should default to empty list
|
|
|
|
# ------------------------------------------------------------------
|
|
# _build_decisions: stored strategy_decisions_json path (M6 wiring)
|
|
# ------------------------------------------------------------------
|
|
|
|
Scenario: _build_decisions uses stored strategy_decisions_json when present
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan with stored strategy_decisions_json in error_details
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call build_decisions
|
|
Then the cov2 built decisions should match the stored JSON decisions
|
|
|
|
Scenario: _build_decisions falls back to definition_of_done on corrupt JSON
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan with corrupt strategy_decisions_json in error_details
|
|
And a cov2 PlanExecutor without execution context
|
|
When I cov2 call build_decisions
|
|
Then the cov2 built decisions should have "2" entries
|
|
|
|
Scenario: run_strategize passes resources from project_links to actor
|
|
Given a cov2 mock lifecycle service
|
|
And a cov2 plan in Strategize-Queued state with project links
|
|
And a cov2 PlanExecutor with a capturing strategize actor
|
|
When I cov2 run strategize
|
|
Then the cov2 strategize actor received resources from project links
|
|
|
|
Scenario: StrategizeStubActor.execute accepts extra keyword arguments
|
|
When I cov2 call StrategizeStubActor execute with extra kwargs
|
|
Then the cov2 stub result should have decisions
|