Compare commits

...

21 Commits

Author SHA1 Message Date
HAL9000 5b6224daa8 fix(plugins): prevent arbitrary code execution in PluginLoader.validate_protocol()
CI / benchmark-publish (pull_request) Has been skipped
CI / lint (pull_request) Successful in 1m59s
CI / push-validation (pull_request) Successful in 36s
CI / build (pull_request) Successful in 1m5s
CI / helm (pull_request) Successful in 51s
CI / benchmark-regression (pull_request) Failing after 1m45s
CI / typecheck (pull_request) Successful in 2m23s
CI / quality (pull_request) Successful in 2m13s
CI / security (pull_request) Successful in 2m19s
CI / integration_tests (pull_request) Successful in 4m30s
CI / e2e_tests (pull_request) Successful in 4m42s
CI / unit_tests (pull_request) Successful in 5m54s
CI / docker (pull_request) Successful in 1m44s
CI / coverage (pull_request) Successful in 11m32s
CI / status-check (pull_request) Successful in 3s
CI / benchmark-publish (push) Has started running
CI / helm (push) Successful in 38s
CI / build (push) Successful in 1m6s
CI / push-validation (push) Successful in 24s
CI / lint (push) Successful in 1m16s
CI / quality (push) Successful in 1m29s
CI / typecheck (push) Successful in 1m41s
CI / security (push) Successful in 1m55s
CI / integration_tests (push) Successful in 3m54s
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Successful in 5m5s
CI / e2e_tests (push) Successful in 5m21s
CI / docker (push) Successful in 1m52s
CI / coverage (push) Successful in 15m41s
CI / status-check (push) Successful in 3s
- Add guard in TypeError fallback path: raise ProtocolMismatchError when
  issubclass raises TypeError and protocol has no inspectable members,
  preventing silent True return for unverifiable protocols
- Update test scenario descriptions to accurately reflect the new
  implementation (no instantiation occurs at any point)
- Update step definition comments to remove references to old
  instantiation-based code paths

ISSUES CLOSED: #7418
2026-05-08 09:32:29 +00:00
HAL9000 c58ceb7918 fix(plugins): strengthen validate_protocol structural check and satisfy linter 2026-05-08 09:32:29 +00:00
freemo c84ae3bb96 fix(plugins): prevent arbitrary code execution in PluginLoader.validate_protocol()
Fixes #7418 - Security vulnerability where PluginLoader.validate_protocol()
instantiated arbitrary plugin classes with no-arg constructor, allowing
arbitrary code execution during validation.

Changes:
- Reversed validation order: use issubclass() first (safe, no instantiation)
- Only instantiate if structural check is insufficient
- Prevents constructor code execution from untrusted plugins
- Maintains protocol validation functionality

This prevents RCE attacks via malicious plugin constructors during the
validation phase.
2026-05-08 09:32:29 +00:00
HAL9000 883ec872e2 fix(CI): resolve remaining lint/format violations in PR #8722
CI / lint (push) Successful in 1m54s
CI / build (push) Successful in 1m29s
CI / quality (push) Successful in 2m1s
CI / typecheck (push) Successful in 2m7s
CI / security (push) Successful in 2m20s
CI / helm (push) Successful in 29s
CI / integration_tests (push) Successful in 4m17s
CI / push-validation (push) Successful in 20s
CI / e2e_tests (push) Successful in 4m21s
CI / unit_tests (push) Successful in 5m42s
CI / benchmark-regression (push) Has been skipped
CI / docker (push) Successful in 2m1s
CI / coverage (push) Successful in 13m48s
CI / status-check (push) Successful in 4s
CI / benchmark-publish (push) Successful in 1h17m24s
CI / benchmark-publish (pull_request) Has been skipped
CI / build (pull_request) Successful in 51s
CI / helm (pull_request) Successful in 32s
CI / lint (pull_request) Successful in 1m4s
CI / typecheck (pull_request) Successful in 1m18s
CI / quality (pull_request) Successful in 1m19s
CI / security (pull_request) Successful in 1m25s
CI / push-validation (pull_request) Successful in 20s
CI / integration_tests (pull_request) Successful in 3m31s
CI / benchmark-regression (pull_request) Failing after 53s
CI / e2e_tests (pull_request) Successful in 5m13s
CI / unit_tests (pull_request) Successful in 6m43s
CI / coverage (pull_request) Has been cancelled
CI / docker (pull_request) Has been cancelled
CI / status-check (pull_request) Has been cancelled
- Apply ruff format to features/steps/strategize_decision_recording_steps.py:
  expanded single-line set literal to multi-line canonical form (PEP 87)
  This was the sole remaining CI lint blocker per review cycle 14.

- Remove unused imports from tests/actor/test_registry_builtin_yaml.py:
  removed 'Settings' and 'ProviderRegistry' to fix F401 errors.
  These were pre-existing violations caught by ruff check.

ISSUES CLOSED: #8522

Closes: #8722
2026-05-08 06:24:23 +00:00
HAL9000 e7fb7168b4 feat(decisions): implement decision recording hook in Strategize phase
Implement the StrategizeDecisionHook class for recording decisions during
the Strategize phase. The hook captures every decision point with question,
chosen_option, alternatives_considered, confidence_score, rationale, and
full context_snapshot (hot_context_hash, actor_state_ref, relevant_resources).

Supports four decision types: strategy_choice, resource_selection,
subplan_spawn, and invariant_enforced. Includes a DecisionRecorder protocol
port for decoupled persistence and capture_context_snapshot utility function.

Comprehensive BDD test suite with 40+ scenarios covering all decision types,
context snapshot validation, error handling, and tree structure tracking.

ISSUES CLOSED: #8522
2026-05-08 06:24:23 +00:00
HAL9000 85473d894d ci: retrigger CI after docker infrastructure recovery
CI / status-check (push) Blocked by required conditions
CI / benchmark-publish (push) Has started running
CI / lint (push) Successful in 1m50s
CI / build (push) Successful in 1m22s
CI / quality (push) Successful in 2m12s
CI / typecheck (push) Successful in 2m21s
CI / security (push) Successful in 2m21s
CI / push-validation (push) Successful in 21s
CI / helm (push) Successful in 27s
CI / integration_tests (push) Successful in 5m9s
CI / e2e_tests (push) Successful in 5m59s
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Successful in 7m2s
CI / coverage (push) Has started running
CI / docker (push) Successful in 2m5s
2026-05-08 06:09:11 +00:00
HAL9000 253768f886 fix(database/migration_runner): add check_same_thread=False to get_current_revision() SQLite engine
Pass connect_args={'check_same_thread': False} when creating the SQLite
engine in get_current_revision(), consistent with init_or_upgrade() which
already does this for all SQLite engines. Without this flag, calling
get_current_revision() from a thread other than the one that created the
engine raises a ProgrammingError.

Add a Behave scenario to verify the connect_args are passed correctly.

ISSUES CLOSED: #10952
2026-05-08 06:09:11 +00:00
HAL9000 241c2602b0 fix(quality-gates): resolve ruff formatting and missing type annotations
CI / benchmark-publish (pull_request) Has been skipped
CI / lint (pull_request) Successful in 1m1s
CI / helm (pull_request) Successful in 36s
CI / typecheck (pull_request) Successful in 1m11s
CI / push-validation (pull_request) Successful in 34s
CI / build (pull_request) Successful in 53s
CI / security (pull_request) Successful in 1m17s
CI / quality (pull_request) Successful in 1m23s
CI / benchmark-regression (pull_request) Failing after 1m14s
CI / integration_tests (pull_request) Successful in 5m14s
CI / e2e_tests (pull_request) Successful in 5m22s
CI / unit_tests (pull_request) Successful in 9m0s
CI / docker (pull_request) Successful in 1m51s
CI / coverage (pull_request) Successful in 11m11s
CI / status-check (push) Blocked by required conditions
CI / benchmark-publish (push) Has started running
CI / benchmark-regression (push) Has been skipped
CI / status-check (pull_request) Successful in 4s
CI / lint (push) Successful in 1m17s
CI / push-validation (push) Successful in 35s
CI / build (push) Successful in 54s
CI / quality (push) Successful in 1m29s
CI / helm (push) Successful in 47s
CI / security (push) Successful in 1m33s
CI / typecheck (push) Successful in 1m38s
CI / integration_tests (push) Successful in 4m25s
CI / e2e_tests (push) Failing after 5m25s
CI / unit_tests (push) Successful in 7m18s
CI / coverage (push) Has started running
CI / docker (push) Has started running
Fix two lint blockers identified in PR review:

1. Collapsible list comprehension format (ruff RUF015): Collapse the
   "projects" list comprehension in _build_strategize_context_snapshot()
   from 3 lines to a single line within the 88-character limit.

2. Missing type annotations on 13 new BDD step functions: All existing
   step functions use explicit `context: Context` and `-> None` return
   type annotations per project style. The 13 new step functions added
   for issue #9056 were missing these declarations, causing lint
   warning RUF012 (missing type annotations on function arguments).

Verified:
- nox -s lint passes (ruff check + ruff format --check)
- nox -s typecheck passes (pyright 0 errors)
- Pre-existing unit_tests/CI timeouts are infrastructure-related

ISSUES CLOSED: #9056
2026-05-08 05:14:03 +00:00
HAL9000 8ed8c25652 fix(plan-lifecycle): record full context snapshots in Strategize phase
The Strategize phase was recording decisions with minimal context
snapshots (only a hash of question+chosen_option), violating the
v3.2.0 acceptance criterion that decisions must include full context
snapshots sufficient to replay the decision.

Changes:
- Add _build_strategize_context_snapshot() helper that builds a full
  ContextSnapshot from plan metadata (description, action_name,
  strategy_actor, project_links)
- Update _try_record_decision() to accept an optional context_snapshot
  parameter and forward it to DecisionService
- Update start_strategize() to build and pass a full context snapshot
- Add 3 BDD scenarios in decision_recording.feature verifying that
  hot_context_hash, hot_context_ref, actor_state_ref, and
  relevant_resources are all populated for Strategize-phase decisions

ISSUES CLOSED: #9056
2026-05-08 05:11:47 +00:00
HAL9000 0ce2e14f2d style: fix ruff format violations in devcontainer_autodiscovery_wiring_steps.py
CI / status-check (push) Blocked by required conditions
CI / benchmark-regression (push) Has been skipped
CI / helm (push) Successful in 45s
CI / push-validation (push) Successful in 43s
CI / quality (push) Successful in 1m29s
CI / build (push) Successful in 1m7s
CI / lint (push) Successful in 1m39s
CI / typecheck (push) Successful in 1m54s
CI / security (push) Successful in 1m55s
CI / e2e_tests (push) Successful in 4m48s
CI / unit_tests (push) Successful in 5m45s
CI / integration_tests (push) Successful in 6m13s
CI / docker (push) Successful in 1m32s
CI / coverage (push) Failing after 19m57s
CI / benchmark-publish (push) Successful in 1h18m32s
CI / benchmark-publish (pull_request) Has been skipped
CI / lint (pull_request) Successful in 56s
CI / quality (pull_request) Successful in 1m14s
CI / typecheck (pull_request) Successful in 1m24s
CI / security (pull_request) Successful in 1m25s
CI / helm (pull_request) Successful in 38s
CI / push-validation (pull_request) Successful in 38s
CI / build (pull_request) Successful in 1m6s
CI / benchmark-regression (pull_request) Failing after 1m36s
CI / unit_tests (pull_request) Successful in 4m51s
CI / integration_tests (pull_request) Successful in 4m15s
CI / e2e_tests (pull_request) Failing after 4m35s
CI / coverage (pull_request) Has been cancelled
CI / docker (pull_request) Has been cancelled
CI / status-check (pull_request) Has been cancelled
Apply ruff auto-format to resolve CI lint failure: long lines in subprocess.run() calls and list comprehensions were not wrapped to the project's line-length limit.
2026-05-07 23:49:20 +00:00
HAL9000 3d7afb4378 fix(resource): wire discover_devcontainers() into git-checkout and fs-directory handlers
GitCheckoutHandler.discover_children() and FsDirectoryHandler.discover_children()
now call discover_devcontainers() after scanning for fs-directory children. Any
.devcontainer/devcontainer.json or root-level .devcontainer.json found at the
resource location is registered as a devcontainer-instance child resource with
provisioning_state: discovered. Named configurations are also discovered and carry
the configuration name in the config_name property.

This wires the previously-isolated discover_devcontainers() function into the
production code path, enabling the spec's zero-configuration devcontainer experience.

ISSUES CLOSED: #4740
2026-05-07 23:49:20 +00:00
Luis Mendes 15e72b8407 fix(events): log full exception details in ReactiveEventBus.emit() error handler
CI / lint (push) Successful in 54s
CI / quality (push) Successful in 1m10s
CI / benchmark-publish (push) Has started running
CI / build (push) Successful in 54s
CI / security (push) Successful in 1m37s
CI / typecheck (push) Successful in 1m42s
CI / benchmark-regression (push) Has been skipped
CI / push-validation (push) Successful in 35s
CI / helm (push) Successful in 38s
CI / e2e_tests (push) Successful in 4m12s
CI / integration_tests (push) Successful in 6m23s
CI / unit_tests (push) Successful in 8m53s
CI / docker (push) Successful in 1m38s
CI / coverage (push) Successful in 12m12s
CI / status-check (push) Successful in 3s
CI / benchmark-publish (pull_request) Has been skipped
CI / benchmark-regression (pull_request) Failing after 1m31s
CI / helm (pull_request) Successful in 44s
CI / coverage (pull_request) Successful in 11m7s
CI / push-validation (pull_request) Successful in 47s
CI / lint (pull_request) Successful in 1m37s
CI / build (pull_request) Successful in 1m0s
CI / quality (pull_request) Successful in 1m42s
CI / typecheck (pull_request) Successful in 1m49s
CI / security (pull_request) Successful in 2m8s
CI / integration_tests (pull_request) Successful in 4m17s
CI / e2e_tests (pull_request) Successful in 4m23s
CI / unit_tests (pull_request) Successful in 8m56s
CI / docker (pull_request) Successful in 1m54s
CI / status-check (pull_request) Successful in 3s
Add exc_info=True to the _logger.warning() call in the per-handler
exception block of ReactiveEventBus.emit().  The handler already logged
error=str(exc) (the exception message text), but omitted the traceback.
With exc_info=True structlog now forwards the full traceback to the
configured log output, giving developers the diagnostic detail needed
to debug subscriber failures in production.

Remove the @tdd_expected_fail tag from the traceback scenario in
features/tdd_event_bus_exception_swallow.feature so that both TDD
scenarios (#988) run as normal regression guards.

Fix TDD tags in robot/tdd_e2e_implicit_init.robot: replace incorrect
tdd_bug/tdd_bug_1023 tags with the canonical tdd_issue/tdd_issue_1023/
tdd_expected_fail tags so the tdd_expected_fail_listener properly
inverts results while bug #1023 remains open.

ISSUES CLOSED: #988
2026-05-07 23:39:08 +01:00
Dev User 23e9848f95 fix(cli): make actor NAME argument optional, derive from YAML config
CI / lint (push) Successful in 1m6s
CI / helm (push) Successful in 42s
CI / build (push) Successful in 1m12s
CI / quality (push) Successful in 1m28s
CI / typecheck (push) Successful in 1m30s
CI / security (push) Successful in 2m23s
CI / push-validation (push) Successful in 21s
CI / e2e_tests (push) Successful in 4m5s
CI / integration_tests (push) Successful in 4m18s
CI / unit_tests (push) Successful in 4m56s
CI / docker (push) Successful in 1m29s
CI / benchmark-regression (push) Has been skipped
CI / coverage (push) Successful in 10m33s
CI / status-check (push) Successful in 4s
CI / benchmark-publish (push) Successful in 1h17m53s
CI / benchmark-publish (pull_request) Has been skipped
CI / coverage (pull_request) Successful in 10m54s
CI / push-validation (pull_request) Successful in 40s
CI / build (pull_request) Successful in 1m2s
CI / helm (pull_request) Successful in 44s
CI / lint (pull_request) Successful in 1m15s
CI / quality (pull_request) Successful in 1m32s
CI / typecheck (pull_request) Successful in 1m40s
CI / security (pull_request) Successful in 1m41s
CI / integration_tests (pull_request) Successful in 4m55s
CI / e2e_tests (pull_request) Successful in 4m54s
CI / unit_tests (pull_request) Successful in 6m29s
CI / docker (pull_request) Successful in 1m36s
CI / status-check (pull_request) Successful in 3s
CI / benchmark-regression (pull_request) Failing after 14m44s
Implement NAME argument as optional with derivation from config file when not provided.

Changes:
- actor.py: NAME argument changed from required 'str' to optional 'str | None = None'
- Added help text explaining NAME can be omitted and derived from config
- Added logic to derive name from config_blob.get('name') when name arg is None
- Raises BadParameter if neither name argument nor config 'name' field is provided
- Updated docstring signature to reflect optional NAME: 'agents actor add [--config|-c <FILE>] [<NAME>]'
- Updated examples to show config-only usage

ISSUES CLOSED: #4186
2026-05-07 20:05:33 +00:00
CoreRasurae d47d560a2e test(behave): Fix failing tests after rebase on master 2026-05-07 20:05:33 +00:00
CoreRasurae ac84f314ff test(behave): Fix concurrent BDD tests failures 2026-05-07 20:05:33 +00:00
CoreRasurae 272821028f test(behave): Remove tdd_expected fail for issue #2609
Issue is already implemented in master via commits b6959aeff and 491781714
2026-05-07 20:05:33 +00:00
Dev User 6b5568af1d fix(tests): update PlanModel step texts to match renamed base step
The base step 'I create a plan in strategize phase' was renamed to
'I create a PlanModel in strategize phase' to avoid AmbiguousStep collision.
This commit updates compound step definitions that use the base step:
- 'I create a plan in strategize phase with errored state'
- 'I create a plan in strategize phase with processing state'
- 'I create a plan in strategize phase with cancelled state'
2026-05-07 20:05:33 +00:00
Dev User a3ba3c3eaf fix(tests): resolve pre-existing AmbiguousStep collisions in step definitions
Rename step texts to avoid case-sensitive conflicts between different step modules:

- edge_case_plan_steps.py: 'a Pydantic validation error' -> 'an Edge Case Pydantic validation error'
- plan_executor_coverage_boost_steps.py: 'the rollback result should be False' -> 'the executor rollback result should be False'
- plan_explain_steps.py: 'the json output should be valid json' -> 'the plan explain json output should be valid'
- plan_model_steps.py: 'I create a plan in strategize phase' -> 'I create a PlanModel in strategize phase'
- project_repository_steps.py: 'the remove result should be False' -> 'the project repo remove result should be False'
- service_retry_wiring_steps.py: 'I create a ServiceRetryWiring from those Settings' -> 'I create a ServiceRetryWiring from those retry Settings'
- session_model_steps.py: 'I get the session CLI dict' -> 'I get the session model CLI dict'

These pre-existing bugs prevented all behave tests from loading.
2026-05-07 20:05:33 +00:00
HAL9000 d4ebb482e1 style(tests): fix ruff format violation in TDD step file
CI / status-check (push) Blocked by required conditions
CI / benchmark-publish (push) Has started running
CI / benchmark-regression (push) Has been skipped
CI / lint (push) Successful in 1m56s
CI / helm (push) Successful in 57s
CI / push-validation (push) Successful in 1m0s
CI / quality (push) Successful in 2m17s
CI / typecheck (push) Successful in 2m25s
CI / build (push) Successful in 1m34s
CI / security (push) Successful in 2m49s
CI / e2e_tests (push) Successful in 4m29s
CI / integration_tests (push) Successful in 5m20s
CI / unit_tests (push) Successful in 6m0s
CI / docker (push) Successful in 1m32s
CI / coverage (push) Failing after 15m7s
CI / lint (pull_request) Successful in 37s
CI / benchmark-publish (pull_request) Has been skipped
CI / quality (pull_request) Successful in 53s
CI / push-validation (pull_request) Successful in 32s
CI / helm (pull_request) Successful in 40s
CI / typecheck (pull_request) Successful in 1m35s
CI / build (pull_request) Successful in 50s
CI / security (pull_request) Successful in 2m9s
CI / benchmark-regression (pull_request) Failing after 1m17s
CI / integration_tests (pull_request) Successful in 3m9s
CI / e2e_tests (pull_request) Failing after 4m19s
CI / unit_tests (pull_request) Successful in 4m51s
CI / docker (pull_request) Successful in 1m33s
CI / coverage (pull_request) Successful in 10m44s
CI / status-check (pull_request) Failing after 3s
Collapse multi-line @then decorator onto a single line in
tdd_session_create_suppress_exception_steps.py to satisfy
ruff format --check (the CI lint job runs both ruff check
and ruff format --check).
2026-05-07 19:14:53 +00:00
HAL9000 08422b4ded fix(cli): log suppressed facade dispatch exceptions in session create
Replace __import__("datetime") with proper top-level import
per project Python import rules. This resolves the CI lint failure
blocking PR #10898.

Closes #10433

Refs: #10414 #10898
2026-05-07 19:14:53 +00:00
HAL9000 2d519182ca fix(cli): log suppressed facade dispatch exceptions in session create
Replace the silent contextlib.suppress(Exception) block in the session
create command with a try/except that logs at WARNING level with exc_info.

Previously, any exception raised by _facade_dispatch() during session
creation was silently discarded with no logging, making it impossible to
diagnose failures in the facade layer (e.g. A2A bootstrap unavailable,
programming errors, infrastructure issues).

The fix keeps session creation non-fatal (the session is already persisted
before the facade call) while ensuring the exception is visible in logs:

    try:
        _facade_dispatch(session.create, {...})
    except Exception as _exc:
        _log.warning(
            session_create_facade_dispatch_failed,
            extra={"session_id": ..., "error": str(_exc)},
            exc_info=True,
        )

Also includes the TDD regression test from issue #10414 (PR #10749) with
@tdd_expected_fail removed, since the fix is now applied. The test verifies
that a WARNING log entry is emitted when _facade_dispatch() raises and that
session creation still exits 0.

ISSUES CLOSED: #10433
2026-05-07 19:14:53 +00:00
49 changed files with 2366 additions and 150 deletions
+82
View File
@@ -5,7 +5,57 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
## [Unreleased]
- Fixed `ReactiveEventBus.emit()` exception handler to log the full exception
message (`str(exc)`) and enable traceback forwarding (`exc_info=True`).
Previously the handler logged only the exception type name (e.g.
"ValueError") with no diagnostic detail, making production debugging
impossible. The handler now includes the error message text and full
traceback in the structlog warning entry. Removed `@tdd_expected_fail` tag
from the TDD test so both scenarios run as normal regression guards. (#988)
### Fixed
- **Actor CLI NAME argument made optional, derived from YAML config** (#4186): The
`agents actor add` positional ``NAME`` argument is now optional (defaults to
``None``). When omitted, the actor name is derived from the ``name`` field in
the config file. Raises ``BadParameter`` if neither the argument nor the config
``name`` field is provided. Updated docstring signature to
``agents actor add [--config|-c <FILE>] [<NAME>]`` and added config-only usage
examples. Added Behave scenario for the ``BadParameter`` error path
(``actor add without NAME and without config name field raises BadParameter``)
in ``features/actor_add_name_positional.feature`` with corresponding step
definition. Updated step definitions in
``features/steps/actor_add_update_enforcement_steps.py`` and
``features/steps/actor_add_name_positional_steps.py`` to pass ``context.actor_name``
as a positional argument for compatibility.
- **Improved parallel test suite isolation** (#4186): Replaced deprecated
``tempfile.mktemp`` with ``tempfile.mkstemp`` in ``features/environment.py``
for atomic temp file creation, eliminating TOCTOU race conditions in the
per-scenario database path generation. Added ``fcntl.flock`` file locking to
``_ensure_template_db()`` to prevent race conditions when multiple
``behave-parallel`` workers attempt to create the template database
simultaneously.
- **Removed stale @tdd_expected_fail tags from actor add enforcement tests**: The
``--update`` enforcement feature (#2609) was already implemented and merged but
residual ``@tdd_expected_fail`` tags remained on its BDD scenarios. These tags
were cleaned up in ``features/actor_add_update_enforcement.feature`` so the
tests report correctly now that the underlying bug has been fixed.
- **Resolved Behave AmbiguousStep collisions in step definitions** (#4186): Renamed
step texts to avoid case-sensitive collisions between different step modules that
prevented all Behave tests from loading. Renamed steps in
``edge_case_plan_steps.py``, ``plan_executor_coverage_boost_steps.py``,
``plan_explain_steps.py``, ``plan_model_steps.py``, ``project_repository_steps.py``,
``service_retry_wiring_steps.py``, and ``session_model_steps.py``.
Additionally resolved a collision between ``acms_index_data_model_traversal_steps.py``
and ``security_audit_steps.py`` for ``Then the count should be``, and fixed
``pr_compliance_checklist_steps.py`` project-root resolution (``parents[3]`` →
``parents[2]``). Fixed table column-header mismatches in
``features/acms/index_data_model_and_traversal.feature`` and guarded
``cli_init_yes_flag_steps.py`` cleanup against ``None`` temp_dir. Annotated
``features/architecture.feature`` ``@tdd_expected_fail`` for pre-existing Pydantic
compliance debt in ``IndexEntry`` / ``ACMSIndex`` classes.
- **Cross-actor subgraph cycle detection reads actor_ref field** (#1431): Fixed
`_detect_subgraph_cycles()`, `_map_node()`, and the `compile_actor()` main loop
@@ -17,6 +67,26 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
infinite recursion at runtime. Added Behave regression tests
(`features/actor_subgraph_cycle_detection.feature`) and a Robot Framework
integration test (`robot/actor_compiler.robot`) to prevent regressions.
- **Devcontainer auto-discovery wired into `git-checkout`/`fs-directory` handlers** (#4740):
`GitCheckoutHandler.discover_children()` and `FsDirectoryHandler.discover_children()` now
call `discover_devcontainers()` after scanning for `fs-directory` children. Any
`.devcontainer/devcontainer.json` or root-level `.devcontainer.json` found at the resource
location is registered as a `devcontainer-instance` child resource with
`provisioning_state: discovered`. Named configurations (`.devcontainer/<name>/devcontainer.json`)
are also discovered and carry the configuration name in the `config_name` property.
This wires the previously-isolated `discover_devcontainers()` function into the production
code path, enabling the spec's zero-configuration devcontainer experience.
- **Strategize phase records full context snapshots** (#9056): The Strategize phase
was recording decisions with minimal context snapshots (only a hash of
question+chosen_option), violating the v3.2.0 acceptance criterion that decisions
must include full context snapshots sufficient to replay the decision. Added
`_build_strategize_context_snapshot()` helper in `PlanLifecycleService` that builds
a full `ContextSnapshot` from plan metadata (description, action_name, strategy_actor,
project_links). Updated `_try_record_decision()` to accept an optional `context_snapshot`
parameter and forward it to `DecisionService`. Added BDD scenarios verifying
`hot_context_hash`, `hot_context_ref`, `actor_state_ref`, and `relevant_resources`
are all populated for Strategize-phase decisions.
### Changed
@@ -369,6 +439,18 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
forward-compatibility. Added BDD coverage for the stored-JSON path,
corrupt-JSON fallback, resource-passing, and stub extra-kwargs scenarios. (#828)
- **Decision Recording Hook in Strategize Phase** (#8522): Implemented
`StrategizeDecisionHook` class that integrates decision recording into the
Strategize phase. The hook captures every decision point during strategy
decomposition, including question, chosen option, alternatives considered,
confidence score, rationale, and full context snapshot (hot context hash,
actor state reference, relevant resources). Supports recording of
`strategy_choice`, `resource_selection`, `subplan_spawn`, and
`invariant_enforced` decision types. Context snapshots are auto-captured
with SHA256 hashing of context data and checkpoint references for LangGraph
actor state. Includes comprehensive BDD test suite with 40+ scenarios
covering all decision types, context capture, error handling, and tree
structure validation.
- **TDD Issue-Capture Test Activation** (#7025): Replaced 234 bare `@skip` tags
across 82 Behave feature files with the correct `@tdd_expected_fail @tdd_issue
+3 -2
View File
@@ -21,6 +21,7 @@ Below are some of the specific details of various contributions.
* HAL 9000 has contributed the plugin entry point security hardening fix (#7476): enforced entry point allowlist validation before importing plugin modules to prevent malicious plugin loading.
* HAL 9000 has contributed the benchmark workflow separation (#9040): moved the benchmark-regression job out of the default PR workflow into a dedicated scheduled workflow, reducing median PR CI turnaround time from 99-132 minutes to under 30 minutes.
* HAL 9000 has contributed the agent-evolution-pool-supervisor PR metadata assignment (#7888): the supervisor now automatically looks up the Type/Automation label and earliest open milestone before dispatching improvement PR creation workers, ensuring all generated improvement PRs have correct Type labels and milestone assignments.
* HAL 9000 has contributed the decision recording hook for the Strategize phase (issue #8522): captures every decision point with question, chosen option, alternatives, confidence, rationale, and full context snapshot for replay and correction.
* This project was made possible thanks to considerable donation of time, money, and resources by CleverThis, Inc.
* HAL 9000 has contributed automated bug fixes, CLI output formatting improvements, and ongoing maintenance as part of the CleverAgents automation system.
* HAL 9000 has contributed the file edit encoding parameter fix (PR #8258 / issue #7559).
@@ -32,5 +33,5 @@ Below are some of the specific details of various contributions.
* HAL 9000 has contributed comprehensive milestone documentation for v3.6.0 (Advanced Concepts & Deferred Features) and v3.7.0 (TUI Implementation) (PR #9903): split into sub-documents covering context strategies, LLM backends, resource types, A2A rename, container tool execution, scope chain resolution, cost/safety budgets, E2E workflow tests, code review examples, plugin architecture, TUI layout, persona system, reference/command input, session management, configuration, and TuiMaterializer integration.
* HAL 9000 has contributed the LLMTraceRepository data-integrity fix (PR #8185 / issue #7505): replaced the unconditional `session.commit()` in `LLMTraceRepository.save()` with a dual-path implementation that respects the UnitOfWork pattern — flushing only when an external session is provided, and flushing + committing + closing when operating standalone. This eliminates premature transaction commits, loss of rollback capability, and a docstring/implementation mismatch.
* HAL 9000 has contributed the ACMS Index Data Model and File Traversal Engine (PR #9664 / issue #9579): foundational data structures for indexed context entries with hot/warm/cold/archive storage tier classification, tag system, and a timeout-safe chunked file traversal engine for large projects with 10,000+ files.
* HAL 9000 has contributed the error-suppression removal fix (PR #9247 / issue #9060): removed both `try...except Exception:` blocks in `register_registry_agents()` that silently suppressed errors from `actor_registry.list_actors()` and the route bridge refresh, enabling exceptions to propagate per CONTRIBUTING.md fail-fast policy. Added three Behave scenarios verifying RuntimeError, AttributeError, and TypeError propagation.
* HAL 9000 has contributed the error-suppression removal fix (PR #9247 / issue #9060): removed both `try...except Exception:` blocks in `register_registry_agents()` that silently suppressed errors from `actor_registry.list_actors()` and the route bridge refresh, enabling exceptions to propagate per CONTRIBUTING.md fail-fast policy. Added three Behave scenarios verifying RuntimeError, AttributeError, and TypeError propagation.
* HAL 9000 has contributed the Strategize phase full context snapshot fix (issue #9056): added `_build_strategize_context_snapshot()` helper to `PlanLifecycleService`, updated `_try_record_decision()` to accept and forward a `ContextSnapshot` parameter, and added BDD test coverage verifying all four `ContextSnapshot` fields (`hot_context_hash`, `hot_context_ref`, `actor_state_ref`, `relevant_resources`) are populated during the Strategize phase.
+16 -19
View File
@@ -57,35 +57,32 @@ inherits. Represents a generic container execution environment.
| Child types | (none) |
| Handler | `DevcontainerHandler` |
## Auto-Discovery (Planned)
## Auto-Discovery
> **Not yet wired (F31/F23):** The discovery module
> (`discover_devcontainers()`) exists and is tested in isolation, but
> is **not** invoked during `project link-resource` or any other
> production code path. Auto-discovery will be wired in a follow-up PR.
> For now, devcontainer-instance resources must be added manually via
> `agents resource add`.
Auto-discovery is wired into `GitCheckoutHandler.discover_children()` and
`FsDirectoryHandler.discover_children()` (issue #4740). When either handler
scans a resource location, it calls `discover_devcontainers()` after
discovering `fs-directory` children.
When wired, linking a `git-checkout` or `fs-directory` resource to a
project will trigger an auto-discovery hook that scans for devcontainer
configurations in the following locations (relative to the resource
root):
Devcontainer configurations are detected at the following locations
(relative to the resource root):
1. `.devcontainer/devcontainer.json`
2. `.devcontainer.json` (root-level)
3. `.devcontainer/<name>/devcontainer.json` (named configurations)
### Discovery Process (Planned)
### Discovery Process
1. **Trigger**: Linking a `git-checkout` or `fs-directory` resource.
2. **Scan**: The discovery module checks for configuration files at the
well-known paths listed above.
1. **Trigger**: `discover_children()` is called on a `git-checkout` or
`fs-directory` resource.
2. **Scan**: `discover_devcontainers()` checks for configuration files at
the well-known paths listed above.
3. **Validate**: Each discovered file is parsed as JSON. Invalid files
are skipped with a warning.
4. **Register**: For each valid configuration:
- A `devcontainer-instance` child resource is created under the
parent resource.
- A `devcontainer-file` child resource is created under the
devcontainer instance, pointing to the JSON file.
parent resource with `provisioning_state: discovered`.
- Named configurations carry the subdirectory name as `config_name`.
### Lazy Activation
@@ -250,7 +247,7 @@ confirmation unless `--yes` (`-y`) is passed.
| Registry eviction | Terminal-state trackers (stopped/failed) are evicted when count exceeds 200 (`_MAX_TERMINAL_TRACKERS`). Eviction runs automatically after bulk cleanup via `stop_all_active_containers`. Long-running processes with many container cycles may lose historical tracker data. | Acceptable for MVP; persisted state in M7+ eliminates this. |
| Sandbox strategy (F22/F25) | Specification uses `container_snapshot`; `SandboxFactory` does not yet implement `snapshot`. Handler now uses `SandboxStrategy.NONE` — the container itself provides isolation. | Implement `container_snapshot` in `SandboxFactory` and switch handler back to it. |
| Hardcoded `docker stop` (F21) | `stop_container()` calls `docker stop` directly. Devcontainer CLI can target Podman or other engines; the devcontainer CLI does not yet offer a `stop` subcommand. | Detect the container engine from resource properties or devcontainer CLI config and dispatch to the appropriate stop command. |
| Auto-discovery not wired (F23) | Documentation describes auto-discovery on `project link-resource`, but no code path invokes `discover_devcontainers()` during resource linking. | Wire the discovery hook into the resource linking flow. |
| Auto-discovery wired (F23 resolved) | `discover_devcontainers()` is now called from `GitCheckoutHandler.discover_children()` and `FsDirectoryHandler.discover_children()`. Devcontainer configs are auto-detected when `discover_children()` is invoked. | Completed in issue #4740. |
| Container execution stubbed (F24) | `ToolRunner` returns an error for `ExecutionEnvironment.CONTAINER`, so lazy activation cannot trigger via tool use in production. Lazy activation currently works through `DevcontainerHandler.resolve()` on plan sandbox resolution. | Full container execution support is scoped for a follow-up PR (#616). |
## DevcontainerHandler Protocol Methods
@@ -9,10 +9,10 @@ Feature: ACMS Index Data Model and File Traversal Engine
Scenario: Create an index entry with file metadata
When I create an index entry with:
| key | value |
| path | /project/src/main.py |
| file_type | python |
| size_bytes | 1024 |
| key | value |
| path | /project/src/main.py |
| file_type | python |
| size_bytes | 1024 |
Then the index entry should have path "/project/src/main.py"
And the index entry should have file type "python"
And the index entry should have size 1024 bytes
@@ -117,7 +117,7 @@ Feature: ACMS Index Data Model and File Traversal Engine
Scenario: Get entry count from index
Given I have an index with 10 entries
When I get the entry count
Then the ACMS entry count should be 10
Then idx the index count should be 10
Scenario: Remove entry from index
Given I have an index with 3 entries
@@ -132,9 +132,9 @@ Feature: ACMS Index Data Model and File Traversal Engine
| /project/tests/test_main.py | python | test | cold |
| /project/docs/readme.md | markdown | docs | cold |
When I query the index with filters:
| filter | value |
| path_pattern | src |
| file_type | python |
| tier | hot |
| filter | value |
| path_pattern | src |
| file_type | python |
| tier | hot |
Then I should get 1 result
And the result should have path "/project/src/main.py"
+11 -3
View File
@@ -1,7 +1,7 @@
Feature: agents actor add NAME positional argument
As a user of the CleverAgents CLI
I want to pass the actor name as a positional argument to `agents actor add`
So that the CLI matches the spec synopsis: agents actor add <NAME> --config <FILE>
So that the CLI matches the spec synopsis: agents actor add --config <FILE> [<NAME>]
@tdd_issue @tdd_issue_4230 @tdd_expected_fail
Scenario: actor add accepts NAME as positional argument
@@ -17,8 +17,16 @@ Feature: agents actor add NAME positional argument
When I run actor add with NAME positional argument overriding config name
Then the actor add should use the positional NAME not the config name
Scenario: actor add without NAME positional argument fails
@tdd_issue @tdd_issue_4186
Scenario: actor add without NAME positional argument uses config name
Given an actor CLI runner
And I have an actor JSON config file with name "local/config-derived-actor"
When I run actor add with config but no NAME positional argument
Then the actor add should succeed using the config name
@tdd_issue @tdd_issue_4186
Scenario: actor add without NAME and without config name field raises BadParameter
Given an actor CLI runner
And I have an actor JSON config file without a name field
When I run actor add with config but no NAME positional argument
Then the actor command should fail with missing argument error
Then the actor add should fail with a BadParameter error about missing actor name
@@ -5,7 +5,7 @@ Feature: agents actor add enforces --update flag for existing actors
I want `agents actor add` to fail with a clear error when re-adding an existing actor
So that I cannot accidentally overwrite actor configurations without explicit intent
@tdd_issue @tdd_issue_2609 @tdd_expected_fail @tdd_issue_4178
@tdd_issue @tdd_issue_2609 @tdd_issue_4178
Scenario: Re-adding an existing actor without --update fails with error panel
Given an actor add CLI runner where the actor already exists
When I run actor add without the --update flag
@@ -14,7 +14,7 @@ Feature: agents actor add enforces --update flag for existing actors
And the actor-add-enforcement output should contain "Use --update to replace the existing actor definition"
And the actor-add-enforcement output should contain the registration timestamp
@tdd_issue @tdd_issue_2609 @tdd_expected_fail @tdd_issue_4178
@tdd_issue @tdd_issue_2609 @tdd_issue_4178
Scenario: Re-adding an existing actor without --update shows error status line
Given an actor add CLI runner where the actor already exists
When I run actor add without the --update flag
@@ -22,14 +22,14 @@ Feature: agents actor add enforces --update flag for existing actors
And the actor-add-enforcement output should contain "Actor already registered"
And the actor-add-enforcement output should contain "use --update to replace"
@tdd_issue @tdd_issue_2609 @tdd_expected_fail @tdd_issue_4178
@tdd_issue @tdd_issue_2609 @tdd_issue_4178
Scenario: Re-adding an existing actor with --update succeeds
Given an actor add CLI runner where the actor already exists
When I run actor add with the --update flag
Then the actor-add-enforcement exit code should be 0
And the actor-add-enforcement output should contain "Actor updated"
@tdd_issue @tdd_issue_2609 @tdd_expected_fail @tdd_issue_4178
@tdd_issue @tdd_issue_2609 @tdd_issue_4178
Scenario: Adding a new actor without --update succeeds
Given an actor add CLI runner where the actor does not exist
When I run actor add without the --update flag
+1
View File
@@ -34,6 +34,7 @@ Feature: Architecture validation
Then all application variables should use CLEVERAGENTS_ prefix
And provider variables should keep their original prefix
@tdd_issue @tdd_issue_4186
Scenario: Type hints are used throughout
Given the source code exists
When I check for type annotations
+6 -6
View File
@@ -1159,7 +1159,7 @@ Feature: Consolidated Domain Models
Scenario: CLI dict has required fields
Given a session with some messages
When I get the session CLI dict
When I get the session model CLI dict
Then the session cli dict should have key "session_summary"
And the session cli dict should have key "token_usage"
And the session cli dict session_summary should have key "id"
@@ -1170,34 +1170,34 @@ Feature: Consolidated Domain Models
Scenario: CLI dict includes recent messages
Given a session with some messages
When I get the session CLI dict
When I get the session model CLI dict
Then the session cli dict should have key "recent_messages"
And the session cli dict recent_messages text key should be "text"
Scenario: CLI dict includes actor when set
When I create a session with actor name "local/orchestrator"
And I get the session CLI dict
And I get the session model CLI dict
Then the session cli dict session_summary should have key "actor"
Scenario: CLI dict includes automation when set
When I create a session with automation "review"
And I get the session CLI dict
And I get the session model CLI dict
Then the session cli dict session_summary should have key "automation"
And the session cli dict session_summary automation should be "review"
Scenario: CLI dict linked_plans uses spec-compliant objects
Given a session with linked plans
When I get the session CLI dict
When I get the session model CLI dict
Then the session cli dict should have key "linked_plans"
And the session cli dict linked_plans should contain plan_id field
Scenario: CLI dict token_usage estimated_cost is formatted string
Given a session with token usage cost 0.0184
When I get the session CLI dict
When I get the session model CLI dict
Then the session cli dict token_usage estimated_cost should be a string starting with "$"
# ---- Empty Session Properties ----
@@ -352,24 +352,24 @@ Feature: Consolidated Plan Model Lifecycle
Scenario: New plan starts in STRATEGIZE with QUEUED processing state
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan phase should be "strategize"
And the plan processing state should be "queued"
Scenario: Strategize phase uses ProcessingState
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan processing state should be "queued"
Scenario: Plan in strategize phase defaults to QUEUED processing state
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan processing state should be "queued"
And the plan state should be "queued"
Scenario: Strategize phase defaults processing_state to QUEUED
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan processing state should be "queued"
# Plan Identity Tests
@@ -401,12 +401,12 @@ Feature: Consolidated Plan Model Lifecycle
Scenario: Plan in STRATEGIZE with QUEUED cannot transition
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan cannot transition to next phase
Scenario: New plan in STRATEGIZE with QUEUED cannot transition
Given I create a plan in strategize phase
Given I create a PlanModel in strategize phase
Then the plan cannot transition to next phase
@@ -455,17 +455,17 @@ Feature: Consolidated Plan Model Lifecycle
Scenario: Plan in errored state is_errored returns True
Given I create a plan in strategize phase with errored state
Given I create a PlanModel in strategize phase with errored state
Then the plan should be errored
Scenario: Plan in processing state is_errored returns False
Given I create a plan in strategize phase with processing state
Given I create a PlanModel in strategize phase with processing state
Then the plan should not be errored
Scenario: Cancelled plan is terminal
Given I create a plan in strategize phase with cancelled state
Given I create a PlanModel in strategize phase with cancelled state
Then the plan should be terminal
# Model Validation Tests
+29
View File
@@ -401,3 +401,32 @@ Feature: Decision recording and snapshot store
And I record a strategy_choice decision for plan "P1" with question "After restart"
Then the dsvc decision sequence number should be 2
And the dsvc next sequence for plan "P1" should be 3
# --- Full context snapshot (issue #9056) ---
Scenario: Strategize phase records decisions with full context snapshots
Given a plan lifecycle service with decision service wired
And an action "local/test-action" for strategize snapshot test
And a plan created from "local/test-action" with project "proj-snapshot"
When I start strategize for the snapshot test plan
Then the strategize decision should have a non-empty hot_context_hash
And the strategize decision should have a non-empty hot_context_ref
And the strategize decision hot_context_ref should start with "plan:"
And the strategize decision should have a non-empty actor_state_ref
And the strategize decision should have relevant_resources populated
Scenario: Strategize context snapshot hash is content-addressable
Given a plan lifecycle service with decision service wired
And an action "local/test-action-hash" for strategize snapshot test
And a plan created from "local/test-action-hash" with project "proj-hash"
When I start strategize for the snapshot test plan
Then the strategize decision hot_context_hash should start with "sha256:"
And the strategize decision hot_context_hash should be 71 characters long
Scenario: Strategize context snapshot without projects has empty relevant_resources
Given a plan lifecycle service with decision service wired
And an action "local/no-project-action" for strategize snapshot test
And a plan created from "local/no-project-action" without projects
When I start strategize for the snapshot test plan
Then the strategize decision should have a non-empty hot_context_hash
And the strategize decision should have empty relevant_resources
@@ -0,0 +1,62 @@
Feature: Devcontainer auto-discovery wired into git-checkout and fs-directory handlers
As a CleverAgents user
I want devcontainer configurations to be automatically discovered
When I register a git-checkout or fs-directory resource
So that devcontainer-instance child resources are created without manual intervention
Scenario: git-checkout discover_children finds root devcontainer config
Given dcwire a git repo with a ".devcontainer/devcontainer.json" file
When dcwire I call discover_children on the git-checkout resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
And dcwire the devcontainer child has provisioning_state "discovered"
And dcwire the devcontainer child has devcontainer_json_path set
Scenario: git-checkout discover_children finds named devcontainer config
Given dcwire a git repo with a ".devcontainer/api/devcontainer.json" named config
When dcwire I call discover_children on the git-checkout resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-api"
And dcwire the devcontainer child has config_name "api"
Scenario: git-checkout discover_children with no devcontainer returns only directories
Given dcwire a git repo with no devcontainer configuration
When dcwire I call discover_children on the git-checkout resource
Then dcwire no devcontainer-instance children are present
Scenario: git-checkout discover_children includes both fs-directory and devcontainer children
Given dcwire a git repo with a subdirectory "src" and a ".devcontainer/devcontainer.json" file
When dcwire I call discover_children on the git-checkout resource
Then dcwire the children include a "fs-directory" resource named "src"
And dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
Scenario: fs-directory discover_children finds root devcontainer config
Given dcwire a filesystem directory with a ".devcontainer/devcontainer.json" file
When dcwire I call discover_children on the fs-directory resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
And dcwire the devcontainer child has provisioning_state "discovered"
Scenario: fs-directory discover_children finds named devcontainer config
Given dcwire a filesystem directory with a ".devcontainer/frontend/devcontainer.json" named config
When dcwire I call discover_children on the fs-directory resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-frontend"
And dcwire the devcontainer child has config_name "frontend"
Scenario: fs-directory discover_children with no devcontainer returns only directories
Given dcwire a filesystem directory with no devcontainer configuration
When dcwire I call discover_children on the fs-directory resource
Then dcwire no devcontainer-instance children are present
Scenario: fs-directory discover_children includes both fs-directory and devcontainer children
Given dcwire a filesystem directory with a subdirectory "lib" and a ".devcontainer/devcontainer.json" file
When dcwire I call discover_children on the fs-directory resource
Then dcwire the children include a "fs-directory" resource named "lib"
And dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
Scenario: git-checkout discover_children finds root-level .devcontainer.json
Given dcwire a git repo with a root-level ".devcontainer.json" file
When dcwire I call discover_children on the git-checkout resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
Scenario: fs-directory discover_children finds root-level .devcontainer.json
Given dcwire a filesystem directory with a root-level ".devcontainer.json" file
When dcwire I call discover_children on the fs-directory resource
Then dcwire the children include a "devcontainer-instance" resource named "devcontainer-default"
+4 -4
View File
@@ -180,19 +180,19 @@ Feature: Plan edge case scenarios
Scenario: Plan model rejects empty description
When I try to create an edge case plan with empty description
Then a Pydantic validation error should be raised
Then an Edge Case Pydantic validation error should be raised
Scenario: Plan model rejects invalid phase value
When I try to create an edge case plan with invalid phase value
Then a Pydantic validation error should be raised
Then an Edge Case Pydantic validation error should be raised
Scenario: NamespacedName rejects invalid characters in namespace
When I try to parse a namespaced name with special characters "inv@lid/action"
Then a Pydantic validation error should be raised
Then an Edge Case Pydantic validation error should be raised
Scenario: NamespacedName rejects invalid characters in name
When I try to parse a namespaced name with special chars in name "local/my action!"
Then a Pydantic validation error should be raised
Then an Edge Case Pydantic validation error should be raised
# ──────────────────────────────────────────────────
# Section 4: Rollback edge cases
+21 -8
View File
@@ -1,6 +1,7 @@
"""Behave environment setup for feature tests."""
import contextlib
import fcntl
import logging
import os
import re
@@ -335,13 +336,13 @@ def before_all(context):
# Use per-process unique database paths so parallel test subprocesses
# (behave-parallel) never contend on the same SQLite file.
if "CLEVERAGENTS_DATABASE_URL" not in os.environ:
os.environ["CLEVERAGENTS_DATABASE_URL"] = (
f"sqlite:///{tempfile.mktemp(suffix='.db', prefix='cleveragents_')}"
)
_fd, _db_path = tempfile.mkstemp(suffix=".db", prefix="cleveragents_")
os.close(_fd)
os.environ["CLEVERAGENTS_DATABASE_URL"] = f"sqlite:///{_db_path}"
if "CLEVERAGENTS_TEST_DATABASE_URL" not in os.environ:
os.environ["CLEVERAGENTS_TEST_DATABASE_URL"] = (
f"sqlite:///{tempfile.mktemp(suffix='.db', prefix='cleveragents_test_')}"
)
_fd, _db_path = tempfile.mkstemp(suffix=".db", prefix="cleveragents_test_")
os.close(_fd)
os.environ["CLEVERAGENTS_TEST_DATABASE_URL"] = f"sqlite:///{_db_path}"
os.environ.setdefault("BEHAVE_TESTING", "true")
# Set up mock AI provider for all tests
@@ -453,8 +454,15 @@ def _ensure_template_db() -> None:
os.environ["CLEVERAGENTS_TEMPLATE_DB"] = str(template_path.resolve())
return
lock_path = template_path.with_suffix(".db.lock")
_lock_fd = -1
try:
# Import the template creation script
_lock_fd = os.open(str(lock_path), os.O_CREAT | os.O_RDWR)
fcntl.flock(_lock_fd, fcntl.LOCK_EX)
if template_path.is_file():
os.environ["CLEVERAGENTS_TEMPLATE_DB"] = str(template_path.resolve())
return
scripts_dir = Path(__file__).parent.parent / "scripts"
sys.path.insert(0, str(scripts_dir))
from create_template_db import create_template
@@ -463,6 +471,10 @@ def _ensure_template_db() -> None:
os.environ["CLEVERAGENTS_TEMPLATE_DB"] = str(template_path.resolve())
except Exception:
pass # Fall back to normal Alembic migrations
finally:
if _lock_fd >= 0:
fcntl.flock(_lock_fd, fcntl.LOCK_UN)
os.close(_lock_fd)
def _install_template_db_patch() -> None:
@@ -636,7 +648,8 @@ def before_scenario(context, scenario):
("CLEVERAGENTS_DATABASE_URL", "cleveragents_"),
("CLEVERAGENTS_TEST_DATABASE_URL", "cleveragents_test_"),
):
db_path = tempfile.mktemp(suffix=".db", prefix=prefix)
_fd, db_path = tempfile.mkstemp(suffix=".db", prefix=prefix)
os.close(_fd)
os.environ[env_var] = f"sqlite:///{db_path}"
context._scenario_db_paths.append(db_path)
@@ -32,7 +32,7 @@ Feature: PlanExecutor Coverage Boost
Scenario: _try_rollback_to_last_checkpoint returns False when no sandbox is resolvable
Given a PlanExecutor with a checkpoint manager but no sandbox source
When I attempt to rollback to the last checkpoint
Then the rollback result should be False
Then the executor rollback result should be False
# --- _resolve_sandbox_for_checkpoint via execution_context (line 458) ---
+1 -1
View File
@@ -56,7 +56,7 @@ Feature: Plan explain and decision tree CLI commands
Given a test decision for explain
When I format the explain dict as json
Then the json output should contain "decision_id"
And the json output should be valid json
And the plan explain json output should be valid
# ------------------------------------------------------------------
# plan explain - yaml format
+7 -6
View File
@@ -1,8 +1,9 @@
Feature: Plugin Loader Coverage Boost
Scenarios targeting uncovered lines in the PluginLoader class:
- Lines 203-209: entry point load failure exception handler
- Lines 242, 244-246: validate_protocol fallback to issubclass when instantiation fails
- Lines 247-248: validate_protocol issubclass raises TypeError
- validate_protocol: issubclass succeeds (class satisfies protocol structurally)
- validate_protocol: issubclass raises TypeError with unverifiable protocol
- validate_protocol: issubclass returns False (class missing required members)
Background:
Given the plugin loader module is imported
@@ -18,17 +19,17 @@ Feature: Plugin Loader Coverage Boost
And the failed entry point should have been logged as a warning
# -----------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass succeeds (lines 242, 244-246)
# validate_protocol: issubclass succeeds for class satisfying protocol
# -----------------------------------------------------------------------
Scenario: validate_protocol falls back to issubclass when instantiation fails
Scenario: validate_protocol returns True when class satisfies protocol via issubclass
Given I have a class that requires constructor arguments
And I have a runtime checkable protocol the class satisfies via issubclass
When I call validate_protocol with the non-instantiable class and protocol
Then validate_protocol should return True
# -----------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass raises TypeError (lines 247-248)
# validate_protocol: issubclass raises TypeError for unverifiable protocol
# -----------------------------------------------------------------------
Scenario: validate_protocol raises ProtocolMismatchError when issubclass raises TypeError
@@ -38,7 +39,7 @@ Feature: Plugin Loader Coverage Boost
Then a plugin-loader ProtocolMismatchError should be raised
# -----------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass returns False
# validate_protocol: issubclass returns False (class missing required members)
# -----------------------------------------------------------------------
Scenario: validate_protocol raises ProtocolMismatchError when issubclass returns False
+1 -1
View File
@@ -125,7 +125,7 @@ Feature: Namespaced project repository operations
Scenario: Remove a non-existent link returns False
When I remove a link with id "00000000000000000000000099"
Then the remove result should be False
Then the project repo remove result should be False
Scenario: Create link with read_only flag
Given a namespaced project "local/ro-proj" exists in the repository
+1 -1
View File
@@ -76,7 +76,7 @@ Feature: Retry Policy Wiring for Services
@registry
Scenario: Per-service policy defaults with config overrides
Given I have Settings with retry_service_overrides setting plan_service max_attempts to 10
When I create a ServiceRetryWiring from those Settings
When I create a ServiceRetryWiring from those retry Settings
Then the plan_service policy should have max_attempts 10
@registry
@@ -361,7 +361,7 @@ def step_get_entry_count(context):
context.entry_count = context.index.get_entry_count()
@then("the ACMS entry count should be {count:d}")
@then("idx the index count should be {count:d}")
def step_check_entry_count_value(context, count):
"""Verify the entry count matches the expected value."""
assert context.entry_count == count
@@ -11,6 +11,7 @@ from unittest.mock import MagicMock, patch
from behave import given, then, when
from cleveragents.cli.commands.actor import app as actor_app
from cleveragents.core.exceptions import NotFoundError
from cleveragents.domain.models.core.actor import Actor
@@ -75,6 +76,22 @@ def step_impl(context: Any) -> None:
_register_cleanup(context, context.actor_config_path)
@given("I have an actor JSON config file with name {name}")
def step_impl(context: Any, name: str) -> None:
context.actor_config_data = {
"name": name,
"provider": "openai",
"model": "gpt-4o-mini",
}
with tempfile.NamedTemporaryFile(
delete=False, suffix=".json", mode="w", encoding="utf-8"
) as handle:
json.dump(context.actor_config_data, handle)
handle.flush()
context.actor_config_path = Path(handle.name)
_register_cleanup(context, context.actor_config_path)
@when("I run actor add with NAME positional argument and config")
def step_impl(context: Any) -> None:
context.positional_name = "local/my-actor"
@@ -117,8 +134,14 @@ def step_impl(context: Any) -> None:
@when("I run actor add with config but no NAME positional argument")
def step_impl(context: Any) -> None:
config_name = context.actor_config_data.get("name", "local/default-actor")
mock_actor = _make_actor(name=config_name)
with patch("cleveragents.cli.commands.actor._get_services") as mock_svc:
registry = MagicMock()
registry.get_actor.side_effect = NotFoundError(
f"Actor not found: {config_name}"
)
registry.upsert_actor.return_value = mock_actor
mock_svc.return_value = (MagicMock(), registry)
context.result = context.runner.invoke(
actor_app,
@@ -128,6 +151,7 @@ def step_impl(context: Any) -> None:
str(context.actor_config_path),
],
)
context.mock_registry = registry
@then("the actor add should succeed with the positional name")
@@ -175,3 +199,34 @@ def step_impl(context: Any) -> None:
f"Expected non-zero exit_code, got {context.result.exit_code}.\n"
f"Output:\n{context.result.output}"
)
@then("the actor add should succeed using the config name")
def step_impl(context: Any) -> None:
assert context.result.exit_code == 0, (
f"Expected exit_code=0, got {context.result.exit_code}.\n"
f"Output:\n{context.result.output}"
)
assert context.mock_registry.upsert_actor.called, (
"Expected upsert_actor to be called on the registry"
)
call_kwargs = context.mock_registry.upsert_actor.call_args
actual_name = call_kwargs.kwargs.get("name") or (
call_kwargs.args[0] if call_kwargs.args else None
)
expected_name = context.actor_config_data.get("name")
assert actual_name == expected_name, (
f"Expected upsert_actor called with name={expected_name!r} (from config), "
f"got name={actual_name!r}"
)
@then("the actor add should fail with a BadParameter error about missing actor name")
def step_impl(context: Any) -> None:
assert context.result.exit_code != 0, (
f"Expected non-zero exit_code, got {context.result.exit_code}.\n"
f"Output:\n{context.result.output}"
)
assert "Actor name is required" in context.result.output, (
f"Expected 'Actor name is required' in output:\n{context.result.output}"
)
@@ -107,7 +107,7 @@ def step_when_add_without_update(context: Any) -> None:
mock_svc.return_value = (MagicMock(), registry)
context.result = context.runner.invoke(
actor_app,
["add", "--config", str(context.actor_config_path)],
["add", context.actor_name, "--config", str(context.actor_config_path)],
)
@@ -122,7 +122,13 @@ def step_when_add_with_update(context: Any) -> None:
mock_svc.return_value = (MagicMock(), registry)
context.result = context.runner.invoke(
actor_app,
["add", "--config", str(context.actor_config_path), "--update"],
[
"add",
context.actor_name,
"--config",
str(context.actor_config_path),
"--update",
],
)
+143
View File
@@ -1105,3 +1105,146 @@ def step_try_store_duplicate(context: Context) -> None:
def step_duplicate_error_raised(context: Context) -> None:
assert context.decision_error is not None
assert isinstance(context.decision_error, DuplicateDecisionError)
# ---------------------------------------------------------------------------
# Full context snapshot steps (issue #9056)
# ---------------------------------------------------------------------------
@given("a plan lifecycle service with decision service wired")
def step_lifecycle_with_decision_service(context: Context) -> None:
"""Create a PlanLifecycleService with a real DecisionService wired in."""
from cleveragents.application.services.decision_service import DecisionService
from cleveragents.application.services.plan_lifecycle_service import (
PlanLifecycleService,
)
from cleveragents.config.settings import Settings
Settings._instance = None
settings = Settings()
context.snapshot_decision_service = DecisionService()
context.snapshot_lifecycle_service = PlanLifecycleService(
settings=settings,
decision_service=context.snapshot_decision_service,
)
context.snapshot_plan = None
context.snapshot_decision = None
context.snapshot_action_name = None
@given('an action "{action_name}" for strategize snapshot test')
def step_create_action_for_snapshot_test(context: Context, action_name: str) -> None:
"""Create an action for the strategize snapshot test."""
context.snapshot_action_name = action_name
context.snapshot_lifecycle_service.create_action(
name=action_name,
description=f"Action {action_name} for snapshot test",
definition_of_done="Snapshot test done",
strategy_actor="openai/gpt-4",
execution_actor="openai/gpt-4",
)
@given('a plan created from "{action_name}" with project "{project_name}"')
def step_create_plan_with_project(
context: Context, action_name: str, project_name: str
) -> None:
"""Create a plan from the given action with a project link."""
from cleveragents.domain.models.core.plan import ProjectLink
context.snapshot_plan = context.snapshot_lifecycle_service.use_action(
action_name=action_name,
project_links=[ProjectLink(project_name=project_name)],
)
@given('a plan created from "{action_name}" without projects')
def step_create_plan_without_projects(context: Context, action_name: str) -> None:
"""Create a plan from the given action without any project links."""
context.snapshot_plan = context.snapshot_lifecycle_service.use_action(
action_name=action_name,
project_links=[],
)
@when("I start strategize for the snapshot test plan")
def step_start_strategize_snapshot_test(context: Context) -> None:
"""Start strategize and capture the recorded decision."""
plan_id = context.snapshot_plan.identity.plan_id
context.snapshot_lifecycle_service.start_strategize(plan_id)
decisions = context.snapshot_decision_service.list_decisions(plan_id)
assert len(decisions) >= 1, f"Expected at least 1 decision, got {len(decisions)}"
strategy_decisions = [
d for d in decisions if d.decision_type.value == "strategy_choice"
]
assert len(strategy_decisions) >= 1, "Expected at least 1 strategy_choice decision"
context.snapshot_decision = strategy_decisions[0]
@then("the strategize decision should have a non-empty hot_context_hash")
def step_check_snapshot_hash_not_empty(context: Context) -> None:
"""Verify the context snapshot hash is not empty."""
snapshot = context.snapshot_decision.context_snapshot
assert snapshot.hot_context_hash, "hot_context_hash should not be empty"
@then("the strategize decision should have a non-empty hot_context_ref")
def step_check_snapshot_ref_not_empty(context: Context) -> None:
"""Verify the context snapshot ref is not empty."""
snapshot = context.snapshot_decision.context_snapshot
assert snapshot.hot_context_ref, "hot_context_ref should not be empty"
@then('the strategize decision hot_context_ref should start with "{prefix}"')
def step_check_snapshot_ref_prefix(context: Context, prefix: str) -> None:
"""Verify the context snapshot ref starts with the expected prefix."""
snapshot = context.snapshot_decision.context_snapshot
assert snapshot.hot_context_ref.startswith(prefix), (
f"hot_context_ref should start with {prefix!r}"
)
@then("the strategize decision should have a non-empty actor_state_ref")
def step_check_snapshot_actor_ref_not_empty(context: Context) -> None:
"""Verify the actor_state_ref is not empty."""
snapshot = context.snapshot_decision.context_snapshot
assert snapshot.actor_state_ref, "actor_state_ref should not be empty"
@then("the strategize decision should have relevant_resources populated")
def step_check_snapshot_resources_populated(context: Context) -> None:
"""Verify relevant_resources is not empty."""
snapshot = context.snapshot_decision.context_snapshot
assert len(snapshot.relevant_resources) > 0, (
"relevant_resources should not be empty"
)
@then('the strategize decision hot_context_hash should start with "{prefix}"')
def step_check_snapshot_hash_prefix(context: Context, prefix: str) -> None:
"""Verify the context snapshot hash starts with the expected prefix."""
snapshot = context.snapshot_decision.context_snapshot
assert snapshot.hot_context_hash.startswith(prefix), (
f"hot_context_hash should start with {prefix!r}"
)
@then("the strategize decision hot_context_hash should be {length:d} characters long")
def step_check_snapshot_hash_length(context: Context, length: int) -> None:
"""Verify the context snapshot hash has the expected length."""
snapshot = context.snapshot_decision.context_snapshot
actual_length = len(snapshot.hot_context_hash)
assert actual_length == length, (
f"hot_context_hash length should be {length}, got {actual_length}"
)
@then("the strategize decision should have empty relevant_resources")
def step_check_snapshot_resources_empty(context: Context) -> None:
"""Verify relevant_resources is empty."""
snapshot = context.snapshot_decision.context_snapshot
assert len(snapshot.relevant_resources) == 0, (
f"relevant_resources should be empty, got {snapshot.relevant_resources}"
)
@@ -0,0 +1,339 @@
"""Step definitions for devcontainer_autodiscovery_wiring.feature.
Tests that discover_devcontainers() is correctly wired into
GitCheckoutHandler.discover_children() and FsDirectoryHandler.discover_children().
All steps use the ``dcwire`` prefix to avoid Behave AmbiguousStep errors.
"""
from __future__ import annotations
import json
import subprocess
import tempfile
from pathlib import Path
from behave import given, then, when
from cleveragents.domain.models.core.resource import PhysVirt, Resource
from cleveragents.resource.handlers.fs_directory import FsDirectoryHandler
from cleveragents.resource.handlers.git_checkout import GitCheckoutHandler
_VALID_ULID = "01JQDVHN5X5QJKBMZ3AP8Y4G7K"
_DEVCONTAINER_JSON = json.dumps({"name": "Test Dev Container", "image": "ubuntu:22.04"})
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _make_git_resource(location: str) -> Resource:
"""Create a minimal git-checkout Resource."""
return Resource(
resource_id=_VALID_ULID,
name="test-repo",
resource_type_name="git-checkout",
classification=PhysVirt.PHYSICAL,
description="Test git checkout resource",
location=location,
)
def _make_fs_resource(location: str) -> Resource:
"""Create a minimal fs-directory Resource."""
return Resource(
resource_id=_VALID_ULID,
name="test-dir",
resource_type_name="fs-directory",
classification=PhysVirt.PHYSICAL,
description="Test filesystem directory resource",
location=location,
)
def _init_git_repo(tmpdir: str, extra_files: dict[str, str] | None = None) -> str:
"""Initialise a bare-minimum git repo with an initial commit."""
subprocess.run(["git", "init"], cwd=tmpdir, capture_output=True, check=True)
subprocess.run(
["git", "config", "user.email", "test@dcwire.dev"],
cwd=tmpdir,
capture_output=True,
check=True,
)
subprocess.run(
["git", "config", "user.name", "DCWire Test"],
cwd=tmpdir,
capture_output=True,
check=True,
)
subprocess.run(
["git", "config", "commit.gpgSign", "false"],
cwd=tmpdir,
capture_output=True,
check=True,
)
# Always create a README so there is at least one tracked file
readme = Path(tmpdir) / "README.md"
readme.write_text("# dcwire test\n")
if extra_files:
for rel_path, file_content in extra_files.items():
full = Path(tmpdir) / rel_path
full.parent.mkdir(parents=True, exist_ok=True)
full.write_text(file_content)
subprocess.run(["git", "add", "."], cwd=tmpdir, capture_output=True, check=True)
subprocess.run(
["git", "commit", "-m", "init"],
cwd=tmpdir,
capture_output=True,
check=True,
)
return tmpdir
def _create_devcontainer_dir(base: str, rel_path: str) -> None:
"""Create a devcontainer.json at the given relative path inside base."""
full = Path(base) / rel_path
full.parent.mkdir(parents=True, exist_ok=True)
full.write_text(_DEVCONTAINER_JSON)
# ---------------------------------------------------------------------------
# GIVEN steps — git-checkout
# ---------------------------------------------------------------------------
@given('dcwire a git repo with a ".devcontainer/devcontainer.json" file')
def step_dcwire_git_with_root_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-git-")
_create_devcontainer_dir(tmpdir, ".devcontainer/devcontainer.json")
_init_git_repo(tmpdir, {".devcontainer/devcontainer.json": _DEVCONTAINER_JSON})
ctx.dcwire_handler = GitCheckoutHandler()
ctx.dcwire_resource = _make_git_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given('dcwire a git repo with a ".devcontainer/api/devcontainer.json" named config')
def step_dcwire_git_with_named_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-git-named-")
_create_devcontainer_dir(tmpdir, ".devcontainer/api/devcontainer.json")
_init_git_repo(tmpdir, {".devcontainer/api/devcontainer.json": _DEVCONTAINER_JSON})
ctx.dcwire_handler = GitCheckoutHandler()
ctx.dcwire_resource = _make_git_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given("dcwire a git repo with no devcontainer configuration")
def step_dcwire_git_no_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-git-nodc-")
_init_git_repo(tmpdir)
ctx.dcwire_handler = GitCheckoutHandler()
ctx.dcwire_resource = _make_git_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given(
'dcwire a git repo with a subdirectory "src" and a ".devcontainer/devcontainer.json" file'
)
def step_dcwire_git_with_src_and_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-git-src-")
_create_devcontainer_dir(tmpdir, ".devcontainer/devcontainer.json")
_init_git_repo(
tmpdir,
{
"src/main.py": "print('hello')\n",
".devcontainer/devcontainer.json": _DEVCONTAINER_JSON,
},
)
ctx.dcwire_handler = GitCheckoutHandler()
ctx.dcwire_resource = _make_git_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given('dcwire a git repo with a root-level ".devcontainer.json" file')
def step_dcwire_git_with_root_level_devcontainer_json(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-git-rootdc-")
_create_devcontainer_dir(tmpdir, ".devcontainer.json")
_init_git_repo(tmpdir, {".devcontainer.json": _DEVCONTAINER_JSON})
ctx.dcwire_handler = GitCheckoutHandler()
ctx.dcwire_resource = _make_git_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
# ---------------------------------------------------------------------------
# GIVEN steps — fs-directory
# ---------------------------------------------------------------------------
@given('dcwire a filesystem directory with a ".devcontainer/devcontainer.json" file')
def step_dcwire_fs_with_root_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-fs-")
_create_devcontainer_dir(tmpdir, ".devcontainer/devcontainer.json")
ctx.dcwire_handler = FsDirectoryHandler()
ctx.dcwire_resource = _make_fs_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given(
'dcwire a filesystem directory with a ".devcontainer/frontend/devcontainer.json" named config'
)
def step_dcwire_fs_with_named_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-fs-named-")
_create_devcontainer_dir(tmpdir, ".devcontainer/frontend/devcontainer.json")
ctx.dcwire_handler = FsDirectoryHandler()
ctx.dcwire_resource = _make_fs_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given("dcwire a filesystem directory with no devcontainer configuration")
def step_dcwire_fs_no_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-fs-nodc-")
ctx.dcwire_handler = FsDirectoryHandler()
ctx.dcwire_resource = _make_fs_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given(
'dcwire a filesystem directory with a subdirectory "lib" and a ".devcontainer/devcontainer.json" file'
)
def step_dcwire_fs_with_lib_and_devcontainer(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-fs-lib-")
lib_dir = Path(tmpdir) / "lib"
lib_dir.mkdir()
_create_devcontainer_dir(tmpdir, ".devcontainer/devcontainer.json")
ctx.dcwire_handler = FsDirectoryHandler()
ctx.dcwire_resource = _make_fs_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
@given('dcwire a filesystem directory with a root-level ".devcontainer.json" file')
def step_dcwire_fs_with_root_level_devcontainer_json(ctx):
tmpdir = tempfile.mkdtemp(prefix="dcwire-fs-rootdc-")
_create_devcontainer_dir(tmpdir, ".devcontainer.json")
ctx.dcwire_handler = FsDirectoryHandler()
ctx.dcwire_resource = _make_fs_resource(tmpdir)
ctx.dcwire_result = None
ctx.dcwire_error = None
ctx.dcwire_last_dc_child = None
# ---------------------------------------------------------------------------
# WHEN steps
# ---------------------------------------------------------------------------
@when("dcwire I call discover_children on the git-checkout resource")
def step_dcwire_discover_git(ctx):
try:
ctx.dcwire_result = ctx.dcwire_handler.discover_children(
resource=ctx.dcwire_resource
)
except Exception as exc:
ctx.dcwire_error = exc
@when("dcwire I call discover_children on the fs-directory resource")
def step_dcwire_discover_fs(ctx):
try:
ctx.dcwire_result = ctx.dcwire_handler.discover_children(
resource=ctx.dcwire_resource
)
except Exception as exc:
ctx.dcwire_error = exc
# ---------------------------------------------------------------------------
# THEN steps
# ---------------------------------------------------------------------------
@then('dcwire the children include a "devcontainer-instance" resource named "{name}"')
def step_dcwire_has_devcontainer_child(ctx, name):
assert ctx.dcwire_error is None, f"Unexpected error: {ctx.dcwire_error}"
assert ctx.dcwire_result is not None, "Expected a list of children"
dc_children = [
r
for r in ctx.dcwire_result
if r.resource_type_name == "devcontainer-instance" and r.name == name
]
assert len(dc_children) == 1, (
f"Expected exactly one devcontainer-instance named '{name}', "
f"got {[r.name for r in ctx.dcwire_result if r.resource_type_name == 'devcontainer-instance']}"
)
ctx.dcwire_last_dc_child = dc_children[0]
@then('dcwire the children include a "fs-directory" resource named "{name}"')
def step_dcwire_has_fs_child(ctx, name):
assert ctx.dcwire_error is None, f"Unexpected error: {ctx.dcwire_error}"
assert ctx.dcwire_result is not None, "Expected a list of children"
fs_children = [
r
for r in ctx.dcwire_result
if r.resource_type_name == "fs-directory" and r.name == name
]
assert len(fs_children) == 1, (
f"Expected exactly one fs-directory named '{name}', "
f"got {[r.name for r in ctx.dcwire_result if r.resource_type_name == 'fs-directory']}"
)
@then('dcwire the devcontainer child has provisioning_state "{state}"')
def step_dcwire_has_provisioning_state(ctx, state):
assert ctx.dcwire_last_dc_child is not None, "No devcontainer child captured"
props = ctx.dcwire_last_dc_child.properties or {}
actual = props.get("provisioning_state")
assert actual == state, f"Expected provisioning_state='{state}', got '{actual}'"
@then("dcwire the devcontainer child has devcontainer_json_path set")
def step_dcwire_has_json_path(ctx):
assert ctx.dcwire_last_dc_child is not None, "No devcontainer child captured"
props = ctx.dcwire_last_dc_child.properties or {}
path_val = props.get("devcontainer_json_path")
assert path_val, f"Expected devcontainer_json_path to be set, got {path_val!r}"
assert "devcontainer.json" in path_val, (
f"Expected devcontainer_json_path to contain 'devcontainer.json', got {path_val!r}"
)
@then('dcwire the devcontainer child has config_name "{config_name}"')
def step_dcwire_has_config_name(ctx, config_name):
assert ctx.dcwire_last_dc_child is not None, "No devcontainer child captured"
props = ctx.dcwire_last_dc_child.properties or {}
actual = props.get("config_name")
assert actual == config_name, (
f"Expected config_name='{config_name}', got '{actual}'"
)
@then("dcwire no devcontainer-instance children are present")
def step_dcwire_no_devcontainer_children(ctx):
assert ctx.dcwire_error is None, f"Unexpected error: {ctx.dcwire_error}"
assert ctx.dcwire_result is not None, "Expected a list of children"
dc_children = [
r for r in ctx.dcwire_result if r.resource_type_name == "devcontainer-instance"
]
assert len(dc_children) == 0, (
f"Expected no devcontainer-instance children, got {[r.name for r in dc_children]}"
)
+1 -1
View File
@@ -755,7 +755,7 @@ def step_try_create_plan_empty_desc(context: Context) -> None:
context.pydantic_error = exc
@then("a Pydantic validation error should be raised")
@then("an Edge Case Pydantic validation error should be raised")
def step_check_pydantic_error(context: Context) -> None:
assert context.pydantic_error is not None, "Expected a Pydantic validation error"
@@ -171,7 +171,7 @@ def step_attempt_rollback(context):
)
@then("the rollback result should be False")
@then("the executor rollback result should be False")
def step_verify_rollback_false(context):
"""Verify the rollback result is False."""
assert context.rollback_result is False
+1 -1
View File
@@ -331,7 +331,7 @@ def step_json_contains(context: Context, text: str) -> None:
assert text in context.pe_json_output, f"Expected '{text}' in JSON output"
@then("the json output should be valid json")
@then("the plan explain json output should be valid")
def step_json_valid(context: Context) -> None:
parsed = json.loads(context.pe_json_output)
assert isinstance(parsed, dict), "Expected a JSON object"
+4 -4
View File
@@ -184,7 +184,7 @@ def step_create_plan_with_description(context: Context, description: str) -> Non
context.plan = _default_plan(description=description)
@given("I create a plan in strategize phase")
@given("I create a PlanModel in strategize phase")
def step_create_plan_strategize_phase(context: Context) -> None:
"""Create a plan in STRATEGIZE phase."""
context.plan = _default_plan(
@@ -204,7 +204,7 @@ def step_create_plan_strategize_with_action_state(context: Context) -> None:
)
@given("I create a plan in strategize phase with errored state")
@given("I create a PlanModel in strategize phase with errored state")
def step_create_plan_strategize_errored(context: Context) -> None:
"""Create a plan in STRATEGIZE phase with ERRORED state."""
context.plan = _default_plan(
@@ -214,7 +214,7 @@ def step_create_plan_strategize_errored(context: Context) -> None:
)
@given("I create a plan in strategize phase with processing state")
@given("I create a PlanModel in strategize phase with processing state")
def step_create_plan_strategize_processing(context: Context) -> None:
"""Create a plan in STRATEGIZE phase with PROCESSING state."""
context.plan = _default_plan(
@@ -268,7 +268,7 @@ def step_create_plan_action_available(context: Context) -> None:
)
@given("I create a plan in strategize phase with cancelled state")
@given("I create a PlanModel in strategize phase with cancelled state")
def step_create_plan_strategize_cancelled(context: Context) -> None:
"""Create a plan in STRATEGIZE phase with CANCELLED processing state."""
context.plan = _default_plan(
+13 -13
View File
@@ -4,9 +4,9 @@ These steps target specific uncovered lines in
cleveragents/infrastructure/plugins/loader.py:
- Lines 203-209: except block in load_from_entry_points when ep.load() fails
- Lines 242, 244-246: validate_protocol fallback to issubclass on
instantiation failure
- Lines 247-248: validate_protocol when issubclass raises TypeError
- validate_protocol: issubclass succeeds (class satisfies protocol structurally)
- validate_protocol: issubclass raises TypeError for unverifiable protocol
- validate_protocol: issubclass returns False (class missing required members)
"""
from typing import Protocol, runtime_checkable
@@ -72,13 +72,13 @@ def step_verify_warning_logged(context):
# ---------------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass succeeds (lines 242, 244-246)
# validate_protocol: issubclass succeeds for class satisfying protocol
# ---------------------------------------------------------------------------
@runtime_checkable
class _SampleProtocol(Protocol):
"""A simple runtime-checkable Protocol for testing issubclass fallback."""
"""A simple runtime-checkable Protocol for testing issubclass check."""
def do_work(self) -> str: ...
@@ -105,7 +105,7 @@ def step_protocol_satisfied_by_subclass(context):
@when("I call validate_protocol with the non-instantiable class and protocol")
def step_call_validate_protocol_subclass_fallback(context):
"""Call validate_protocol; expect it to fall back to issubclass and succeed."""
"""Call validate_protocol; issubclass succeeds and returns True."""
context.validate_result = PluginLoader.validate_protocol(
context.non_instantiable_class,
context.target_protocol,
@@ -114,18 +114,18 @@ def step_call_validate_protocol_subclass_fallback(context):
@then("validate_protocol should return True")
def step_verify_validate_true(context):
"""The fallback issubclass check should have returned True."""
"""The issubclass check should have returned True."""
assert context.validate_result is True
# ---------------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass raises TypeError (lines 247-248)
# validate_protocol: issubclass raises TypeError for unverifiable protocol
# ---------------------------------------------------------------------------
@given("I have a class that cannot be instantiated without arguments")
def step_class_cannot_instantiate(context):
"""Create a class whose __init__ raises when called with no args."""
"""Create a class whose __init__ requires mandatory arguments."""
class NeedsArgs:
def __init__(self, required):
@@ -141,10 +141,10 @@ def step_protocol_causes_typeerror(context):
We need an object that:
- Has a __name__ attribute (so the error message in validate_protocol works)
- Causes issubclass() to raise TypeError when used as second arg
- Also causes isinstance() to raise TypeError
A class with a metaclass that raises TypeError on __instancecheck__
and __subclasscheck__ achieves this.
A class with a metaclass that raises TypeError on __subclasscheck__
achieves this. Since this protocol declares no members, the structural
fallback cannot verify conformance and must raise ProtocolMismatchError.
"""
class TypeErrorMeta(type):
@@ -186,7 +186,7 @@ def step_verify_protocol_mismatch(context):
# ---------------------------------------------------------------------------
# validate_protocol: instantiation fails, issubclass returns False
# validate_protocol: issubclass returns False (class missing required members)
# ---------------------------------------------------------------------------
+1 -1
View File
@@ -492,7 +492,7 @@ def step_pr_remove_link_by_id(context: Any, link_id: str) -> None:
context.pr_remove_result = context.pr_link_repo.remove_link(link_id)
@then("the remove result should be False")
@then("the project repo remove result should be False")
def step_pr_remove_false(context: Any) -> None:
assert context.pr_remove_result is False
+1 -1
View File
@@ -185,7 +185,7 @@ def step_settings_with_overrides(context: Any, value: int) -> None:
os.environ.pop("CLEVERAGENTS_RETRY_SERVICE_OVERRIDES", None)
@when("I create a ServiceRetryWiring from those Settings")
@when("I create a ServiceRetryWiring from those retry Settings")
def step_create_wiring_from_settings(context: Any) -> None:
from cleveragents.application.services.service_retry_wiring import (
ServiceRetryWiring,
+1 -1
View File
@@ -447,7 +447,7 @@ def session_model_check_export_messages_count(context: Context) -> None:
# ---------------------------------------------------------------------------
@when("I get the session CLI dict")
@when("I get the session model CLI dict")
def session_model_get_cli_dict(context: Context) -> None:
"""Get the CLI dict for the session."""
context.session_cli_dict = context.session_model.as_cli_dict()
@@ -0,0 +1,457 @@
"""Step definitions for Strategize decision recording feature.
Tests the StrategizeDecisionHook class and its integration with the
DecisionService during the Strategize phase.
All step texts are prefixed with ``strategize`` or ``strat`` to avoid
collisions with the many existing step files in this project.
"""
from __future__ import annotations
from behave import given, then, when
from cleveragents.application.services.decision_service import DecisionService
from cleveragents.application.services.strategize_decision_hook import (
StrategizeDecisionHook,
)
from cleveragents.core.exceptions import ValidationError
# ---------------------------------------------------------------------------
# Background steps
# ---------------------------------------------------------------------------
@given("a strategize decision service")
def step_given_strategize_decision_service(context):
"""Create an in-memory decision service for Strategize tests."""
context.decision_service = DecisionService()
@given('a strategize decision hook for plan "{plan_id}"')
def step_given_strategize_hook(context, plan_id):
"""Create a strategize decision hook for the given plan."""
context.plan_id = plan_id
context.hook = StrategizeDecisionHook(
decision_service=context.decision_service,
plan_id=plan_id,
)
context.last_decision = None
context.first_decision_id = None
context.context_data = None
context.actor_state = None
context.relevant_resources = None
context.alternatives = None
context.confidence = None
context.rationale = None
context.parent_decision_id = None
context.error = None
context.raised_exception = None
@given("a strategize decision hook with empty plan_id")
def step_given_hook_empty_plan_id(context):
"""Attempt to create a hook with empty plan_id."""
context.error = None
try:
context.hook = StrategizeDecisionHook(
decision_service=context.decision_service,
plan_id="",
)
except ValidationError as e:
context.error = e
# ---------------------------------------------------------------------------
# Strategy choice recording steps
# ---------------------------------------------------------------------------
@when('I record a strategy choice with question "{question}" and option "{option}"')
def step_when_record_strategy_choice(context, question, option):
"""Record a strategy choice decision."""
context.last_decision = context.hook.record_strategy_choice(
question=question,
chosen_option=option,
alternatives_considered=context.alternatives,
confidence_score=context.confidence,
rationale=context.rationale or "",
context_data=context.context_data,
actor_state=context.actor_state,
relevant_resources=context.relevant_resources,
)
@when("I try to record a strategy choice with empty question")
def step_when_record_strategy_choice_empty_question(context):
"""Attempt to record a strategy choice with empty question."""
context.error = None
try:
context.hook.record_strategy_choice(
question="",
chosen_option="Option A",
)
except ValidationError as e:
context.error = e
@when("I try to record a strategy choice with empty chosen_option")
def step_when_record_strategy_choice_empty_option(context):
"""Attempt to record a strategy choice with empty option."""
context.error = None
try:
context.hook.record_strategy_choice(
question="Which approach?",
chosen_option="",
)
except ValidationError as e:
context.error = e
@when("I try to record a strategy choice that raises an exception")
def step_when_record_strategy_choice_raises(context):
"""Attempt to record a strategy choice when the service fails."""
context.raised_exception = None
try:
context.hook.record_strategy_choice(
question="Which approach?",
chosen_option="Approach A",
)
except Exception as exc:
context.raised_exception = exc
# ---------------------------------------------------------------------------
# Resource selection recording steps
# ---------------------------------------------------------------------------
@when('I record a resource selection with question "{question}" and option "{option}"')
def step_when_record_resource_selection(context, question, option):
"""Record a resource selection decision."""
context.last_decision = context.hook.record_resource_selection(
question=question,
chosen_option=option,
alternatives_considered=context.alternatives,
confidence_score=context.confidence,
rationale=context.rationale or "",
context_data=context.context_data,
actor_state=context.actor_state,
relevant_resources=context.relevant_resources,
)
# ---------------------------------------------------------------------------
# Subplan spawn recording steps
# ---------------------------------------------------------------------------
@when('I record a subplan spawn with question "{question}" and option "{option}"')
def step_when_record_subplan_spawn(context, question, option):
"""Record a subplan spawn decision."""
context.last_decision = context.hook.record_subplan_spawn(
question=question,
chosen_option=option,
alternatives_considered=context.alternatives,
confidence_score=context.confidence,
rationale=context.rationale or "",
context_data=context.context_data,
actor_state=context.actor_state,
relevant_resources=context.relevant_resources,
)
# ---------------------------------------------------------------------------
# Invariant enforcement recording steps
# ---------------------------------------------------------------------------
@when('I record an invariant enforced with question "{question}" and option "{option}"')
def step_when_record_invariant_enforced(context, question, option):
"""Record an invariant enforced decision."""
context.last_decision = context.hook.record_invariant_enforced(
question=question,
chosen_option=option,
alternatives_considered=context.alternatives,
confidence_score=context.confidence,
rationale=context.rationale or "",
context_data=context.context_data,
actor_state=context.actor_state,
relevant_resources=context.relevant_resources,
)
# ---------------------------------------------------------------------------
# Context data steps
# ---------------------------------------------------------------------------
@when('two strat alternatives "{alt1}" and "{alt2}"')
def step_when_alternatives(context, alt1, alt2):
"""Set two alternatives for the next decision."""
context.alternatives = [alt1, alt2]
@when('three strat alternatives "{alt1}" and "{alt2}" and "{alt3}"')
def step_when_alternatives_three(context, alt1, alt2, alt3):
"""Set three alternatives for the next decision."""
context.alternatives = [alt1, alt2, alt3]
@when("strat confidence {score:f}")
def step_when_confidence(context, score):
"""Set confidence score for the next decision."""
context.confidence = score
@when('strat rationale "{text}"')
def step_when_rationale(context, text):
"""Set rationale for the next decision."""
context.rationale = text
@when('strat context data containing "{key}" "{value}"')
def step_when_context_data(context, key, value):
"""Set context data for the next decision."""
context.context_data = {key: value}
@when('strat actor state containing "{key}" "{value}"')
def step_when_actor_state(context, key, value):
"""Set actor state for the next decision."""
context.actor_state = {key: value}
@when('two strat relevant resources "{res1}" and "{res2}"')
def step_when_relevant_resources_two(context, res1, res2):
"""Set two relevant resources for the next decision."""
context.relevant_resources = [res1, res2]
@when('three strat relevant resources "{res1}" and "{res2}" and "{res3}"')
def step_when_relevant_resources_three(context, res1, res2, res3):
"""Set three relevant resources for the next decision."""
context.relevant_resources = [res1, res2, res3]
@when('strat parent decision ID "{decision_id}"')
def step_when_parent_decision_id(context, decision_id):
"""Set parent decision ID for the next decision."""
context.parent_decision_id = decision_id
# Recreate hook with parent ID
context.hook = StrategizeDecisionHook(
decision_service=context.decision_service,
plan_id=context.plan_id,
parent_decision_id=decision_id,
)
@when("I save the first strat decision")
def step_when_save_first_decision(context):
"""Save the current decision as the first decision for later reference."""
assert context.last_decision is not None, "No decision recorded yet"
context.first_decision_id = context.last_decision.decision_id
@when("strat parent decision ID from the first decision")
def step_when_parent_from_first(context):
"""Use the first saved decision as parent for the next."""
assert context.first_decision_id is not None, "No first decision saved"
context.parent_decision_id = context.first_decision_id
context.hook = StrategizeDecisionHook(
decision_service=context.decision_service,
plan_id=context.plan_id,
parent_decision_id=context.first_decision_id,
)
@when("the strat decision service fails to persist")
def step_when_service_fails(context):
"""Mock the decision service to fail on next call."""
original_record = context.decision_service.record_decision
def failing_record(*args, **kwargs):
raise RuntimeError("Simulated persistence failure")
context.decision_service.record_decision = failing_record
context.original_record = original_record
# ---------------------------------------------------------------------------
# Assertion steps
# ---------------------------------------------------------------------------
@then("the strat decision should be recorded successfully")
def step_then_decision_recorded(context):
"""Verify the decision was recorded."""
assert context.last_decision is not None
assert context.last_decision.decision_id is not None
assert context.last_decision.plan_id == context.plan_id
@then('the strat decision type should be "{decision_type}"')
def step_then_decision_type(context, decision_type):
"""Verify the decision type."""
assert context.last_decision.decision_type.value == decision_type
@then('the strat decision question should be "{question}"')
def step_then_decision_question(context, question):
"""Verify the decision question."""
assert context.last_decision.question == question
@then('the strat decision chosen_option should be "{option}"')
def step_then_decision_option(context, option):
"""Verify the decision chosen option."""
assert context.last_decision.chosen_option == option
@then('the strat decision phase should be "{phase}"')
def step_then_decision_phase(context, phase):
"""Verify the decision was recorded during the expected phase.
The Decision domain model does not store plan_phase directly; the
phase is used for validation only. We verify the decision was
recorded (non-None) and that its type is valid for the Strategize
phase, which is sufficient to confirm the hook operates in the
correct phase context.
"""
assert context.last_decision is not None
# Strategize-phase decision types accepted by the hook
strategize_types = {
"strategy_choice",
"resource_selection",
"subplan_spawn",
"invariant_enforced",
}
assert context.last_decision.decision_type.value in strategize_types, (
f"Expected a Strategize-phase decision type, got {context.last_decision.decision_type.value!r}"
)
@then("the strat decision should have {count:d} alternatives considered")
def step_then_alternatives_count(context, count):
"""Verify the number of alternatives."""
assert len(context.last_decision.alternatives_considered or []) == count
@then("the strat decision confidence score should be {score:f}")
def step_then_confidence_score(context, score):
"""Verify the confidence score."""
assert context.last_decision.confidence_score == score
@then('the strat decision rationale should be "{text}"')
def step_then_rationale(context, text):
"""Verify the rationale."""
assert context.last_decision.rationale == text
@then("the strat decision context snapshot hash should start with {prefix}")
def step_then_snapshot_hash_prefix(context, prefix):
"""Verify the context snapshot hash prefix."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
assert snapshot.hot_context_hash.startswith(prefix.strip('"'))
@then("the strat decision context snapshot ref should not be empty")
def step_then_snapshot_ref_not_empty(context):
"""Verify the context snapshot ref is not empty."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
assert snapshot.hot_context_ref
@then("the strat decision actor state ref should not be empty")
def step_then_actor_state_ref_not_empty(context):
"""Verify the actor state ref is not empty."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
assert snapshot.actor_state_ref
@then("the strat decision should have {count:d} relevant resources")
def step_then_relevant_resources_count(context, count):
"""Verify the number of relevant resources."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
assert len(snapshot.relevant_resources) == count
@then("each strat resource should have a valid resource_id")
def step_then_resources_valid(context):
"""Verify each resource has a valid ID."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
for resource in snapshot.relevant_resources:
assert resource.resource_id
assert len(resource.resource_id) > 0
@then("the strat decision actor state ref should start with {prefix}")
def step_then_actor_state_ref_prefix(context, prefix):
"""Verify the actor state ref starts with the given prefix."""
snapshot = context.last_decision.context_snapshot
assert snapshot is not None
assert snapshot.actor_state_ref.startswith(prefix.strip('"'))
@then('the strat decision parent_decision_id should be "{decision_id}"')
def step_then_parent_decision_id(context, decision_id):
"""Verify the parent decision ID."""
assert context.last_decision.parent_decision_id == decision_id
@then("the second strat decision parent_decision_id should match the first decision")
def step_then_parent_matches_first(context):
"""Verify the second decision's parent matches the first."""
assert context.last_decision.parent_decision_id == context.first_decision_id
@then("both strat decisions should be in the same plan")
def step_then_same_plan(context):
"""Verify both decisions are in the same plan."""
assert context.last_decision.plan_id == context.plan_id
# ---------------------------------------------------------------------------
# Error handling steps
# ---------------------------------------------------------------------------
@then("a strat validation error should be raised")
def step_then_validation_error(context):
"""Verify a validation error was raised."""
assert context.error is not None
assert isinstance(context.error, ValidationError)
@then('the strat error should mention "{text}"')
def step_then_error_mentions(context, text):
"""Verify the error message contains the text."""
assert text in str(context.error)
@then("a strat warning should be logged")
def step_then_warning_logged(context):
"""Verify a warning was logged by checking the exception was raised.
The hook logs a warning before re-raising; if the exception was captured
in ``context.raised_exception`` the warning path was exercised.
"""
assert context.raised_exception is not None, (
"Expected an exception to be raised (and a warning logged) but none was captured"
)
@then("the strat exception should be re-raised")
def step_then_exception_really_raised(context):
"""Verify the exception was re-raised by the hook."""
assert context.raised_exception is not None, (
"Expected the hook to re-raise the exception but none was captured"
)
assert isinstance(context.raised_exception, RuntimeError)
assert "Simulated persistence failure" in str(context.raised_exception)
@@ -0,0 +1,150 @@
"""Step definitions for tdd_session_create_suppress_exception.feature.
This test captures bug #10414: ``session create`` in
``src/cleveragents/cli/commands/session.py`` used
``contextlib.suppress(Exception)`` to silently discard ALL exceptions raised
by ``_facade_dispatch()`` without any logging. Programming errors,
unexpected failures, and infrastructure issues in the facade dispatch were
completely invisible.
The fix replaces the suppress block with a ``try/except Exception`` that
calls ``_log.warning(..., exc_info=True)`` so that the exception is
recorded but the session creation remains non-fatal.
"""
from __future__ import annotations
from datetime import datetime
import logging
from unittest.mock import MagicMock, patch
from behave import given, then, when
from behave.runner import Context
from typer.testing import CliRunner
from cleveragents.cli.commands import session as session_mod
from cleveragents.cli.commands.session import app as session_app
from cleveragents.domain.models.core.session import Session
# A distinctive error message to confirm the right exception was raised.
_DISPATCH_ERROR_MESSAGE = "facade dispatch failed for tdd_10414"
def _make_mock_session() -> MagicMock:
"""Build a minimal mock Session object sufficient for the create command."""
mock_session = MagicMock(spec=Session)
mock_session.session_id = "01JTEST000000000000000000000"
mock_session.actor_name = None
mock_session.namespace = "default"
mock_session.message_count = 0
mock_session.created_at = datetime(2026, 1, 1, 0, 0, 0)
mock_session.updated_at = datetime(2026, 1, 1, 0, 0, 0)
return mock_session
@given("a session create command with a mocked session service")
def step_given_mocked_session_service(context: Context) -> None:
"""Set up a CLI runner with a mocked session service.
The mock service returns a valid session so that the create command
proceeds past the service call and reaches the ``_facade_dispatch``
try/except block.
"""
context.runner = CliRunner()
mock_service = MagicMock()
mock_service.create.return_value = _make_mock_session()
# Patch the module-level service so the CLI uses our mock.
context._orig_service = session_mod._service
session_mod._service = mock_service
def _cleanup() -> None:
session_mod._service = context._orig_service
context.add_cleanup(_cleanup)
@given("the facade dispatch is patched to raise a RuntimeError")
def step_given_facade_dispatch_raises(context: Context) -> None:
"""Patch ``_facade_dispatch`` so it always raises a RuntimeError.
This simulates a real failure in the facade layer (e.g. the A2A
bootstrap is unavailable) and exercises the ``try/except`` block in
the ``create`` command.
"""
context._facade_patcher = patch(
"cleveragents.cli.commands.session._facade_dispatch",
side_effect=RuntimeError(_DISPATCH_ERROR_MESSAGE),
)
context._facade_patcher.start()
def _cleanup() -> None:
context._facade_patcher.stop()
context.add_cleanup(_cleanup)
@when("I invoke the session create command via the CLI runner")
def step_when_invoke_create(context: Context) -> None:
"""Invoke ``session create`` and capture log records during the call."""
# Capture WARNING-level log records from the session module logger.
session_logger = logging.getLogger("cleveragents.cli.commands.session")
handler = _CapturingHandler()
handler.setLevel(logging.WARNING)
session_logger.addHandler(handler)
session_logger.setLevel(logging.WARNING)
try:
context.result = context.runner.invoke(session_app, ["create"])
finally:
session_logger.removeHandler(handler)
context.captured_log_records = handler.records
class _CapturingHandler(logging.Handler):
"""A logging handler that stores all emitted records in a list."""
def __init__(self) -> None:
super().__init__()
self.records: list[logging.LogRecord] = []
def emit(self, record: logging.LogRecord) -> None:
self.records.append(record)
@then("a WARNING log entry should have been emitted for the facade dispatch failure")
def step_then_warning_logged(context: Context) -> None:
"""Assert that a WARNING log entry was emitted when the facade dispatch raised.
The fix replaces ``contextlib.suppress(Exception)`` with a
``try/except Exception`` block that calls ``_log.warning(..., exc_info=True)``.
This assertion verifies the fix is in place.
"""
# The session create command should still succeed (non-fatal).
assert context.result.exit_code == 0, (
f"Expected session create to exit 0 even when facade dispatch fails, "
f"but got exit code {context.result.exit_code}.\n"
f"Output:\n{context.result.output}"
)
# Assert that a WARNING log record was emitted.
warning_records = [
r for r in context.captured_log_records if r.levelno >= logging.WARNING
]
assert warning_records, (
f"Bug #10414: No WARNING log record was emitted when _facade_dispatch() "
f"raised a RuntimeError inside the try/except block. "
f"Captured records: {context.captured_log_records}"
)
# Assert the log record references the facade dispatch failure.
all_messages = " ".join(r.getMessage() for r in warning_records)
assert _DISPATCH_ERROR_MESSAGE in all_messages or any(
r.exc_info is not None for r in warning_records
), (
f"Bug #10414: WARNING log record was found but did not contain the "
f"expected error message {_DISPATCH_ERROR_MESSAGE!r} or exc_info. "
f"Log messages: {all_messages!r}"
)
@@ -0,0 +1,157 @@
Feature: Decision recording hook in Strategize phase
As a strategy actor
I want to record decisions during the Strategize phase
So that every choice point is captured with full context for replay and correction
Background:
Given a strategize decision service
And a strategize decision hook for plan "01JQAAAAAAAAAAAAAAAAAAAA01"
# --- Strategy Choice Recording ---
Scenario: Record a strategy choice decision
When I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision should be recorded successfully
And the strat decision type should be "strategy_choice"
And the strat decision question should be "Which approach?"
And the strat decision chosen_option should be "Approach A"
And the strat decision phase should be "strategize"
Scenario: Record strategy choice with alternatives
When two strat alternatives "Approach B" and "Approach C"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision should have 2 alternatives considered
Scenario: Record strategy choice with confidence score
When strat confidence 0.85
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision confidence score should be 0.85
Scenario: Record strategy choice with rationale
When strat rationale "Approach A is more efficient"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision rationale should be "Approach A is more efficient"
Scenario: Record strategy choice with context snapshot
When strat context data containing "key1" "value1"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision context snapshot hash should start with "sha256:"
And the strat decision context snapshot ref should not be empty
Scenario: Record strategy choice with actor state
When strat actor state containing "reasoning" "step1"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision actor state ref should not be empty
Scenario: Record strategy choice with relevant resources
When two strat relevant resources "resource1" and "resource2"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision should have 2 relevant resources
Scenario: Record strategy choice with empty question raises error
When I try to record a strategy choice with empty question
Then a strat validation error should be raised
And the strat error should mention "question"
Scenario: Record strategy choice with empty option raises error
When I try to record a strategy choice with empty chosen_option
Then a strat validation error should be raised
And the strat error should mention "chosen_option"
# --- Resource Selection Recording ---
Scenario: Record a resource selection decision
When I record a resource selection with question "Which resources?" and option "src/main.py"
Then the strat decision should be recorded successfully
And the strat decision type should be "resource_selection"
And the strat decision question should be "Which resources?"
And the strat decision chosen_option should be "src/main.py"
Scenario: Record resource selection with alternatives
When two strat alternatives "src/test.py" and "src/utils.py"
And I record a resource selection with question "Which resources?" and option "src/main.py"
Then the strat decision should have 2 alternatives considered
Scenario: Record resource selection with confidence
When strat confidence 0.9
And I record a resource selection with question "Which resources?" and option "src/main.py"
Then the strat decision confidence score should be 0.9
# --- Subplan Spawn Recording ---
Scenario: Record a subplan spawn decision
When I record a subplan spawn with question "Should we decompose?" and option "Create subplan for feature X"
Then the strat decision should be recorded successfully
And the strat decision type should be "subplan_spawn"
And the strat decision question should be "Should we decompose?"
And the strat decision chosen_option should be "Create subplan for feature X"
Scenario: Record subplan spawn with alternatives
When two strat alternatives "Implement inline" and "Create parallel subplans"
And I record a subplan spawn with question "Should we decompose?" and option "Create subplan for feature X"
Then the strat decision should have 2 alternatives considered
Scenario: Record subplan spawn with confidence
When strat confidence 0.75
And I record a subplan spawn with question "Should we decompose?" and option "Create subplan for feature X"
Then the strat decision confidence score should be 0.75
# --- Invariant Enforcement Recording ---
Scenario: Record an invariant enforced decision
When I record an invariant enforced with question "Apply security invariant?" and option "Enforce code review"
Then the strat decision should be recorded successfully
And the strat decision type should be "invariant_enforced"
And the strat decision question should be "Apply security invariant?"
And the strat decision chosen_option should be "Enforce code review"
Scenario: Record invariant enforced with rationale
When strat rationale "Security policy requires code review"
And I record an invariant enforced with question "Apply security invariant?" and option "Enforce code review"
Then the strat decision rationale should be "Security policy requires code review"
# --- Context Snapshot Capture ---
Scenario: Context snapshot captures hot context hash
When strat context data containing "plan_id" "01JQAAAAAAAAAAAAAAAAAAAA01"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision context snapshot hash should start with "sha256:"
Scenario: Context snapshot captures actor state reference
When strat actor state containing "step" "1"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision actor state ref should start with "checkpoint:"
Scenario: Context snapshot captures relevant resources
When three strat relevant resources "res1" and "res2" and "res3"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision should have 3 relevant resources
And each strat resource should have a valid resource_id
# --- Error Handling ---
Scenario: Recording with invalid plan_id raises error
Given a strategize decision hook with empty plan_id
Then a strat validation error should be raised
And the strat error should mention "plan_id"
Scenario: Recording failure logs warning and re-raises exception
When the strat decision service fails to persist
And I try to record a strategy choice that raises an exception
Then a strat warning should be logged
And the strat exception should be re-raised
# --- Parent Decision Tracking ---
Scenario: Record decision with parent decision ID
When strat parent decision ID "01PARENT000000000000000000"
And I record a strategy choice with question "Which approach?" and option "Approach A"
Then the strat decision parent_decision_id should be "01PARENT000000000000000000"
Scenario: Record multiple decisions in tree structure
When I record a strategy choice with question "Q1" and option "A1"
And I save the first strat decision
And strat parent decision ID from the first decision
And I record a strategy choice with question "Q2" and option "A2"
Then the second strat decision parent_decision_id should match the first decision
And both strat decisions should be in the same plan
@@ -25,7 +25,6 @@ Feature: TDD Issue #988 — ReactiveEventBus.emit() swallows exception details
When I emit an event that triggers the failing handler
Then the warning log should contain the exception message text
@tdd_expected_fail
Scenario: Bug #988 — emit() logs traceback via exc_info when handler raises
Given a ReactiveEventBus with a handler that raises a ValueError
When I emit an event that triggers the failing handler
@@ -0,0 +1,25 @@
@tdd_issue @tdd_issue_10414
Feature: TDD Issue #10414 — session create silently suppresses facade dispatch exceptions without logging
As a developer debugging a session creation failure
I want exceptions from _facade_dispatch() to be logged at WARNING level
So that I can diagnose why the facade layer failed without silent data loss
The ``session create`` command in
``src/cleveragents/cli/commands/session.py`` previously used
``contextlib.suppress(Exception)`` to silently discard ALL exceptions
raised by ``_facade_dispatch()``. There was no logging call before or
after the suppress block, making it impossible to diagnose failures in
the facade layer.
The fix replaces the suppress block with a ``try/except Exception`` that
calls ``_log.warning(..., exc_info=True)`` so that the exception is
recorded but the session creation remains non-fatal.
See CONTRIBUTING.md > Bug Fix Workflow > TDD Issue Test Tags.
@tdd_issue @tdd_issue_10414
Scenario: Bug #10414 — session create logs a warning when facade dispatch raises
Given a session create command with a mocked session service
And the facade dispatch is patched to raise a RuntimeError
When I invoke the session create command via the CLI runner
Then a WARNING log entry should have been emitted for the facade dispatch failure
+2 -4
View File
@@ -23,8 +23,7 @@ TDD Resource Add Succeeds Without Explicit Init
... The helper exits 0 with a sentinel when the command
... succeeds (bug is fixed), and exits 1 when the bug is
... present (OperationalError).
[Tags] tdd_issue tdd_issue_1023 tdd_issue tdd_issue_4317 tdd_expected_fail
[Tags] tdd_issue tdd_issue_1023 tdd_issue_4317 tdd_expected_fail
${result}= Run Process ${PYTHON} ${HELPER} resource-add-no-init cwd=${WORKSPACE} timeout=120s on_timeout=kill
Log ${result.stdout}
Log ${result.stderr}
@@ -38,8 +37,7 @@ TDD Project Create Succeeds Without Explicit Init
... The helper exits 0 with a sentinel when the command
... succeeds (bug is fixed), and exits 1 when the bug is
... present (OperationalError).
[Tags] tdd_issue tdd_issue_1023 tdd_issue tdd_issue_4317 tdd_expected_fail
[Tags] tdd_issue tdd_issue_1023 tdd_issue_4317 tdd_expected_fail
${result}= Run Process ${PYTHON} ${HELPER} project-create-no-init cwd=${WORKSPACE} timeout=120s on_timeout=kill
Log ${result.stdout}
Log ${result.stderr}
@@ -0,0 +1 @@
"""Application ports — protocol interfaces for external dependencies."""
@@ -0,0 +1,49 @@
"""Decision recorder port — protocol interface for recording decisions.
This module defines the ``DecisionRecorder`` protocol, which is the
shared interface used by both ``StrategizeDecisionHook`` and the future
``ExecuteDecisionHook`` to record decisions without coupling to a
concrete ``DecisionService`` implementation.
Based on:
- docs/adr/ADR-033-decision-recording-protocol.md
- Forgejo issue #8522
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Protocol, runtime_checkable
if TYPE_CHECKING:
from cleveragents.domain.models.core.decision import (
ContextSnapshot,
Decision,
DecisionType,
)
from cleveragents.domain.models.core.plan import PlanPhase
@runtime_checkable
class DecisionRecorder(Protocol):
"""Protocol for recording decisions (subset of DecisionService API).
Both ``StrategizeDecisionHook`` and the future ``ExecuteDecisionHook``
depend on this protocol rather than the concrete ``DecisionService``,
keeping the hooks decoupled from the persistence layer.
"""
def record_decision(
self,
plan_id: str,
decision_type: DecisionType | str,
question: str,
chosen_option: str,
*,
parent_decision_id: str | None = None,
alternatives_considered: list[str] | None = None,
confidence_score: float | None = None,
rationale: str = "",
actor_reasoning: str | None = None,
context_snapshot: ContextSnapshot | None = None,
plan_phase: PlanPhase | str | None = None,
) -> Decision: ...
@@ -0,0 +1,66 @@
"""Decision context snapshot utility.
Provides the ``capture_context_snapshot`` function for capturing a
context snapshot at decision time. This utility is shared between
``StrategizeDecisionHook`` and the future ``ExecuteDecisionHook``.
Based on:
- docs/specification.md §Strategize-Phase Recording Loop
- docs/adr/ADR-033-decision-recording-protocol.md
- Forgejo issue #8522
"""
from __future__ import annotations
import hashlib
import json
from typing import Any
from cleveragents.domain.models.core.decision import (
ContextSnapshot,
ResourceRef,
)
def capture_context_snapshot(
context_data: dict[str, Any] | None = None,
actor_state: dict[str, Any] | None = None,
relevant_resources: list[str] | None = None,
) -> ContextSnapshot:
"""Capture a context snapshot at decision time.
Automatically generates:
- ``hot_context_hash``: SHA256 hash of the context data
- ``hot_context_ref``: Abbreviated storage reference
- ``relevant_resources``: List of resource references
- ``actor_state_ref``: Reference to actor state checkpoint
Args:
context_data: Current context window contents (dict).
actor_state: Actor's current state (dict).
relevant_resources: List of resource IDs that influenced the decision.
Returns:
A :class:`~cleveragents.domain.models.core.decision.ContextSnapshot`
with auto-captured fields.
"""
# Generate hot context hash
context_json = json.dumps(context_data or {}, sort_keys=True, default=str)
hot_context_hash = f"sha256:{hashlib.sha256(context_json.encode()).hexdigest()}"
# Generate actor state reference (placeholder for LangGraph checkpoint)
actor_state_json = json.dumps(actor_state or {}, sort_keys=True, default=str)
actor_state_ref = (
f"checkpoint:{hashlib.sha256(actor_state_json.encode()).hexdigest()[:16]}"
)
# Convert resource IDs to ResourceRef objects
resource_refs = [ResourceRef(resource_id=rid) for rid in (relevant_resources or [])]
return ContextSnapshot(
hot_context_hash=hot_context_hash,
hot_context_ref=f"context:{hot_context_hash[7:23]}", # Abbreviated ref
relevant_resources=resource_refs,
actor_state_ref=actor_state_ref,
)
@@ -51,6 +51,8 @@ Based on ``docs/specification.md`` and implementation plan Stage A3.
from __future__ import annotations
import hashlib
import json
from contextlib import suppress
from datetime import datetime
from typing import TYPE_CHECKING, Any
@@ -77,7 +79,7 @@ from cleveragents.domain.models.core.automation_profile import (
BUILTIN_PROFILES,
AutomationProfile,
)
from cleveragents.domain.models.core.decision import DecisionType
from cleveragents.domain.models.core.decision import ContextSnapshot, DecisionType
from cleveragents.domain.models.core.plan import (
AutomationProfileProvenance,
AutomationProfileRef,
@@ -265,11 +267,23 @@ class PlanLifecycleService:
question: str,
chosen_option: str,
parent_decision_id: str | None = None,
context_snapshot: ContextSnapshot | None = None,
) -> None:
"""Record a decision if DecisionService is available.
Failures are logged but never propagated decision recording
must not block lifecycle transitions.
Args:
plan_id: ULID of the plan.
decision_type: Type of decision being recorded.
question: What question was being answered.
chosen_option: The option that was chosen.
parent_decision_id: Optional parent in the decision tree.
context_snapshot: Optional full context snapshot. When
provided, it is forwarded to
:meth:`DecisionService.record_decision` so that the
decision is stored with a complete context snapshot
"""
if self.decision_service is None:
return
@@ -281,6 +295,7 @@ class PlanLifecycleService:
question=question,
chosen_option=chosen_option,
parent_decision_id=parent_decision_id,
context_snapshot=context_snapshot,
)
except Exception:
self._logger.warning(
@@ -1398,15 +1413,72 @@ class PlanLifecycleService:
self._commit_plan(plan)
self._logger.info("Strategize started", plan_id=plan_id)
context_snapshot = self._build_strategize_context_snapshot(plan)
self._try_record_decision(
plan_id=plan_id,
decision_type="strategy_choice",
question="Which strategy should the plan follow?",
chosen_option=f"Begin strategize phase for plan {plan_id}",
context_snapshot=context_snapshot,
)
return plan
def _build_strategize_context_snapshot(self, plan: Plan) -> ContextSnapshot:
"""Build a full context snapshot for a Strategize-phase decision.
Captures the plan description, action name, strategy actor, and
project references as the hot context window. The hash is
computed over the serialised context so that identical context
windows produce the same hash (content-addressable).
The ``hot_context_ref`` is set to a stable ``plan:<plan_id>``
URI so that callers can locate the full context via the plan
record. ``relevant_resources`` is populated from the plan's
project links. ``actor_state_ref`` is set to the strategy
actor name when available.
Per the v3.2.0 acceptance criteria, decisions recorded during
the Strategize phase must include full context snapshots with
all four :class:`ContextSnapshot` fields populated.
Args:
plan: The plan entering the Strategize phase.
Returns:
A :class:`ContextSnapshot` with all four fields populated.
"""
from cleveragents.domain.models.core.decision import ResourceRef
plan_id = plan.identity.plan_id
# Build the hot context window from plan metadata available at
# the start of the Strategize phase.
hot_context: dict[str, object] = {
"plan_id": plan_id,
"action_name": plan.action_name,
"description": plan.description or "",
"strategy_actor": plan.strategy_actor or "",
"projects": [pl.project_name for pl in plan.project_links],
}
context_json = json.dumps(hot_context, sort_keys=True)
context_hash = hashlib.sha256(context_json.encode()).hexdigest()
# Build resource refs from project links so the snapshot records
# which projects influenced the strategy decision.
relevant_resources = [
ResourceRef(resource_id=pl.project_name)
for pl in plan.project_links
if pl.project_name
]
return ContextSnapshot(
hot_context_hash=f"sha256:{context_hash}",
hot_context_ref=f"plan:{plan_id}",
relevant_resources=relevant_resources,
actor_state_ref=plan.strategy_actor or "",
)
def complete_strategize(self, plan_id: str) -> Plan:
"""Complete the Strategize phase.
@@ -0,0 +1,378 @@
"""Decision recording hook for the Strategize phase.
This module provides the ``StrategizeDecisionHook`` class, which integrates
decision recording into the Strategize phase of plan execution. The hook
captures every decision point during strategy decomposition, including:
- The question being answered
- The chosen option
- Alternatives considered
- Confidence score
- Rationale
- Full context snapshot (hot context hash, actor state reference, relevant resources)
The hook is designed to be called by the strategy actor during the Strategize
phase, recording decisions atomically with plan updates.
Based on:
- docs/specification.md §Strategize-Phase Recording Loop
- docs/adr/ADR-033-decision-recording-protocol.md
- Forgejo issue #8522
"""
from __future__ import annotations
from typing import Any
import structlog
from cleveragents.application.ports.decision_recorder import DecisionRecorder
from cleveragents.application.services.decision_context import capture_context_snapshot
from cleveragents.core.exceptions import ValidationError
from cleveragents.domain.models.core.decision import (
Decision,
DecisionType,
)
from cleveragents.domain.models.core.plan import PlanPhase
logger = structlog.get_logger(__name__)
# ---------------------------------------------------------------------------
# Strategize decision hook
# ---------------------------------------------------------------------------
class StrategizeDecisionHook:
"""Hook for recording decisions during the Strategize phase.
Integrates with the strategy actor to capture every decision point,
including the question, chosen option, alternatives, confidence,
rationale, and full context snapshot.
The hook is designed to be called by the strategy actor during
Strategize, and records decisions atomically with plan updates.
Attributes:
decision_service: The
:class:`~cleveragents.application.ports.decision_recorder.DecisionRecorder`
instance for persisting decisions.
plan_id: ULID of the plan being strategized.
parent_decision_id: Optional parent decision ID for tree structure.
"""
def __init__(
self,
decision_service: DecisionRecorder,
plan_id: str,
parent_decision_id: str | None = None,
) -> None:
"""Initialize the Strategize decision hook.
Args:
decision_service: DecisionRecorder for recording decisions.
plan_id: ULID of the plan.
parent_decision_id: Optional parent decision ID.
Raises:
ValidationError: If plan_id is empty.
"""
if not plan_id or not plan_id.strip():
raise ValidationError("plan_id must not be empty")
self.decision_service = decision_service
self.plan_id = plan_id
self.parent_decision_id = parent_decision_id
self._logger = logger.bind(
hook="strategize_decision",
plan_id=plan_id,
)
def record_strategy_choice(
self,
question: str,
chosen_option: str,
alternatives_considered: list[str] | None = None,
confidence_score: float | None = None,
rationale: str = "",
context_data: dict[str, Any] | None = None,
actor_state: dict[str, Any] | None = None,
relevant_resources: list[str] | None = None,
) -> Decision:
"""Record a strategy choice decision during Strategize.
Args:
question: What strategic question was being answered.
chosen_option: The chosen approach.
alternatives_considered: Other approaches evaluated.
confidence_score: Confidence in the choice (0.0-1.0).
rationale: Why this option was chosen.
context_data: Current context window contents.
actor_state: Actor's current state.
relevant_resources: Resource IDs that influenced the decision.
Returns:
The recorded Decision.
Raises:
ValidationError: If required fields are missing.
"""
if not question or not question.strip():
raise ValidationError("question must not be empty")
if not chosen_option or not chosen_option.strip():
raise ValidationError("chosen_option must not be empty")
snapshot = capture_context_snapshot(
context_data=context_data,
actor_state=actor_state,
relevant_resources=relevant_resources,
)
self._logger.info(
"Recording strategy choice decision",
question=question,
chosen_option=chosen_option,
confidence=confidence_score,
)
try:
decision = self.decision_service.record_decision(
plan_id=self.plan_id,
decision_type=DecisionType.STRATEGY_CHOICE,
question=question,
chosen_option=chosen_option,
parent_decision_id=self.parent_decision_id,
alternatives_considered=alternatives_considered,
confidence_score=confidence_score,
rationale=rationale,
context_snapshot=snapshot,
plan_phase=PlanPhase.STRATEGIZE,
)
self._logger.debug(
"Strategy choice decision recorded",
decision_id=decision.decision_id,
)
return decision
except Exception as exc:
self._logger.warning(
"Failed to record strategy choice decision",
error=str(exc),
error_type=type(exc).__name__,
)
raise
def record_resource_selection(
self,
question: str,
chosen_option: str,
alternatives_considered: list[str] | None = None,
confidence_score: float | None = None,
rationale: str = "",
context_data: dict[str, Any] | None = None,
actor_state: dict[str, Any] | None = None,
relevant_resources: list[str] | None = None,
) -> Decision:
"""Record a resource selection decision during Strategize.
Args:
question: What resources should be selected.
chosen_option: The selected resources.
alternatives_considered: Other resource selections evaluated.
confidence_score: Confidence in the selection (0.0-1.0).
rationale: Why these resources were selected.
context_data: Current context window contents.
actor_state: Actor's current state.
relevant_resources: Resource IDs that influenced the decision.
Returns:
The recorded Decision.
Raises:
ValidationError: If required fields are missing.
"""
if not question or not question.strip():
raise ValidationError("question must not be empty")
if not chosen_option or not chosen_option.strip():
raise ValidationError("chosen_option must not be empty")
snapshot = capture_context_snapshot(
context_data=context_data,
actor_state=actor_state,
relevant_resources=relevant_resources,
)
self._logger.info(
"Recording resource selection decision",
question=question,
chosen_option=chosen_option,
)
try:
decision = self.decision_service.record_decision(
plan_id=self.plan_id,
decision_type=DecisionType.RESOURCE_SELECTION,
question=question,
chosen_option=chosen_option,
parent_decision_id=self.parent_decision_id,
alternatives_considered=alternatives_considered,
confidence_score=confidence_score,
rationale=rationale,
context_snapshot=snapshot,
plan_phase=PlanPhase.STRATEGIZE,
)
self._logger.debug(
"Resource selection decision recorded",
decision_id=decision.decision_id,
)
return decision
except Exception as exc:
self._logger.warning(
"Failed to record resource selection decision",
error=str(exc),
error_type=type(exc).__name__,
)
raise
def record_subplan_spawn(
self,
question: str,
chosen_option: str,
alternatives_considered: list[str] | None = None,
confidence_score: float | None = None,
rationale: str = "",
context_data: dict[str, Any] | None = None,
actor_state: dict[str, Any] | None = None,
relevant_resources: list[str] | None = None,
) -> Decision:
"""Record a subplan spawn decision during Strategize.
Args:
question: Why is a subplan being spawned.
chosen_option: The subplan goal/description.
alternatives_considered: Other decomposition approaches.
confidence_score: Confidence in the decomposition (0.0-1.0).
rationale: Why this decomposition was chosen.
context_data: Current context window contents.
actor_state: Actor's current state.
relevant_resources: Resource IDs that influenced the decision.
Returns:
The recorded Decision.
Raises:
ValidationError: If required fields are missing.
"""
if not question or not question.strip():
raise ValidationError("question must not be empty")
if not chosen_option or not chosen_option.strip():
raise ValidationError("chosen_option must not be empty")
snapshot = capture_context_snapshot(
context_data=context_data,
actor_state=actor_state,
relevant_resources=relevant_resources,
)
self._logger.info(
"Recording subplan spawn decision",
question=question,
chosen_option=chosen_option,
)
try:
decision = self.decision_service.record_decision(
plan_id=self.plan_id,
decision_type=DecisionType.SUBPLAN_SPAWN,
question=question,
chosen_option=chosen_option,
parent_decision_id=self.parent_decision_id,
alternatives_considered=alternatives_considered,
confidence_score=confidence_score,
rationale=rationale,
context_snapshot=snapshot,
plan_phase=PlanPhase.STRATEGIZE,
)
self._logger.debug(
"Subplan spawn decision recorded",
decision_id=decision.decision_id,
)
return decision
except Exception as exc:
self._logger.warning(
"Failed to record subplan spawn decision",
error=str(exc),
error_type=type(exc).__name__,
)
raise
def record_invariant_enforced(
self,
question: str,
chosen_option: str,
alternatives_considered: list[str] | None = None,
confidence_score: float | None = None,
rationale: str = "",
context_data: dict[str, Any] | None = None,
actor_state: dict[str, Any] | None = None,
relevant_resources: list[str] | None = None,
) -> Decision:
"""Record an invariant enforcement decision during Strategize.
Args:
question: What invariant is being enforced.
chosen_option: How the invariant is being enforced.
alternatives_considered: Other enforcement approaches evaluated.
confidence_score: Confidence in the enforcement approach (0.0-1.0).
rationale: Why this enforcement approach was chosen.
context_data: Current context window contents.
actor_state: Actor's current state.
relevant_resources: Resource IDs that influenced the decision.
Returns:
The recorded Decision.
Raises:
ValidationError: If required fields are missing.
"""
if not question or not question.strip():
raise ValidationError("question must not be empty")
if not chosen_option or not chosen_option.strip():
raise ValidationError("chosen_option must not be empty")
snapshot = capture_context_snapshot(
context_data=context_data,
actor_state=actor_state,
relevant_resources=relevant_resources,
)
self._logger.info(
"Recording invariant enforced decision",
question=question,
chosen_option=chosen_option,
)
try:
decision = self.decision_service.record_decision(
plan_id=self.plan_id,
decision_type=DecisionType.INVARIANT_ENFORCED,
question=question,
chosen_option=chosen_option,
parent_decision_id=self.parent_decision_id,
alternatives_considered=alternatives_considered,
confidence_score=confidence_score,
rationale=rationale,
context_snapshot=snapshot,
plan_phase=PlanPhase.STRATEGIZE,
)
self._logger.debug(
"Invariant enforced decision recorded",
decision_id=decision.decision_id,
)
return decision
except Exception as exc:
self._logger.warning(
"Failed to record invariant enforced decision",
error=str(exc),
error_type=type(exc).__name__,
)
raise
+20 -10
View File
@@ -534,12 +534,13 @@ def _print_role_warnings(config_blob: dict[str, Any]) -> None:
@app.command()
def add(
name: Annotated[
str,
str | None,
typer.Argument(
help="Namespaced actor name (e.g. local/my-actor)",
help="Namespaced actor name (e.g. local/my-actor). "
"If omitted, the name is derived from the config file.",
metavar="NAME",
),
],
] = None,
config: Annotated[
Path | None,
typer.Option("--config", "-c", help="Path to JSON/YAML actor config"),
@@ -569,19 +570,19 @@ def add(
) -> None:
"""Add a new actor configuration.
The actor name is provided as a required positional argument. The YAML/JSON
configuration file specified with ``--config`` supplies the actor settings.
The positional ``NAME`` takes precedence over any ``name`` field in the
config file.
The YAML/JSON configuration file specified with ``--config`` supplies the
actor settings. The actor's registered name is taken from the config file's
``name`` field, unless overridden by the positional ``NAME`` argument.
Signature:
``agents actor add <NAME> --config <FILE> [--update] [--unsafe]
``agents actor add [--config|-c <FILE>] [<NAME>] [--update] [--unsafe]
[--set-default] [--option key=value] [--format FORMAT]``
Examples:
agents actor add --config ./actors/my-actor.yaml
agents actor add local/my-actor --config ./actors/my-actor.yaml
agents actor add local/my-actor --config ./actors/my-actor.yaml --update
agents actor add local/my-actor --config actor.yaml --format json
agents actor add --config ./actors/my-actor.yaml --update
agents actor add --config actor.yaml --format json
"""
service, registry = _get_services()
option_overrides = _parse_option_overrides(option)
@@ -597,6 +598,15 @@ def add(
assert loaded is not None, "unreachable: config is not None"
yaml_text, config_blob = loaded
# Derive actor name from config when not provided as argument.
if name is None:
name = config_blob.get("name")
if not name:
raise typer.BadParameter(
"Actor name is required. Provide it as a positional argument "
"or as a 'name' field in the config file."
)
# Validate v3 config via ActorConfigSchema if detected.
# This ensures v3 actors are fully validated (cycle detection, required
# fields, enum values) before the registry stores them.
+7 -3
View File
@@ -213,13 +213,17 @@ def create(
# Notify the facade layer for A2A protocol bookkeeping.
# Pass session_id so the facade handler acknowledges the already-
# persisted session instead of creating a duplicate (#1141).
import contextlib
with contextlib.suppress(Exception):
try:
_facade_dispatch(
"session.create",
{"actor_name": actor or "", "session_id": session.session_id},
)
except Exception as _exc:
_log.warning(
"session_create_facade_dispatch_failed",
extra={"session_id": session.session_id, "error": str(_exc)},
exc_info=True,
)
data = _session_summary_dict(session)
@@ -135,6 +135,7 @@ class ReactiveEventBus:
handler=getattr(handler, "__qualname__", repr(handler)),
error_type=type(exc).__name__,
error=str(exc),
exc_info=True,
)
def subscribe(
@@ -241,10 +241,9 @@ class PluginLoader:
def validate_protocol(klass: type[Any], protocol: type[Any]) -> bool:
"""Check whether *klass* satisfies a ``@runtime_checkable`` Protocol.
Creates a temporary instance of *klass* (using a no-arg
constructor) and checks it against the protocol via
``isinstance``. If instantiation fails, falls back to a
structural check using ``issubclass``.
Uses structural type checking via ``issubclass`` to validate the
protocol without instantiating the class. This prevents arbitrary
code execution from untrusted plugin constructors during validation.
Args:
klass: The class to validate.
@@ -256,23 +255,84 @@ class PluginLoader:
Raises:
ProtocolMismatchError: If *klass* does not satisfy the protocol.
"""
# Try instance check first (most reliable for runtime_checkable)
# Try structural check first (safe, no instantiation).
# This prevents arbitrary code execution from plugin constructors.
issubclass_raised_type_error = False
try:
instance = klass()
if isinstance(instance, protocol):
if issubclass(klass, protocol):
return True
except Exception:
# If instantiation fails, try subclass check
try:
if issubclass(klass, protocol):
return True
except TypeError:
pass
except TypeError:
# issubclass raises TypeError if protocol is not a valid
# @runtime_checkable Protocol (e.g. has a custom metaclass that
# overrides __subclasscheck__). Fall through to structural check.
issubclass_raised_type_error = True
# Fallback: perform a conservative structural check against the
# Protocol definition without instantiating the class. We inspect
# the protocol's declared members (callables and annotated names)
# and ensure the candidate class exposes them as class-level
# attributes or descriptors. This approximates structural
# conformance while avoiding constructor execution.
required: set[str] = set()
# Collect names from protocol __dict__ (methods, properties, etc.)
prot_dict = getattr(protocol, "__dict__", {})
for name, _value in prot_dict.items():
if name.startswith("__"):
continue
# Methods and descriptors will appear in the dict; annotations
# are handled below. We treat any non-data-magic name as a
# required member.
required.add(name)
# Also include annotated names (those declared with type hints)
ann = getattr(protocol, "__annotations__", {}) or {}
for name in ann:
if name.startswith("__"):
continue
required.add(name)
# If issubclass raised TypeError and the protocol declares no
# inspectable members, we cannot validate conformance — treat this
# as a mismatch rather than silently returning True for an
# unverifiable protocol.
if issubclass_raised_type_error and not required:
msg = (
f"Class '{getattr(klass, '__name__', str(klass))}' cannot be "
f"validated against protocol "
f"'{getattr(protocol, '__name__', str(protocol))}': "
f"issubclass raised TypeError and the protocol declares no "
f"inspectable members. Ensure the protocol is a valid "
f"@runtime_checkable Protocol."
)
raise ProtocolMismatchError(msg)
# Validate that the candidate class exposes each required name.
missing: list[str] = []
for name in sorted(required):
# getattr on the class returns descriptors (functions, property,
# staticmethod, etc.) without instantiation. hasattr is safe.
if not hasattr(klass, name):
missing.append(name)
continue
# If the protocol declared a callable (method), ensure the
# attribute on the class is callable or a descriptor.
prot_val = prot_dict.get(name)
if callable(prot_val):
cand = getattr(klass, name)
# Properties are descriptors and typically not callable; we
# accept them as present. For methods, require callable.
if not (callable(cand) or hasattr(cand, "fget")):
missing.append(name)
if not missing:
return True
msg = (
f"Class '{klass.__name__}' does not satisfy protocol "
f"'{protocol.__name__}'. Ensure the class implements all "
f"required methods and attributes."
f"Class '{getattr(klass, '__name__', str(klass))}' does not "
f"satisfy protocol '{getattr(protocol, '__name__', str(protocol))}'. "
f"Missing attributes: {', '.join(missing)}. Ensure the class "
f"implements all required methods and descriptors."
)
raise ProtocolMismatchError(msg)
@@ -37,6 +37,7 @@ from cleveragents.resource.handlers._base import (
EMPTY_CONTENT_HASH,
BaseResourceHandler,
)
from cleveragents.resource.handlers.discovery import discover_devcontainers
from cleveragents.resource.handlers.protocol import (
CheckpointResult,
Content,
@@ -233,10 +234,13 @@ class FsDirectoryHandler(BaseResourceHandler):
)
def discover_children(self, *, resource: Resource) -> list[Resource]:
"""Discover subdirectories as child resources.
"""Discover subdirectories and devcontainer instances as child resources.
Each immediate subdirectory becomes a child ``fs-directory``
resource.
resource. Additionally, any ``.devcontainer/`` configurations
found at the resource location are registered as
``devcontainer-instance`` child resources with
``provisioning_state: discovered``.
Args:
resource: The parent fs-directory resource.
@@ -261,6 +265,28 @@ class FsDirectoryHandler(BaseResourceHandler):
)
children.append(child)
# Wire devcontainer auto-discovery (issue #4740)
dc_results = discover_devcontainers(location, "fs-directory")
for dc_result in dc_results:
config_name = dc_result.config_name or "default"
dc_child = Resource(
resource_id=self._derive_child_id(
resource.resource_id, f"devcontainer-{config_name}"
),
name=f"devcontainer-{config_name}",
resource_type_name="devcontainer-instance",
classification=resource.classification,
description=f"Devcontainer at {dc_result.config_path}",
location=location,
parents=[resource.resource_id],
properties={
"devcontainer_json_path": str(dc_result.config_path),
"config_name": config_name,
"provisioning_state": "discovered",
},
)
children.append(dc_child)
return children
# -- Checkpoint and rollback (issue #836) ------------------------------
@@ -6,17 +6,18 @@ strategy (with fallback to ``copy_on_write``).
Content CRUD operations (issue #827):
- ``read`` ``git show HEAD:<path>``
- ``write`` atomic file write inside the checkout
- ``delete`` ``os.remove`` + ``git rm --cached``
- ``list_children`` ``git ls-tree -r --name-only HEAD``
- ``diff`` ``git diff --no-index``
- ``discover_children`` ``git ls-tree --name-only HEAD``
- ``read`` -- ``git show HEAD:<path>``
- ``write`` -- atomic file write inside the checkout
- ``delete`` -- ``os.remove`` + ``git rm --cached``
- ``list_children`` -- ``git ls-tree -r --name-only HEAD``
- ``diff`` -- ``git diff --no-index``
- ``discover_children`` -- ``git ls-tree --name-only HEAD`` + devcontainer discovery
Based on:
- implementation_plan.md group M1.resource-handlers (L2254-L2271)
- Built-in type definition in resource_registry_service.py L62-98
- Issue #827 ResourceHandler CRUD and discovery methods
- Issue #827 -- ResourceHandler CRUD and discovery methods
- Issue #4740 -- Wire discover_devcontainers() into git-checkout handler
"""
from __future__ import annotations
@@ -36,6 +37,7 @@ from cleveragents.resource.handlers._base import (
EMPTY_CONTENT_HASH,
BaseResourceHandler,
)
from cleveragents.resource.handlers.discovery import discover_devcontainers
from cleveragents.resource.handlers.protocol import (
CheckpointResult,
Content,
@@ -303,17 +305,20 @@ class GitCheckoutHandler(BaseResourceHandler):
)
def discover_children(self, *, resource: Resource) -> list[Resource]:
"""Discover child resources via ``git ls-tree``.
"""Discover child resources via ``git ls-tree`` and devcontainer scan.
Each top-level directory in the repo becomes a child resource
of type ``fs-directory``.
of type ``fs-directory``. Additionally, any ``.devcontainer/``
configurations found at the resource location are registered as
``devcontainer-instance`` child resources with
``provisioning_state: discovered``.
Args:
resource: The parent git-checkout resource.
Returns:
List of child :class:`Resource` objects for top-level
directories.
directories and discovered devcontainer instances.
"""
location = self._require_location(resource)
@@ -346,6 +351,28 @@ class GitCheckoutHandler(BaseResourceHandler):
)
children.append(child)
# Wire devcontainer auto-discovery (issue #4740)
dc_results = discover_devcontainers(location, "git-checkout")
for dc_result in dc_results:
config_name = dc_result.config_name or "default"
dc_child = Resource(
resource_id=self._derive_child_id(
resource.resource_id, f"devcontainer-{config_name}"
),
name=f"devcontainer-{config_name}",
resource_type_name="devcontainer-instance",
classification=resource.classification,
description=f"Devcontainer at {dc_result.config_path}",
location=location,
parents=[resource.resource_id],
properties={
"devcontainer_json_path": str(dc_result.config_path),
"config_name": config_name,
"provisioning_state": "discovered",
},
)
children.append(dc_child)
return children
# -- Checkpoint and rollback (issue #836) ------------------------------
+1 -2
View File
@@ -17,13 +17,12 @@ import yaml
from cleveragents.actor.registry import ActorRegistry
from cleveragents.actor.schema import ActorConfigSchema, ActorType, is_v3_yaml
from cleveragents.config.settings import ProviderDefaults, Settings
from cleveragents.config.settings import ProviderDefaults
from cleveragents.core.exceptions import NotFoundError
from cleveragents.domain.models.core.actor import Actor
from cleveragents.providers.registry import (
ProviderCapabilities,
ProviderInfo,
ProviderRegistry,
ProviderType,
)