Files
cleveragents-core/CHANGELOG.md
T
HAL9000 262087ca3e feat(context): implement semantic context search strategy using embeddings
Fix ruff format lint on plugin.py by removing the out-of-scope stub and
its main.py registration. Fix bandit B324 security finding by annotating
the MockEmbeddingProvider MD5 call with usedforsecurity=False. Add
CHANGELOG entry under [Unreleased].

ISSUES CLOSED: #5254
2026-06-18 01:17:23 -04:00

1403 lines
106 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Changelog
All notable changes to this project will be documented in this file.
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
Changed `wf10_batch.robot` to be less likely to create files, and
`plan_generation_graph.robot` to give more test answers.
## [Unreleased]
- **docs(spec): fix checkpoint config key path and trigger name defaults** (#5009 / PR #5163): Corrects the Configuration Reference table entry `sandbox.checkpoint.auto-create-on``core.checkpoints.auto-create-on`, matching the implementation in `config_service.py`. Aligns the default trigger-name values (`before_tool_execute`, `after_tool_execute`) with the implementation in `tool/runner.py`, resolving specimplementation discrepancies identified in issue #5009.
- **docs: module guides for Sandbox & Checkpoint, Correction Attempts, and Invariant Reconciliation** (#4848): Added three comprehensive module guides covering purpose, core classes, lifecycle diagrams, exception hierarchies, CLI usage, and ADR links for `SandboxManager`, `CorrectionAttemptManager`, and `InvariantReconciliationActor`. Includes security callouts for `NoSandbox` bypass (permanent writes, no rollback), `guidance` prompt-injection risk, `archived_artifacts_path` provenance, and `non_overridable` global invariant access control.
- **feat(context): PriorityContextStrategy** (#9997 / PR #10772): Implements a priority-based context strategy that ranks context fragments by configurable priority scores — default role-based rules (system > tool > user > assistant), exponential recency decay (7-day half-life), and explicit priority tag boost. Supports custom scoring function injection and custom PriorityRule list injection. Registered in the ACMS pipeline under key `priority_context`. `PriorityRule` uses Pydantic `BaseModel` for architecture conformance. Includes 18 BDD scenarios covering all acceptance criteria.
- **docs(timeline): verify timeline status for 2026-04-16 Cycle 2** (#8519): Updated `docs/timeline.md` with Days 104-106 Cycle 2 milestone snapshot. No changes detected since Cycle 1. All M3-M7 milestones remain overdue. Timeline verification performed by AUTO-TIME-3 supervisor agent. Refs #8519.
- **feat(context): implement semantic context search strategy using embeddings** (#5254): Added `EmbeddingProvider` abstract base class and `SimpleWordEmbeddingProvider` (deterministic word-vector implementation) to `application/services/embedding_provider.py`. Introduced `SemanticContextSearchStrategy` that retrieves context chunks by cosine-similarity scoring against query embeddings. Includes BDD scenarios covering embedding generation, dimension constraints, cosine similarity computation, vocabulary overflow warnings, and semantic context ranking.
- **Virtual Resource Type Base Class** (#8610): Implemented `VirtualResource` base class with two example concrete implementations (`MetricResource`, `APIEndpointResource`) for abstract/computed resources that are derived rather than mapped to physical files. Virtual resources are computed on demand via a `compute_fn` callable. Includes Behave BDD scenarios in `features/resource_virtual_types.feature` exercising construction, computation, name validation, kwargs passthrough, exception handling, string representation, and subclassing. Resource names are validated against `^[a-zA-Z][a-zA-Z0-9_-]*$` (must start with a letter; alphanumeric, hyphens, and underscores otherwise).
- **test(e2e): restore complete M2 acceptance test** (#11191): Restored the truncated M2 full actor compiler and LLM integration e2e acceptance test to its complete 10-step form. Added dynamic LLM provider selection via `Resolve LLM Actor` (falls back to Anthropic when OpenAI is unavailable or quota-exhausted), replacing hardcoded `gpt-4` / `openai/gpt-4` references in the actor config and action YAML. Added explicit return-code validation (`Should Be Equal As Integers ${r_actor.rc} 0`) for the actor registration step.
- **docs(a2a): ACP to A2A migration guide** (#10230): Added migration guide documenting how to upgrade from the ACP module to the A2A module introduced in v3.6.0, including symbol renames, field renames, operation-name mappings, and YAML configuration updates.
- **Plan Prompt JSON Timing Field** (#9353): `agents plan prompt --format json` now
includes `timing.started` as an ISO 8601 UTC timestamp in the JSON envelope,
matching the spec (§CLI Commands — `agents plan prompt`). Extended
`cleveragents.cli.formatting.format_output` (and `_build_envelope`) with an
optional `started_at: datetime` parameter; when provided, the envelope's
`timing` dict includes a `started` field alongside `duration_ms`. Refactored
`prompt_plan_cmd` to delegate envelope construction to `format_output` so the
envelope keys (`command`, `status`, `data`, `timing.started`, `messages`) are
populated correctly at the JSON root rather than nested under a synthetic
inner `data` field.
- **fix(plan): NamespacedName digit-start validation** (#2145, #2147): `NamespacedName` field validators now reject `namespace` and `name` components whose first character is a digit, raising `pydantic.ValidationError` with message `"must start with a letter"`. BDD constructor scenarios updated to use the `"a Pydantic ValidationError should be raised"` step so the assertion correctly matches the exception type raised by Pydantic model construction.
- **fix(cli/plan): plan correct JSON output envelope fix and BDD test coverage** (#8584 / PR #8662): Restructured `agents plan correct --format json` output to nest correction fields under `data.correction` (e.g., `data.correction.mode`) and populate the spec-required CLI envelope with `command="plan correct"`, `status`, `exit_code`, `timing`, and `messages` fields. Added three BDD scenarios in `features/tdd_plan_correct_json_output.feature` validating the envelope structure for both revert and append modes.
- **fix(cli): add --url flag to resource add for git resource type** (#6322): Added support for the `--url` flag on `agents resource add git` command, allowing users to specify a remote URL for git resources. The flag is validated to only apply to git resource types. Includes Behave BDD tests in `features/resource_cli_git_url_flag.feature` and Robot Framework integration tests verifying correct URL validation and CLI behavior.
- **Session create JSON envelope** (#6441): Fixed `agents session create --format json` returning a flat `data` dict instead of the spec-required nested structure with `data.session`, `data.settings`, and `data.actor_details` sub-objects. The `command` field is now populated correctly. Extended JSON envelope coverage to `agents session list`, `show`, `delete --format json`, `export --output-format json`, and `import --format json` so all session commands emit a structured `messages[].text` field (`"0 sessions listed"`, `"Session details loaded"`, `"Session deleted"`, `"Export completed"`, `"Import completed"`).
- **feat(tui): conversation content pruning** (#6350): Added `ConversationStream` to the TUI layer implementing hysteresis-based line-count pruning. When the rendered conversation exceeds `trigger_line_count` (`prune_low_mark + prune_excess`, defaults 1 500 + 1 000 = 2 500 lines), the oldest non-protected blocks are removed until total lines falls back to `prune_low_mark`. A styled pruning note is inserted at the head of the visible conversation. Pruning thresholds are configurable via `~/.config/cleveragents/tui-settings.json` (`ui.prune_low_mark`, `ui.prune_excess`). Includes Behave BDD tests, Robot Framework integration tests, and ASV performance benchmarks.
- **fix(resources): remove unsupported executable resource type and fix resource list columns** (#3077 / PR #3248): Removed `executable` from `LSP_RESOURCE_TYPES` and `BUILTIN_TYPE_NAMES` (the specification defines no such built-in type). Updated `agents resource list` CLI table columns from `[ID, Name, Type, Status, Kind, Location, Description]` to the spec-required `[Name, ID, Type, Phys/Virt, Children, Projects]`. Deleted orphaned `examples/resource-types/executable.yaml`. Lifecycle state for container resources is now displayed as a note below the resource table.
- **fix(cli): add Read-Only and Writes columns to tool list output** (#1476): Rewrote
`list_tools()` in `src/cleveragents/cli/commands/tool.py` to render exactly the 5
spec-required columns (Name, Type, Source, Read-Only, Writes), removed the legacy
Description and Timeout columns, computes read-only/writes status from `capability`
metadata (rendering `✓`/`—`), and added a Summary panel showing Total, Tools,
Validations, Read-Only, Writes, and Namespaces counts. Adds Behave BDD tests in
`features/tool_cli.feature` verifying correct column names, capability rendering, and
Summary panel presence.
- **Fix actor compiler to read LSP bindings from typed `lsp_binding` field** (#1488): Fixed
`_extract_lsp_bindings()` in `actor/compiler.py` to read from `node.lsp_binding` (the typed
`NodeLspBinding` Pydantic field on `NodeDefinition`) as the primary path, with
backward-compatible fallback to the legacy `lsp_bindings` config dict key. Per-node LSP
bindings specified via the `lsp_binding:` YAML key are no longer silently dropped. Includes
Behave BDD and Robot Framework regression tests (#1432).
- **TDD regression tests for automation_profile DI bypass** (#1031): Added BDD scenarios
and Robot Framework integration tests verifying that ``_get_service()`` in
``automation_profile.py`` resolves ``AutomationProfileService`` through the DI container
rather than directly calling ``create_engine`` or ``sessionmaker``. Bug #990 was fixed by
PR #1181 before this TDD test PR merged; these tests serve as permanent regression guards
confirming the fix remains in place.
- Added team collaboration features: multi-user connection handling with
user identity tracking (TeamMember with owner/admin/member/viewer roles),
role-based access control (TeamPermission with read/write/admin/manage_members),
concurrent session support (SessionRegistry with thread-safe locking),
and optimistic-locking conflict resolution (VersionStamp with last-writer-wins,
reject, and merge strategies). Includes TeamCollaborationService orchestrating
all collaboration operations. Behave BDD tests and Robot Framework integration
tests included. (#863)
- Added FastAPI-based ASGI server endpoint served by uvicorn for the
CleverAgents server mode. Includes health check endpoint (`/health`),
A2A Agent Card discovery (`/.well-known/agent.json`), A2A JSON-RPC 2.0
routing (`/a2a`), configurable host:port binding via Settings, graceful
shutdown on SIGTERM/SIGINT, and `agents server start` CLI command.
Behave BDD tests and Robot Framework integration tests included. (#862)
- **SubplanExecutionService lazy wiring in CLI** (#10268): Wired `subplan_service`
from the DI container into `_get_plan_executor()` so `PlanExecutor` can spawn child
plans during the Execute phase. When `SubplanService` is available but
`SubplanExecutionService` is not explicitly injected, `_execute_subplans()` now
lazily creates a `SubplanExecutionService` using the parent plan's `subplan_config`
and the `_execute_child_plan` callback. Added recursion guard and strategize result
validation to `_execute_child_plan` to prevent re-entrant or orphaned child plan
execution.
- **Actor namespace/name disambiguation** (#11254): When action YAML references
actors using `namespace/name` format (e.g. `strategy_actor: local/my-strategist`),
the `_parse_actor_name()` functions no longer mistake the namespace prefix for an
LLM provider name. A new `_is_known_provider()` utility checks whether the first
slash-separated segment matches a known `ProviderType` (e.g. `openai`, `anthropic`);
if not, the input is treated as namespace/name. `PlanLifecycleService` now provides
`resolve_actor_provider_model()` to resolve namespaced references to their underlying
`provider/model` via the actor registry. All affected call sites — `StrategyActor`,
`LLMStrategizeActor`, `LLMExecuteActor`, and `SessionWorkflow._resolve_llm()` —
pre-resolve actor names before passing them to the LLM provider, fixing the
`ValueError: Unknown provider type` crash when using namespaced actor references.
Data integrity fix: ValidationAttachmentRepository argument swap (#7492): Fixed
a critical data integrity issue in `ValidationAttachmentRepository.attach` where
`validation_name` and `resource_id` arguments were being silently swapped based on a
fragile heuristic (`"/" in resource_id`). Arguments are now passed in the correct order,
ensuring data is stored with proper parameter values.
- Hardened the TDD bug-fix quality gate for issue #629: PR parsing now
requires whole-word closing keywords (avoids false positives like
"prefixes #12"), TDD bug tag discovery now uses exact token matching
(avoids `@tdd_bug_42` matching `@tdd_bug_420`), and gate evaluation now
requires expected-fail tag removal to be present in the PR diff. CI
integration now fetches the PR base branch for diff analysis and includes
`tdd_quality_gate` in `status-check` requirements for pull requests.
Review-round fixes: `check_expected_fail_removed` now uses word-boundary
matching via `_contains_tag_token` (avoids false positives on partial tag
names); diff expected-fail removal detection now tracks flags at file level
instead of per-hunk (fixes false negatives when tags span different hunks);
`parse_bug_refs` filters out issue number zero; redundant double error
reporting eliminated; regex compilation cached via `lru_cache`; nox session
no longer installs the full project (script uses stdlib only); CI checkout
uses `fetch-depth: 0` for reliable merge-base resolution.
Review-round 2 fixes: diff expected-fail removal detection now requires the
removed line to contain both the expected-fail tag and the specific bug tag
(fixes false positives when two bugs' TDD tests reside in the same file);
`check_expected_fail_removed` error messages now use the correct tag prefix
per file type (`@tdd_bug_N` for `.feature`, `tdd_bug_N` for `.robot`);
`bool` values are now rejected by bug-number validation guards; file-read
error handling in `find_tdd_tests` and `check_expected_fail_removed` now
catches `UnicodeDecodeError` (root-safe unreadable-file handling); temp
directory cleanup added to `after_scenario` hook; 8 new Behave scenarios
covering bool guards, co-located bug false-positive, `run_quality_gate`
argument validation, and `main()` CLI entry point.
Review-round 3 fixes: synthetic PR diff helper now auto-detects `.robot`
vs `.feature` file type and generates the matching diff format (fixes
under-tested robot-format diff code path); `check_expected_fail_removed`
test step now filters files by bug tag via `find_tdd_tests` before
checking (matches the production code path); `after_scenario` temp
directory cleanup no longer sets `context.temp_dir = None` (fixes
cleanup conflict with `cli_init_yes_flag_steps.py`); 2 new Behave
scenarios covering multi-line PR description parsing and non-string
`pr_diff` type guard.
- Added `TokenAuthMiddleware` and DI wiring to emit missing auth domain
events (`AUTH_SUCCESS`, `AUTH_FAILURE`) during token checks, including
spec-aligned audit details (`user_identity`, `attempted_identity`,
`ip_address`, `token_prefix`, `failure_reason`) and best-effort publish
behavior. Auth token prefixes now persist to audit logs without
redaction-key collisions, and short tokens are masked (`***...`) to
prevent full token disclosure. Added Behave coverage and Robot
audit-pipeline integration tests for auth event persistence. (#714)
- Added TDD bug-capture E2E tests for bug #1028 — ACMS indexing pipeline not
wired into CLI. Four Robot Framework E2E tests prove ContextTierService starts
empty on every CLI invocation. Tests use ``@tdd_expected_fail`` until the bug
fix is merged. (#1029)
- Added Fix-then-Revalidate orchestration loop for required validations:
bounded retry with configurable limits (0--100 per Safety Profile),
strategy revision escalation via `auto_strategy_revision` float
threshold, user escalation via `needs_user_escalation` result flag,
and domain events (`VALIDATION_FIX_ATTEMPTED`, `VALIDATION_FIX_SUCCEEDED`,
`VALIDATION_FIX_EXHAUSTED`). Validation errors are treated as required
failures regardless of mode. Includes `auto_validation_fix` threshold,
per-resource retry tracking, early-exit signalling via `None` return from
`FixCallback`, event bus circuit breaker with lock-protected failure
counter, spec-required `validation_summary` and
`final_validation_results` fields on the result model, DI container
- **fix(tui): rename ActorSelectionOverlay._render to _refresh_display (issue #11039)** — `ActorSelectionOverlay._render()` shadows Textual's `Widget._render()` which must return a `Strip`. In textual >=1.0, layout calls `get_content_height()` `self._render()` gets `None` `AttributeError: 'NoneType' object has no attribute 'get_height'`. Renamed the method to `_refresh_display()` and updated all four internal call sites (`show()`, `move_up()`, `move_down()`, `set_search()`) to use the new name.
- **Structural Component Output Validation** (#8164): Replaces exact character matching with structural component checking for output validation. Implements three validators covering plan tree output, decision CLI dicts, and structured session snapshots. The `validate_plan_tree` function validates node dicts for required keys (`decision_id`, `type`, `sequence`, `question`, `children`), ULID format, correct types, and sibling ordering. The `validate_decision_dict` function validates decision CLI output against the `Decision.as_cli_dict()` schema with field presence, type, ULID pattern, confidence range [0..1], and boolean field checks. The `validate_structured_output` function validates the StructuredOutput envelope for `command`, `session_id` (ULID), status membership, `exit_code`, and elements integrity. A unified dispatcher (`validate_structured_component_output`) enables routing by target_type. BDD test coverage added in `features/structural_validation.feature`. [Epic #8137](https://git.cleverthis.com/cleveragents/cleveragents-core/issues/8137)
- **Fixed `agents actor add --config` crash with nested `actors:` map and `config.actor` combined shorthand** (#11189): The
CLI command `agents actor add --config` now correctly parses spec-canonical YAML using the
nested `actors:` map format with `config.actor: "<provider>/<model>"` combined shorthand.
Fixed two defects in `ActorConfiguration._extract_v3_actor()`: `type` is now detected at
the actor-entry level (sibling of `config`), and `config.actor` is parsed as fallback
when separate `provider`/`model` keys are absent. Added validation to reject malformed
combined values with empty provider or model halves.
- **Fixed plan tree reporting zero decision nodes after strategize** (#10813): The `plan tree` command showed no ``decision_id`` fields even though planning completed successfully. Root cause: the PlanExecutor's ``run_strategize()`` method produced strategy decisions via the StrategyActor but never persisted them as domain ``Decision`` objects through the DecisionService wiring. Added ``decision_service`` parameter to the PlanExecutor constructor, wired it from the CLI dependency-injection container in ``_get_plan_executor()``, and added ``_persist_strategy_decisions()`` that converts each strategy decision into a domain Decision with correct type mapping (prompt_definition, strategy_choice, subplan_spawn) so they appear in plan tree output.
- **`task-implementor` posts work-started notification comments** (#11031): Both
the `issue_impl` and `pr_fix` procedures now post an informational "work
started" comment to the Forgejo issue/PR before beginning implementation.
The comment includes the issue/PR title, procedure type, and expected
duration. Posted asynchronously ("fire and move on") so it does not block
the workflow. Step numbering in both procedures has been re-numbered to
accommodate the new step.
- **`agents session tell` invokes real LLM orchestrator actor** (#5784): Replaced the
M3 echo-stub with real actor invocation via `SessionWorkflow`, routing through
`LangChainSessionCaller` → `ToolCallingRuntime.run_tool_loop()`. The user prompt
is sent to the session's bound orchestrator actor (or `--actor` override), the
assistant's response is persisted via `SessionService.append_message()`, and token
usage is tracked via `SessionService.update_token_usage()`. Output includes a
**Usage** panel (Rich/Plain) or `usage` object (JSON/YAML) with input tokens,
output tokens, estimated cost, and duration. The `--stream` flag produces real
LLM streaming output. A `SessionActorNotConfiguredError` is raised with exit
code 1 when no actor is configured.
### Added
- **ACMS Context Policy Configuration Loader and Plan Execution Integration**:
Implemented `ContextPolicyConfigurationLoader` for loading YAML/TOML policy
configurations with full schema validation, and `PlanExecutionACMSIntegration`
for wiring ACMS-assembled context into the plan execution engine. Enables
flexible, per-view context policy configuration with scope rules, priority
weights, and budget overrides. `PlanExecutor` now accepts an optional
`acms_integration` parameter (dependency injection) and passes it to
`RuntimeExecuteActor`, which uses it to assemble context via ACMS policies
before LLM calls instead of passing raw file dumps. (#9584)
- **Automated CLI Docstring Example Validation** (#9106): Added `DocstringExampleValidator`
in `src/cleveragents/cli/docstring_validator.py` that introspects Typer command signatures
and validates `Examples:` sections to ensure positional arguments appear before option
flags. Validation runs automatically via `nox -s unit_tests` through the new Behave feature
`features/cli_docstring_example_validation.feature`. Fixed `rollback_plan` docstring in
`src/cleveragents/cli/commands/plan.py` to show correct positional argument order.
CONTRIBUTING.md updated with the required CLI docstring example style guide.
### Documentation
- **`context_tier_hydrator` module documented in ACMS architecture section** (#9208): Added
a new **Context Tier Hydration** subsection to the ACMS Architecture section of the
specification (`docs/specification.md`), documenting the `context_tier_hydrator` module's
public interface (`hydrate_tiers_for_plan`, `hydrate_tiers_from_project`), file listing
strategy (git ls-files for git-checkout, os.walk fallback), budget limits (256 KB per file,
10 MB total per project), and fragment structure (`TieredFragment` with HOT tier placement
and metadata keys `path`, `detail_depth`, `relevance_score`). Closes #6175.
### Added
- **ACMS budget enforcement for max_file_size and max_total_size constraints** (#9673, #9583):
Implemented `BudgetEnforcer` in ``src/cleveragents/acms/budget_enforcement.py`` with three
core dataclasses — ``BudgetEnforcer``, ``BudgetViolation``, and ``ContextFile``. The enforcer
validates per-file size limits (``max_file_size``) and cumulative context budgets
(``max_total_size``), gracefully excluding files that exceed constraints while tracking
violations with clear, actionable error messages including filename, file size, and limit
metadata. Full type annotations throughout, ruff linting compliant, BDD test coverage
(11 Behave scenarios + Robot Framework integration) ensures correctness across boundary
conditions, empty file handling, multi-byte UTF-8 measurement, and file ordering semantics.
### Fixed
- **fileConfig error handling in alembic env.py** (#7874): Wrapped the `fileConfig()`
call in `alembic/env.py` with a `try/except` block that catches malformed INI logging
configuration and emits a clear, user-actionable error message to stderr (including the
config file path and guidance on the `[loggers]` section) before exiting with code 1.
Includes Behave scenarios for the error-handling logic.
- **`agents project context show` JSON/YAML output** (#6323): Fixed structured output to include spec-required envelope (`command`, `status`, `exit_code`, `data`, `timing`, `messages`) and all four data sections (`context_policy`, `limits`, `summarization`, `current_usage`). Rich output now renders four panels including the new `Current Usage` panel.
### Added
- **A2A module rename BDD test suite** (#8615): Comprehensive Behave tests validating that the ACP→A2A module rename is complete — verifying all 22 A2A symbols are properly exported, no legacy ACP references remain in `.py` files under `cleveragents.a2a/`, and the module docstring uses current A2A naming. The step definitions include self-contained symbol lookups to avoid cross-scenario dependency failures.
### Added
- **Project switch command** (#8675): Added `agents project switch <project>` subcommand
to the project CLI that allows switching between registered namespaced projects
without changing filesystem directories. The command validates the target project
exists in the NamespacedProject registry, sets it as the active project via the
``CLEVERAGENTS_PROJECT`` environment variable (persisted to a session helper file
for shell sourcing), and displays previous/current project info across rich, json,
yaml, plain, and table output formats. Includes BDD test coverage and confirmation
prompt (bypassable with ``--yes``/``-y``).
### Security
- **PyYAML declared as explicit runtime dependency** (#11012 / #13605): Added `pyyaml>=6.0.3` as a direct runtime dependency in `pyproject.toml`. PyYAML was previously only transitive (pulled via langchain ecosystem), listed solely as type stubs (`types-pyyaml>=6.0.0`) in the dev extras group, while being used at runtime in `src/cleveragents/actor/yaml_loader.py` for actor configuration YAML loading with Jinja2 template support and environment variable interpolation. This change explicitly pins the dependency to `>=6.0.3` to mitigate CVE-2025-8045 (arbitrary remote code execution via crafted YAML payloads) and prevents silent breakage if upstream transitive dependencies change their PyYAML requirements in future releases. The version floor ensures vulnerable versions (<6.0.3) cannot be installed even if upstream transitive dependencies have loose version constraints.
- **plan explain output uses structured alternatives objects** (#9166): The `alternatives_considered` field is replaced with a structured `alternatives` array where each element contains an `"index"` (1-based), `"description"`, and `"chosen"` (bool) key.
### Added
- **Cloud Infrastructure Resource Types** (#10592): Implemented cloud infrastructure
resource type stubs for AWS, GCP, and Azure providers. This includes Pydantic model
definitions in `src/cleveragents/resource/handlers/cloud.py` with credential resolution
(`resolve_credentials`) and validation (`validate_credentials`), built-in type
registry entries (`aws-account`, `gcp`, `azure`) with hierarchical CLI argument
definitions, provider extraction from type names via prefix matching, three BDD
feature files (`cloud_resources.feature`, `cloud_handler_coverage.feature`,
`cloud_handler_coverage_r3.feature`) covering credential masking, inheritance
chains, sandbox strategies, and full ResourceHandler Protocol conformance testing
for CRUD and lifecycle stub methods (read, write, delete, list_children, diff,
discover_children, create_sandbox, create_checkpoint, rollback_to, project_access).
### Fixed
- Fixed `ReactiveEventBus.emit()` exception handler to log the full exception
message (`str(exc)`) and enable traceback forwarding (`exc_info=True`).
Previously the handler logged only the exception type name (e.g.
"ValueError") with no diagnostic detail, making production debugging
impossible. The handler now includes the error message text and full
traceback in the structlog warning entry. Removed `@tdd_expected_fail` tag
from the TDD test so both scenarios run as normal regression guards. (#988)
- **`agents plan apply --format json` now returns spec-required JSON envelope** (#9449):
Fixed `lifecycle_apply_plan` in `src/cleveragents/cli/commands/plan.py` to wrap
non-rich format output in the spec-required JSON envelope instead of a raw plan
dictionary. The envelope includes `command: "plan apply"`, `status: "ok"`,
`exit_code: 0`, and `timing` fields. The `data` field contains `artifacts` (count
of sandbox refs), `changes` (insertions/deletions from git diff), `project` (first
project link), `applied_at` (ISO timestamp), `validation` (test/lint/type_check
results with duration), `sandbox_cleanup` (worktree/branch/checkpoint status derived
from actual plan terminal state), and `lifecycle` (phase, state, total_duration,
total_cost, decisions_made, child_plans). The `Plan as LifecyclePlan` import was moved
to the module top-level. The new `_apply_output_dict` helper is scoped exclusively to
`plan apply` call sites; all other commands (`plan status`, `plan cancel`, `plan
revert`, `plan use`) continue to use `_plan_spec_dict` unchanged.
### Added
- **`pr-review-worker` review-started notification** (#11028): The `first_review`
and `re_review` modes now post a "review started" notification comment to the
PR at the beginning of the review, giving PR authors immediate visibility
that a review is in progress. Posted asynchronously so it does not block
the workflow.
- **Plan Rollback Command** (#8557): Implemented `agents plan rollback <plan-id> [<checkpoint-id>]` for checkpoint-based plan state restoration in Epic #8493. The command restores a plan's sandbox to the state captured at a given checkpoint, discarding all decisions made after that checkpoint. The checkpoint can be specified as an optional positional second argument or via the `--to-checkpoint` named option. Supports `--yes/-y` flag to skip confirmation prompts and `--format/-f` for output format selection (rich/plain/json/yaml). Included with comprehensive BDD test coverage (>= 97%) and spec-aligned output formatting showing rollback summary, changes reverted, impact analysis, and post-rollback state panels.
### Changed
- **Expanded ruff lint scope to cover `.opencode/`** (#10848): The `lint` nox session
now includes `.opencode/` as a ruff check target, ensuring Python scripts placed
under `.opencode/` are covered by the lint gate.
### Fixed
- **ContextTierService defaults aligned with TierBudget model** (#1443): Fixed three budget default constants in ``context_tier_settings.py`` that did not match the canonical ``TierBudget`` model defaults. ``DEFAULT_MAX_TOKENS_HOT`` was 8000 but should match the ``TierBudget`` default of 16000; ``DEFAULT_MAX_DECISIONS_WARM`` was 500 but should be 100; and ``DEFAULT_MAX_DECISIONS_COLD`` was 5000 but should be 500. These incorrect values caused wrong budget enforcement when settings were not provided (i.e. ``budget_from_settings(None)``).
- **ACMS execute-phase assembler respects project-level hot_max_tokens** (#11035): Fixed
``_resolve_hot_max_tokens()`` to read ``hot_max_tokens`` from
``context_policy_json["acms_config"]["hot_max_tokens"]`` — the correct sub-key written
by ``agents project context set --hot-max-tokens``. The previous implementation read
from the top-level key (``config_dict.get("hot_max_tokens")``), which was always
``None``, causing the assembler to silently fall back to the global 16K default even
when a project-level override was configured. Also adds two Behave regression scenarios
with ``@tdd_issue @tdd_issue_11035`` tags that exercise the DB query code path and
verify the project-level budget is applied to ``CoreContextBudget`` and
``ContextRequest``.
- **Actor add `--config` crashes with combined-format ``config.actor`` YAML** (#11189): Fixed a bug where `agents actor add --config test/actor.yaml` raised ``click.BadParameter: "provider is required"`` when the config file used the spec-compliant combined ``config.actor`` format (both the compact string form ``config:\n actor: "<provider>/<model>"`` and the nested-dict form ``config:\n actor:\n type: llm\n provider: gcp\n model: gemini``). Added ``_detect_nested_config_actor()``, ``_flatten_config_actor()``, and corresponding handling in ``ActorConfiguration.from_blob()`` to transparently flatten the nesting so downstream v3 detection, schema validation, and canonicalisation see a flat dictionary. Includes new BDD scenarios and Python unit tests.
- **Guard cleanup_stale against execute/processing and execute/complete plans** (#11121):
``_create_sandbox_for_plan()`` in ``src/cleveragents/cli/commands/plan.py`` now
skips ``GitWorktreeSandbox.cleanup_stale()`` when the plan is in
``execute/processing`` (execution in progress) or ``execute/complete`` (execution
finished, awaiting apply) state. Previously, re-invoking ``agents plan execute``
on a completed plan would silently destroy the ``cleveragents/plan-<id>`` git
worktree branch, causing ``plan apply`` to merge zero artifacts. The guard
preserves the branch per spec (§sandbox.cleanup defaults to ``on_apply``).
- **Global CLI options ``--data-dir``, ``--config-path``, and ``-v`` now work correctly**
(#6785): These spec-required flags were absent from ``main_callback()`` in
``src/cleveragents/cli/main.py``, causing any invocation with these flags to crash
with ``NoSuchOption: No such option``. All three options are now implemented per
ADR-021 §Global CLI Flags and ADR-024 §Resolution Chain. ``--data-dir`` sets
``CLEVERAGENTS_DATA_DIR`` before any ``Settings``-reading code runs, ``--config-path``
sets ``CLEVERAGENTS_CONFIG_PATH`` for ``ConfigService`` to pick up, and ``-v``
(repeatable count) maps to the appropriate ``structlog`` log level.
- **Add regression test for ActorRegistry.add() nested provider/model extraction** (#4321):
Added BDD regression test confirming that provider and model are correctly
extracted from nested `actors.<name>.config` blocks when `type` is only at
the nested level and `name` uses the `local/` namespace prefix. This
scenario was previously fixed in #4300; the test ensures the fix remains
in place and prevents future regressions.
- **Actor configuration validation incorrectly requires top-level provider field** (#4300):
Actor configuration in V3 is now obtained from the nested configuration
parameter, according to the specification.
Removed the legacy V2 fallback support and the tests affected by that
removal. Mocked existing steps to allow remaining V2 features to be
covered/tested.
- **Automation profile threshold gates fully respect spec semantics** (#4328): Added
``_should_auto_progress_for_threshold()`` helper in ``PlanLifecycleService`` that
implements the full spec Threshold Semantics: ``0.0`` = always auto, ``1.0`` =
always human approval, ``0.0 < v < 1.0`` = proceed only if confidence >= threshold.
Integrated with ``AutonomyController.should_proceed_automatically()`` for
intermediate thresholds. Updated ``should_auto_progress()``, ``try_auto_run()``,
``execute_async_job()``, ``try_auto_revert_from_apply()``, and
``try_auto_revert_from_execute()`` to use the helper. Previously the service
only checked ``< 1.0`` (treated all intermediates as auto). Added BDD regression
tests in ``features/tdd_automation_profile_gates_4328.feature`` covering all 8
built-in profiles including the ``cautious`` profile's intermediate thresholds
(e.g. ``create_tool=0.7``).
- **ProviderType.GEMINI added to fallback chains** (#10906): Added `ProviderType.GEMINI` to
`ProviderRegistry.FALLBACK_ORDER` (after GOOGLE) and `"gemini"` to
`FallbackSelector.DEFAULT_FALLBACK_ORDER` (after "google"), ensuring Gemini is
selectable via auto-discovery when only GEMINI_API_KEY is configured.
- **TUI Prompt Symbol Mode Awareness** (#6431): The prompt widget now displays a
mode-dependent symbol (`` normal, `/` command, `$` shell, `` multi-line),
implemented via `_PromptSymbolMixin` and `InputMode.MULTILINE`. The widget uses
a `_TextualPromptInput` composite (Horizontal + Static + Input) when Textual is
available, and a `_FallbackPromptInput` otherwise. Zero `# type: ignore`
suppressions — all typing uses Protocol definitions and `cast()`.
- **InvariantService thread safety** (#7524): Added `threading.RLock` to protect
`_invariants` (dict) and `_enforcement_records` (list) from concurrent access
during parallel plan execution. Prevents `RuntimeError: dictionary changed size
during iteration` and data corruption when multiple subplans simultaneously
call `add_invariant()`, `list_invariants()`, `remove_invariant()`,
`get_effective_invariants()`, or `enforce_invariants()`. Also adds three
thread-safe helper read methods: `get_enforcement_records()`, `get_invariant()`,
and `get_invariants_snapshot()`. Includes comprehensive BDD tests for concurrent
access patterns (adds, lists, removes, enforcement, and mixed operations).
- **Actor CLI NAME argument made optional, derived from YAML config** (#4186): The
`agents actor add` positional `NAME` argument is now optional (defaults to
`None`). When omitted, the actor name is derived from the `name` field in
the config file. Raises `BadParameter` if neither the argument nor the config
`name` field is provided. Updated docstring signature to
`agents actor add [--config|-c <FILE>] [<NAME>]` and added config-only usage
examples. Added Behave scenario for the `BadParameter` error path
(`actor add without NAME and without config name field raises BadParameter`)
in `features/actor_add_name_positional.feature` with corresponding step
definition. Updated step definitions in
`features/steps/actor_add_update_enforcement_steps.py` and
`features/steps/actor_add_name_positional_steps.py` to pass `context.actor_name`
as a positional argument for compatibility.
- **Improved parallel test suite isolation** (#4186): Replaced deprecated
`tempfile.mktemp` with `tempfile.mkstemp` in `features/environment.py`
for atomic temp file creation, eliminating TOCTOU race conditions in the
per-scenario database path generation. Added `fcntl.flock` file locking to
`_ensure_template_db()` to prevent race conditions when multiple
`behave-parallel` workers attempt to create the template database
simultaneously.
- **Removed stale @tdd_expected_fail tags from actor add enforcement tests**: The
`--update` enforcement feature (#2609) was already implemented and merged but
residual `@tdd_expected_fail` tags remained on its BDD scenarios. These tags
were cleaned up in `features/actor_add_update_enforcement.feature` so the
tests report correctly now that the underlying bug has been fixed.
- **Fixed `merge_invariants()` missing ACTION scope — 4-tier precedence restored** (#9126): Updated ``merge_invariants()`` and ``InvariantSet.merge()`` to accept a fourth parameter ``action_invariants`` alongside plan, project, and global tiers. The module docstring, ``InvariantScope`` docstring, and ``InvariantService.get_effective_invariants()`` now all reflect the correct precedence chain: ``plan > action > project > global``. Added ``action_name`` parameter to ``get_effective_invariants()`` so action-scoped invariants are collected and passed through the merge pipeline instead of silently dropped. All docstrings across both files were corrected from ``plan > project > global`` to ``plan > action > project > global``. Comprehensive Behave scenarios added covering four-tier merge precedence, action-before-project ordering, action override of project with same text, and effective invariant computation with all four scopes.
step texts to avoid case-sensitive collisions between different step modules that
prevented all Behave tests from loading. Renamed steps in
`edge_case_plan_steps.py`, `plan_executor_coverage_boost_steps.py`,
`plan_explain_steps.py`, `plan_model_steps.py`, `project_repository_steps.py`,
`service_retry_wiring_steps.py`, and `session_model_steps.py`.
Additionally resolved a collision between `acms_index_data_model_traversal_steps.py`
and `security_audit_steps.py` for `Then the count should be`, and fixed
`pr_compliance_checklist_steps.py` project-root resolution (`parents[3]` →
`parents[2]`). Fixed table column-header mismatches in
`features/acms/index_data_model_and_traversal.feature` and guarded
`cli_init_yes_flag_steps.py` cleanup against `None` temp_dir. Annotated
`features/architecture.feature` `@tdd_expected_fail` for pre-existing Pydantic
compliance debt in `IndexEntry` / `ACMSIndex` classes.
- **PR Review Pool Supervisor Tracking Prefix** (#7891): Updated `pr-review-pool-supervisor.md`,
`docs/development/automation-tracking.md`, and `docs/development/agent-system-specification.md`
to replace all occurrences of the outdated `AUTO-REV-POOL` tracking prefix with the correct
`AUTO-REV-SUP` prefix. The agent was already using `AUTO-REV-SUP` in production; this change
aligns the documentation with the actual runtime behaviour.
- **BDD Feature File Tag Coverage** (#9124): Added required `@a2a`, `@session`, and `@cli` Gherkin tags to all A2A, session, and CLI feature files (30 files) to enable tag-based test filtering via `behave --tags=a2a,session,cli`. This restores the ability to selectively run test categories and enables CI to execute targeted test suites without running the full suite.
- **`ProviderRegistry.FALLBACK_ORDER` missing `ProviderType.GEMINI`** (#10906): Added
`ProviderType.GEMINI` to the fallback provider order list so that when only a
Gemini API key is configured, the registry correctly selects it as the default
provider instead of falling through to no provider. The enum value, capabilities,
models, and the `_create_provider_instance()` factory already supported Gemini — this
fix closes the gap in the fallback chain. Includes BDD regression scenarios in
`features/fallback_gemini_provider.feature`.
- **Cross-actor subgraph cycle detection reads actor_ref field** (#1431): Fixed
`_detect_subgraph_cycles()`, `_map_node()`, and the `compile_actor()` main loop
in `src/cleveragents/actor/compiler.py` to read `actor_ref` from the top-level
`NodeDefinition.actor_ref` field instead of `node.config.get("actor_ref", "")`.
Because `actor_ref` is a typed, validated Pydantic field (not a key inside the
untyped `config` dict), the old code always returned an empty string, causing
cross-actor cycle detection to silently fail and leaving the system vulnerable to
infinite recursion at runtime. Added Behave regression tests
(`features/actor_subgraph_cycle_detection.feature`) and a Robot Framework
integration test (`robot/actor_compiler.robot`) to prevent regressions.
- **ActorLoader.list_actors TOCTOU race condition** (#8588): Moved the namespace
filter inside the ``with self._lock:`` block so that the actors dictionary is read
and filtered atomically. Previously an actor ``.clear()`` could run between the
dictionary read and the filter step, causing the loader to return stale results or
raise ``RuntimeError: dictionary changed size during iteration``. Added a
``@unit @actor @concurrency`` Behave scenario in ``features/actor_loading.feature``
that exercises ``list_actors(namespace=...)`` and ``clear()`` on multiple threads
via ``threading.Barrier``.
- **Devcontainer auto-discovery wired into `git-checkout`/`fs-directory` handlers** (#4740):
`GitCheckoutHandler.discover_children()` and `FsDirectoryHandler.discover_children()` now
call `discover_devcontainers()` after scanning for `fs-directory` children. Any
`.devcontainer/devcontainer.json` or root-level `.devcontainer.json` found at the resource
location is registered as a `devcontainer-instance` child resource with
`provisioning_state: discovered`. Named configurations (`.devcontainer/<name>/devcontainer.json`)
are also discovered and carry the configuration name in the `config_name` property.
This wires the previously-isolated `discover_devcontainers()` function into the production
code path, enabling the spec's zero-configuration devcontainer experience.
- **Strategize phase records full context snapshots** (#9056): The Strategize phase
was recording decisions with minimal context snapshots (only a hash of
question+chosen_option), violating the v3.2.0 acceptance criterion that decisions
must include full context snapshots sufficient to replay the decision. Added
`_build_strategize_context_snapshot()` helper in `PlanLifecycleService` that builds
a full `ContextSnapshot` from plan metadata (description, action_name, strategy_actor,
project_links). Updated `_try_record_decision()` to accept an optional `context_snapshot`
parameter and forward it to `DecisionService`. Added BDD scenarios verifying
`hot_context_hash`, `hot_context_ref`, `actor_state_ref`, and `relevant_resources`
are all populated for Strategize-phase decisions.
- **`auto_debug` node functions return partial state updates instead of mutating state** (#10494, #10496): The `_analyze_error()`, `_generate_fix()`, `_validate_fix()`, and `_finalize()` methods in `src/cleveragents/agents/graphs/auto_debug.py` now return new `dict[str, Any]` objects containing only the keys they update rather than mutating the input state in-place or returning full state objects. This respects the LangGraph node contract (all `StateGraph` node functions are passed the current state and must return a partial dict representing only their incremental updates). Callers rely on LangGraph's built-in state delta-merge logic to combine these partial updates across successive nodes.
- **A2aEventQueue thread-safety (#7604)**: Guarded ``_subscriptions`` and ``_events`` with a
``threading.Lock`` to prevent ``RuntimeError: dictionary changed size during iteration`` when
``publish()``, ``subscribe_local()``, and ``unsubscribe()`` are called concurrently from multiple
threads. The ``publish()`` method snapshots the subscriber dict inside the lock and invokes
callbacks outside the lock to prevent deadlocks. The ``_is_closed`` check in ``publish()`` is
performed inside the lock to eliminate a TOCTOU race with concurrent ``close()`` calls, and the
``is_closed`` property reads ``_is_closed`` under the lock for memory-visibility guarantees.
### Changed
- Fixed stale `AUTO-BUG-POOL` tracking prefix references in automation-tracking.md documentation and agent-system-specification.md spec document, replaced with correct `AUTO-BUG-SUP` prefix used by the bug-hunt-pool-supervisor agent (#7875).
- **`agents session list` now displays full 26-character session ULIDs** (#10970): The Rich table
and Summary panel ("Most Recent" / "Oldest") previously showed only the first 8 characters of
each session ULID. This made the output unusable for copy-paste into `session tell`,
`session show`, `session delete`, and `session export`, all of which require the full
26-character identifier. The full ULID is now displayed in all output formats (Rich, plain,
JSON, YAML, table).
- **`pr-creator` now applies State/In Review and Priority labels** (#8520): Extended
`pr-creator` step 4 to apply three labels on every new PR: the `Type/` label (from
the caller's `type_label` parameter), `State/In Review` (always applied), and the
`Priority/` label matching the linked issue. Updated Rule 1 to enumerate all required
labels. Addresses the 53% missing-State-label rate observed across open PRs of
2026-04-13.
- **Suppress passing BDD scenario output in `unit_tests` by default** (#10987): Implemented
`PassSuppressFormatter`, a custom Behave formatter (in `scripts/behave_pass_suppress_formatter.py`)
that buffers all per-scenario output and only flushes it to stdout when a scenario fails or
errors. An all-passing `nox -s unit_tests` run now produces ≤ 10 lines (the summary block
only), eliminating ~100,000 lines of noise that previously made CI logs unreadable. The
formatter is embedded in the `behave-parallel` in-process runner; coverage mode
(`BEHAVE_PARALLEL_COVERAGE=1`) is unaffected.
### Security
- **PyYAML upgraded to >=6.0.3 to address known security vulnerability** (#9055):
Added an explicit `pyyaml>=6.0.3` dependency constraint to `pyproject.toml` to
prevent installation of vulnerable older versions with known YAML parsing security
issues.
- **aiohttp upgraded to >=3.13.4 to remediate CVE-2026-34513 and CVE-2026-34515** (#1549, #1544):
Added an explicit `aiohttp>=3.13.4` dependency constraint to `pyproject.toml` to remediate
two high-severity open redirect vulnerabilities. Both CVEs affect the CleverAgents platform's
HTTP infrastructure including A2A server communication, tool source fetching (MCP servers,
Agent Skills), and agent-to-agent protocol handlers. The version floor ensures vulnerable
versions (<3.13.4) cannot be installed even if upstream transitive dependencies have loose
version constraints.
### Fixed
- **Concurrent ValidationPipeline stdout/stderr restoration** (#7623): Fixed a race
condition where concurrent `ValidationPipeline.run()` calls could permanently leave
`sys.stdout` and `sys.stderr` wrapped in `_ThreadLocalStream` objects. Pipeline B
could capture Pipeline A's wrapper as its "original" stream, then restore that wrapper
in its `finally` block — permanently leaving the global streams wrapped. The fix
introduces a reference-counted shared wrapper manager (`_install_thread_local_streams`
/ `_release_thread_local_streams`) protected by a `threading.Lock`. The first caller
saves the true original streams and installs the wrappers; subsequent concurrent callers
increment the reference count and reuse the same wrappers; the last caller restores the
saved originals. Per-pipeline stdout/stderr capture remains fully isolated via
thread-local buffers.
- **Error suppression removed from `reactive_registry_adapter.py`** (#9060): Removed two `try...except Exception:` blocks in `register_registry_agents()` that silently suppressed errors, violating the CONTRIBUTING.md fail-fast policy. Exceptions from `actor_registry.list_actors()` and the route bridge refresh now propagate to the caller instead of being swallowed. Added Behave scenarios verifying RuntimeError, AttributeError, and TypeError propagation.
- **Implementation Supervisor PR Compliance Checklist** (#9824): Added a mandatory
8-item PR Compliance Checklist to the worker prompt body in `implementation-supervisor.md`
that every implementation worker must complete before creating a PR. Checklist covers:
CHANGELOG.md update, CONTRIBUTORS.md update, commit footer (`ISSUES CLOSED: #N`),
CI verification, BDD tests, Epic reference, label application via `forgejo-label-manager`,
and milestone assignment. This eliminates systemic PR merge blockers caused by workers
omitting required items.
- **Implementation Pool Supervisor PR Compliance Checklist** (#9824): Added a mandatory
8-item PR Compliance Checklist to the new `implementation-pool-supervisor.md` agent definition.
Supervisors must enforce that workers complete all 8 checklist items (CHANGELOG.md update,
CONTRIBUTORS.md update, commit footer, CI verification, BDD tests, Epic reference, label
application, and milestone assignment) before creating any PR. Includes concrete markdown
examples for each subsection and compliance verification pseudocode to ensure reproducible
adherence.
- **ACMS context path matching now handles absolute fragment paths** (#10972): Fixed
`_path_matches()` in `execute_phase_context_assembler.py` and `_matches_pattern()` in
`context_phase_analysis.py` to correctly match absolute paths (e.g. `/app/.opencode/skills/SKILL.md`)
against relative glob patterns (e.g. `.opencode/**`, `docs/*`). Previously
`PurePath.full_match()` required the entire path to match the pattern, so relative
include/exclude filters were silently ineffective for absolute paths in fragment metadata.
Updated each pattern to be tried as-is via `full_match()`, then with a `**/` prefix so that
relative globs also match absolute paths. Added BDD regression tests in
`execute_phase_context_assembler_coverage.feature` and `project_context_phase_analysis.feature`.
### Added
- **CI Dockerfile.server security scan with Trivy** (#1927): Added Trivy vulnerability
scan step to the CI docker job. The scan runs against the built `Dockerfile.server`
image and fails the build on HIGH or CRITICAL severity findings. Trivy is installed
at a pinned version with checksum verification to prevent supply chain attacks. Scan
results are surfaced in CI output. Robot Framework integration tests verify the
workflow configuration.
### Changed
- Restored `benchmark-regression` CI job to `master.yml` with `pull_request` trigger guard
(`if: forgejo.event_name == 'pull_request'`). The job was previously absent from
`master.yml`, causing benchmark regression testing to never run on PRs. The job is
informational only and is not in `status-check`'s required needs list. (Closes #10716)
- **CI coverage job now waits for unit_tests** (#10714): Added `unit_tests` to the
`needs` list of the `coverage` job in `ci.yml`. Previously the coverage job ran
in parallel with unit tests, which could produce misleading pass results when
tests were still in-flight or had already failed. Coverage now only starts after
unit tests succeed, eliminating redundant parallel test execution and ensuring
coverage results are always meaningful.
- **Bandit B608 f-string SQL in plan phases migration** (#10777): Replaced f-string
SQL construction in `a5_005_rebaseline_plan_phases.py` with plain string
concatenation. The `INSERT INTO _v3_plans_new ... SELECT ... FROM v3_plans`
statement used f-strings to interpolate `_ALL_DATA_COLUMNS`, which Bandit
flags as B608 (SQL injection risk). The constant is hardcoded and safe, but
the f-string pattern blocks tightening the bandit severity gate from HIGH to
MEDIUM (issue #9945). Replaced with `"INSERT INTO _v3_plans_new (" +
_ALL_DATA_COLUMNS + ") " "SELECT " + _ALL_DATA_COLUMNS + " FROM v3_plans"`.
- **Diagnostics spec examples expanded to all 9 providers** (#5320): Updated the
`agents diagnostics` command examples in the specification to show all 9 supported
providers (OpenAI, Anthropic, Google, Gemini, Azure, OpenRouter, Cohere, Groq,
Together), matching the implementation from PR #3469. Rich, plain, JSON, and YAML
example outputs now reflect comprehensive provider coverage with accurate warning
counts and per-provider recommendations.
### Added
- **Automation Profile Precedence Chain** (#8234): Implemented and validated the
four-level automation profile precedence chain (plan > action > project > global) as a
dedicated `automation_profile_precedence` module. All 16 combinations of
plan/action/project/global configurations are tested with BDD scenarios.
Resolution chain is logged at debug level via the
`PrecedenceResolution` dataclass and `PrecedenceSource` enum.
- **CLI showcase: version/info/diagnostics basics (#4211)**: Added a verified CLI showcase example
for the `version`, `info`, and `diagnostics` introspection commands. The guide walks through
fast-path eager flag behavior (`--help`, `--version`), rich format output with Rich panels,
machine-readable JSON envelope structure shared across all CLI commands, and CI-friendly
`diagnostics --check` health monitoring. Registered in `docs/showcase/examples.json`. Closes #7592.
- **TuiMaterializer A2A integration layer** (#5326): Implemented the
``TuiMaterializer`` class that bridges the Output Rendering Framework to
Textual UI widgets, enabling all CLI command producers to render in the TUI
without modification. Implements the ``MaterializationStrategy`` protocol and
maps ``ElementHandle`` events (Panel, Table, Status, Progress, Tree, Code,
Diff, Separator, ActionHint, Text) to plain-text renderings for TUI display.
Includes A2A event routing for ``PermissionRequest`` and ``ThoughtBlock`` events.
Thread-safe event accumulation with thread-lock guards on all state mutations.
Supports real-time streaming updates via the ``on_event`` callback pattern.
Comprehensive Behave BDD test suite covering all element types, callback
invocation, rendered output accumulation, A2A routing, and concurrent thread safety.
- `agents actor context clear` command to reset actor message history and
state while preserving the underlying context directory via `ContextManager`
(#6370).
- **Quick Start Guide** (PR #9245): Added `docs/quickstart.md` with an end-to-end quick start guide covering prerequisites, installation, project creation, resource registration, plan/apply workflow, and troubleshooting. Updated `mkdocs.yml` navigation to include the Quick Start page.
- **Plan checkpoint management CLI commands** (#8683): Added `agents plan checkpoint-list <plan-id>` and `agents plan checkpoint-delete <checkpoint-id>` commands. Listing output now highlights checkpoint ID, type, created timestamp, reason, phase, and decision linkage with a concise field summary footer across rich/table/json/yaml formats. Deletion supports batch IDs, interactive confirmation (skip with `--yes`), and structured JSON/YAML responses for automation-friendly scripting.
- **Invariant Remove CLI Command** (#8530): Implemented `agents invariant remove <id>` command that soft-deletes an invariant by ID. The command displays a confirmation prompt before removal (bypassable with `--yes`/`-y`), outputs the removed invariant ID on success, and shows a clear error message when the invariant ID does not exist. Supports `--format` flag for JSON and YAML output. Full BDD test coverage and Robot Framework integration tests included.
- **Semgrep guard for broad exception suppression** (#9185): Added two new Semgrep rules
(`python-no-suppressed-exception` and `python-no-suppress-exception`) to `.semgrep.yml`
to automate enforcement of the CONTRIBUTING.md guideline against suppressing `Exception`
or `BaseException`. Rules detect `except Exception`, `except BaseException`,
`contextlib.suppress(Exception)`, and `contextlib.suppress(BaseException)` patterns.
Supports an opt-out escape hatch via Semgrep's native `# nosemgrep` comment combined with
`# error-propagation: allow` audit annotation. Integrated Semgrep into the `nox -s lint`
session in audit mode (with `success_codes=[0,1]`) during phased rollout to prevent CI
failures from ~337 existing violations while they are triaged. Added pre-commit hook for
local enforcement and comprehensive BDD test coverage across all rule patterns and escape
hatch scenarios. Closes #9103.
- **TDD: MCPToolAdapter.infer_resource_slots() TypeError with null properties** (#10470):
Added a TDD issue-capture Behave scenario that reproduces the bug where
`MCPToolAdapter.infer_resource_slots()` raises `TypeError` when the input schema
contains `{"properties": None}`. The test is tagged `@tdd_expected_fail` and will
pass (by inversion) until the underlying bug is fixed.
- **Architecture Pool Supervisor Milestone Assignment** (#7521): Added a "PR Workflow
for Major Changes" section to the `architecture-pool-supervisor` agent definition
documenting the milestone assignment step for spec PRs. The agent now has
`forgejo_update_pull_request` permission to assign PRs to the current active
milestone after creation, improving traceability of specification changes within
project milestone planning. Includes BDD test coverage for the new workflow
documentation and permission configuration.
- **Git Worktree TOCTOU Race Condition** (#7507): Fixed a Time-Of-Check-To-Time-Of-Use
(TOCTOU) race condition in `git_worktree.py` that could cause `git worktree add`
operations to fail under concurrent execution. The fix replaces the unsafe
`mkdtemp()` + `rmdir()` pattern with a parent-directory approach that maintains
the OS-level uniqueness guarantee throughout the entire operation. The parent
temporary directory is now persisted and properly cleaned up on both success and
failure paths. Comprehensive BDD test coverage validates the fix under concurrent
execution and confirms proper cleanup behavior.
- **Database resource types (PostgreSQL, SQLite) with transaction-based sandbox strategy** (#8608):
Implemented comprehensive database resource support enabling users to interact with
PostgreSQL and SQLite backends through a unified resource interface. Introduces
`DatabaseResourceHandler` providing full CRUD operations (`read`, `write`, `delete`,
`list_children`), connection validation with automatic credential masking via
:mod:`cleveragents.shared.redaction`, and transaction-based sandbox strategy using
BEGIN/COMMIT/ROLLBACK wrappers for safe, isolated database operations. SQLite-specific
checkpoint and rollback support with SAVEPOINT semantics. Support for multiple backends (PostgreSQL, SQLite, MySQL, DuckDB) via unified "DatabaseResourceHandler" and type-specific routing. BDD test
coverage in `features/database_resources.feature` (connection validation, CRUD workflows,
transaction/rollback behavior, error handling, credential masking verification) and
Robot Framework integration tests in `robot/database_resources.robot`.
- **TransactionSandbox infrastructure for database resource isolation** (#8608):
Implemented `TransactionSandbox` class with BEGIN/COMMIT/ROLLBACK lifecycle
management for transaction-based sandbox strategy. Wired into `SandboxFactory`
as the strategy resolver for database resource types. Added `database` resource type
registration in bootstrap builtin types and updated `_resource_registry_data.py`
to recognize database resource categories.
### Fixed
- **UAT tester docs PR duplicate detection and examples.json conflict fix** (#5768): Fixed two bugs in the
`create_documentation_pr()` function of the uat-tester agent that caused
documentation PRs to be skipped: branch-existence guard incorrectly treated
error responses as success, and open-PR duplicate check missed owner prefix.
Removed `update_examples_json()` from `create_documentation_pr()` and moved
it to a single batch call after all docs PRs are created per cycle, eliminating
merge conflicts on `examples.json` when parallel UAT workers run simultaneously.
- **`invariant_enforced` decisions not propagated to child plans on subplan spawn** (#9131):
Fixed `SubplanService.spawn()` to propagate all `invariant_enforced` decisions from the
parent plan's decision tree to each child plan's decision tree. Previously, child plans
started Strategize with a completely empty invariant set, violating the spec requirement:
"recorded as `invariant_enforced` decisions that propagate to child plans." The fix adds
a `_propagate_invariant_decisions()` helper that re-records each parent
`invariant_enforced` decision on the child plan, including `non_overridable` global
invariants. BDD regression coverage added in
`features/tdd_invariant_propagation_subplan.feature`.
- **fix(repositories): derive PlanResult.success from result_success column instead of error_message** (#7501):
Fixed a critical bug in `PlanRepository._to_domain` where `PlanResult.success` was incorrectly
derived from `error_message is None`. Because `error_message` is shared between the build phase
and the result phase, a plan with a historical build error would be marked as failed even after
successfully completing and being applied. The fix introduces a dedicated `result_success` boolean
column in the `plans` table (migration `m9_003_plan_result_success_column`) and updates the
repository read path to use it. For backward compatibility, when `result_success` is NULL
(pre-migration records), the legacy `error_message is None` heuristic is preserved.
- **`LLMTraceRepository.save()` premature commit breaks UnitOfWork transactions** (#7505):
Replaced the unconditional `session.commit()` in `LLMTraceRepository.save()` with a
dual-path implementation that respects the UnitOfWork (UoW) pattern. When an external
session is provided (UoW mode), the method now calls only `session.flush()`, leaving
transaction control to the caller. When no session is provided (standalone mode), the
method creates its own session, flushes, commits, and closes it to ensure durable
persistence. This eliminates three data-integrity violations: premature commit of outer
UoW transactions, loss of rollback capability for subsequent failures, and a mismatch
between the class docstring ("Callers are responsible for commit") and the implementation.
Input validation for the `trace` argument was also added. Two new BDD scenarios verify
the session contract: `Repository save() calls flush not commit` and `LLM trace rolled
back when UnitOfWork transaction rolls back`.
- **engine_cache MEMORY_ENGINES TOCTOU Race Condition** (#7566): Fixed a
Time-Of-Check-To-Time-Of-Use (TOCTOU) race condition in the in-memory SQLite
engine cache where two concurrent threads could both observe a cache miss on
`MEMORY_ENGINES` and each create a separate `Engine` instance for the same
`sqlite:///:memory:` URL, violating the single-shared-engine contract and
causing duplicate database connections and inconsistent transaction state.
The fix adds a module-level `MEMORY_ENGINES_LOCK: threading.Lock` in
`engine_cache.py` (exported for use by dependants) and wraps the
check-and-set in `UnitOfWork.engine` with `with MEMORY_ENGINES_LOCK:`.
Also fixed a cache-hit bug where `self._engine` was only assigned inside the
`if url not in MEMORY_ENGINES` block, leaving it `None` on a cache hit;
the assignment is now unconditional inside the `with` block so every call
that reaches the lock exits with a valid engine reference. Four new BDD
scenarios in `features/tdd_engine_cache_toctou.feature` (with step
definitions in `features/steps/tdd_engine_cache_toctou_steps.py`) verify
lock export, cache-hit correctness, lock acquisition, and thread safety
under 10 concurrent threads.
- **git_tools.\_get_base_env() TOCTOU Race Condition** (#7619): Fixed a
Time-Of-Check-To-Time-Of-Use race condition in `git_tools._get_base_env()`
where two concurrent threads could both observe `_BASE_ENV is None`, both
snapshot `os.environ`, and write potentially different snapshots. The fix
adds a module-level `_BASE_ENV_LOCK: threading.Lock` and replaces the bare
`if _BASE_ENV is None` assignment with double-checked locking: the outer
check keeps the warm-cache path lock-free; the inner check inside
`with _BASE_ENV_LOCK` prevents duplicate initialisation on the very first
concurrent call. Three new BDD scenarios in `features/git_tools.feature`
(with step definitions in
`features/steps/git_tools_thread_safety_steps.py`) verify caching identity,
content correctness, and thread safety under 20 concurrent threads.
- **Unified provider factory: eliminate divergence between `create_llm()` and `create_ai_provider()`** (#10949):
Introduced `_create_provider_instance()` as the single internal factory so that
both public methods delegate to one place. Creating a new provider now
requires changes in exactly one method.
- **Fixed API key regression**: the unified factory now explicitly passes
the validated API key to all LangChain constructors (OpenAI, Anthropic,
Google / Gemini, Azure, Groq, Together, Cohere, and OpenRouter). Users
who configure providers via `CLEVERAGENTS_`-prefixed variables are no
longer silently failed when LangChain falls back to raw environment
variable lookup. Pre-validated keys are forwarded through the
`api_key` kwarg to avoid a second settings lookup in the factory
closure. (Closes #10949)
- **Fixed mock provider accessibility in production**: `ProviderType.MOCK`
is now gated by the `CLEVERAGENTS_ALLOW_MOCK_PROVIDER=true` sentinel
environment variable. Without this flag, both `create_llm()` and
`create_ai_provider()` raise `ValueError` when MOCK is requested,
preventing accidental or malicious use of the fake LLM in production.
`resolve_provider_by_name("mock")` now also respects the guard and
`is_provider_configured(ProviderType.MOCK)` returns `True` as expected.
- **Fixed type annotation**: `create_llm()` now declares `**kwargs: Any`
instead of `**kwargs: object`, restoring correct Pyright inference for
forwarded keyword arguments.
- **`create_llm()` raises `Unsupported provider type: openrouter`** (#10948): Fixed
`ProviderRegistry._create_provider_llm()` missing an `OPENROUTER` branch, which
caused `agents actor run openrouter/<model>` to fail with `ValueError`. Added a
`ProviderType.OPENROUTER` branch that creates a `ChatOpenAI` instance configured
with `openai_api_base="https://openrouter.ai/api/v1"` and the OpenRouter API key,
matching the behavior of `create_ai_provider("openrouter")`. Supports optional
`default_headers` kwarg with automatic string coercion for non-string keys/values.
- **LoadingThrobber Widget Restored** (#6357): Restored `LoadingThrobber` widget
- **Built-in actors v3 YAML format** (#10883): Fixed `agents actor run` failing for
built-in actors (e.g., `openai/gpt-4`, `anthropic/claude-3-opus`) due to missing
v3 `type` field in stored configuration. `ActorRegistry.ensure_built_in_actors()`
now generates and persists v3 YAML text with `type: llm` and `description` fields,
ensuring built-in actors work identically to custom actors. The
`_generate_builtin_actor_yaml()` helper creates spec-compliant YAML that passes
`ReactiveConfigParser._is_v3_format()` validation. Includes BDD scenarios and unit
tests covering YAML generation, schema validation, and multiple provider handling.
- **Atomic `server_connect` config writes** (#993): Fixed `server_connect` in
`cli/commands/server.py` to write all three config values (`server.url`,
`server.namespace`, `server.tls-verify`) atomically. A snapshot of the config
file is taken before any writes; if any `set_value()` call fails, the snapshot is
restored and compensating `CONFIG_CHANGED` events are emitted for already-applied
keys so the audit trail reflects the rollback. Added `emit_config_changed()` helper
to `ConfigService` for decoupled event emission in rollback flows. Added
`close()` method to `ReactiveEventBus` for proper resource cleanup in tests.
Resolved merge conflict in `config_service.py` integrating the PR's
`emit_config_changed()` helper with master's scoped config infrastructure.
Removed `# type: ignore[assignment]` by introducing a typed `_AutoDiscover`
sentinel class. BDD regression coverage in
`features/tdd_server_connect_atomic_writes.feature`.
- **Atomic `load_from_metadata` for Autonomy Guardrails** (#7504): Fixed
`AutonomyGuardrailService.load_from_metadata()` to validate both
`AutonomyGuardrails` and `GuardrailAuditTrail` models before writing either
to state, ensuring atomic updates. Previously, a validation failure on the
audit trail after guardrails were already written would leave the system in
an inconsistent state with partial updates. The method now uses a two-phase
validate-then-write approach: all model validation occurs in Phase 1, and
state mutations only happen in Phase 2 after all validations succeed.
- **`agents actor run` empty response for built-in LLM actors** (#10861): Fixed
`resolve_config_files` in `cli/commands/_resolve_actor.py` silently returning
empty output when invoked with a built-in actor name (e.g.
`anthropic/claude-sonnet-4-20250514`). Built-in actors generated from the
provider registry have a `config_blob` with `provider` and `model` fields but
no `type` field. Serialising this blob as-is produced YAML that
`ReactiveConfigParser` could not interpret (no agents, no routes → empty
`ReactiveConfig` → empty response). Fix: `_synthesize_llm_yaml()` now
synthesises a minimal v3 `type: llm` YAML when the actor has no `yaml_text`
and the `config_blob` has `provider` and `model` but no `type` field, allowing
the reactive config parser to create a working agent and graph route. BDD
regression coverage in `features/tdd_actor_run_response.feature`.
- **ReactiveConfigParser route synthesis for v3 actors** (#10807): Fixed
`agents actor run` silently returning empty output for v3 `type:llm` actors.
`_build_from_v3()` and `_build()` now synthesise a default single-node
graph route when agents are created without explicit routes, ensuring
`run_single_shot()` can invoke the LLM via `GraphExecutor`. The nested
`actors:` map format also translates the v3 `actor: "provider/model"` key
into separate `provider` and `model` keys so the correct LLM provider is
instantiated.
- **ActorRegistry.add() spec-compliant YAML support** (#4466): The registry now
accepts actor YAML using the spec's `actors:` map format with nested `config:`
blocks, in addition to the legacy top-level `provider`/`model` format. The
`unsafe` flag and graph descriptor from nested config are now correctly
preserved during registration. Multi-actor YAML (>1 entry in `actors:`/`agents:`
map) is now rejected by `add()` with a `ValidationError`. Nested
`config.options` are now correctly preserved. The `unsafe` coercion now uses
strict `is True or == 1` instead of `bool()` to prevent truthy non-boolean
YAML values (e.g. `unsafe: "no"`) from being treated as unsafe.
- **UKO Runtime Layer 2 (Paradigm) Indexing** (#9351): Added missing `rdf:type
uko-oo:Class` triple emission in `PythonAnalyzer._extract_class()` so that
Python class definitions are now correctly classified at layer 2 (paradigm/OO)
in addition to layer 3 (technology). Added the corresponding Behave scenario
`Indexing a Python file populates layer 2 (paradigm)` to
`features/uko_runtime.feature`, completing four-layer guarantee verification
for the UKO runtime.
- **Actor CLI v3 YAML Schema Support** (#6283): Fixed three components to add
full v3 `ActorConfigSchema` support to the actor CLI registration and
execution paths. `ActorConfiguration.from_blob()` now detects v3 format
(top-level `type` key of `llm`/`graph`/`tool`) and correctly extracts
provider, model, and graph descriptors — including `type: tool` actors
without a `model` field. `ActorRegistry.add()` validates against the full
Pydantic v2 schema, persists `skills`/`lsp`/`description` in the config
blob, and compiles graph actors with proper metadata.
`ReactiveConfigParser._build_from_v3()` now uses correct `source`/`target`
edge keys (fixing `KeyError` in `to_graph_config()`), handles `config: null`
nodes without crashing, propagates `context_view`/`memory`/`context`/
`env_vars`/`response_format`/`lsp_capabilities`/`lsp_context_enrichment`
into agent configs, and validates `entry_node` against the nodes map.
Exception handling narrowed from broad `except Exception` to specific
`NotFoundError` and `ActorCompilationError`. v3 registration logic
extracted to `v3_registry.py` to keep `registry.py` under the 500-line
limit. 19 BDD scenarios cover all v3 paths including tool actors,
update mode, LSP dict bindings, and field propagation.
- **TDD Non-AssertionError Guard Visibility** (#8294): `apply_tdd_inversion` in
`features/environment.py` now emits its non-assertion exception guard warning to
both the structured logger and `stderr` via a new `_warning_with_stderr` helper.
This makes the guard firing visible in standard Behave console output and CI log
snippets where the structured logging sink may not be displayed. BDD infrastructure
coverage added: a new scenario in `tdd_expected_fail_infrastructure.feature`
asserts that the warning is emitted to stderr when a non-AssertionError exception
is encountered in an `@tdd_expected_fail` scenario, and a second scenario asserts
the warning is NOT emitted when the exception is an `AssertionError`. The
`CONTRIBUTING.md` now documents that `@tdd_expected_fail` step definitions must
signal expected failures via `AssertionError`.
- **Parallel Behave Runner Log Noise Reduction** (#8351): The parallel behave
runner now suppresses captured stdout/stderr for passing worker chunks and
only replays diagnostics for failed, errored, or crashed chunks. This makes
failure output significantly easier to spot in CI and local runs. A worker
crash (unhandled exception) is detected via an all-zero summary and the
captured traceback is always surfaced.
- **Bug Hunt Pool Supervisor Non-Blocking Tracking**: Updated `bug-hunt-pool-supervisor` to make the automation tracking step non-blocking. The `automation-tracking-manager` call in step 5 is now best-effort — if it does not complete within a reasonable time or fails, the supervisor skips it and continues to the next cycle. Added explicit rule 9 clarifying that tracking must never block the main loop. Core functionality (module scanning and worker dispatch) takes priority over status reporting.
- **Name Validator Server-Qualified Format** (#9074): Updated actor, skill, and tool name
validators to accept the spec-required `[[server:]namespace/]name` format. Previously,
server-qualified names like `dev:freemo/custom-analysis` were incorrectly rejected.
Added BDD scenarios for server-qualified name acceptance and rejection. All three
validators (`ActorConfigSchema.validate_name`, `NAMESPACED_NAME_RE`,
`_TOOL_NAME_PATTERN`) now correctly support optional server prefixes while maintaining
backward compatibility with existing `namespace/name` names.
- **Legacy CLI command removal** (#4181): Removed all legacy plan lifecycle CLI
commands (`tell`, `build`, `new`, `current`, `cd`, `continue`) and their
associated tests to support V3 Plan Lifecycle exclusively. Removed `tell` and
`build` CLI shortcuts from `main.py` that delegated to the deprecated commands.
Removed orphaned `_tell_streaming` dead code from `plan.py`. Updated help text
and command validation to recognize only V3 commands. Added `--format` option
to `agents session tell` for consistency with other session commands. Fixed
MCP logger thread-safety in `session.py` using a threading lock. Created
migration guide for users transitioning from legacy to V3 workflow.
- **Plan Tree JSON/YAML Command Envelope** (#9163): `agents plan tree --format json/yaml`
now wraps output in the spec-required command envelope with `command`, `status`,
`exit_code`, `data`, `timing`, and `messages` fields. The `data` field contains
`plan_id`, `tree`, `summary` (nodes, depth, child_plans, invariants, superseded),
`child_plans` list, and `decision_ids` mapping. Timing now reflects actual elapsed
milliseconds from command start to envelope construction.
- **Automation Profile Silent Fallback** (#8232): `_resolve_profile_for_plan` in
`PlanLifecycleService` now raises a clear `ValidationError` when a plan's
automation profile name is not a known built-in profile, instead of silently
falling back to `"manual"`. Users who configured custom automation profiles
(e.g. `"semi-auto"`, `"acme/strict"`) will now receive an actionable error
message listing available built-in profiles. The resolved profile name is also
logged at debug level for observability.
- **DoD Gating in Apply Phase** (#7927): `PlanLifecycleService.apply_plan` now
evaluates the plan's `definition_of_done` criteria before transitioning to the
Apply phase. If any required criteria fail, a `DoDGatingError` is raised and
the plan remains in Execute/COMPLETE state. The evaluation result is stored in
`plan.validation_summary` with `dod_evaluated=True` and `dod_all_passed`
reflecting the outcome. Plans with no DoD text skip evaluation and proceed
normally.
### Added
- **ACMS Index Data Model and File Traversal Engine** (#9579): Implements the
foundational ACMS index data model with structured fields for file metadata
(path, size, last modified, type), tag system, and hot/warm/cold/archive
storage tier assignment. Introduces a timeout-safe large-project file traversal
engine capable of handling 10,000+ files without memory exhaustion through
chunked processing. Provides a complete index entry pipeline for creation,
storage, and retrieval with full queryability by path, tag, type, and recency.
- **ACMS Large-Project Indexing BDD Coverage** (#8726): Added 7 Behave scenarios
covering walk-based indexing of 10,000+ files without timeout, binary-file
skipping, oversized-file skipping, git-checkout indexing, fallback to walk when
`git ls-files` is unavailable on a non-git directory, and total-bytes budget
enforcement. Optimised fixture setup to pre-create subdirectories (99x fewer
syscalls). Added `timeout=120` to git subprocess calls to prevent CI hangs.
Cached `get_scoped_view` results in `When` steps to avoid redundant re-queries
in `Then` steps.
- **Agent Evolution Pool Supervisor PR Metadata Assignment** (#7888): The
agent-evolution-pool-supervisor now automatically looks up the Type/Automation
label and the earliest open milestone from the repository before dispatching
improvement PR creation workers. Label and milestone IDs are passed to workers
via the dispatch context, ensuring all generated improvement PRs have correct
Type labels and milestone assignments. Graceful error handling skips label or
milestone assignment when either is unavailable. Added comprehensive BDD test
suite (7 scenarios) covering label lookup, milestone lookup, worker dispatch,
PR creation with metadata, and error handling for missing labels/milestones.
- **Tool and Validation Management Showcase** (#4565): Added
`docs/showcase/cli-tools/tool-and-validation-management.md`, a step-by-step
CLI guide covering the complete lifecycle for managing tools and validations
(register, list, inspect, update, remove). Demonstrates the unified registry,
capability flags, validation modes and wrapped validations. Indexed the new
example in `docs/showcase/examples.json` and added a callout explaining why
capability flags display as `(default)` in the detail view.
- Wired `StrategyActor` into the real plan execution path: `_get_plan_executor`
in `plan.py` now resolves the strategy actor via `resolve_strategy_actor()`
(reading the `actor.default.strategy` config key) instead of always
constructing `LLMStrategizeActor`. `run_strategize` in `PlanExecutor` now
passes `resources` (derived from `plan.project_links`) and `project_context`
to the actor so the LLM prompt receives full project context. Strategy
decisions are serialised as JSON in `plan.error_details["strategy_decisions_json"]`
so `_build_decisions` can reconstruct the full hierarchy (dependency ordering,
parent/child structure) during Execute instead of rebuilding from
`definition_of_done`. `StrategizeStubActor.execute` accepts `**kwargs` for
forward-compatibility. Added BDD coverage for the stored-JSON path,
corrupt-JSON fallback, resource-passing, and stub extra-kwargs scenarios. (#828)
- **Decision Recording Hook in Strategize Phase** (#8522): Implemented
`StrategizeDecisionHook` class that integrates decision recording into the
Strategize phase. The hook captures every decision point during strategy
decomposition, including question, chosen option, alternatives considered,
confidence score, rationale, and full context snapshot (hot context hash,
actor state reference, relevant resources). Supports recording of
`strategy_choice`, `resource_selection`, `subplan_spawn`, and
`invariant_enforced` decision types. Context snapshots are auto-captured
with SHA256 hashing of context data and checkpoint references for LangGraph
actor state. Includes comprehensive BDD test suite with 40+ scenarios
covering all decision types, context capture, error handling, and tree
structure validation.
- **TDD Issue-Capture Test Activation** (#7025): Replaced 234 bare `@skip` tags
across 82 Behave feature files with the correct `@tdd_expected_fail @tdd_issue
@tdd_issue_<N>` tag system. Scenarios whose referenced bugs were already fixed
had `@tdd_expected_fail` removed and now run as permanent regression guards.
Net result: 629 features active in CI (up from ~545), zero `@skip` tags remain.
- **Spec alignment for `agents project delete` output** (#7872): Updated
`docs/specification.md` to document the correct JSON/YAML output structure
for the project delete command, replacing the legacy `deletion_summary`
object with the `deleted`, `success`, and `deleted_at` fields that match
the actual implementation introduced in PR #6639. Added BDD feature file
and step definitions to validate the output format.
- **Git Worktree Sandbox Apply** (#4454): The `plan apply` command now merges
LLM-generated changes via `git merge` from an isolated worktree branch
instead of flat `shutil.copy2`. Displays spec-aligned Apply Summary
(plan ID, artifacts, insertions/deletions, project, timestamp), Sandbox
Cleanup panel, and `✓ OK Changes applied` footer. Non-git projects fall
back to the original flat file copy.
- **Context Hydration Fix** (#4454): Fixed `ContextFragment` metadata types
(`detail_depth` and `relevance_score` must be strings, not int/float) that
caused Pydantic validation errors during context assembly, resulting in the
LLM receiving zero file context.
- **Automation Tracking System**: Replaced shared session state issue tracking with
individual per-agent tracking issues. Each agent now creates its own `[AUTO-<PREFIX>]`
titled issues with standardized headers, reporting intervals, and health indicators.
Agents: `session-persister`, `implementation-orchestrator`, `system-watchdog`,
`backlog-groomer`, `human-liaison`. Documentation at
`docs/development/automation-tracking.md`.
- **Automated Health Monitoring and Recovery**: The `system-watchdog` now runs
`audit_automation_tracking_health()` every 5 minutes, detecting stalled agents when
tracking issues are >20% overdue from their declared reporting interval. On detection,
it terminates stalled sessions via the OpenCode Server API, performs root-cause analysis,
creates high-priority diagnostic issues, and closes stale tracking issues with recovery
notes.
- **Centralized Label Management** (`forgejo-label-manager`): A new specialized subagent
centralizes all Forgejo label operations across the agent system. Enforces the
organization-level label system, prohibits label creation, and validates label compliance.
Agents `backlog-groomer`, `human-liaison`, `project-owner`, `epic-planner`,
`new-issue-creator`, and `issue-state-updater` now delegate all label operations to this
subagent.
- **PR-Issue Label Synchronization**: PRs now inherit `Priority/`, `MoSCoW/`, `Points/`,
and `State/` labels from their associated issues at creation time
(`pr-api-creator`). The `backlog-groomer` adds a continuous Pass 19 for ongoing
PR-issue label synchronization. The `issue-state-updater` syncs PR state labels whenever
issue states change.
- **Automation Tracking Announcements**: Extended `automation-tracking-manager` with
announcement issue support (`CREATE_ANNOUNCEMENT_ISSUE`, `CLOSE_ANNOUNCEMENT_ISSUE`,
`LIST_TRACKING_ISSUES`, `READ_ANNOUNCEMENTS`, `REVIEW_OWN_ANNOUNCEMENTS`). Supervisors
and workers now read critical announcements before each cycle for cross-agent awareness.
Priority-based filtering (Critical/High/Medium/Low) reduces noise. Backlog-groomer
performs intelligent cleanup with age thresholds by priority.
- **PR Agent Reorganization**: All PR-related agents renamed and reorganized to follow
the `*-pool-supervisor` naming pattern. New agents added: `pr-editor` (safe PR editing
with description preservation), `pr-manager` (unified PR interface), and
`pr-merge-pool-supervisor` (automated PR merging supervisor). Renamed:
`pr-api-creator` to `pr-creator`, `pr-checker` to `pr-ci-test-fixer`,
`pr-status-checker` to `pr-status-analyzer`, `pr-self-reviewer` to `pr-reviewer`,
`pr-fix-orchestrator` to `pr-fix-pool-supervisor`.
- **Automated PR Merging** (`pr-merge-pool-supervisor`): New supervisor continuously
monitors for merge-ready PRs and merges them automatically when all criteria are met
(approvals, CI passing, no conflicts). Supports both formal reviews and comment-based
approvals (LGTM, ready to merge, etc.).
- **Implementation Worker Workflow Completion**: `implementation-worker` now implements
work claiming protocols with conflict detection, comprehensive review feedback handling
with intelligent parsing, sophisticated merge conflict resolution with multiple
strategies, and parallel subtask execution with wave-based dependency analysis.
Pass rate improved from 48.15% to 84.8%.
- **Container Resource Stop Support**: `agents resource stop` now correctly stops
`container-instance` and `devcontainer-instance` resource types.
- **Centralized Automation Tracking Manager** (`automation-tracking-manager`): The
subagent is now the single interface for all tracking issue operations
(`CREATE_TRACKING_ISSUE`, `UPDATE_TRACKING_ISSUE`, `CLOSE_TRACKING_ISSUE`,
`READ_TRACKING_STATE`, `GET_NEXT_CYCLE_NUMBER`). Agents delegate to the manager rather
than calling the Forgejo API directly, ensuring sequential cycle numbers across
restarts, preventing duplicate issues, and enforcing consistent label application.
Migrated agents include `system-watchdog`, `implementation-orchestrator`,
`timeline-updater`, `project-owner`, `product-builder`, `backlog-groomer`,
`implementation-pool-supervisor`, `timeline-update-pool-supervisor`, and
`project-owner-pool-supervisor`. The legacy `shared/automation_tracking.md` module was
removed.
- **Documentation Writer Tracking** (`docs-writer`): The documentation writer now
participates in the automation tracking system by creating individual `[AUTO-DOCS]
Documentation Report (Cycle N)` issues every 10 cycles (~3.3 hours). The manager applies
the mandatory `Automation Tracking` label automatically, while teams may add additional
workflow labels as needed. See `docs/development/automation-tracking.md` and the new
`docs/development/docs-writer.md` reference.
- **ACMS / UKO API Documentation** (`docs/api/acms.md`): Added comprehensive API
reference for the `cleveragents.acms` package covering the four-layer UKO ontology
hierarchy, `VocabularyRegistry`, `ProvenanceInfo`, `UKOClass`, `UKOProperty`,
`UKOVocabulary`, `Layer2Dependency`, `ParadigmVocabulary`, `DetailLevelMapBuilder`,
and all Layer 3 language vocabulary types (Python, TypeScript, Rust, Java).
The new page is linked from the API Reference index and the MkDocs navigation.
- **Comprehensive Worker Tracking System**: All 16 supervisors now provide detailed
visibility into worker activities and health via the OpenCode API. Enhanced
`product-builder`, `implementation-orchestrator`, `continuous-pr-reviewer`, and
`uat-tester` with detailed session monitoring, real cycle-time calculations, stale
worker detection and restart, and proper tracking issue lifecycle management (delete
previous, create new each cycle). Tracking now extends to previously uncovered
supervisors such as `architect`, `timeline-updater`, `docs-writer`, and
`architecture-guard`.
- **Plan Action Argument Upsert**: `PlanLifecycleService` now upserts action arguments
during `plan use` to avoid `UNIQUE` constraint violations when reusing actions.
Includes batch-delete updates with identity-map eviction, invariants unique constraint,
and an Alembic migration. (#4174)
### Changed
- **`product-builder` Worker Allocation Tier Comments** (#8169): Clarified the
`N_FULL` tier comment to explicitly document that PR fixing is handled by
`implementation-pool-supervisor` via its PR-First Priority rule. Updated the
`N_QUARTER` comment to enumerate the pools it covers (UAT, bug hunting, test
infra). Prevents confusion about which supervisor handles PR fix work.
- **Actor CLI Exception Handling Refinement** (#8567): `_get_services()` in
`actor.py` now carries an explicit `tuple[Any, Any | None]` return type
annotation. `_load_config_text()` now catches `yaml.YAMLError` and `TypeError`
(instead of `ValueError`/`AttributeError`) around `yaml.safe_load()`, preserving
the user-friendly `typer.BadParameter` message for malformed YAML.
`_compute_actor_impact()` defensive guards now also catch `CleverAgentsError`,
SQLAlchemy `OperationalError`, and `ValidationError` in addition to
`AttributeError`/`RuntimeError`, keeping actor removal resilient when the
database layer is unavailable.
- **Decision Tree Full ULID Display** (#5825): The `agents plan tree` command now
displays full 26-character ULIDs for all decisions instead of truncating them to
8 characters. This enables users to copy decision IDs directly from tree output
and use them in follow-up CLI commands like `agents plan correct` without manual
ID reconstruction. Added "Decision IDs (for correction)" section with human-readable
labels for easy reference. Applies to both table and rich/plain text output formats.
- **Automation Tracking Format**: All automation tracking issues now use a standardized
header format with mandatory `Reporting Interval: <interval> (Next report expected: <ts>)`
declarations, enabling precise staleness detection.
- **PR Review Policy**: Reduced PR review requirement from 2 approvals to 1. Self-approval
is now permitted including for automated bot PRs. Approval can be a formal review OR an
approval comment (LGTM, Approved, ready to merge).
- **Label Delegation Enforcement**: `automation-tracking-manager` now enforces delegation
to `forgejo-label-manager` for all label operations, preventing "invalid label ID" errors
and ensuring label application uses correct name-to-ID mapping.
- **Automation Tracking Label Guidance**: Documentation now clarifies that the manager
automatically applies the `Automation Tracking` label and that additional labels such as
`Type/Automation`, `State/In Progress`, or `Priority/Medium` remain optional workflow
choices rather than mandatory.
- **Automation Tracking Agent Prefix Registry**: Expanded from 5 agents to 18 agents.
New prefixes include `AUTO-DOCS`, `AUTO-REV-POOL`, `AUTO-UAT-POOL`, `AUTO-BUG-POOL`,
`AUTO-INF-POOL`, `AUTO-ARCH`, `AUTO-EPIC`, `AUTO-EVLV`, `AUTO-GUARD`, `AUTO-SPEC`,
`AUTO-TIME`, `AUTO-PROJ-OWN`, and `AUTO-PROD-BLDR`.
- **ACMS Context Hydration**: Fixed ACMS indexing pipeline not wired into CLI —
`ContextTierService` started empty on every CLI invocation so LLM received zero file
context during plan execution. Added `context_tier_hydrator.py` that reads files from
linked project resources (via `git ls-files` or `os.walk`) and stores them as
`TieredFragment` objects in the tier service. Hydration runs automatically before context
assembly in `LLMExecuteActor.execute()`. Respects max file size (256KB), total budget
(10MB), binary file exclusion, and `.git`/`node_modules`/`__pycache__` directory
skipping. (#1028)
- **Product-Builder Tracking Migration**: `product-builder` now creates individual
per-cycle tracking issues (prefix `AUTO-PROD-BLDR`) instead of a long-running shared
session state issue. Each cycle closes the previous tracking issue and creates a fresh
one, providing better isolation and traceability.
- **Implementation Orchestrator Scaling**: Scaled to 32 parallel workers. Reduced
dispatch loop sleep from 10s to 2s, simplified worker verification, reduced retry
delays from 15s to 2s, and reduced idle sleep from 60s to 10s for dramatically
faster throughput.
- **Specification — Validation Gate Empty-Run Guard** (#8146): Updated `docs/specification.md`
to document the security invariant introduced in PR #7786 (fixing issue #7508). The spec now
explicitly states that `ApplyValidationSummary.all_required_passed` returns `False` when no
validations have been run (empty summary), blocking apply. Added a prominent danger admonition
block, updated the validation process results section, the `final_validation_results` data
model description, and two milestone acceptance criteria to reflect the corrected blocking
behavior for empty validation summaries and no-attachment runs.
### Fixed
- **Plan Concurrency Race Condition** (#7989): Fixed critical race condition in `execute_plan()` and
`apply_plan()` where concurrent CLI/worker sessions could simultaneously modify the same plan,
corrupting plan state. `LockService` is now wired into the plan lifecycle with plan-level advisory
locking. Each invocation generates a unique caller identity (UUID) to prevent re-entrant lock
acquisition by concurrent sessions on the same plan. Concurrent attempts now raise `LockConflictError`
instead of silently racing. Lock is acquired before phase transition and released in a `finally`
block to ensure cleanup even on error.
- **`--format color` ANSI Output** (#7910): Fixed `format_output` routing the `color` format
option to `_format_plain`, which produced plain uncoloured text instead of ANSI escape
sequences. The `color` format is now routed to `format_output_session` which uses the
`ColorMaterializer` to emit proper ANSI-coloured output. `--format plain` and all other
formats remain unaffected.
- **ContextTierService Thread Safety** (#7547): Added `threading.RLock` to
`ContextTierService` to prevent `RuntimeError: dictionary changed size during
iteration` and data corruption under concurrent plan execution. All public
methods (`store`, `get`, `promote`, `demote`, `evict_lru`, `enforce_staleness`,
`get_metrics`, `get_all_fragments`, `get_hot_fragments`, `get_for_actor`,
`get_scoped_view`, `get_scoped_by_resource`, `get_scoped_metrics`) now acquire
the reentrant lock before accessing the hot/warm/cold tier dicts. The service
was previously documented as single-threaded but registered as a DI Singleton,
causing potential data corruption when parallel subplans shared the same
instance. The `TierRuntimeMixin.enforce_staleness()` and
`ScopedTierMixin.get_scoped_by_resource()` / `get_scoped_metrics()` methods
are also protected. The DI container registration as `providers.Singleton`
is now correct and safe.
- **TOCTOU Race Condition in Git Worktree Sandbox** (#7507): Fixed Time-Of-Check-To-Time-Of-Use race condition in `GitWorktreeSandbox.create()` by replacing unsafe mkdtemp+rmdir pattern with persistent parent directory approach. Parent directory is now held throughout operation lifetime and properly cleaned up in all error paths (timeout, CalledProcessError, OSError) and in the cleanup() method, eliminating race window where another process could claim the worktree path. Comprehensive BDD coverage added for all error-path cleanup branches.
- **Validation Gate Empty-Run Guard** (#7508): Fixed `ApplyValidationSummary.all_required_passed`
returning `True` when zero validations were run, silently bypassing the apply gate. The property
now returns `False` when the validation result set is empty (`is_empty` is `True`), ensuring
that apply is blocked unless at least one validation was actually executed. Also added
`required_total` property for completeness. Updated `consolidated_validation.feature` scenarios
to reflect the corrected blocking behavior for empty summaries and no-attachment runs.
- **ACMS context tier hydration**: `ContextTierService` no longer starts empty
on every CLI invocation. A new `context_tier_hydrator.py` reads files from
linked project resources (via `git ls-files` or `os.walk`), creates
`TieredFragment` objects, and stores them in the tier service before context
assembly in `LLMExecuteActor.execute()`. The LLM now receives real file
context during plan execution. Respects max file size (256 KB), total budget
(10 MB), binary file exclusion, and `.git`/`node_modules`/`__pycache__`
directory skipping. (#1028)
- **Sandbox root wiring**: `_get_plan_executor()` now passes
`sandbox_root=.cleveragents/sandbox/`, so LLM file output (`FILE:` blocks)
is written to disk during the execute phase. (#4222)
- **SubplanExecutionService fail_fast cancellation** (#7582): Fixed a race condition where
already-running parallel subplans were not cancelled when `fail_fast` fired. Previously,
`Future.cancel()` only prevented queued futures from starting but had no effect on
in-flight futures that completed after `stop_flag` was set -- their `COMPLETE` results
were incorrectly included in the merge output. The fix adds a post-completion guard that
overrides any non-`ERRORED`/non-`CANCELLED` result to `CANCELLED` when `stop_flag` is
active, and clears the associated output to prevent it from entering the merge. Also
replaces the O(n) linear `status` lookup in the `as_completed()` loop with an O(1)
`status_map` dict pre-computed before the executor block.
- **JSON/YAML envelope `messages[].text` content** (#6457): `agents session create`,
`list`, `show`, `delete`, `export`, and `import` commands now populate the
`messages[].text` field with human-readable text (`"Session created"`,
`"N sessions listed"`, `"Session details loaded"`, `"Session deleted"`,
`"Export completed"`, `"Import completed"`) instead of the generic `"ok"` fallback.
The `export` command gains `--output-format` and the `import` command gains `--format`
to select the output envelope format independently of the export/import file format.
- **Robot Framework TDD Listener Guards** (#5436): Added three guard conditions to the
`tdd_expected_fail_listener` `end_test()` function to prevent blindly inverting ALL test
failures to passes, which was masking infrastructure errors and causing flaky CI behavior.
Guards: setup/teardown error detection, non-assertion failure detection (infrastructure
errors), and dry-run mode detection. Also fixed `Variable Should Exist` syntax errors in
e2e test files and removed `tdd_expected_fail` from 4 context assembly e2e tests where
bugs were already fixed.
- **PluginLoader entry point prefix validation** (#7476): Parse entry point targets before
import, enforce the module allowlist ahead of loading, and add Behave plus Robot Framework
regression coverage to ensure disallowed prefixes never execute untrusted module-level code
in either unit or integration flows.
- **`issue-state-updater` Bash Script Errors**: Removed problematic bash script examples
that tried to invoke `task forgejo-label-manager` as a bash command (the Task tool cannot
be invoked from bash). Replaced with clear step-by-step operational instructions and
direct label management via API.
- **`automation-tracking-manager` Label Delegation Syntax**: Fixed incorrect delegation
syntax when calling `forgejo-label-manager`. The manager now uses correct natural language
requests (e.g., "Apply labels to issue #123: Automation Tracking") instead of structured
parameters, ensuring tracking issues receive proper labels.
- **`product-builder` Missing Supervisors**: Added missing `pr-fix-pool-supervisor` and
`pr-merge-pool-supervisor` to the product-builder's supervisor launch list (18 total
supervisors). Updated all numeric references, pre-flight checklists, and validation logic.
- `ActionRepository.update()` now uses explicit bulk `sa_delete()` + `session.flush()`
before re-inserting child rows for `action_arguments` and `action_invariants`, fixing
a `sqlite3.IntegrityError: UNIQUE constraint failed` crash when `agents plan use` was
called on an action that already had arguments registered via `action create`. (#4197)
- **ACMS Indexing Pipeline CLI Wiring**: `ContextTierService` was starting empty on
every CLI invocation, causing the LLM to receive zero file context during plan
execution. Added `context_tier_hydrator.py` that reads files from linked project
resources (via `git ls-files` or `os.walk`) and stores them as `TieredFragment`
objects in the tier service. Hydration runs automatically before context assembly
in `LLMExecuteActor.execute()`. Respects max file size (256KB), total budget (10MB),
binary file exclusion, and `.git`/`node_modules`/`__pycache__` directory skipping.
(#1028)
- **CI Lint**: Resolved 51 ruff violations in `scripts/validate_automation_tracking.py`
(import ordering, deprecated `typing` generics, unused imports, line-length, whitespace).
- **CI Integration Tests**: Removed stale `tdd_expected_fail` tag from
`robot/coverage_threshold.robot` — the underlying bug (issue #4305) is resolved and
the tag was inverting a passing test to a failure. (#5266)
- **Orchestrator Worker Dispatch**: Fixed `verify_worker_started()` to handle the dict
response format from the OpenCode API `/session/status` endpoint instead of an array.
Workers now dispatch and verify correctly, preventing incorrect session deletion.
- **Actor compiler ignores `actor_ref` field on SUBGRAPH nodes** (#1429): Fixed the
actor compiler (`src/cleveragents/actor/compiler.py`) to read `actor_ref` from the
top-level `NodeDefinition.actor_ref` field instead of `node.config.get("actor_ref")`.
Before this fix, `CompiledActor.metadata.subgraph_refs` was always empty and
`NodeConfig.subgraph` on every SUBGRAPH node was always `None`, silently breaking all
hierarchical/nested actor graph compositions. Added Behave regression tests covering
subgraph compilation with `actor_ref` fields and Robot Framework integration tests
verifying that `subgraph_refs` is correctly populated after compilation.
- **Resource Removal Guard** (#6886): Fixed `agents resource remove` silently deleting
parent resources that still have linked children. The guard now queries
`ResourceLinkModel` (the active DAG link table) instead of the legacy
`ResourceEdgeModel`, so the child-link check correctly blocks deletion.
### Fixed
- **TUI — Shell safety integration** (#6361): The Textual prompt now uses
`ShellSafetyService` to analyse shell commands, highlight dangerous input
with `$error` styling, and display an advisory warning banner. Removed the
inline environment-variable gate, added `shell.warn_dangerous` configuration,
and wired the warning indicator into the TUI layout.
---
### Fixed
- **CLI (`agents actor remove`)** (#6491): Restores output parity with the
other actor commands by honoring `--format`/`-f` for JSON/YAML/plain/Rich
envelopes. Adds a Robot Framework regression test to assert the JSON
envelope structure and updates the CLI synopsis in `docs/specification.md`
to document the option.
---
## [3.8.0] -- 2026-04-05
### Added
- Wired Invariant Reconciliation Actor auto-invocation into
`PlanLifecycleService` phase transitions (`start_strategize`,
`execute_plan`, `apply_plan`). Reconciliation failures now block
the transition with `ReconciliationBlockedError` and emit
`INVARIANT_VIOLATED` events. Post-correction reconciliation runs
via `CORRECTION_APPLIED` event subscription (best-effort). Added
`InvariantService` Singleton provider in the DI container.
- **TUI -- Shell danger detection**: The TUI shell mode (`!` prefix) now detects
dangerous command patterns before execution. A configurable pattern registry
classifies commands by danger level (warning, critical) and surfaces a user
warning overlay before proceeding. Patterns cover destructive filesystem
operations, privilege escalation, network exfiltration, and more. (#1003)
- **TUI -- Permission Question Widget**: A new inline `PermissionQuestionWidget`
renders permission requests directly in the conversation stream for single-key
operations. Users can allow/reject with single-key shortcuts (`a`/`A`/`r`/`R`),
navigate with arrow keys, confirm with `Enter`, or press `v` to open the full