7f5a330f23
CI / lint (push) Failing after 20s
CI / helm (push) Successful in 29s
CI / quality (push) Successful in 34s
CI / security (push) Failing after 51s
CI / build (push) Successful in 1m0s
CI / typecheck (push) Failing after 1m2s
CI / coverage (push) Has been skipped
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Failing after 2m11s
CI / docker (push) Has been skipped
CI / e2e_tests (push) Failing after 15m58s
CI / integration_tests (push) Failing after 20m51s
CI / status-check (push) Failing after 1s
CI / benchmark-publish (push) Has been cancelled
- tool add: handle tool: wrapper key (#1471) - session export: sha256: prefix checksum (#1450) - ThoughtBlockWidget: correct background color (#1448) ISSUES CLOSED: #1471 #1450 #1448
1887 lines
127 KiB
Markdown
1887 lines
127 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to this project will be documented in this file.
|
||
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||
|
||
## [Unreleased]
|
||
|
||
### Added
|
||
|
||
- **TUI — First-run experience with actor selection overlay**: On first launch
|
||
(no personas configured), a centred `ActorSelectionOverlay` widget guides the
|
||
user to select an actor from a curated list (`anthropic/claude-4-sonnet`,
|
||
`anthropic/claude-4-opus`, `openai/gpt-4o`, `openai/o3`, `google/gemini-2`).
|
||
Supports keyboard navigation (`j`/`k`), fuzzy search (`/`), and confirmation
|
||
(`enter`). Selecting an actor creates a `"default"` persona and dismisses the
|
||
overlay. Subsequent launches restore the last active persona. (#1391)
|
||
|
||
- **TUI — Session export to Markdown transcript**: `agents session export` gains
|
||
a `--format` flag accepting `json` (default, canonical, re-importable) and
|
||
`md` (human-readable Markdown transcript via `Session.as_export_markdown()`).
|
||
The Markdown format renders a header with session metadata, full message
|
||
history with role/timestamp/content, and linked plan references. The
|
||
`/session:export [--format md] [path]` and `/session:import <path>` slash
|
||
commands are also wired in the TUI command router. (#1004)
|
||
|
||
- **ACMS — UKO provenance tracking and temporal versioning**: The UKO runtime
|
||
now records provenance metadata (`sourceResource`, `validFrom`, `isCurrent`)
|
||
on every typed triple produced by `UKOIndexer.index_graph()`. A revision chain
|
||
tracks ontology state across indexing runs, enabling temporal queries over
|
||
historical UKO snapshots. `UKOGraphPersistence` serialises/restores the full
|
||
graph (including provenance) via JSON-file or in-memory backends, satisfying
|
||
the "persists across restarts" requirement. 47 new BDD scenarios cover typed
|
||
triples, temporal queries, and provenance. (#891)
|
||
|
||
### Fixed
|
||
|
||
- `agents tool add` now accepts YAML configs using the spec-required `tool:` wrapper key
|
||
(e.g. `tool:\n name: ...`). The `cleveragents:` version header is silently ignored.
|
||
Flat format (without wrapper) remains supported for backward compatibility. (#1471)
|
||
|
||
- Session export checksum format corrected from raw hex to `sha256:`-prefixed format as
|
||
required by the specification. (#1450)
|
||
|
||
- `ThoughtBlockWidget` background corrected from `$primary 20%` to `$primary-muted 20%`,
|
||
making thought blocks visually lighter and more subtle per spec §29811. (#1448)
|
||
|
||
## [3.7.0] — 2026-04-02
|
||
|
||
### Added
|
||
|
||
- **TUI — Interactive Terminal UI** (`agents tui`): Full-screen Textual-based
|
||
application with multi-session tabs, real-time plan monitoring, and rich
|
||
conversation with actors. Requires the optional `cleveragents[tui]` extra.
|
||
([ADR-044](docs/adr/ADR-044-tui-architecture-and-framework.md))
|
||
|
||
- **TUI — Persona system**: YAML-backed personas stored in
|
||
`~/.config/cleveragents/personas/`. Each persona binds an actor, optional
|
||
argument presets, and scope references to a named identity. Personas are
|
||
managed via `/persona:*` slash commands or the `PersonaRegistry` API.
|
||
Per-session state (active persona, active preset) is tracked in memory and
|
||
the last-used persona is persisted to `tui-state.yaml`.
|
||
([ADR-045](docs/adr/ADR-045-tui-persona-system.md))
|
||
|
||
- **TUI — Input mode routing**: The prompt auto-detects three input modes from
|
||
the first character — Normal (message + `@reference` expansion), Command
|
||
(`/` slash commands), and Shell (`!` subprocess passthrough). The
|
||
`InputModeRouter` dispatches each mode to the appropriate handler.
|
||
([ADR-046](docs/adr/ADR-046-tui-reference-and-command-system.md))
|
||
|
||
- **TUI — Slash command catalog**: 67 slash commands across 14 groups
|
||
(Session, Persona, Scope, Plan, Project, Actor, Resource, Config, Tool,
|
||
Skill, Invariant, Profile, Context, Utility) exposed via the
|
||
`SlashCommandOverlay` widget and `SLASH_COMMAND_SPECS` catalog.
|
||
|
||
- **TUI — Context-sensitive help panel (F1)**: `HelpPanelOverlay` toggled by
|
||
`F1` with content that adapts to the current prompt context (Main Screen,
|
||
Slash Commands, Reference Picker, Shell Mode). Global key bindings are
|
||
always shown alongside context-specific shortcuts. (#1013)
|
||
|
||
- **TUI — Persona bar**: Bottom status bar (`PersonaBar`) displaying the
|
||
active persona name, bound actor, current argument preset, and scope
|
||
reference count. Cycle presets with `Ctrl+T`.
|
||
|
||
- **Session management**: `agents session` command group with `create`,
|
||
`list`, `show`, `delete`, `export`, `import`, and `tell` subcommands.
|
||
Sessions are persisted via the DI-wired `SessionService` and routed through
|
||
the A2A local facade. Fixed DI container wiring (`session_service` provider)
|
||
that previously caused `AttributeError` on all session commands. (#554, #570)
|
||
|
||
- **Server mode**: `agents server connect` persists server URL and token to
|
||
configuration. `agents server status` reports connection state. Kubernetes
|
||
Helm chart added in `k8s/` for production deployment with Deployment,
|
||
Service, Ingress (TLS), ConfigMap, ServiceAccount, Secrets, and optional
|
||
Redis subchart for multi-instance session affinity. `Dockerfile.server`
|
||
provides a multi-stage ASGI container image. (#928)
|
||
|
||
- **A2A integration**: A2A local facade handlers wired to live application
|
||
services — `session.create`/`close`, `plan.create`/`execute`/`status`/
|
||
`diff`/`apply`, `registry.list_tools`/`list_resources`, `event.subscribe`.
|
||
Domain-to-A2A error code mapping (`NOT_FOUND`, `VALIDATION_ERROR`,
|
||
`INVALID_STATE`, `PLAN_ERROR`). (#501)
|
||
|
||
- **TUI — PermissionsScreen with diff view**: Full-screen overlay for tool
|
||
permission requests. Displays a file list on the left and a diff view on
|
||
the right. Supports three diff display modes (unified, side-by-side,
|
||
context) toggled with `d`. Keyboard bindings: `a` allow-once, `A`
|
||
allow-always, `r` reject-once, `R` reject-always. Permission decisions are
|
||
persisted via `PermissionService`. (#996)
|
||
|
||
- **TUI — Actor thought block rendering**: `ThoughtBlock` domain model
|
||
represents actor reasoning traces with configurable `max_lines` (default
|
||
10), expanded/collapsed state, and helper methods for truncated/full
|
||
content. `ThoughtBlockWidget` renders thought blocks with muted styling and
|
||
a space-bar toggle. (#1001)
|
||
|
||
- **ACMS — UKO runtime operationalized**: Three new services complete the
|
||
Universal Knowledge Ontology runtime per spec §185:
|
||
- `UKOQueryInterface` — typed interface for ACMS context strategies to
|
||
query UKO classification data (layer, primary type, implicit relations).
|
||
- `UKOInferenceEngine` — semantic analysis producing implicit triples
|
||
(`uko:implicitSiblingOf`, `uko:implicitContains`, `uko:implicitDependsOn`)
|
||
with confidence 0.7.
|
||
- `UKOGraphPersistence` — serialises/restores UKO graph state via JSON or
|
||
in-memory backends, satisfying the "persists across restarts" requirement.
|
||
`UKOIndexer.index_graph()` now runs inference and populates `uko:layer`
|
||
triples for all four ontology layers. (#891)
|
||
|
||
- **ACMS — Pipeline Phase 2 protocol aliases**: Spec-aligned
|
||
`Protocol` type aliases added for all Phase 2 (Fragment Fusion) pipeline
|
||
component interfaces: `FragmentDeduplicatorProtocol`,
|
||
`DetailDepthResolverProtocol`, `FragmentScorerProtocol`,
|
||
`BudgetPackerProtocol`, `FragmentOrdererProtocol`. (#540)
|
||
|
||
- **Resource — DevcontainerHandler protocol completion**: Four previously
|
||
missing protocol methods implemented on `DevcontainerHandler`:
|
||
`delete()` (uses `devcontainer exec rm -rf`), `list_children()` (uses
|
||
`devcontainer exec ls -1`), `diff()` (content-hash comparison), and
|
||
`create_sandbox()` (delegates to `BaseResourceHandler` with lazy
|
||
activation). All methods return graceful failure results for missing or
|
||
stopped containers rather than raising. (#1242)
|
||
|
||
- **Resource — DatabaseResourceHandler CRUD and checkpoint methods**:
|
||
Full CRUD and checkpoint implementation for both SQLite and remote
|
||
database types. `read()` queries `sqlite_master` for schema; `write()`
|
||
executes SQL statements; `delete()` executes `DROP TABLE IF EXISTS`;
|
||
`list_children()` lists tables/views; `diff()` compares schemas via
|
||
content hash; `create_checkpoint()` creates a SQLite `SAVEPOINT`;
|
||
`rollback_to()` executes `ROLLBACK TO SAVEPOINT`. Remote database
|
||
operations return not-supported results gracefully. (#1241)
|
||
|
||
- **Estimation lifecycle hook**: `actor.default.estimation` config key
|
||
wired as fallback for plan estimation actor selection. The
|
||
Strategize-to-Estimate lifecycle hook in `PlanLifecycleService` now
|
||
invokes the estimation actor, stores `EstimationResult` on the plan,
|
||
populates `plan.cost_estimate_usd`, and emits a
|
||
`PLAN_ESTIMATION_COMPLETE` domain event. Estimation failures are
|
||
informational only and never block the Execute transition. (#1310)
|
||
|
||
- **Events — `user_identity` field on `DomainEvent`**: All domain events
|
||
now carry an optional `user_identity` field propagated through the full
|
||
event pipeline (event bus → audit subscriber → audit log). (#1257)
|
||
|
||
- **Events — `PLAN_APPLIED` enriched with changeset statistics**:
|
||
`PlanApplyService.apply_with_validation_gate()` now computes
|
||
`files_changed`, `lines_added`, `lines_removed`, `resources_modified`,
|
||
and `apply_duration_seconds` from `SpecChangeSet.summary()` and passes
|
||
them to `PlanLifecycleService.complete_apply()`, which includes them in
|
||
the `PLAN_APPLIED` event details for SEC7 audit logging. (#716)
|
||
|
||
- **Events — `PLAN_CANCELLED` enriched with progress context**: The
|
||
`PLAN_CANCELLED` event now includes progress percentage, completed/total
|
||
action counts, and resource cleanup context. (#1301)
|
||
|
||
- **Server — `agents server serve` subcommand**: BDD scenarios added to
|
||
verify that `Dockerfile.server` uses `python -m cleveragents` as its
|
||
`ENTRYPOINT` and `server serve` as its `CMD`. (#1088)
|
||
|
||
### Changed (original 3.7.0)
|
||
|
||
- `agents actor run` now takes positional `<NAME>` and `<PROMPT>` arguments,
|
||
aligning with the specification. The previous `--prompt/-p` option is
|
||
removed. (#901)
|
||
|
||
- `resource_selection` decision type reclassified from Execute-only to
|
||
phase-agnostic (valid in both Strategize and Execute), aligning with
|
||
ADR-007 and ADR-033. (#931)
|
||
|
||
- `SandboxManager.commit_all()` is now an all-or-nothing atomic operation.
|
||
On partial failure, already-committed sandboxes are rolled back in reverse
|
||
(LIFO) order. (#925)
|
||
|
||
### Fixed
|
||
|
||
- `agents session list`, `session create`, and other session subcommands no
|
||
longer raise `AttributeError: 'DynamicContainer' object has no attribute
|
||
'db'` after `agents init`. (#554, #570, #680)
|
||
|
||
- `plan execute` no longer fails with stale-cache state when run in a
|
||
separate CLI process from `plan use`. (#960)
|
||
|
||
- `project context set` now commits changes via `session.commit()` instead
|
||
of `session.flush()`, preventing silently lost data. (#745)
|
||
|
||
- `agents action create` now accepts `--format`/`-f` flag, matching all
|
||
other action subcommands. (#959)
|
||
|
||
## Unreleased (pre-3.7.0)
|
||
|
||
- Eliminated redundant fields (`excluded_decisions`, `rollback_tier_depth`,
|
||
`child_plans_to_rollback`) from `CorrectionDryRunReport` that duplicated data
|
||
already present in the embedded `CorrectionImpact` object. Consumers now
|
||
access these values via `report.impact.excluded_decisions`,
|
||
`report.impact.rollback_tier_depth`, and `report.impact.affected_child_plans`
|
||
respectively. Updated `CorrectionService.generate_dry_run_report()`, Behave
|
||
step definitions, and Robot Framework helpers to use the canonical `impact`
|
||
sub-object. (#1087)
|
||
- Added a context-sensitive TUI help panel overlay toggled by `F1`, with
|
||
help content that varies for main-screen, slash-command, reference, and
|
||
shell prompt modes. Updated Behave and Robot coverage for help-panel
|
||
rendering and mode switching. (#1013)
|
||
- Added direct BDD coverage for `_fast_init_or_upgrade` early-return behavior,
|
||
template-copy/fallback delegation paths, and the existing-empty-DB branch in
|
||
`features/fast_init_upgrade.feature`. Uses race-safe temp-path allocation
|
||
(`mkstemp`/`mkdtemp`) throughout new fast-init test steps. (#733)
|
||
- Added `get_hover` and `get_definitions` methods to `LspClient` and
|
||
`LspRuntime`, completing the functional LSP runtime (Epic #824).
|
||
`LspClient.get_hover()` sends `textDocument/hover` and returns the
|
||
hover result dict. `LspClient.get_definitions()` sends
|
||
`textDocument/definition` and handles Location, Location[], and
|
||
LocationLink[] responses. `LspRuntime` wrappers add input validation,
|
||
file reading, language detection, and 1-based to 0-based line/column
|
||
conversion. Tool adapter now dispatches HOVER and DEFINITIONS
|
||
capabilities to the runtime instead of raising `LspNotAvailableError`.
|
||
Includes 10 Behave BDD scenarios. (#824)
|
||
- Expanded the TUI slash command overlay catalog to include 67 commands across
|
||
14 groups, aligned with the specification command reference for session,
|
||
persona, scope, plan, project, registry/config, context, and utility flows.
|
||
Added Behave coverage for command and group cardinality plus representative
|
||
command presence checks. (#1002)
|
||
- Strengthened WF02 trusted-profile integration coverage for automated test
|
||
generation. The helper now validates and incorporates mocked provider output
|
||
(instead of discarding it), asserts all WF02 invariant conventions (test-only
|
||
paths, `test_<module>.py` naming, and fixture/conftest usage), verifies
|
||
user-facing `plan artifacts` dispatch behavior, and adds explicit negative
|
||
guardrail tests for absolute/traversal/non-tests destinations. Also narrowed
|
||
Ruff suppression scope in the WF02 helper. (#766)
|
||
- Added TDD bug-capture tests for bug #1025 — ``plan correct`` auto-resolve
|
||
fails in isolated E2E environments. Two Behave BDD scenarios
|
||
(``@tdd_bug @tdd_bug_1025 @tdd_expected_fail``) verify that
|
||
``_resolve_active_plan_id()`` finds an ``Execute/COMPLETE`` plan when
|
||
``--plan`` is omitted. Two Robot Framework integration tests exercise
|
||
the same path in a subprocess context. Tests simulate the divergent-
|
||
container condition (fresh ``CLEVERAGENTS_HOME`` with empty database).
|
||
ASV benchmark measures active-plan filtering overhead. (#1035)
|
||
- Added missing `LspServerConfig` model fields per specification:
|
||
`description` (max 1000 chars), `transport` (`LspTransport` enum with
|
||
`stdio`/`tcp`, default `stdio`), `initialization` (dict for LSP
|
||
`initializationOptions`), and `workspace_settings` (dict for
|
||
`workspace/didChangeConfiguration`). Updated `agents lsp show` Rich
|
||
output to display new fields. All fields have defaults for backward
|
||
compatibility. Includes 20 Behave scenarios and 5 Robot tests. (#835)
|
||
- Added E2E test for Workflow Example 18: Container with Remote Repo Clone
|
||
(trusted profile). Introduces the new `--clone-into` CLI flag on
|
||
`resource add` for container-instance and devcontainer-instance resources
|
||
(format: `REPO_URL:CONTAINER_PATH`), with input validation and type
|
||
restriction. Exercises two-step project creation and linking, plan-level
|
||
`--execution-environment` with `--execution-env-priority fallback`, and full
|
||
plan lifecycle including container commit/push verification on apply.
|
||
(`robot/e2e/wf18_container_clone.robot`,
|
||
`src/cleveragents/cli/commands/resource.py`) (#764)
|
||
- Fixed execution environment resolution to honour project-level override
|
||
(precedence level 2). Threaded `plan_env` and `project_env` through
|
||
`ToolCallRouter`, `ToolCallingRuntime`, and `PlanExecutionContext` so
|
||
the resolver receives project-level execution environment values stored
|
||
in `ContextConfig.execution_environment`. (#1080)
|
||
- Added E2E test for Workflow Example 12 — large-scale hierarchical feature
|
||
implementation (supervised profile). Covers 4-project setup with per-project
|
||
invariants, spec-compliant action YAML (estimation_actor, invariant_actor,
|
||
automation_profile: cautious, action-level invariants), all-project plan use,
|
||
hierarchical tree inspection, plan correct (append mode) on non-root decision,
|
||
phased lifecycle-apply, and terminal-state verification via JSON status.
|
||
Dynamic actor selection and UUID-suffixed names for CI safety.
|
||
Known limitations: `plan prompt` not yet implemented as CLI subcommand,
|
||
action `--arg` omitted due to UNIQUE constraint bug, validation registration
|
||
omitted pending independent validation. (#758)
|
||
- Added E2E test for Workflow Example 17: explicit container with directory
|
||
mount using trusted automation profile. Exercises container-instance resource
|
||
registration, project link-resource, execution environment setting via
|
||
project context set, plan-level execution-env-priority override via plan use,
|
||
and full plan lifecycle with dynamic actor selection. Includes TDD
|
||
bug-capture tests for deferred acceptance criteria: dual mount registration
|
||
(#1078), project-level execution-env-priority (#1079), and precedence
|
||
level 2 resolution (#1080).
|
||
(`robot/e2e/wf17_explicit_container.robot`) (#763)
|
||
- Added `correction_attempts` table per specification DDL with
|
||
`CorrectionAttemptModel` ORM, `CorrectionAttemptRecord` domain model,
|
||
`CorrectionAttemptRepository` CRUD layer, Alembic migration, and
|
||
`CorrectionAttemptState` enum. Repository `update_state()` accepts
|
||
typed `CorrectionAttemptState` enum and `datetime` parameters and
|
||
enforces the spec lifecycle (`pending → executing → complete|failed`)
|
||
via `InvalidCorrectionStateTransitionError`.
|
||
`CorrectionAttemptRecord.guidance` validates non-empty with
|
||
`max_length=10_000`. `created_at` column includes spec-aligned
|
||
`server_default`; `to_domain()` normalises naive timestamps to UTC.
|
||
State transition validation extracted to domain-level
|
||
`validate_correction_state_transition()` function with
|
||
`CORRECTION_ATTEMPT_VALID_TRANSITIONS` and
|
||
`CORRECTION_ATTEMPT_TERMINAL_STATES` constants.
|
||
`update_state()` rejects `completed_at` on non-terminal transitions.
|
||
Improved FK-violation error messages in `create()` and `update_state()`.
|
||
Normalised timestamp format in `from_domain()` to millisecond precision
|
||
(`SS.mmm`) matching SQLite `server_default` `strftime('%f')` output
|
||
for consistent string-based ordering.
|
||
`CORRECTION_ATTEMPT_VALID_TRANSITIONS` and
|
||
`CORRECTION_ATTEMPT_TERMINAL_STATES` now use typed
|
||
`CorrectionAttemptState` enum keys/values.
|
||
Updated repository module docstring tables.
|
||
`from_domain()` normalises timestamps to UTC via `astimezone(UTC)`
|
||
before formatting, preventing silent data loss for non-UTC datetimes.
|
||
`update_state()` auto-sets `completed_at` when transitioning to
|
||
terminal states if not explicitly provided.
|
||
`CorrectionAttemptRecord` `plan_id` and `original_decision_id`
|
||
validators now return stripped values, preventing whitespace-padded
|
||
IDs from causing FK lookup failures.
|
||
Added new domain exports to `__init__.py` `__all__`.
|
||
Aligned `update_state()` `completed_at` timestamp to millisecond
|
||
precision for consistency.
|
||
Improved FK-violation error message in `update_state()` to avoid
|
||
misleading reference when `new_decision_id` is `None`.
|
||
Removed unnecessary `session.rollback()` in read-only repository
|
||
methods (`get()`, `list_by_plan()`) for consistency with other repos.
|
||
Added `created_at` and `completed_at` Pydantic validators on
|
||
`CorrectionAttemptRecord` to normalise naive datetimes to UTC,
|
||
preventing `ValueError` in `from_domain()` `astimezone()` calls.
|
||
Added defensive enum coercion in `CorrectionAttemptModel.to_domain()`
|
||
with warning-level logging for invalid `mode`/`state` DB values,
|
||
consistent with `LifecyclePlanModel.to_domain()` pattern.
|
||
Moved `CorrectionAttemptState` from `TYPE_CHECKING`-only to runtime
|
||
import in the repository module, removing redundant in-method import.
|
||
Changed `original_decision_id` FK from `CASCADE` to `RESTRICT`
|
||
matching the spec DDL default and the codebase convention for
|
||
non-dependency FK references to decisions, preserving correction
|
||
audit trail when decisions are cleaned up.
|
||
Added `new_decision_id` strip-and-validate field validator matching
|
||
the pattern used for `plan_id` and `original_decision_id`.
|
||
Added input validation for `new_decision_id` in `update_state()`
|
||
rejecting empty and whitespace-only values per CONTRIBUTING.md
|
||
argument validation guidelines.
|
||
Fixed BDD mode-validation scenario to use dedicated `Then` step
|
||
with field-level assertion instead of reusing guidance error step.
|
||
Moved `new_decision_id` and `archived_artifacts_path` argument
|
||
validation in `update_state()` before any ORM row mutations per
|
||
CONTRIBUTING.md early-validation guidelines, preventing dirty
|
||
session state on validation failure.
|
||
Added `archived_artifacts_path` empty/whitespace-only rejection
|
||
in `update_state()` matching the `new_decision_id` validation
|
||
pattern per CONTRIBUTING.md argument validation guidelines.
|
||
43 BDD scenarios and 5 Robot integration tests including cascade
|
||
deletion, terminal-state rejection, failed-path transition, guidance
|
||
validation, max-length boundary, min-length boundary, not-found
|
||
update, FK-violation update, completed_at guard, timezone
|
||
normalization, archived_artifacts_path round-trip, delete-in-complete-
|
||
state, cross-plan list isolation, auto-set completed_at on terminal
|
||
transition, FK violation on create, invalid mode rejection,
|
||
whitespace/empty `new_decision_id` rejection, combined field
|
||
update, self-transition rejection, and empty/whitespace
|
||
`archived_artifacts_path` rejection.
|
||
Fixed `update_state()` bug where `archived_artifacts_path` was
|
||
stored without stripping leading/trailing whitespace, unlike
|
||
`new_decision_id` which correctly used the stripped value.
|
||
Extracted `SQLITE_TIMESTAMP_MS_LEN` constant and
|
||
`format_sqlite_timestamp()` helper for millisecond-precision
|
||
timestamp formatting, used by both `from_domain()` and
|
||
`update_state()`.
|
||
Changed `InvalidCorrectionStateTransitionError` base class from
|
||
`DatabaseError` to `BusinessRuleViolation` per CONTRIBUTING.md
|
||
exception semantics (state transition is a business rule, not a
|
||
database error).
|
||
Changed `new_decision_id` FK from `SET NULL` to `RESTRICT` matching
|
||
the spec DDL default (no ON DELETE clause) and consistent with
|
||
`original_decision_id`.
|
||
Changed `update_state()` input validation for `new_decision_id` and
|
||
`archived_artifacts_path` from `DatabaseError` to `ValueError` per
|
||
CONTRIBUTING.md argument validation guidelines.
|
||
Defensive `to_domain()` coercion now defaults corrupted state to
|
||
`failed` (terminal) instead of `pending`, preventing re-execution
|
||
of completed/failed corrections with corrupted DB values.
|
||
45 BDD scenarios (was 43) with new RESTRICT FK test for
|
||
`original_decision_id` and stronger cross-plan isolation test.
|
||
Added ORM-level `relationship(cascade="all, delete-orphan")` on
|
||
`LifecyclePlanModel` for `CorrectionAttemptModel`, consistent with
|
||
all other `v3_plans` child tables, ensuring ORM-level cascade
|
||
deletes work even when SQLite FK enforcement is disabled.
|
||
Added defensive `to_domain()` coercion for corrupted `guidance`
|
||
column (defaults to `"[corrupted]"` with warning log), consistent
|
||
with existing mode/state coercion pattern.
|
||
Added `ValueError` guard in `format_sqlite_timestamp()` rejecting
|
||
naive datetimes per CONTRIBUTING.md fail-fast argument validation.
|
||
Fixed `update_state()` to defensively handle corrupted DB state
|
||
values (coerces to `failed` terminal with warning log), consistent
|
||
with `to_domain()` defensive coercion pattern.
|
||
Strengthened RESTRICT FK BDD assertion to verify exception type
|
||
(`IntegrityError`/`DatabaseError`) instead of only checking
|
||
presence.
|
||
Split multi-When/Then cross-plan isolation BDD scenario into
|
||
idiomatic single-When/Then scenarios.
|
||
53 BDD scenarios (was 45) with new defensive `to_domain()` coercion
|
||
tests (corrupted mode/state/guidance), `format_sqlite_timestamp()`
|
||
naive datetime rejection, domain model naive datetime normalisation,
|
||
and corrupted DB state handling in `update_state()`.
|
||
(#920)
|
||
- Hardened automation profile configuration validation after the task-flag
|
||
rename: `AutomationProfile` now rejects unknown top-level fields instead of
|
||
silently ignoring them, and raises an actionable ``ValueError`` listing the
|
||
required renames when legacy ``auto_*`` keys are supplied (instead of a
|
||
generic Pydantic "Extra inputs are not permitted" error). Updated
|
||
automation profile schema and documentation references (`specification`,
|
||
ADRs, and reference docs) to the task-type field names, added BDD coverage
|
||
for rejecting legacy threshold keys in `automation-profile add`, restored
|
||
docs-schema parity for optional `guards` profile configs, aligned the
|
||
`m5_001` migration header metadata text with its actual revision chain,
|
||
updated M6 fixture files to use the spec-defined field names, added
|
||
phase-transition semantic bridge comments in `PlanLifecycleService`.
|
||
Restored categorised CLI ``automation-profile show`` output to match the
|
||
specification (Phase Transitions / Decision Automation / Self-Repair /
|
||
Execution Controls), added missing ``access_network`` field to spec ``show``
|
||
output examples, aligned ADR-017 and reference doc descriptions with the
|
||
specification's Automatable Tasks table (all 11 fields), and extended the
|
||
repository roundtrip test to assert all 11 threshold fields. Fixed
|
||
benchmark ``_make_profile()`` helper passing safety fields as top-level
|
||
kwargs instead of via ``SafetyProfile`` sub-model (incompatible with
|
||
``extra="forbid"``). Aligned CLI JSON/YAML output structure for
|
||
``automation-profile show`` with the specification's grouped format
|
||
(``phase_transitions``, ``decision_automation``, ``self_repair``,
|
||
``execution_controls``). Moved safety boolean fields into the Execution
|
||
Controls section of Rich output per spec examples. Reverted ``auto``
|
||
profile description to "Fully automatic except apply" per specification.
|
||
Added comprehensive field-name mapping table to ``automation_profile.py``
|
||
module docstring documenting the old phase-transition names to new
|
||
spec task-type names correspondence and added cross-reference in
|
||
``AutonomyController._get_threshold()`` docstring.
|
||
(#902)
|
||
- Added TDD bug-capture tests for bug #1141 — session create does not persist
|
||
into subsequent session list output. Added a Behave scenario and Robot E2E
|
||
test with required tags (`@tdd_bug`, `@tdd_bug_1141`, `@tdd_expected_fail` /
|
||
`tdd_bug`, `tdd_bug_1141`, `tdd_expected_fail`) to assert create→list should
|
||
show one session. The underlying assertion currently fails and is intentionally
|
||
inverted until bug #1141 is fixed. (#1142)
|
||
- Fixed `project context set` missing `--execution-env-priority` flag.
|
||
Setting is persisted and displayed by `project context show`.
|
||
Project-level priority propagates to `plan use` when no plan-level
|
||
override is specified. (#1079)
|
||
- Added Robot Framework integration test for Specification Workflow Example 5:
|
||
Database Schema Migration with Safety Nets. 7 test cases exercising review
|
||
automation profile, custom `local/postgres-db` resource type with
|
||
`transaction_rollback` sandbox and spec-matching cli_args (host/port/database/
|
||
schema), project-resource linking, custom skill with 3 database tools,
|
||
action with 4 typed args and 4 invariants, checkpoint creation/rollback via
|
||
CheckpointManager, 5-phase sequential SubplanService.spawn with fail-fast,
|
||
and plan lifecycle through strategize-to-execute. (#769)
|
||
- Implemented `--mount` flag on `resource add container-instance`. Supports
|
||
resource-reference mounts (`--mount local/api-repo:/workspace`) and
|
||
host-path mounts (`--mount /var/config:/config:ro`). Multiple `--mount`
|
||
flags can be specified. Mount info is persisted as JSON in resource
|
||
properties and displayed by `resource show`. (#1078)
|
||
- Added TDD bug-capture tests for #1078 — resource add container-instance
|
||
missing --mount flag. Three Behave scenarios prove the --mount flag is not
|
||
recognised. Uses @tdd_expected_fail until fix is merged. (#1099)
|
||
- Added WF03 plan prompt and confidence-threshold pausing tests.
|
||
Behave BDD scenarios and Robot Framework integration tests exercise
|
||
`plan prompt` via the A2A facade dispatch path (S15822) -- verifying
|
||
guidance propagation -- and verify cautious-profile confidence-threshold
|
||
pausing (S37262-37367) including a pause-and-resume flow. Facade stub
|
||
updated to echo guidance text. (#961)
|
||
- Added TDD bug-capture tests for bug #1023: CLI commands fail without explicit
|
||
`agents init` when `CLEVERAGENTS_AUTO_APPLY_MIGRATIONS=true` is set. Two
|
||
Behave BDD scenarios and two Robot Framework integration tests verify that
|
||
`resource add` and `project create` succeed in a fresh environment without
|
||
prior init. Tests use `@tdd_expected_fail` until the bug fix is merged.
|
||
(#1033)
|
||
- Added TDD bug-capture tests for bug #1079 — `project context set` missing
|
||
`--execution-env-priority` flag. Six Behave BDD scenarios exercise the CLI
|
||
path (override, fallback, rejection without `--execution-environment`, default
|
||
value, invalid value, and round-trip via `context show`). Three Robot
|
||
Framework integration tests cover override acceptance, fallback acceptance,
|
||
and full persistence round-trip. Tests use `@tdd_expected_fail` until the
|
||
fix is merged. (#1100)
|
||
- Fixed Robot Framework test mocks for ``plan correct`` dry-run and correction
|
||
subplan helpers to use ``container.decision_service()`` instead of the
|
||
non-existent ``container.resolve()``, matching corrected production code.
|
||
Activated regression-guard BDD scenarios for ``plan tree``, ``plan explain``,
|
||
and ``plan correct``. (#647)
|
||
- Added the production ACMS skeleton compression stage via
|
||
`DepthReductionCompressor`. The pipeline now re-renders inherited parent
|
||
fragments to overview depths 0-1 using the UKO detail-level map chain,
|
||
exposes the compressor as the configured builtin, and covers the behavior
|
||
with BDD scenarios for compressor output and default pipeline wiring. (#919)
|
||
- Added TDD bug-capture tests for bug #1076 — `use_action()` does not
|
||
propagate `automation_profile` to Plan. Three Behave BDD scenarios
|
||
(`@tdd_bug @tdd_bug_1076 @tdd_expected_fail`) verify the full precedence
|
||
chain (action, project-scoped config, global default) for automation
|
||
profile resolution at `plan use` time. Tests prove the bug exists: the
|
||
Plan's `automation_profile` is always `None` regardless of the Action's
|
||
profile, project config, or global default. The `@tdd_expected_fail` tag
|
||
inverts this to a CI pass until the fix is merged. (#1098)
|
||
- Added TDD bug-capture tests for bug #1022 — InvariantService in-memory
|
||
storage only. Four Behave BDD scenarios and three Robot Framework
|
||
integration tests verify invariant persistence across simulated CLI
|
||
process restarts. Tests use `@tdd_expected_fail` until bug #1022 is
|
||
fixed. (#1032)
|
||
- Added TDD bug-capture test for bug #988 — ReactiveEventBus.emit() swallows
|
||
exception details. Behave BDD scenario (`@tdd_bug @tdd_bug_988
|
||
@tdd_expected_fail`) captures the missing exception message and traceback in
|
||
the emit() exception handler. The test subscribes a handler that raises
|
||
ValueError with a distinctive message and asserts the message appears in the
|
||
structlog warning log — which currently fails, confirming the bug. The
|
||
`@tdd_expected_fail` tag inverts this to a CI pass until the fix is merged.
|
||
(#1093)
|
||
- Added ResourceHandler sandbox and checkpoint lifecycle methods:
|
||
`create_sandbox` (idempotent, delegates to SandboxManager),
|
||
`create_checkpoint`, `rollback_to`, and `project_access`. Frozen
|
||
dataclass result types (SandboxResult, CheckpointResult, RollbackResult,
|
||
AccessResult) added to the handler protocol. GitCheckoutHandler uses
|
||
`git tag` for checkpoint and `git checkout` for rollback.
|
||
FsDirectoryHandler uses `shutil.copytree` snapshot and clear-and-restore.
|
||
Default `project_access` delegates to PermissionService (local mode =
|
||
always permit). (#836)
|
||
- Added 5 missing LSP capabilities to `LspCapability` enum: `HOVER`,
|
||
`DEFINITIONS`, `SIGNATURE_HELP`, `DOCUMENT_SYMBOLS`, `WORKSPACE_SYMBOLS`.
|
||
Renamed `TYPE_INFO` -> `HOVER`, `SYMBOLS` -> `DOCUMENT_SYMBOLS`,
|
||
`FORMAT` -> `FORMATTING` for spec alignment. Updated tool adapter with
|
||
11 capability mappings, RENAME schema with `new_name` parameter,
|
||
workspace-symbols query schema, and defensive schema validation.
|
||
Extended `initialize()` to advertise all 11 capabilities. Fixed
|
||
`workspace_symbols` runtime handler to accept query-only input. (#834)
|
||
- Added per-phase ACMS context analysis summaries for
|
||
`agents project context inspect/simulate` with human-readable metrics
|
||
(fragment/resource counts, size/tokens, budget utilization), explicit
|
||
strategize->execute->apply narrowing diagnostics, Robot acceptance
|
||
verification upgrades, and Behave edge-case coverage for empty,
|
||
single-resource, multi-resource, and budget-constrained contexts. (#849)
|
||
- Added 10,000-file ACMS indexing reliability improvements: configurable
|
||
runtime timeout bounds propagated through the repository indexing service,
|
||
utility walker, and CLI (`agents repo index --timeout-seconds`), plus
|
||
large-scale verification coverage in Behave and Robot and a dedicated
|
||
10K-file indexing benchmark path for regression tracking. (#851)
|
||
- Added TDD bug-capture E2E tests for bug #1028 — ACMS indexing pipeline not
|
||
wired into CLI. Four Robot Framework E2E tests prove ContextTierService starts
|
||
empty on every CLI invocation. Tests use `@tdd_expected_fail` until the bug
|
||
fix is merged. (#1029)
|
||
- Added Fix-then-Revalidate orchestration loop for required validations:
|
||
bounded retry with configurable limits (0--100 per Safety Profile),
|
||
strategy revision escalation via `auto_strategy_revision` float
|
||
threshold, user escalation via `needs_user_escalation` result flag,
|
||
and domain events (`VALIDATION_FIX_ATTEMPTED`, `VALIDATION_FIX_SUCCEEDED`,
|
||
`VALIDATION_FIX_EXHAUSTED`). Validation errors are treated as required
|
||
failures regardless of mode. Includes `auto_validation_fix` threshold,
|
||
per-resource retry tracking, early-exit signalling via `None` return from
|
||
`FixCallback`, event bus circuit breaker with lock-protected failure
|
||
counter, spec-required `validation_summary` and
|
||
`final_validation_results` fields on the result model, DI container
|
||
registration per ADR-003, and structured logging via `structlog`.
|
||
(#583)
|
||
- Implemented real revert-mode re-execution from decision point.
|
||
`CorrectionService.execute_revert` now performs checkpoint restoration via
|
||
`CheckpointService`, extracts `actor_state_ref` from the target decision's
|
||
context snapshot for reasoning rollback, generates a `user_intervention`
|
||
decision ID for guidance injection, and signals phase transition to
|
||
Strategize. Added `checkpoint_restored`, `actor_state_ref`,
|
||
`user_intervention_decision_id`, and `phase_transition_target` fields to
|
||
`CorrectionResult`. Includes 16 Behave BDD scenarios and 7 Robot Framework
|
||
integration tests. (#844)
|
||
- Added LSP resource types: `executable`, `lsp-server`, `lsp-workspace`,
|
||
`lsp-document` with parent/child hierarchy, auto-discovery rules, and
|
||
handler references. Registered in bootstrap, with YAML configurations,
|
||
Behave BDD tests (21 scenarios), and Robot integration tests (6 tests). (#832)
|
||
- Implemented functional LSP runtime replacing local-mode stubs with real
|
||
LSP protocol support. `StdioTransport` manages server subprocesses via
|
||
JSON-RPC over stdin/stdout. `LspClient` implements initialize/shutdown/
|
||
diagnostics/completions. `LspLifecycleManager` provides reference-counted
|
||
instances, health checks, and crash restart. `LanguageDiscovery` implements
|
||
4-layer detection (extension, shebang, UKO, project config). Includes
|
||
27 Behave BDD scenarios and 6 Robot integration tests. (#826)
|
||
- Added ResourceHandler CRUD and discovery methods: read, write, delete,
|
||
list_children, diff, and discover_children. Frozen dataclass result types
|
||
(Content, WriteResult, DeleteResult, DiffResult) added to the handler
|
||
protocol. GitCheckoutHandler implements all six methods via git plumbing
|
||
and filesystem operations. FsDirectoryHandler implements all six via
|
||
pathlib/os/difflib. DevcontainerHandler implements read, write, and
|
||
discover_children via `devcontainer exec`. DatabaseResourceHandler
|
||
inherits NotImplementedError stubs pending connection management. (#827)
|
||
- Implemented ACMS context tier runtime promotion/demotion/eviction:
|
||
auto-promotion on access with configurable threshold (default: 5),
|
||
time-based staleness enforcement (hot/warm TTL, default: 24h each),
|
||
budget-based LRU eviction on hot-tier overflow, and tier transition
|
||
event emission (TIER_PROMOTED, TIER_DEMOTED, TIER_EVICTED) via
|
||
EventBus. Added `context_tier_promotion_threshold`,
|
||
`context_tier_hot_ttl_hours`, and `context_tier_warm_ttl_hours`
|
||
settings with DI wiring of event_bus into ContextTierService.
|
||
Oversized fragments that exceed the entire hot-tier budget are now
|
||
redirected to the warm tier with a TIER_DEMOTED event. Promotion
|
||
to hot falls back to warm when the promoted fragment is evicted by
|
||
budget enforcement. Event emission is best-effort; a failing event
|
||
bus no longer breaks tier operations. Added `CLEVERAGENTS_CTX_HOT_HOURS`
|
||
env var alias for `context_tier_hot_ttl_hours` for consistency with
|
||
the warm-tier alias. Demotion now resets `access_count` to zero so
|
||
that demoted fragments must accumulate fresh accesses before
|
||
re-promotion, preventing staleness enforcement from being immediately
|
||
undone by a single access. (#821)
|
||
- Added byte-size budget enforcement for the ACMS context assembly
|
||
pipeline. `enforce_size_budget()` filters context fragments against
|
||
`max_file_size` (per-fragment) and `max_total_size` (cumulative)
|
||
limits defined in a `ContextView`. New domain models
|
||
`BudgetViolation` and `BudgetEnforcementResult` provide structured
|
||
violation reporting. Pipeline integration in `ACMSPipeline.assemble()`
|
||
applies enforcement as a pre-filter when a `context_view` is
|
||
provided. (#847)
|
||
- Added E2E test for Workflow Example 16: devcontainer-driven development
|
||
with supervised automation profile. Exercises devcontainer auto-detection
|
||
during resource registration, lazy container build during plan execution,
|
||
tool invocation routing to container workspace, and apply writing changes
|
||
back to host filesystem via bind mount. Uses dynamic actor selection
|
||
(Anthropic/OpenAI) and UUID-suffixed names for parallel CI safety.
|
||
(`robot/e2e/wf16_devcontainer.robot`) (#762)
|
||
- **Breaking (behavioral):** `SandboxManager.commit_all()` is now an
|
||
all-or-nothing atomic operation per specification line 45938. (#925)
|
||
- On partial failure, already-committed sandboxes are rolled back in
|
||
reverse (LIFO) order following the standard transaction-log undo
|
||
pattern.
|
||
- Non-`SandboxError` exceptions are wrapped in `AtomicCommitError`
|
||
(chaining the original as `__cause__`) with `rolled_back_ids` and
|
||
`failed_rollback_ids` attributes; `SandboxError` exceptions return
|
||
a `CommitResult` with `rolled_back` / `rollback_failed` metadata.
|
||
- Non-rollbackable sandboxes (`NoSandbox`, `TransactionSandbox`) are
|
||
committed last in the batch so rollbackable sandboxes can be undone
|
||
if they fail first. A warning is logged when these types are present.
|
||
- `TransactionSandbox.rollback()` from `COMMITTED` now raises
|
||
`SandboxRollbackError` (database commits are irreversible) instead
|
||
of silently reporting success.
|
||
- Extracted shared `_fs_utils` module (`backup_directory`,
|
||
`safe_restore`, `compute_diff`) with symlink, permission, and
|
||
timestamp preservation; replaces duplicated per-class methods.
|
||
- `safe_restore` uses rename-based swap (both renames are O(1) on
|
||
the same filesystem) to prevent data loss during restore.
|
||
- Pre-commit backup is created on the same filesystem as the original
|
||
(avoids cross-device copy overhead); assigned only after
|
||
`backup_directory()` succeeds; skipped when no changes detected.
|
||
- `backup_directory` defers directory permissions and timestamps to a
|
||
bottom-up post-walk pass, fixing POSIX mtime overwrite; skips
|
||
non-regular files (FIFOs, sockets, device files) with a warning.
|
||
- `CopyOnWriteSandbox` and `OverlaySandbox` commit error handler
|
||
restores original from pre-commit backup; catches `Exception`
|
||
(not just `OSError`) so unexpected errors also trigger restoration.
|
||
- Rollback from `COMMITTED` with no backup (no changes applied) is
|
||
a no-op instead of raising `SandboxRollbackError`.
|
||
- `CopyOnWriteSandbox` and `OverlaySandbox` rollback from `COMMITTED`
|
||
resets sandbox copy/merged directory from restored original,
|
||
preventing stale data from being exposed on re-activation.
|
||
- `OverlaySandbox` rollback from `COMMITTED` properly remounts
|
||
OverlayFS (or re-copies for userspace fallback); raises
|
||
`SandboxRollbackError` if unmount fails. No longer double-wraps
|
||
`SandboxRollbackError` — inner errors are re-raised directly.
|
||
- `OverlaySandbox.get_path()` now transitions `ROLLED_BACK → ACTIVE`
|
||
for consistency with `CopyOnWriteSandbox` and the protocol
|
||
transition table.
|
||
- `GitWorktreeSandbox.get_path()` now accepts `ROLLED_BACK` status
|
||
for consistency with all other sandbox types and the protocol
|
||
transition table (`ROLLED_BACK → ACTIVE`).
|
||
- `rollback_all` now also handles sandboxes in `COMMITTED` status
|
||
and catches `Exception` (not just `SandboxError`) to ensure all
|
||
rollbacks are attempted.
|
||
- `cleanup_all` now catches `Exception` (not just `SandboxError`)
|
||
to prevent a single unexpected error from aborting cleanup of
|
||
remaining sandboxes.
|
||
- `CopyOnWriteSandbox` and `OverlaySandbox` rollback from `ACTIVE`
|
||
now uses `dirs_exist_ok=True` to prevent `FileExistsError` when
|
||
`rmtree` silently fails.
|
||
- `cleanup_abandoned` now catches `Exception` (not just
|
||
`SandboxError`) so that unexpected errors do not crash the loop
|
||
and prevent remaining abandoned sandboxes from being cleaned up,
|
||
consistent with `cleanup_all`, `rollback_all`, and
|
||
`_rollback_committed`.
|
||
- `OverlaySandbox._mount_overlay()` and `_unmount_overlay()` now
|
||
catch `subprocess.TimeoutExpired` (in addition to
|
||
`CalledProcessError` and `OSError`), preventing `create()` from
|
||
leaving the sandbox in `PENDING` status and `cleanup()` from
|
||
leaving it in a zombie state when `mount`/`umount` hangs.
|
||
- `OverlaySandbox._mount_overlay()` validates that overlay paths
|
||
do not contain commas, which would corrupt the OverlayFS mount
|
||
options string.
|
||
- `GitWorktreeSandbox.commit()` now checks `git diff` return code
|
||
so that a failed diff command does not silently skip the merge.
|
||
- `safe_restore` cleanup of the temporary rollback container now
|
||
runs in a `finally` block, preventing a temp directory leak when
|
||
the rename fails and the exception is re-raised.
|
||
- `AtomicCommitError` exported from `sandbox` package `__init__.py`.
|
||
- Aligned plan lifecycle model with specification: ERRORED is now
|
||
terminal in `is_terminal`, per-phase state validation enforces
|
||
APPLIED/CONSTRAINED to APPLY-only and COMPLETE to
|
||
STRATEGIZE/EXECUTE-only via model validator, COMPLETE docstring
|
||
clarified as phase-level terminal. Added defensive coercion in
|
||
database deserialization for legacy invalid phase/state
|
||
combinations with warning-level logging. Fixed assignment ordering
|
||
in `execute_plan()` for consistency with phase-state validator.
|
||
Updated `PlanResumeService` docstring to reflect ERRORED
|
||
terminality. (#918)
|
||
- Aligned `v3_plans` table schema with specification DDL: added
|
||
`effective_profile_snapshot` column (TEXT NOT NULL, validated as JSON),
|
||
made `root_plan_id` NOT NULL with self-referencing for root plans and
|
||
explicit `RESTRICT` FK policy, made `automation_profile` NOT NULL with
|
||
default `"balanced"`, and documented the intentional `phase` default
|
||
deviation (`"action"` vs spec `"strategize"`). Migration backfill
|
||
correctly resolves root ancestors for child plans at arbitrary hierarchy
|
||
depth via level-by-level propagation with safety bound. Hardened
|
||
`automation_profile` deserialization to catch `RecursionError`.
|
||
Hardened `effective_profile_snapshot` deserialization in `to_domain()`
|
||
to gracefully fall back to `'{}'` on corrupted JSON, preventing a
|
||
single corrupted row from crashing plan reads. Added `RecursionError`
|
||
to `effective_profile_snapshot` Pydantic validator for consistency
|
||
with `automation_profile` deserialization. Validator error message
|
||
now uses length-only to avoid potential information disclosure.
|
||
Documented intentional column naming conventions vs spec DDL.
|
||
Migration orphan-row fallback now logs affected row count.
|
||
Documented FK ondelete policy drift between ORM and migrated schemas.
|
||
Documented `automation_profile` dual-format storage semantic.
|
||
Added `TypeError` to `effective_profile_snapshot` deserialization
|
||
exception list in `to_domain()` for consistency with the Pydantic
|
||
validator. Extracted default automation profile name to a
|
||
module-level constant (`DEFAULT_AUTOMATION_PROFILE`) to reduce
|
||
sentinel duplication across `models.py` and `repositories.py`.
|
||
Migration cycle-detection now logs affected `plan_id` values
|
||
(truncated to first 50) before the orphan fallback runs.
|
||
Migration SQL statements
|
||
uniformly use `sa.text()` for consistency. Centralised
|
||
automation-profile serialisation into
|
||
`LifecyclePlanModel._serialize_automation_profile()` to
|
||
eliminate duplication between `from_domain()` and
|
||
`LifecyclePlanRepository.update()`.
|
||
Moved `root_plan_id` self-reference resolution from
|
||
`from_domain()` into a `PlanIdentity` `model_validator` so
|
||
the domain model is consistent with the DB ``NOT NULL``
|
||
constraint before and after persistence.
|
||
Used explicit ``is not None`` check in
|
||
`_serialize_automation_profile()` for consistency with the
|
||
explicit-None-check convention used elsewhere in this commit.
|
||
(#921)
|
||
- Fixed `shell=True` subprocess usage in `cli_coverage_steps.py` by replacing
|
||
with `shlex.split()` and `shell=False` for defense-in-depth command injection
|
||
prevention, consistent with the existing pattern in
|
||
`cli_plan_context_commands_steps.py`. (#734)
|
||
- Added TDD bug-capture test for bug #987: AutomationProfileRepository session
|
||
leak. Four Behave BDD scenarios verify that `upsert()` and `delete()` close
|
||
the database session in `auto_commit` mode, capturing the missing
|
||
`session.close()` in a `finally` block. Tests use `@tdd_expected_fail` until
|
||
the bug fix is merged. (#1092)
|
||
- Added ACMS Backend Abstraction Layer (BAL) protocol definitions and
|
||
in-memory stub implementations. Defines `TextBackend`, `VectorBackend`,
|
||
and `GraphBackend` protocols with frozen result dataclasses (`TextResult`,
|
||
`VectorResult`, `GraphResult`). In-memory stubs (`InMemoryTextBackend`,
|
||
`InMemoryVectorBackend`, `InMemoryGraphBackend`) validate arguments and
|
||
return empty results for development and testing. Backends registered as
|
||
configurable singletons in the DI container with provider selection via
|
||
`override_providers()`. Includes Behave BDD tests (35 scenarios), Robot
|
||
Framework smoke tests, ASV benchmarks, and reference documentation. (#498)
|
||
- Added TDD bug-capture tests for #1024 — SQLite DB URL resolves to CWD
|
||
instead of CLEVERAGENTS_HOME. Behave BDD scenarios
|
||
(`@tdd_bug @tdd_bug_1024 @tdd_expected_fail`) verify that the default
|
||
`database_url` resolves inside `CLEVERAGENTS_HOME`, not the current
|
||
working directory. Includes Robot Framework integration tests with a
|
||
helper script exercising the same resolution path via subprocess. (#1034)
|
||
- Added integration Robot Framework test for Specification Workflow Example 7:
|
||
CI/CD Integration — Automated PR Review and Fix. Exercises the `ci`
|
||
automation profile (headless, non-interactive) covering: ci-profile
|
||
configuration (automation-profile, format, log level), idempotent
|
||
resource and project registration with duplicate-detection assertions,
|
||
three validation tools (`ci-lint`, `ci-typecheck`, `ci-tests`) registration
|
||
and resource attachment via `ToolRegistryService`, action creation with
|
||
typed arguments and invariants per spec Step 2, plan lifecycle with
|
||
explicit phase/state transition assertions across all phases
|
||
(strategize, execute, apply) plus terminal `applied` and `cancelled`
|
||
path checks, and JSON output structure verification
|
||
including `plan_id`, `phase`, `state`, `action`, `projects`, and
|
||
`arguments` fields.
|
||
(`robot/wf07_cicd_integration.robot`, `robot/helper_wf07_cicd.py`) (#771)
|
||
- Added integration Robot Framework test for Specification Workflow Example 14:
|
||
Server Mode — Team Collaboration. Exercises server mode configuration,
|
||
config-registry diagnostics, namespace management, action publishing with
|
||
namespaced actor references and supervised profile metadata, shared action
|
||
consumption via `use_action()` with required arguments and project links,
|
||
and namespace + phase plan monitoring using mocked LLM providers and
|
||
in-memory domain services (`CLEVERAGENTS_TESTING_USE_MOCK_AI=true`).
|
||
(`robot/wf14_server_mode_integration.robot`,
|
||
`robot/helper_wf14_server_mode.py`) (#778)
|
||
- Added volatile in-memory `audit_log` to `ReactiveEventBus` — every emitted
|
||
`DomainEvent` is appended to a volatile in-memory log accessible via the
|
||
`audit_log` property (defensive copy). Emit ordering now follows the
|
||
specification: RxPY stream push, then audit append, then handler dispatch.
|
||
Reactive and logging event buses now isolate stream/handler failures so one
|
||
subscriber cannot block audit recording or downstream subscribers. The
|
||
reactive stream now returns a read-only observable view (preventing direct
|
||
`on_next()` bypass), and `ReactiveEventBus` supports explicit in-memory
|
||
retention controls via `max_audit_log_size` and `clear_audit_log()`. Includes
|
||
expanded Behave and Robot integration coverage, 5 ASV benchmark suites, and
|
||
`vulture_whitelist.py` entry. (#587)
|
||
- **Breaking (CLI):** `agents actor run` now takes positional `<NAME>` and
|
||
`<PROMPT>` arguments, aligning the command signature with the specification.
|
||
The previous `--prompt/-p` option is removed. `--config/-c` is preserved as
|
||
an optional fallback that overrides registry-based name resolution (spec
|
||
deviation documented in `_resolve_actor.py`). Shared resolution logic
|
||
extracted to `_resolve_actor.py` with `yaml.safe_dump`, early name
|
||
validation, input sanitisation, resilient `atexit`-based temp file cleanup,
|
||
and graceful error handling for missing actors, empty config data, and
|
||
non-serialisable config blobs (without exposing serializer internals in
|
||
user-facing errors). Comprehensive BDD and Robot Framework
|
||
tests cover all resolution paths and edge cases. (#901)
|
||
- Fixed `list_actions()` to query the database when persistence is enabled
|
||
so that actions created by previous CLI invocations are visible. Falls
|
||
back to the in-memory cache on database errors or when no Unit of Work
|
||
is wired. Added `ActionRepository.list_all()` for unfiltered action
|
||
listing. (#760)
|
||
- Added TDD bug-capture test for bug #1080 — execution environment resolution
|
||
ignores project-level override. Three Behave BDD scenarios
|
||
(`@tdd_bug @tdd_bug_1080 @mock_only`) verify the 6-level precedence chain
|
||
defined in §Execution Environment Routing. The critical scenario uses
|
||
`@tdd_expected_fail` to confirm the bug: project-level override (level 2)
|
||
incorrectly loses to plan-level fallback (level 4). Two regression guard
|
||
scenarios verify existing correct behaviour (plan override vs project
|
||
override, project override vs host default). (#1101)
|
||
- Added TDD bug-capture tests for bug #1038 — `agents validation add`
|
||
missing `--required`/`--informational` flags. Four Behave BDD scenarios
|
||
(`@tdd_bug @tdd_bug_1038 @tdd_expected_fail`) verify that the `add`
|
||
command accepts `--required` and `--informational` flags and that
|
||
`--required` overrides the YAML config mode. Tests use
|
||
`@tdd_expected_fail` until the bug fix is merged. (#1102)
|
||
- Fixed execution environment resolution to honour project-level override
|
||
(precedence level 2). Threaded `plan_env` and `project_env` through
|
||
`ToolCallRouter`, `ToolCallingRuntime`, and `PlanExecutionContext` so
|
||
the resolver receives project-level execution environment values stored
|
||
in `ContextConfig.execution_environment`. (#1080)
|
||
- Added BuiltinAdapter class and MCP automatic resource slot creation.
|
||
BuiltinAdapter wraps register_file_tools/register_git_tools/register_subplan_tool
|
||
into a unified adapter interface. McpAdapter.infer_resource_slots() analyzes
|
||
tool input schemas to detect file/directory/repository parameters. (#882)
|
||
- Added large-project scaling performance tests with ASV benchmarks for
|
||
context assembly, execution throughput, and project scaling. Includes
|
||
Behave BDD scenarios (78 scenarios, 200 steps), Robot Framework tests,
|
||
and baseline threshold fixtures. (#576)
|
||
- Added overlay filesystem sandbox strategy with real OverlayFS support
|
||
and userspace copy-tree fallback. Includes lifecycle management, diff
|
||
computation, and sandbox factory registration. (#880)
|
||
- Added Kubernetes Helm chart in `k8s/` directory for server deployment.
|
||
Chart includes Deployment, Service, Ingress (with TLS termination),
|
||
ConfigMap, ServiceAccount, Secrets, and optional Redis subchart for
|
||
multi-instance session affinity. Added `Dockerfile.server` for ASGI
|
||
server containerization with multi-stage build, non-root user, and
|
||
pinned uv version. Includes deployment README with configuration
|
||
reference and quick-start instructions. (#928)
|
||
- Added CLI polish infrastructure: shared constants.py (exit codes, format
|
||
constants), centralized errors.py (cli_error, cli_not_found, cli_warning),
|
||
and completion command for shell tab-completion generation. (#861)
|
||
- Added tool-level execution environment preferences with NONE, REQUIRED,
|
||
PREFERRED, and SPECIFIC modes. ToolRunner routes tool execution based on
|
||
preference mode with caller-override precedence. (#879)
|
||
- Added Robot Framework integration test suite for Specification Workflow
|
||
Example 4: Multi-Project Dependency Update. 8 test cases exercising
|
||
supervised automation profile with 4 projects, child plan spawning,
|
||
dependency-ordered execution, and coordinated apply. Uses mocked LLM
|
||
providers via `CLEVERAGENTS_TESTING_USE_MOCK_AI`. (#768)
|
||
- Added TDD bug-capture tests for #969 — `plan correct` expects `decision_id`
|
||
but M3 acceptance test passes `plan_id`. Behave BDD scenarios (revert and
|
||
append modes) and Robot Framework integration tests verify that
|
||
`request_correction` is called with the root decision ID when a plan_id is
|
||
given as the first positional argument. Tests use `@tdd_expected_fail` until
|
||
the bug fix is merged. Shared mock fixtures extracted to
|
||
`features/mocks/tdd_plan_correct_plan_id_fixtures.py`. (#979)
|
||
- Added TDD bug-capture tests for bug #968: `plan explain` expects a
|
||
decision_id but the M3 acceptance test passes a plan_id. Two Behave BDD
|
||
scenarios (`@tdd_bug @tdd_bug_968 @tdd_expected_fail`) verify the fixed
|
||
behaviour — `plan explain <plan_id>` succeeds (rc=0) and displays
|
||
decision details. Includes Robot Framework integration tests with a
|
||
helper script exercising the same CLI path via subprocess, and step
|
||
definitions following established patterns. (#978)
|
||
- Added TDD bug-capture tests for bug #967 — `plan execute` phase processing.
|
||
Tests exercise the CLI orchestration layer via CliRunner (Behave) and
|
||
replicated CLI logic (Robot) to verify that `plan execute` correctly
|
||
handles plans in Strategize/QUEUED state by running `run_strategize()`
|
||
before transitioning. Includes four Behave scenarios and four Robot
|
||
integration test cases covering CLI execute from QUEUED, full lifecycle
|
||
orchestration, positive control, and auto-discovery of QUEUED plans.
|
||
(`features/tdd_plan_execute_phase_processing.feature`,
|
||
`robot/tdd_plan_execute_phase_processing.robot`) (#977)
|
||
- **Breaking (behavioral):** `resource_selection` decision type reclassified from
|
||
Execute-only to phase-agnostic (valid in both Strategize and Execute).
|
||
`DecisionType.RESOURCE_SELECTION` now appears in both `STRATEGIZE_TYPES` and
|
||
`EXECUTE_TYPES`. Code relying on `is_strategize_type` or `is_execute_type`
|
||
returning `False` for `resource_selection` will see different results.
|
||
Reclassification aligns with ADR-007 L72 and ADR-033 L74 which permit
|
||
resource selection during planning. (#931)
|
||
- Added E2E test for Workflow Example 4: Multi-Project Dependency Update
|
||
(supervised profile). Exercises the full supervised plan lifecycle across 4
|
||
git repositories (common-lib + 3 services), validates child plan spawning,
|
||
dependency-ordered execution and apply, per-project validation attachment,
|
||
and automation profile enforcement via the `agents plan use` CLI. (#750)
|
||
- Added ResourceHandler CRUD and discovery methods: read, write, delete,
|
||
list_children, diff, and discover_children. Frozen dataclass result types
|
||
(Content, WriteResult, DeleteResult, DiffResult) added to the handler
|
||
protocol. GitCheckoutHandler implements all six methods via git plumbing
|
||
and filesystem operations. FsDirectoryHandler implements all six via
|
||
pathlib/os/difflib. DevcontainerHandler implements read, write, and
|
||
discover_children via `devcontainer exec`. DatabaseResourceHandler
|
||
inherits NotImplementedError stubs pending connection management. (#827)
|
||
- Added E2E test for Workflow Example 5: Database Schema Migration with Safety
|
||
Nets (review automation profile). Exercises custom resource type registration
|
||
(`resource type add`), custom skill creation with spec-aligned database tools
|
||
(`query_db`, `execute_migration`, `backfill_column`), phased child plan
|
||
execution verification via `plan tree`, checkpoint-based rollback via
|
||
`plan rollback`, and post-apply migration content verification.
|
||
(`robot/e2e/wf05_db_migration.robot`) (#751)
|
||
- Implemented 6-level execution environment precedence chain per spec
|
||
lines 19324-19386. Plan/project environments now support `override` vs
|
||
`fallback` priority modes. Level 3 (nearest-ancestor devcontainer)
|
||
auto-detection integrated. Added `execution_env_priority` field to
|
||
`ContextConfig`. New `resolve_with_precedence()` API on
|
||
`ExecutionEnvironmentResolver`. Legacy 4-level `resolve()` preserved
|
||
for backward compatibility. Includes 13 new Behave scenarios. (#877)
|
||
- Added built-in deferred virtual resource types: `remote`, `submodule`, and
|
||
`symlink` with equivalence metadata rules for cross-repo and cross-layer
|
||
identity tracking. Registry bootstrap includes deferred virtual types but
|
||
hides them from `resource add` scaffolding (user_addable: false). Includes
|
||
YAML configurations, Behave BDD tests (47 scenarios), Robot Framework
|
||
integration tests, ASV benchmarks, and reference documentation. (#331)
|
||
- Enhanced `CorrectionService` subtree isolation: `analyze_impact()` now
|
||
populates `excluded_decisions` and `rollback_tier_depth`; added
|
||
`compute_rollback_tier()`, `validate_subtree_isolation()`, and dry-run
|
||
report enhancements with tier-0 root-targeted warnings. Fixed status
|
||
state-machine regression in `execute_revert()` where `analyze_impact()`
|
||
overwrote status back to ANALYZING; `execute_revert()` now transitions
|
||
through ANALYZING before EXECUTING for correct lifecycle ordering.
|
||
Fixed `validate_subtree_isolation()` to check structural-only BFS for
|
||
sibling invariant so influence-DAG-caused sibling reachability is not
|
||
misreported as a violation. Fixed false-positive cycle-detection warnings
|
||
from convergent (diamond) topologies by using a global enqueued set in
|
||
BFS instead of per-node seen_this_round. Added `dry_run` enforcement
|
||
guard in `_assert_executable()` to prevent execution of dry-run-only
|
||
corrections per spec (§ plan correct --dry-run). Added terminal-state
|
||
guard in `analyze_impact()` to reject re-analysis after execution.
|
||
Added mode validation in `execute_revert()`/`execute_append()` to
|
||
prevent mode-mismatched execution. Fixed `generate_dry_run_report()`
|
||
to preserve request status (dry-run is non-mutating). Fixed tier-0
|
||
warning to only trigger when target is genuinely in the structural
|
||
tree. Fixed `_collect_all_decisions()` to always include the target
|
||
decision in the universe. Improved cycle-detection log message
|
||
accuracy. Extracted cost/time estimation constants. Fixed
|
||
`generate_dry_run_report()` to use try/finally for status restoration
|
||
so that an exception during `analyze_impact()` does not leave the
|
||
request stuck in ANALYZING status. Promoted terminal-status set to
|
||
a module-level `_TERMINAL_STATUSES` frozenset constant. Review-cycle
|
||
fixes: fixed `generate_dry_run_report()` to also restore `_impacts`
|
||
dict (not only status) so dry-run is fully non-mutating; fixed tier-0
|
||
warning to compare against the actual root via `_find_root()` rather
|
||
than key-in-tree heuristic, preventing false warnings for forest
|
||
topologies; fixed BFS cycle detection to log at enqueue time so the
|
||
warning is reachable (previously dead code due to redundant visited
|
||
vs enqueued sets); added `_EXECUTABLE_STATUSES` module-level
|
||
frozenset for `_assert_executable()` and `cancel_correction()`;
|
||
added `_MAX_TREE_NODES` guard in `analyze_impact()` to reject
|
||
pathologically large inputs; `execute_revert()` now reuses cached
|
||
impact from `_impacts` when available to avoid redundant O(V+E)
|
||
recomputation; `execute_append()` now transitions through ANALYZING
|
||
before EXECUTING for consistent lifecycle across both modes; event
|
||
emission (`_emit_correction_applied`) now includes `attempt_id` and
|
||
logs failures at error level; added `max_length=10000` to
|
||
`CorrectionRequest.guidance` field; moved status transition after
|
||
attempt creation in both execution paths. Includes Behave BDD
|
||
scenarios (influence DAG, append mode, negative isolation validation,
|
||
dry-run enforcement, execute-revert end-to-end, status guard, mode
|
||
mismatch, single-node tree, terminal state guard, exact-match
|
||
affected count, convergent diamond DAG topology, dry-run exception
|
||
recovery, execute-revert with influence edges, DAG-only nodes in
|
||
excluded set), Robot Framework integration tests, and updated
|
||
dry-run report model fields. (#845)
|
||
- Added deferred physical resource types for git object taxonomy
|
||
(`git`, `git-remote`, `git-branch`, `git-tag`, `git-commit`, `git-tree`,
|
||
`git-tree-entry`, `git-stash`, `git-submodule`) and filesystem link types
|
||
(`fs-symlink`, `fs-hardlink`). All types are built-in, physical, and
|
||
auto-discovered with bounded scan_depth. Updated `fs-directory` child types
|
||
and auto-discovery. Updated `git-checkout` child types. Includes YAML
|
||
configs, Behave BDD tests, Robot tests, and ASV benchmarks. (#330)
|
||
- Added TDD bug-capture tests for #932 (plan apply missing --yes flag). (#950)
|
||
- Modified `auto_progress()` to complete the Apply phase immediately after
|
||
transitioning from Execute to Apply, since Apply is a metadata transition
|
||
with no LLM processing. This ensures `plan execute` drives the plan to
|
||
the terminal `applied` state when the automation profile permits (ci,
|
||
full-auto profiles with `auto_apply < 1.0`).
|
||
Extracted `_complete_apply_if_queued()` helper that consolidates the
|
||
Apply-completion pattern (start_apply + complete_apply) into a single
|
||
method with error recovery (calls `fail_apply` on failure) and async-job
|
||
guard (skips inline completion when async execution is enabled to avoid
|
||
orphaning enqueued jobs). Used by `auto_progress()`,
|
||
`lifecycle_apply_plan()`, and `try_auto_run()`.
|
||
Added `PlanLifecycleService.try_auto_run()` that drives plans through all
|
||
lifecycle phases (Strategize → Execute → Apply) when automation-profile
|
||
thresholds allow automatic progression; a threshold of 1.0 stops the plan
|
||
at that phase boundary for human approval.
|
||
Fixed `lifecycle-apply` CLI leaving plans stuck in `apply/queued` without
|
||
completing. The command now calls `_complete_apply_if_queued()` when the
|
||
plan is in Apply/queued, driving it to the terminal `applied` state.
|
||
Fixed stale RICH output in `lifecycle_apply_plan` that printed
|
||
"Plan is now in Apply phase (queued)" after the plan had already reached
|
||
terminal `applied` state; now branches on `plan.is_terminal`.
|
||
Fixed SQLite UNIQUE constraint violation in
|
||
`LifecyclePlanRepository.update()`: added `session.flush()` after
|
||
`clear()` on child collections (project_links, arguments, invariants)
|
||
before re-inserting rows.
|
||
Added `state` alias in `_plan_spec_dict()` JSON output for spec §Example 7
|
||
`jq` compatibility.
|
||
Updated plan execute and lifecycle-apply reference documentation.
|
||
(`src/cleveragents/application/services/plan_lifecycle_service.py`,
|
||
`src/cleveragents/cli/commands/plan.py`,
|
||
`src/cleveragents/infrastructure/database/repositories.py`,
|
||
`docs/reference/plan_cli.md`) (#753)
|
||
- Fixed `plan execute` CLI failing with "Plan is not in an executable state
|
||
(current: strategize/queued)" after strategize completed successfully.
|
||
Root cause: `_get_plan_executor()` created a second `PlanLifecycleService`
|
||
Factory instance with its own in-memory `_plans` cache. After the executor's
|
||
`run_strategize()` advanced the plan to `execute/queued` (via `auto_progress`),
|
||
the CLI handler's separate service instance returned stale `strategize/queued`
|
||
state from its cache. Fix: `_get_plan_executor()` now accepts an optional
|
||
`lifecycle_service` parameter; the `plan execute` handler passes its own
|
||
service instance so both share the same cache.
|
||
(`src/cleveragents/cli/commands/plan.py`)
|
||
- Improved type safety: `_get_plan_executor()` parameter
|
||
`lifecycle_service` now typed as `PlanLifecycleService | None` instead
|
||
of `Any | None`.
|
||
(`src/cleveragents/cli/commands/plan.py`)
|
||
- Added BDD regression test verifying that `plan execute` CLI handler
|
||
passes its lifecycle service instance to `_get_plan_executor()`,
|
||
preventing stale-cache regressions.
|
||
(`features/plan_lifecycle_cli_coverage.feature`)
|
||
- Added M5 (v3.4.0) E2E acceptance test suite `robot/e2e/m5_acceptance.robot`
|
||
with 21 zero-mock test cases covering context assembly, context policy
|
||
configuration, budget enforcement, context analysis, 10,000+ file scaling,
|
||
and plan execution with real LLM calls (`openai/gpt-4o-mini`). (#745)
|
||
- Fixed `project context set` writing policy changes via `session.flush()`
|
||
instead of `session.commit()`, causing silently lost data on
|
||
`session.close()`. (#745)
|
||
- Added `session_factory` DI provider to `Container` for CLI project-context
|
||
commands. The four `project context` subcommands (`set`, `show`, `inspect`,
|
||
`simulate`) previously called `container.session_factory()` which did not
|
||
exist, causing `AttributeError` at runtime. (#745)
|
||
- Added Google/Gemini API key pattern (`AIzaSy...`) to secret redaction in
|
||
`redaction.py`. (#745)
|
||
- Added `--skill <SKILL>` repeatable flag to `agents actor run` and
|
||
`actor-run` CLI commands. The flag resolves named skills from the
|
||
Skill Registry at runtime and merges their tools into agents that
|
||
already have configured tools, enabling ad-hoc skill injection
|
||
without modifying YAML configuration. Skill resolution uses the
|
||
DI-provided `SkillService` singleton; unknown or invalid skill names
|
||
produce a clear error and exit code 2. (#887)
|
||
- Added `--execution-env-priority` flag to `agents plan use` command, accepting
|
||
`fallback` (default) or `override` to control execution environment routing
|
||
precedence per ADR-043. Includes `ExecutionEnvPriority` StrEnum on the domain
|
||
model, domain-level model validation (priority requires environment),
|
||
`as_cli_dict()` support for both `execution_environment` and
|
||
`execution_env_priority`, database persistence via new columns on
|
||
`LifecyclePlanModel` with Alembic migration, and a `save_plan()` service
|
||
method to re-persist CLI overrides after plan creation. (#886)
|
||
- Added estimation actor support and role-aware actor validation for issue #650.
|
||
- **Schema/validation:** introduced `role_hint` and `response_format` fields,
|
||
plus role-aware compatibility warnings through shared validation helpers.
|
||
- **Preflight/CLI wiring:** aligned actor registration and preflight warning
|
||
paths to use the same warning logic and resolved estimation actor configs
|
||
before preflight compatibility checks.
|
||
- **Examples/docs/tests:** added `examples/actors/estimator.yaml`, updated
|
||
actor example docs, and expanded Behave/Robot coverage for estimator schema
|
||
and warning scenarios.
|
||
- **E2E helper behavior:** aligned M1/M2/M3/M6 integration helper handling so
|
||
missing OpenAI provider keys in local environments are controlled non-crash
|
||
outcomes while tracebacks/unexpected internal failures still fail.
|
||
- Runtime enforcement of `response_format` in provider invocation remains
|
||
planned and tracked via TODO comments in runtime code. (#650)
|
||
- Added interactive TUI persona and input-mode support with a dedicated
|
||
`agents tui` entry point and Textual app scaffolding. Personas are now
|
||
managed as local YAML configs with per-session binding/state, and input
|
||
handling supports Normal (`@` references), Command (`/` slash commands),
|
||
and Shell (`!` passthrough) flows with picker/overlay components and fuzzy
|
||
matching primitives. Includes Behave feature coverage for TUI persona/input
|
||
behavior, Robot TUI smoke validation, and an ASV fuzzy-reference benchmark.
|
||
(#695)
|
||
- Added TDD regression tests for bug #647 (`Container.resolve()` crash in
|
||
`plan tree`, `plan explain`, and `plan correct`) using real DI wiring in
|
||
Behave and Robot. Added regression-guard assertions, cache/singleton
|
||
cleanup hardening, and targeted issue-648 review follow-ups. (#648)
|
||
- Added four CLI-based integration test cases to M5 E2E verification suite
|
||
for v3.4.0 milestone acceptance criteria validation. Tests exercise
|
||
`project create`, `resource add git-checkout`, `project link-resource`, and
|
||
`project show` via real subprocess calls to `python -m cleveragents` with
|
||
per-test workspace isolation. (#496)
|
||
- Fixed `ProjectResourceLinkRepository.create_link()` and `remove_link()`
|
||
only calling `session.flush()` without `session.commit()`, causing linked
|
||
resource data to be lost between sessions. Added `finally: session.close()`
|
||
to both methods to match the session-factory lifecycle pattern used by all
|
||
other mutating repository methods. (#496)
|
||
- Fixed `agents plan execute` always using local-only stub actors that returned
|
||
empty changesets instead of invoking real LLM providers. The CLI command only
|
||
performed phase transitions (Strategize → Execute) without ever running the
|
||
`PlanExecutor` to drive the strategize or execute actors. Added
|
||
`_get_plan_executor()` helper that resolves `ProviderRegistry` from the DI
|
||
container and constructs `LLMStrategizeActor` / `LLMExecuteActor` for real
|
||
LLM calls. Updated `execute_plan` CLI to detect plan phase/state and
|
||
automatically invoke the appropriate actor: strategize actor when the plan is
|
||
in `Strategize/queued`, phase transition for `Strategize/complete`, and
|
||
execute actor for `Execute/queued`. Existing mock-based tests remain
|
||
backward-compatible via duck-typing fallback. New `llm_actors.py` module
|
||
provides `LLMStrategizeActor` (task decomposition) and `LLMExecuteActor`
|
||
(code generation) that resolve `provider/model` actor names to LangChain LLM
|
||
instances. `PlanExecutor.__init__` now accepts optional `strategize_actor`
|
||
and `execute_actor` parameters with stub defaults. (#960)
|
||
- Fixed `agents action create` missing the `--format`/`-f` flag. All other
|
||
action subcommands (`list`, `show`, `archive`) already accepted `--format`
|
||
and routed through `_print_action()`, but `create` was the only one omitted.
|
||
Running `action create --config action.yaml --format plain` previously failed
|
||
with a Typer unrecognized-option error. Added the `fmt` parameter to the
|
||
`create()` function signature and wired it to `_print_action()`. (#959)
|
||
- Added E2E acceptance test for M2 (v3.1.0): Actor Compiler + Full LLM
|
||
Integration. Robot Framework test suite `robot/e2e/m2_acceptance.robot`
|
||
exercises actor YAML compilation into functional graphs, skill registry,
|
||
tool lifecycle, and plan execution with a custom actor using real LLM API
|
||
keys. Test flow: create temp git repo → register custom actor → register
|
||
resource and project → create action → run full plan lifecycle (use →
|
||
execute strategize → execute → diff → apply) → verify actor compilation
|
||
and plan integrity. Uses `[Tags] E2E`, `Skip If No LLM Keys`, and
|
||
flexible structural assertions with `expected_rc=None` for LLM-dependent
|
||
commands. (#742)
|
||
- Fixed `plan execute` failing with `Error [500] INTERNAL` when run in a
|
||
separate CLI process from `plan use`. Root cause: `start_strategize()`
|
||
built its action registry from the in-memory `_actions` dict only,
|
||
missing DB-persisted actions created by prior CLI invocations. The
|
||
preflight guardrail then rejected the plan with a `PreflightRejection`
|
||
that escaped the CLI error handler (extends bare `Exception`, not
|
||
`CleverAgentsError`). Fixes: (1) `start_strategize()` now loads the
|
||
plan's action from the persistence layer before preflight checks,
|
||
(2) `execute_plan` CLI catches `PreflightRejection` for user-friendly
|
||
errors, (3) `plan execute` runs the execute phase inline so the plan
|
||
progresses through execute/queued → execute/complete in a single CLI
|
||
invocation, (4) `lifecycle-apply` handles plans already auto-progressed
|
||
to apply/queued by `complete_execute()`.
|
||
(`src/cleveragents/application/services/plan_lifecycle_service.py`,
|
||
`src/cleveragents/cli/commands/plan.py`) (#746)
|
||
- Added E2E Robot Framework acceptance test for M6 (v3.5.0) autonomy hardening
|
||
milestone. Exercises session CRUD lifecycle, automation-profile list/show/set,
|
||
project init with git-checkout resource, A2A plan lifecycle (use,
|
||
lifecycle-list, status, execute, lifecycle-apply), guard enforcement via
|
||
automation profiles, and a full autonomy acceptance flow — all via real CLI
|
||
invocations. LLM-dependent tests skip gracefully when API keys are absent.
|
||
Hardened shared E2E keywords: safe JSON parsing with multi-object fallback,
|
||
git return-code checks, special-character-safe API-key detection, `IF`/`ELSE`
|
||
migration from deprecated `Run Keyword If`, per-test teardowns, and
|
||
`Force Tags`. Profile list now verifies all 8 built-in profiles. Session
|
||
delete confirms removal via re-list. Apply step verifies phase transition.
|
||
Execute step asserts plan_id in output. JSON-quoted assertions for short
|
||
profile names (`"ci"`, `"auto"`) prevent false-positive substring matches.
|
||
Added four new E2E tests covering remaining acceptance criteria: guard
|
||
enforcement with custom profile (denylist, budget caps, tool-call limits),
|
||
profile precedence resolution (plan-level overrides global), event queue
|
||
pub/sub via plan lifecycle state transitions, and hierarchical decomposition
|
||
verification via `plan tree`.
|
||
Post-review hardening (PR #803): LLM-dependent tests now Fail instead of
|
||
Skip when API keys are present but `plan use` returns non-zero. Event Queue
|
||
test (AC-3) uses hard assertions for state transition verification.
|
||
Hierarchical Decomposition test (AC-6) asserts at least one decision node
|
||
exists after execution. Guard Enforcement Assertions verify the resolved
|
||
profile name matches the expected value. Extracted `Setup Plan Test
|
||
Resources` keyword to eliminate repeated boilerplate and bring the file
|
||
under the 500-line limit. `Verify Plan In List` and `Full Flow Apply Step`
|
||
keywords use hard assertions instead of WARN fallbacks. Profile Precedence
|
||
test documents that action > global precedence requires production wiring
|
||
not yet present in `PlanLifecycleService.use_action`.
|
||
(`robot/e2e/m6_acceptance.robot`, `robot/e2e/common_e2e.resource`) (#746)
|
||
- Added E2E Robot Framework test for Specification Workflow Example 7: CI/CD
|
||
Integration — Automated PR Review and Fix. Exercises the `ci` automation
|
||
profile (headless, non-interactive) with JSON output and log-level
|
||
configuration, idempotent resource and project registration with `--branch`
|
||
and `--description` flags, three-validation registration (source/mode/code)
|
||
with project attachment and `project show` verification, action creation
|
||
with spec-aligned name (`local/review-pr`), complete `definition_of_done`,
|
||
`invariants`, and `arguments`, plan launch with `--arg` flags, explicit
|
||
`plan execute` for lifecycle progression, `plan status` terminal-state
|
||
assertion, plan diff JSON validation, and JSON output verification.
|
||
Resource/project naming follows spec convention (`local/ci-workspace` project,
|
||
`local/ci-main` resource). Entity creation commands tolerate "already
|
||
exists" for CI re-runnability. `Extract JSON Field` keyword handles CLI
|
||
debug log lines preceding JSON via `JSONDecoder.raw_decode(strict=False)`.
|
||
Fail-fast `expected_rc` only where the spec mandates error suppression
|
||
(`2>/dev/null || true`); first `resource add` and `project create` now
|
||
assert `expected_rc=${0}`. Config assertions use stdout-only matching and
|
||
exact equality for the `ci` profile value. Project idempotency verified
|
||
with occurrence count. Empty plan-diff stdout logged as warning.
|
||
Validation naming aligned with spec (`local/ci-lint` per §Example 7).
|
||
All `Run Process` calls include `on_timeout=kill` per codebase CI
|
||
stability standard. Added `on_timeout=kill` to `Run CleverAgents
|
||
Command` and `Create Temp Git Repo` keywords in `common_e2e.resource`
|
||
for consistent timeout handling across all E2E suites.
|
||
Dynamic actor selection based on available API keys (same pattern as
|
||
`m6_acceptance.robot`) avoids runtime failure when only one provider
|
||
key is set. `Poll Plan Until Terminal` keyword now integrated into
|
||
the CI Plan Launch test case per spec Step 3 polling loop. Replaced
|
||
local `Extract JSON Field` with shared `Safe Parse Json Field` from
|
||
`common_e2e.resource`. Added `Force Tags E2E` and per-test
|
||
`[Teardown]` blocks. Added `WF07 Suite Setup` keyword for database
|
||
initialisation.
|
||
Added `robot/common_vars.py` module placeholder for shared Robot
|
||
Framework variables.
|
||
(`robot/e2e/wf07_cicd.robot`, `robot/e2e/common_e2e.resource`,
|
||
`robot/common_vars.py`) (#753)
|
||
- Added E2E Robot Framework test for Specification Workflow Example 14: Server
|
||
Mode — Team Collaboration. Exercises server mode configuration (server URL,
|
||
token, namespace), diagnostics, action publishing to team namespace with
|
||
namespace-scoped listing, actor registration, plan list smoke test, and
|
||
`supervised` automation profile verification with threshold field assertions
|
||
via real CLI with zero mocking. (`robot/e2e/wf14_server_mode.robot`) (#760)
|
||
- Fixed `agents session list`, `agents session create`, and other session
|
||
subcommands raising `AttributeError: 'DynamicContainer' object has no
|
||
attribute 'db'` after `agents init`. Root cause: `_get_session_service()`
|
||
called `container.db()` but no `db` provider existed. Added a
|
||
`session_service` DI provider in `container.py` that builds the engine,
|
||
sessionmaker, and auto-committing repositories. Rewrote
|
||
`_get_session_service()` to resolve via the container with module-level
|
||
caching. Added `auto_commit` parameter to `SessionRepository` and
|
||
`SessionMessageRepository` to prevent resource leaks in CLI context while
|
||
preserving Unit-of-Work semantics. Unified error handling across all 7
|
||
session subcommands. Includes Behave BDD regression scenarios, Robot
|
||
Framework integration smoke tests, and structlog isolation for parallel
|
||
test execution. (#554, #570, #680)
|
||
- Added Robot Framework E2E acceptance test for M1 (v3.0.0) milestone.
|
||
Tests the complete plan lifecycle (action create → resource add → project
|
||
create → plan use → plan execute strategize → plan execute → plan diff →
|
||
plan apply) with real LLM API keys and no mocking. Gracefully skips when
|
||
API keys are absent. (#741)
|
||
- Added dedicated E2E test infrastructure: new `nox -s e2e_tests` session
|
||
running Robot Framework with `--include E2E` tag filter against `robot/e2e/`
|
||
directory, dedicated CI job with real LLM API key secrets, graceful skip
|
||
when API keys are absent, and `--exclude E2E` on the standard integration
|
||
test session. Includes a minimal smoke test exercising `agents --version`
|
||
and `agents --help`. (#740)
|
||
- Implemented `tdd_expected_fail` tag handling in Robot Framework via a Listener v3
|
||
module (`robot/tdd_expected_fail_listener.py`). Tests tagged `tdd_expected_fail`
|
||
that fail have their result inverted to pass (expected failure); tests that
|
||
unexpectedly pass are reported as failed with guidance to remove the tag. Tag
|
||
validation enforces `tdd_bug` + `tdd_bug_<N>` prerequisites. Includes
|
||
idempotency guard against double-invocation, explicit SKIP status handling,
|
||
and a `close()` hook for clean teardown. Listener is registered in the nox
|
||
`integration_tests` and `slow_integration_tests` sessions. Fixture files are
|
||
excluded from the main pabot runner via `tdd_fixture` tag. Includes 9 Robot
|
||
Framework integration test cases. (#628)
|
||
|
||
- Added TDD-style failing Behave BDD tests for the session list DI container
|
||
missing `db` provider bug. Three scenarios exercise `session list`,
|
||
`_get_session_service()`, and `session list --format json` through the real
|
||
DI path. Includes Robot Framework smoke tests and ASV benchmarks. Tests
|
||
are intentionally failing (`@tdd_expected_fail`) until the bug fix for
|
||
#554 is applied. (#631)
|
||
- Added TDD-style failing Behave BDD tests for the session create DI container
|
||
missing `db` provider bug. Three scenarios exercise `session create`,
|
||
`session create --actor`, and `session create --format json` through the
|
||
real DI path. Includes Robot Framework smoke tests and ASV benchmarks.
|
||
Tests are intentionally failing (`@tdd_expected_fail`) until the bug fix
|
||
for #570 is applied. (#630)
|
||
- Implemented UKO Layer 2 paradigm vocabulary specializations: Object-Oriented
|
||
(`uko-oo:`), Functional (`uko-func:`), and Procedural (`uko-proc:`). Added
|
||
OWL/Turtle class and property definitions for all three paradigms in
|
||
`docs/ontology/uko.ttl`. Implemented `DetailLevelMapBuilder` with insertion
|
||
and integer reassignment logic for extending parent DetailLevelMaps.
|
||
`ParadigmVocabulary`, `VocabularyClass`, `VocabularyProperty`, and
|
||
`VocabularyRegistry` frozen Pydantic models provide the Python API. Includes
|
||
Behave BDD tests (60+ scenarios), Robot Framework integration helper, ASV
|
||
benchmarks for DetailLevelMap operations, and reference documentation.
|
||
**Breaking:** `DetailLevelMap.effective_levels()` now returns
|
||
`MappingProxyType[str, int]` (read-only) instead of `dict[str, int]`;
|
||
callers that mutated the returned mapping must copy to a `dict` first.
|
||
(#575, PR #657)
|
||
- Implemented UKO Layer 3 technology-specific vocabulary extensions for Python
|
||
(`uko-py:`), TypeScript (`uko-ts:`), Rust (`uko-rs:`), and Java (`uko-java:`).
|
||
Each vocabulary defines OWL classes, properties, Layer 2 dependencies, and
|
||
DetailLevelMap insertions per specification lines 44405-44420. Python inserts
|
||
3 new depth levels (DECORATED_SIGNATURES, TYPE_STUBS, WITH_TESTS) producing a
|
||
15-level effective map; TS/RS/Java extend at SIGNATURES level without new
|
||
insertions. Includes OWL/Turtle ontology files, ProvenanceInfo model with 2
|
||
required fields (source_resource, source_path) and 3 defaulted fields
|
||
(source_range, valid_from, is_current), build_detail_level_map/resolve_detail_level
|
||
utilities, and full Behave BDD tests (78 scenarios, 200 steps). (#576)
|
||
- Added `RepoIndexingService` for repository file indexing with incremental
|
||
refresh, extension-based language detection, SHA-256 content hashing, and
|
||
token estimation. Supports policy enforcement via include/exclude globs,
|
||
max file size, and max total size limits from project `ContextConfig`.
|
||
Persists index metadata and per-file records to SQLite via `RepoIndexModel`
|
||
and `IndexedFileModel`. Domain models (`IndexStatus`, `FileRecord`,
|
||
`IndexMetadata`, `RepoIndex`) are frozen Pydantic v2 with ULID IDs and UTC
|
||
datetimes. Wired into the DI container. Includes 28 Behave BDD scenarios,
|
||
3 Robot Framework integration tests, ASV benchmarks (5 time + 2 track), and
|
||
reference documentation. (#195)
|
||
- Wired retry policies and circuit breakers into the service layer.
|
||
`RetryPolicyConfig` and `CircuitBreaker` models govern per-service retry behaviour
|
||
(max attempts, backoff strategy, delay bounds, jitter) and circuit breaker
|
||
protection (failure threshold, recovery timeout, half-open probing, cooldown).
|
||
`ServiceRetryWiring` initialises from `Settings`, creates `CircuitBreaker`
|
||
instances per service, and exposes `execute()`/`async_execute()` helpers.
|
||
`retry_service_operation` decorator adds retry + circuit breaker to any
|
||
service method via a single annotation. Structured logs are emitted on every
|
||
retry attempt and when a circuit breaker opens. Retry amplification is
|
||
prevented by a `contextvars` nesting guard. Per-service overrides are loaded
|
||
from the `retry_service_overrides` JSON config key. Includes Behave BDD unit
|
||
tests, Robot Framework integration tests, and ASV benchmarks. (#313)
|
||
- Added container-aware tool execution and I/O forwarding via `ContainerToolExecutor`
|
||
and `PathMapper`. Tools routed to `execution_environment: container` are executed
|
||
inside a provisioned devcontainer with automatic host↔container path mapping,
|
||
bounded output capture (50 MiB), structured error reporting, and container metadata
|
||
on the `ToolInvocation` audit trail. Includes `ContainerConfig`, `ContainerMetadata`,
|
||
`ContainerExecutionError`, and `ContainerTimeoutError` domain models, `ToolRunner`
|
||
container routing integration, safe environment filtering, symlink/traversal
|
||
protection, and `sync_results_to_host` for file-based result retrieval. Covered by
|
||
Behave BDD scenarios, Robot Framework integration tests, ASV benchmarks, and
|
||
`docs/reference/execution_environment.md`. (#515)
|
||
- Implemented `@tdd_expected_fail` tag handling in Behave environment hooks. Added
|
||
`validate_tdd_tags()` and `should_invert_result()` helper functions in
|
||
`features/environment.py`. Scenarios tagged `@tdd_expected_fail` that fail have their
|
||
result inverted to pass (expected failure); scenarios that unexpectedly pass are reported
|
||
as failed with guidance to remove the tag. Tag validation enforces `@tdd_bug` +
|
||
`@tdd_bug_<N>` prerequisites. Implemented via `Scenario.run()` monkey-patch in `before_all`.
|
||
Includes 34 Behave BDD scenarios (19 tag-validation, 14 infrastructure, and
|
||
1 demo) and 12 Robot Framework integration test cases. (#627)
|
||
- Wired `AuditService.record()` into domain services via EventBus auto-dispatch.
|
||
Created `AuditEventSubscriber` that subscribes to 9 security-relevant event types
|
||
(`plan_applied`, `plan_cancelled`, `resource_modified`, `correction_applied`,
|
||
`config_changed`, `entity_deleted`, `session_created`, `auth_success`,
|
||
`auth_failure`) and persists them via `AuditService.record()` with secret masking
|
||
applied to all audit log details (always `show_secrets=False`). Subscriber enriches
|
||
audit entries with `session_id` and `correlation_id` from the domain event for
|
||
traceability. Wired `PlanLifecycleService` to emit `PLAN_APPLIED` and
|
||
`PLAN_CANCELLED` events with all project names in event details. Added 5 new
|
||
`EventType` enum members. Registered subscriber as eagerly-initialized singleton
|
||
in DI container. `AuditService` now accepts an explicit `database_url` parameter
|
||
so it shares the same database as the rest of the application. All `EventBus.emit()`
|
||
call sites are wrapped in try/except guards with structured logging, and
|
||
`ReactiveEventBus` isolates per-handler failures so one failing subscriber cannot
|
||
block others. `server connect` now emits per-setting `CONFIG_CHANGED` audit events
|
||
via `set_value()`. `SessionService.delete()` emits `ENTITY_DELETED`. Exception
|
||
messages in the DI container bootstrap are redacted before logging.
|
||
`CorrectionService` is registered as a singleton in the DI container. Includes
|
||
23 Behave BDD scenarios, 5 Robot Framework integration tests, and ASV benchmarks.
|
||
(#581)
|
||
|
||
### Added
|
||
|
||
- Resource type single-inheritance via `inherits` field (ADR-042) (#513)
|
||
- Inheritance chain resolution, field merging, and polymorphic type matching
|
||
- `ToolRegistry.find_tools_for_resource()` for polymorphic tool binding
|
||
- Polymorphic handler resolution with ancestor-type fallback
|
||
- CLI: `agents resource type list` shows Inherits column; `type show` displays inheritance chain
|
||
- Alembic migration `m6_004_resource_type_inherits` adds `inherits` column to `resource_types`
|
||
- Fixed `agents actor list` raising a validation error on fresh projects.
|
||
`ActorRegistry._actor_name()` built names via `f"{provider}/{model}"`,
|
||
which produced names with 2+ slashes when providers had models containing
|
||
`/` (e.g. OpenRouter's `anthropic/claude-sonnet-4-20250514`). Now sanitises
|
||
both provider and model names by replacing `/` with `-` and lowercasing to
|
||
satisfy the spec pattern. **Note:** provider/model names are now lowercased;
|
||
existing mixed-case built-in actors will be superseded by lowercased versions
|
||
on the next `ensure_built_in_actors()` call.
|
||
Includes Behave BDD regression scenarios, Robot Framework integration
|
||
smoke tests, and ASV benchmarks. (#592)
|
||
- Added TDD regression tests for `agents session list` DI container wiring
|
||
error (bug #554). `_get_session_service()` calls `container.db()` but the
|
||
`Container` class has no `db` provider, raising `AttributeError`. Includes
|
||
10 Behave BDD scenarios (`@tdd_bug @tdd_bug_554 @tdd_expected_fail`)
|
||
covering empty list, empty-list format validation (JSON/YAML/plain),
|
||
init-then-list lifecycle, post-create list, rich/JSON/plain/YAML output
|
||
formats, and stderr error-path assertions. Robot Framework integration
|
||
smoke tests and ASV service-layer benchmarks. Implements
|
||
`@tdd_expected_fail` infrastructure (Behave `after_scenario` hook and Robot
|
||
listener) and migrates 18 existing TDD scenarios from `@tdd @bugNNN` to
|
||
`@tdd_bug @tdd_bug_NNN` convention. (#554)
|
||
- Added TDD regression tests for `agents session create` DI container wiring
|
||
error (bug #570). `_get_session_service()` calls `container.db()` but the
|
||
`Container` class has no `db` provider, raising `AttributeError`. Same root
|
||
cause as #554. Includes 4 Behave BDD scenarios
|
||
(`@tdd_bug @tdd_bug_570 @tdd_expected_fail`), Robot Framework integration
|
||
smoke tests, and ASV service-layer benchmarks. Tests exercise the real DI
|
||
path with `_service = None` and a file-based SQLite database.
|
||
Also implements the `@tdd_expected_fail` inversion infrastructure:
|
||
a Behave `after_scenario` hook in `features/environment.py` that flips
|
||
pass/fail for `@tdd_expected_fail` scenarios, and a Robot Framework
|
||
Listener API v3 plugin (`robot/tdd_expected_fail_listener.py`) with
|
||
identical semantics. Migrates 18 existing TDD scenarios from the old
|
||
`@tdd @bugNNN` convention to standardised `@tdd_bug @tdd_bug_NNN` tags.
|
||
(#570)
|
||
- Fixed intermittent race condition in M4 validation integration tests when
|
||
running under pabot. Root cause was three-pronged: shared SQLite DB URL,
|
||
shared CLEVERAGENTS_HOME directory, and singleton leaks in chained CLI
|
||
helper invocations. Introduced composable `Setup Database Isolation`
|
||
keyword in `common.resource`, per-suite temp directories, and centralised
|
||
`reset_global_state()` in `robot/helpers_common.py`. Added `timeout=30s`
|
||
to all `Run Process` calls in `m4_e2e_verification.robot`. (#563)
|
||
- Fixed `agents project show` not finding a project immediately after creation.
|
||
Extended the `session.commit()` fix from #589 to also cover `update()` and
|
||
`delete()` in `NamespacedProjectRepository`, and updated the class docstring
|
||
to reflect that all mutating methods now commit within their own session.
|
||
Includes 3 Behave BDD regression scenarios, Robot Framework integration
|
||
smoke tests, and ASV benchmarks. (#590)
|
||
- Fixed `agents project create` not persisting projects to the database.
|
||
`NamespacedProjectRepository.create()` called `session.flush()` but never
|
||
`session.commit()`, so projects were invisible to subsequent
|
||
`agents project list` calls. Added `session.commit()` and a `finally:
|
||
session.close()` guard. Includes 4 Behave BDD regression scenarios,
|
||
Robot Framework integration smoke tests, and ASV benchmarks. (#589)
|
||
- Added TDD-style Behave BDD tests for the built-in `git-checkout` resource type
|
||
bootstrap. Three scenarios: one failing TDD test reproducing bug #524 (no bootstrap
|
||
called during init), and two regression tests verifying `bootstrap_builtin_types()`
|
||
seeds correct data and `agents resource add git-checkout` succeeds. Includes Robot
|
||
Framework regression tests. (Drive-by: corrected bug reference from #523 to #524.) (#553)
|
||
- Added TDD-style Behave BDD tests for the built-in `fs-directory` resource type
|
||
bootstrap. Three scenarios: one failing TDD test reproducing bug #524 (no bootstrap
|
||
called during init), and two regression tests verifying `bootstrap_builtin_types()`
|
||
seeds correct data and `agents resource add fs-directory` succeeds. Includes Robot
|
||
Framework regression tests. (#537)
|
||
- Added TDD-style failing Behave BDD tests for the missing `agents init --yes` flag.
|
||
Five scenarios: four TDD-failing tests (exit code, prompt suppression, `-y` alias,
|
||
output summary) and one regression guard for interactive mode. Includes Robot
|
||
Framework smoke tests and ASV benchmarks. Tests are intentionally failing until
|
||
the bug fix for #522 is applied. (#536)
|
||
- Added Temporal Data Model (Revision-Aware RDF) with 3 storage tiers for
|
||
the ACMS. Temporal metadata fields (`valid_from`, `valid_until`,
|
||
`is_current`, `is_revision_of`) on UKO InformationUnit nodes enable
|
||
revision chain tracking: when code changes, old nodes are marked historical
|
||
and new revision nodes are created with back-links. Three storage tiers
|
||
(hot/warm/cold) filter nodes by temporal scope (current/recent/all) with
|
||
configurable retention (`warm_retention_hours` default 24h,
|
||
`cold_retention_days` default 90d). Includes `TemporalMetadata`,
|
||
`TemporalNode`, `RevisionChain`, `TierQueryResult`, `TierRetentionConfig`
|
||
frozen domain models, `TemporalBackend` protocol,
|
||
`InMemoryTemporalBackend` stub, `TemporalService` with structlog and DI,
|
||
`BackendSet.temporal` typing upgrade from `object | None` to
|
||
`TemporalBackend | None`. 99 Behave scenarios, 8 Robot Framework tests,
|
||
ASV benchmarks, and reference documentation. (#577)
|
||
- Implemented UKO Layer 1 Domain Ontologies (`uko-doc:`, `uko-data:`, `uko-infra:`)
|
||
in the OWL/Turtle ontology file (`docs/ontology/uko.ttl`). Added 17 `uko-doc:` classes
|
||
(Document, Section, Paragraph, Citation, etc.), 13 `uko-data:` classes (Table, Column,
|
||
ForeignKey, View, etc.), and 7 `uko-infra:` classes (Service, Endpoint, Port, etc.) with
|
||
all spec-mandated properties and relationships. Updated all four DetailLevelMap presets
|
||
(`code_detail_map`, `docs_detail_map`, `database_detail_map`, `infra_detail_map`) to
|
||
be spec-complete with every named level. Added `OntologyRegistry` with domain lookup,
|
||
Layer 1 listing, DetailLevelMap inheritance chain building, and Turtle syntax validation.
|
||
Includes 31 Behave BDD scenarios and 6 Robot Framework integration tests. (#574)
|
||
- Added `PostgreSQLAnalyzer` and `DockerComposeAnalyzer` domain-specific analyzers
|
||
(Phase 2 of issue #588). `PostgreSQLAnalyzer` parses DDL content via regex and
|
||
extracts `uko-data:Table`, `uko-data:Column`, `uko-data:ForeignKey`, `uko-data:View`,
|
||
and `uko-data:Schema` triples with column metadata (data type, nullability, primary
|
||
key). `DockerComposeAnalyzer` parses Docker Compose YAML via `yaml.safe_load` and
|
||
extracts `uko-infra:DeploymentUnit`, `uko-infra:Service`, `uko-infra:Port`,
|
||
`uko-infra:EnvironmentVariable`, and `uko-infra:connectsTo` triples. Both satisfy
|
||
`AnalyzerProtocol` and register in `AnalyzerRegistry` by file extension. Includes
|
||
34 Behave BDD scenarios covering all four analyzers (protocol conformance, registry
|
||
operations, triple extraction, error handling, cross-analyzer URI scheme and
|
||
confidence checks), 6 Robot Framework integration smoke tests, and updated
|
||
`__init__.py` exports. (#588)
|
||
- Added TDD-style Behave BDD tests for the built-in `git-checkout` resource type
|
||
bootstrap. Three scenarios: one failing TDD test reproducing bug #524 (no bootstrap
|
||
called during init), and two regression tests verifying `bootstrap_builtin_types()`
|
||
seeds correct data and `agents resource add git-checkout` succeeds. Includes Robot
|
||
Framework regression tests. (#553)
|
||
- Added UKO Indexer for real-time index synchronization. `UKOIndexer` orchestrates
|
||
analysis of resources into UKO triples via pluggable `AnalyzerRegistry` and
|
||
simultaneously indexes into text, vector, and graph backends. Provenance metadata
|
||
(`ProvenanceMetadata`, `ProvenancedTriple`) is attached to every triple, tracking
|
||
source resource, file path, and temporal validity. Write-side index backend protocols
|
||
(`TextIndexBackend`, `VectorIndexBackend`, `GraphIndexBackend`) are distinct from the
|
||
existing read-side query protocols. Graceful degradation: if text or vector backends
|
||
are `None`, the corresponding indexing step is silently skipped. Index lifecycle:
|
||
`index_resource` (add), `remove_resource` (cleanup), `reindex_resource` (change).
|
||
`ContentReader` protocol decouples the indexer from filesystem I/O.
|
||
`IndexLifecycleHook` provides callbacks for indexing events. In-memory stub
|
||
implementations for all three backends. `ResourceFileWatcher` monitors filesystem
|
||
paths via watchdog and triggers re-indexing callbacks on file changes with
|
||
configurable debouncing and optional `RESOURCE_MODIFIED` EventBus emission.
|
||
Includes 166 Behave BDD scenarios, 9 Robot Framework integration tests, ASV
|
||
benchmarks, and reference documentation. (#578)
|
||
- Added built-in virtual core resource types (`file`, `directory`, `commit`,
|
||
`branch`, `tag`, `tree`) under `examples/resource-types/`. All use
|
||
`resource_kind: virtual`, `sandbox_strategy: none`, `user_addable: false`.
|
||
Equivalence metadata defines identity criteria per type (content hash for
|
||
files, git object identity for commits/branches/tags/trees). Extended
|
||
bootstrap registration in `ResourceRegistryService` to include virtual
|
||
types and updated `ResourceTypeSpec.BUILTIN_NAMES`. Added reference
|
||
documentation in `docs/reference/resource_types_builtin.md`. Includes
|
||
Behave BDD tests, Robot integration tests, and ASV benchmarks. (#329)
|
||
- Added general-purpose domain event system under
|
||
`cleveragents.infrastructure.events`. `EventType` StrEnum defines 38 typed
|
||
event identifiers across 9 domains (plan lifecycle, decision, invariant, actor,
|
||
tool, resource, sandbox, context, validation, session, budget). `DomainEvent` is a frozen
|
||
Pydantic model with `event_type`, auto-UTC `timestamp`, auto-ULID
|
||
`correlation_id`, `plan_id`, `root_plan_id`, `session_id`, `actor_name`,
|
||
`project_name`, and `details` fields. `EventBus` is a `@runtime_checkable`
|
||
Protocol with `emit()` and `subscribe()` methods. `ReactiveEventBus` is an
|
||
RxPY `Subject`-backed in-process bus that dispatches synchronously to
|
||
type-filtered handlers and exposes a raw `rx.Observable` stream for advanced
|
||
operators. `LoggingEventBus` is a structlog-based bus for audit logging that
|
||
requires no RxPY dependency. `DecisionService` and `PlanLifecycleService`
|
||
accept an optional `event_bus` parameter (backward-compatible) and emit
|
||
`DECISION_CREATED`, `PLAN_CREATED`, and `PLAN_PHASE_CHANGED` on significant
|
||
state changes. `ReactiveEventBus` is registered as a Singleton in the DI
|
||
container and wired into both services automatically. Includes Behave BDD unit
|
||
tests (27 scenarios, 75 steps, 100% coverage on all new source files), Robot
|
||
Framework smoke tests (9 cases), ASV performance benchmarks (5 suites), and
|
||
reference documentation (`docs/reference/event_bus.md`). (#473)
|
||
- Validated M4 acceptance criteria for v3.3.0 milestone closure. All M4 E2E verification
|
||
tests and correction/subplan smoke tests pass against the final v3.3.0 implementation.
|
||
Added CLI-exercising integration tests for `plan use`, `plan execute`, and `plan tree`
|
||
commands to verify the milestone success criteria through actual Typer CLI invocations.
|
||
Split 1074-line helper into six focused modules (`_common`, `_domain`, `_merge`, `_cli`,
|
||
`_cli_errors`, dispatcher) under the 500-line limit. Added CLI error-path tests for
|
||
read-only plan execute, unavailable action, missing changeset, and empty decision tree.
|
||
Extracted `_make_subplan_status` factory, `_assert_exit_code` and
|
||
`_assert_mock_called_once*` wrappers, frozen timestamp constant, and `shutil.which`
|
||
git pre-check. Removed tautological domain assertions in `plan_tree` and
|
||
`parallel_max`. Fixed CONTRIBUTORS.md alphabetical ordering. (#495)
|
||
- Added minimal LSP server stub with `agents lsp serve` CLI command supporting the
|
||
`initialize`, `shutdown`, and `exit` lifecycle handshake over JSON-RPC stdin/stdout
|
||
transport with Content-Length header framing. Unsupported methods return `MethodNotFound`
|
||
(-32601); requests before `initialize` return `ServerNotInitialized` (-32002); requests
|
||
after `shutdown` return `InvalidRequest` (-32600). Transport hardening includes 10 MB
|
||
max content-length, 32 max header lines, graceful recovery from malformed messages,
|
||
and recursion-depth protection. `LspServer` stores an `A2aLocalFacade` for future
|
||
server-mode wiring. `--log-level` flag controls logging verbosity. Includes 48 Behave
|
||
BDD scenarios, Robot Framework smoke tests, ASV startup latency benchmarks, and
|
||
`docs/reference/lsp_stub.md`. (#203)
|
||
- Validated M3 acceptance criteria for v3.2.0 milestone closure. All 10 E2E
|
||
verification tests pass against the final implementation, exercising real
|
||
CLI command paths (`plan use`, `plan execute`, `plan tree`,
|
||
`plan explain`, project-scoped `invariant add/list`, dry-run and live
|
||
`plan correct`), database-backed persistence, context snapshots, and
|
||
invariant enforcement during strategize. Added acceptance criteria tags and
|
||
milestone documentation to the robot suite. (#494)
|
||
- Added scoped backend view filtering for project-resource isolation in ACMS.
|
||
`ResourceScope` holds the resolved set of resource ULIDs and project names visible
|
||
to a plan (immutable, with `include_paths`/`exclude_paths` glob filtering via
|
||
`PurePath.full_match()`). `ScopedBackendView` wraps text/vector/graph backends to
|
||
auto-inject the `scope` parameter into every query, and filters `TieredFragment`
|
||
visibility by project name and resource ID. `ScopedBackendSet` bundles scoped
|
||
views for all three backend types. `ResourceAliasResolver` translates user-facing
|
||
aliases to canonical resource ULIDs with uniqueness validation.
|
||
`resolve_resource_scope()` builds a `ResourceScope` from projects with
|
||
allowlist/denylist filtering. `validate_project_scope()` and
|
||
`validate_resource_scope()` guard against out-of-scope access. Three enforcement
|
||
hooks added to `ContextTierService`: `get_scoped_by_resource`,
|
||
`validate_fragment_scope`, `store_with_scope_check`. Includes 74 Behave BDD
|
||
scenarios (with full coverage-gap tests), 8 Robot Framework integration tests,
|
||
ASV benchmarks, and reference documentation. (#193)
|
||
- Added `builtin/plan-subplan` tool for strategy actors to emit `SUBPLAN_SPAWN` or
|
||
`SUBPLAN_PARALLEL_SPAWN` decisions. Validates payload via `SubplanPayload` (Pydantic),
|
||
applies defaults (merge strategy, max_parallel, dependencies), generates rationale text,
|
||
and optionally persists the `Decision` via an injected `DecisionService`. Includes
|
||
`register_subplan_tool`, `make_plan_subplan_spec` factory, Behave unit tests, Robot
|
||
Framework integration smoke tests, ASV benchmarks, and an actor YAML example. (#198)
|
||
- Extended `CorrectionService._compute_affected_subtree()` to BFS over both the structural
|
||
decision tree (parent-child) and the influence DAG (`decision_dependencies` edges). Added
|
||
cycle detection guard via visited set. Updated `DecisionService.record_decision()` to accept
|
||
`dependency_decision_ids` parameter for recording influence relationships during decision
|
||
creation. Added `DecisionService.get_influence_edges()` to retrieve influence DAG as
|
||
adjacency list. All public methods on `CorrectionService` (`analyze_impact`,
|
||
`execute_revert`, `execute_correction`, `generate_dry_run_report`) now accept optional
|
||
`influence_edges` parameter. Includes Behave BDD scenarios, Robot Framework integration
|
||
tests, and ASV benchmarks. (#542)
|
||
- Added Semantic Escalation system with `AutonomyController` class implementing
|
||
`should_proceed_automatically()` that computes confidence scores from weighted factors
|
||
(past_success_rate, codebase_familiarity, risk_assessment, invariant_complexity) and
|
||
compares them against automation profile thresholds. Includes `EscalationDecision`,
|
||
`ConfidenceFactors`, `OperationContext`, and `HistoricalOutcome` domain models.
|
||
Historical success tracking records outcomes for future confidence computation.
|
||
Integrates with all 8 built-in automation profiles. DI-wired as singleton
|
||
`autonomy_controller`. Includes 48 Behave BDD scenarios, 10 Robot Framework
|
||
integration tests, and ASV benchmarks. (#546)
|
||
- Added context strategy registry with `ContextStrategy` protocol,
|
||
`StrategyCapabilities`, `BackendSet`, `PlanContext`, `StrategyConfig`,
|
||
`ContextStrategyResult`, and `StrategyRegistryEntry` domain models.
|
||
Six built-in stub strategies (simple-keyword, semantic-embedding,
|
||
breadth-depth-navigator, arce, temporal-archaeology, plan-decision-context)
|
||
with spec-mandated quality scores and backend requirements.
|
||
`StrategyRegistry` supports config-driven registration, per-strategy
|
||
timeout/max-fragment limits, per-project enable/disable overrides,
|
||
plugin discovery from `"module:ClassName"` strings, and validation
|
||
that strategies declare supported resource types. (#191)
|
||
- Fixed `context inspect` to display project-scoped tier fragment counts instead of global
|
||
counts. Added `ContextTierService.get_scoped_metrics()` which returns fragment population
|
||
counts filtered to the target project while keeping hit/miss counters as global service
|
||
metrics. (#499)
|
||
- Fixed `context simulate --focus` to filter fragments by the specified focus URIs during
|
||
dry-run assembly. Previously, focus URIs were passed to `ContextRequest` but not applied
|
||
to fragment selection, making `--focus` a no-op. (#499)
|
||
- Wired project context CLI stubs (`inspect`, `simulate`, `set`, `show`) to live ACMS pipeline
|
||
services. `context inspect` queries `ContextTierService` for tier metrics and per-project
|
||
fragments with optional filtering by strategy, focus area, breadth, and depth. `context
|
||
simulate` performs dry-run context assembly using CRP models with configurable token budget
|
||
and assembly strategies. `context set` gains 12 ACMS pipeline options (`hot_max_tokens`,
|
||
`warm_max_decisions`, `cold_max_decisions`, `summary_max_tokens`, `temporal_scope`,
|
||
`auto_refresh`, `focus_area`, `breadth`, `depth`, `assembly_strategy`, `retrieval_strategy`,
|
||
`summary_strategy`). `context show` displays ACMS pipeline configuration alongside context
|
||
policy. Includes 28 Behave BDD scenarios for wiring coverage, updated Robot Framework
|
||
integration tests, and reference documentation. (#499)
|
||
- Added async command execution infrastructure allowing plan phases (Execute,
|
||
Apply) to run as background jobs processed by a thread pool of workers.
|
||
`AsyncJob` Pydantic v2 domain model with ULID primary key, status state machine
|
||
(`queued -> running -> succeeded/failed/cancelled`), and payload schema versioning.
|
||
`AsyncWorker` service with `ThreadPoolExecutor`-backed concurrent execution,
|
||
race-safe cancellation contract, stuck job detection, and job cleanup.
|
||
`InMemoryJobStore` with atomic `snapshot_counts()` and single-pass `remove_expired()`.
|
||
Plan lifecycle service wired to enqueue jobs when `async.enabled` is True.
|
||
Error messages redacted via `shared/redaction.py` before persisting to the audit
|
||
trail. Configurable via `async.enabled`, `async.max_workers`, `async.poll_interval`,
|
||
`async.job_timeout`, `async.job_ttl`. Includes `AsyncJobModel` SQLAlchemy model,
|
||
Alembic migration, Behave BDD scenarios, Robot Framework integration tests, ASV
|
||
benchmarks, and `docs/reference/async_architecture.md`. (#312)
|
||
- Added hot/warm/cold context tiers with `ContextTier`, `ActorRole`, `TieredFragment`,
|
||
`TierBudget`, `ActorContextView`, `TierMetrics`, and `ScopedBackendView` models.
|
||
`ContextTierService` provides store/get, promotion/demotion with cold-tier summarisation
|
||
hook, LRU eviction, per-actor filtered views (strategist/executor/reviewer), and
|
||
project-scoped isolation. Settings: `context_max_tokens_hot`, `context_max_decisions_warm`,
|
||
`context_max_decisions_cold`. DI-wired as singleton `context_tier_service`. Includes Behave
|
||
BDD scenarios, Robot Framework integration tests, ASV benchmarks, and
|
||
`docs/reference/context_tiers.md`. (#208)
|
||
- Added `ExecutionEnvironment` enum (`host`, `container`) and execution environment
|
||
routing with priority chain (tool > plan > project > default). Includes
|
||
`ExecutionEnvironmentResolver` service, `--execution-environment` CLI flags on
|
||
`agents plan use` and `agents plan execute`, project-level context config support
|
||
via `agents project context set --execution-environment`, tool runner wiring with
|
||
container availability validation, and clear error when container is selected but no
|
||
container resource is linked. Covered by Behave BDD scenarios, Robot Framework smoke
|
||
tests, and ASV benchmarks. (#512)
|
||
- Added multi-project subplan support with `MultiProjectMetadata`, `ProjectScope`,
|
||
`ChangeSetSummary`, `CrossProjectDependency`, and `ProjectScopeResolver` domain models.
|
||
`MultiProjectService` provides scope initialization, context resolution, per-project
|
||
changeset recording, and cross-project constraint validation. Plan model extended with
|
||
`multi_project_metadata` field, `is_multi_project` property, and `get_project_scope()`
|
||
method. CLI `plan status` shows per-project changeset summaries for multi-project plans.
|
||
Includes Behave BDD scenarios, Robot Framework smoke tests, ASV benchmarks, and reference
|
||
documentation. (#199)
|
||
- Added large-project hierarchical decomposition with 4+ levels and bounded context per
|
||
subplan. Includes clustering heuristics (directory, language, size), bounded dependency
|
||
closure with memoization for 10K+ files, DAG execution ordering with cycle detection,
|
||
and DecisionService integration for strategy_choice + subplan_spawn recording.
|
||
Configurable via `planner_max_depth`, `planner_max_files_per_subplan`,
|
||
`planner_max_tokens_per_subplan`, `planner_min_files_per_subplan` settings. (#205)
|
||
- Added `LLMTrace` Pydantic v2 domain model and `llm_traces` database table with
|
||
`LLMTraceRepository` for persisting LLM call telemetry (tokens, cost, latency,
|
||
tool calls, context hash, streaming flag, retry count, error). Defined 14
|
||
`OperationalMetricKey` values with `MetricEntry` / `MetricCollector` for plan-level
|
||
metrics. `TraceService` provides recording, querying, metric computation, and
|
||
optional LangSmith forwarding when `LANGCHAIN_TRACING_V2=true`. (#500)
|
||
- Added `SafetyProfile` domain model with configurable safety constraints (allowed skill
|
||
categories, sandbox/checkpoint requirements, human-approval flag, cost/retry limits) and
|
||
integrated it into the `Action` model via `from_config`/`as_cli_dict`. Persistence backed
|
||
by `safety_profile_json` column on `LifecycleActionModel` with Alembic migration
|
||
`c4_001_safety_profile_column`. Includes `resolve_safety_profile()` stub that raises
|
||
`NotImplementedError` in local mode (real enforcement deferred to server mode). Covered by
|
||
26 Behave BDD scenarios, 6 Robot Framework smoke tests, 5 ASV benchmark suites, and
|
||
`docs/reference/safety_profile.md` reference documentation. (#332)
|
||
- Added `devcontainer-instance`, `devcontainer-file`, and `container-instance` built-in
|
||
resource types with `DevcontainerHandler` and auto-discovery logic that scans for
|
||
`.devcontainer/` directories when `git-checkout` or `fs-directory` resources are linked.
|
||
Includes CLI support, Behave/Robot/ASV tests, and reference documentation. (#511)
|
||
- Added full lifecycle management for `devcontainer-instance` resources: lazy activation
|
||
on first tool target via `DevcontainerHandler.resolve()`, health checking with periodic
|
||
liveness probes, manual `agents resource stop`/`agents resource rebuild` CLI commands
|
||
(rebuild passes `--reset-container` to force container recreation), and automatic
|
||
session-scoped cleanup on session close or plan completion. Lifecycle state tracked via
|
||
`ContainerLifecycleTracker` with six states (`detected`/`building`/`running`/`stopping`/
|
||
`stopped`/`failed`) and validated transitions. Includes Behave BDD scenarios, Robot
|
||
Framework integration tests, ASV benchmarks, and reference documentation. (#514)
|
||
- Added skeleton compressor service (`SkeletonCompressorService`) for ACMS context inheritance.
|
||
Compresses parent plan context fragments by `skeleton_ratio` (0.0–1.0) for propagation to
|
||
child plans. Persists `SkeletonMetadata` (ratio, token counts, source decision IDs) on the
|
||
plan model for auditability. Includes stable fragment ordering, ratio validation with default
|
||
handling, and compression summary. (#194)
|
||
- Wired A2A local facade handlers to live application services. `session.create`/`close`
|
||
delegate to `SessionService`; `plan.create`/`execute`/`status`/`diff`/`apply` delegate to
|
||
`PlanLifecycleService`; `registry.list_tools` and `registry.list_resources` delegate to
|
||
`ToolRegistry` and `ResourceRegistryService`; `event.subscribe` delegates to `A2aEventQueue`.
|
||
`context.get` returns a stub pending ACMS pipeline. Added domain-to-A2A error code mapping
|
||
(`NOT_FOUND`, `VALIDATION_ERROR`, `INVALID_STATE`, `PLAN_ERROR`, etc.) via `map_domain_error()`.
|
||
(#501)
|
||
- Added `plan explain` and `plan tree` CLI commands for decision tree inspection
|
||
with json/yaml/table/rich output formats, and flags for superseded decisions,
|
||
context snapshots, and reasoning details. (#174)
|
||
- Replaced behave-parallel subprocess-per-feature execution model (342 Python interpreter
|
||
startups) with in-process execution via behave's `Runner` API. Sequential mode runs all
|
||
features in a single `Runner.run()` call; parallel mode uses `multiprocessing.Pool` with
|
||
`fork` for COW sharing of heavy modules. Coverage pipeline simplified to a single slipcover
|
||
invocation wrapping the entire process. Unit tests: 24m21s -> 2m05s (91% reduction);
|
||
coverage report: 75m20s -> 3m00s (96% reduction). (#481)
|
||
- Optimized 20 medium-slow BDD features (10-100s tier). Capped `time.sleep` and `asyncio.sleep`
|
||
globally at 10ms in `before_all` to eliminate retry/backoff waits; originals saved as
|
||
`time._original_sleep` / `asyncio._original_sleep` for timing-sensitive tests. Replaced
|
||
`subprocess.run` CLI invocations with `CliRunner` in coverage step files. Switched persistence
|
||
features to in-memory SQLite by default. Total tier runtime reduced from 565s to 21s
|
||
(96%). (#480)
|
||
- Optimized the 8 slowest BDD feature files (100-248s each, 64% of total runtime). Added
|
||
`@mock_only` tag support to skip unnecessary DB setup, extracted shared service-setup helpers
|
||
in `services_coverage_steps.py` (~200 lines of duplicated boilerplate removed), and introduced
|
||
lightweight in-memory plan service for actor-resolution tests. (#479)
|
||
- Added pre-migrated SQLite template database via `scripts/create_template_db.py` to eliminate
|
||
repeated Alembic migrations per BDD scenario. Nox sessions propagate the template via
|
||
`CLEVERAGENTS_TEMPLATE_DB` env var; `features/environment.py` monkey-patches
|
||
`MigrationRunner.init_or_upgrade` to copy the template for fresh scenario temp DBs, falling
|
||
through to real migrations for `:memory:`, existing files, and migration-runner unit
|
||
tests. (#483)
|
||
- Replaced coverage.py (sys.settrace) with slipcover (bytecode instrumentation) for faster
|
||
coverage collection. Each behave-parallel worker now produces per-feature JSON coverage files;
|
||
slipcover --merge combines them. CI workflow JSON key lookups handle both slipcover and
|
||
coverage.py output formats. Documentation updated to reflect slipcover as the coverage
|
||
tool. (#482)
|
||
- Added checkpointing and rollback with `CheckpointService` for creating, listing, pruning,
|
||
and deleting sandbox snapshots, and restoring sandbox state via `plan rollback <plan_id>
|
||
<checkpoint_id>` CLI command. Checkpoint domain models (`Checkpoint`, `CheckpointMetadata`,
|
||
`CheckpointRetentionPolicy`, `RollbackResult`) store sandbox refs, decision alignment,
|
||
checkpoint type (`pre_write`, `post_step`, `manual`), filesystem path, size, and structured
|
||
audit metadata (reason, source tool, phase). Retention policy auto-prunes oldest interior
|
||
checkpoints when exceeding `max_checkpoints` (default 50), preserving the first and most
|
||
recent. Guards reject rollback when plan is applied or sandbox is missing. `CorrectionService`
|
||
accepts an optional `CheckpointService` integration point for future revert delegation.
|
||
`CheckpointRepository` and `CheckpointModel` back persistence via the session-factory pattern
|
||
(ADR-007), with `UnitOfWorkContext.checkpoints` for cross-repository atomicity. Alembic
|
||
migration `m6_001_checkpoint_metadata` adds the `checkpoint_metadata` table. Includes Behave
|
||
BDD scenarios (33 scenarios, 129 steps), Robot Framework integration tests (11 test cases),
|
||
ASV benchmarks, and `docs/reference/checkpointing.md`. (#206)
|
||
- Added semantic validation service with AST-based rules for syntax errors, missing imports,
|
||
broken references, duplicate imports, API misuse, and missing symbols. Includes rule registry,
|
||
file-hash LRU cache, severity mapping, and ValidationPipeline integration. (#448)
|
||
- Added comprehensive M6 autonomy acceptance test suite covering A2A local facade dispatch
|
||
(session/plan/registry/context/event operations), event queue pub/sub with local callbacks
|
||
and close semantics, HTTP transport stub rejection, version negotiation, automation profile
|
||
built-ins (8 profiles), custom profile creation/validation/YAML loading, guard enforcement
|
||
(denylist, allowlist, call-limit, cost-budget, write-approval, apply-approval), and profile
|
||
service 4-level resolution precedence. Includes Behave BDD scenarios (52), Robot Framework
|
||
integration tests (11), ASV performance benchmarks (5 suites), JSON fixtures, and
|
||
documentation update. (#211)
|
||
- Added MCP refresh hooks to wire `notifications/tools/list_changed` events from MCP servers
|
||
to `SkillRegistry`. Introduced `SkillRegistry.refresh(name)` and `refresh_all()` to
|
||
recompute flattened tool sets on demand. Added `MCPRefreshHook` with configurable debounce
|
||
window (default 0.5 s) to coalesce rapid notification bursts into a single refresh call.
|
||
Refresh skips tool-ref validation when no `ToolRegistry` is configured and emits a single
|
||
`WARNING` with recovery steps. Results are summarised as an immutable `SkillRefreshResult`
|
||
(refreshed / failed / skipped counts) for CLI and log output. Includes Behave unit tests
|
||
(19 scenarios), Robot Framework integration tests (10 tests), ASV benchmarks, and
|
||
`docs/reference/skill_refresh.md`. (#168)
|
||
- Added `agents skill refresh <name>|--all` command to recompute tool flattening and sync
|
||
MCP-backed skills. Enhanced `skill list`, `skill show`, and `skill tools` outputs with
|
||
capability summary fields, tool counts, and description columns. Added `--format json/yaml`
|
||
schemas for refresh output. Updated CLI reference documentation with refresh examples and
|
||
caching behavior. (#167)
|
||
- Added UKO Layer 0-3 ontology scaffolding (RDF/TTL) aligned with specification
|
||
Section 14. Layer 0 (`uko:`) defines InformationUnit, Container, Atom,
|
||
Annotation, Boundary plus contains/references/dependsOn relationships,
|
||
content properties (hasRendering, renderingDepth, hasFullContent),
|
||
provenance properties (sourceResource, sourcePath, sourceRange), and
|
||
temporal properties (validFrom, validUntil, isCurrent, isRevisionOf).
|
||
Layer 1 (`uko-code:`) defines Module, Callable, TypeDefinition, TestCase,
|
||
Import plus hasReturnType/hasParameters/testsCallable. Layer 2 (`uko-oo:`)
|
||
defines Class, Interface, Method, Attribute plus inheritsFrom/implements
|
||
with `rdfs:subPropertyOf`. Layer 3 is reserved for DetailLevelMap
|
||
insertions. Loader supports semantic domain prefixes, hyphenated prefix
|
||
names, full-URI layer detection, multi-parent `rdfs:subClassOf` (DAG
|
||
traversal via BFS), `rdfs:domain`/`rdfs:range`/`rdfs:subPropertyOf`
|
||
resolution, and non-existent parent validation. (#189)
|
||
- Added ACMS v1 context assembly pipeline with UKO and CRP integration, three
|
||
fusion strategies (relevance, recency, tiered), budget-constrained assembly,
|
||
and extensible strategy registration. (#188)
|
||
- Added `AgentSkillSpec` loader that parses SKILL.md frontmatter and progressive disclosure
|
||
sections into structured `SkillStep` objects with stable 1-based ordering. Supports
|
||
namespaced naming (`namespace/short_name`), optional `steps`, `version`, `compatibility`,
|
||
`metadata`, and `allowed-tools` frontmatter fields. Explicit validation raises actionable
|
||
errors for missing `name`/`description` and invalid namespace format. Agent Skills are
|
||
mapped to `AgentSkillToolDescriptor` with `source="agent_skill"` and read-only defaults.
|
||
Support directories (`scripts/`, `references/`, `assets/`) are discovered automatically and
|
||
exposed as read-only `AgentSkillResourceSlot` bindings. Includes `docs/reference/agent_skills.md`,
|
||
Behave unit tests, Robot Framework integration tests, and ASV benchmarks. (#160)
|
||
- Added MCP adapter runtime (`MCPToolAdapter`) to connect to external MCP servers via stdio,
|
||
SSE, and streamable-http transports. Supports full connection lifecycle (connect with timeout,
|
||
reconnect, disconnect), tool discovery, input-validated invocation, and bulk registration into
|
||
`ToolRegistry` with `source="mcp"` and `checkpointable=False`. Includes Behave unit tests,
|
||
Robot integration test, ASV benchmarks, and `docs/reference/mcp_adapter.md`.
|
||
|
||
### fix(permissions): address code-review findings for permission system (#448)
|
||
|
||
- Added input validation to `check_permission()` and `get_role_bindings()`:
|
||
empty or whitespace-only `principal` and `scope_id` arguments now raise
|
||
`ValueError` after stripping, preventing silent lookup misses.
|
||
- Aligned module docstring and reference docs to clarify that the
|
||
`enforce_permission` decorator is available but not yet wired into CLI
|
||
or service call sites — integration is deferred to a future pass.
|
||
- Added `permissions.md` to the docs nav in `gen_ref_pages.py`.
|
||
- Added Behave BDD scenarios covering empty, whitespace-only, and
|
||
leading/trailing-whitespace inputs for both `check_permission()` and
|
||
`get_role_bindings()`.
|
||
|
||
### feat(actor): extend hierarchical actor YAML schema and loader
|
||
|
||
- Extended actor YAML schema with hierarchical graph support: per-node LSP bindings (`lsp_binding`), tool-source references (`tool_sources`), and subgraph `actor_ref`.
|
||
- Added graph reachability validation — all nodes must be reachable from `entry_node` via edges or conditional routing targets.
|
||
- Improved loader error reporting with YAML line/column positions and Pydantic field-path hints.
|
||
- Added `docs/reference/actor_config.md` — practical configuration reference with hierarchical examples and error cases.
|
||
- Fixed `examples/actors/graph_workflow.yaml` to use `actor_ref` instead of deprecated `actor_path`.
|
||
- Added Robot smoke test for loading hierarchical actor YAML via `ActorLoader.discover()` (#157).
|
||
- Added decision persistence layer with DecisionRepository, DecisionModel, Alembic migration,
|
||
tree queries (BFS traversal, path-to-root), superseded lookup, and ordered decision path
|
||
retrieval. Includes Behave BDD scenarios, Robot Framework integration tests, and ASV
|
||
benchmarks. (#171)
|
||
- Added token/cost tracking, budget enforcement (per-plan and per-day), provider fallback
|
||
selection with capability filtering, and cost metadata for plan execution. New config keys
|
||
`budget_per_plan`, `budget_per_day`, and `fallback_providers` control spending limits and
|
||
provider ordering. Budget warnings are emitted at 90% usage, and requests are blocked at 100%.
|
||
Per-provider cost table includes default token cost estimates for offline reporting. Budget
|
||
exhaustion events are persisted in plan metadata for auditability. (#324)
|
||
- Added comprehensive Behave, Robot Framework, and ASV test coverage for CLI extension
|
||
features including automation profile resolution, invariant ordering, actor override
|
||
error cases, and output format snapshot assertions.
|
||
- Added comprehensive E2E test suite for M2 (Actor Graphs + Tool Sources) epic covering
|
||
actor YAML loading, skill registry, tool lifecycle, and MCP stub tool discovery with
|
||
Behave BDD scenarios, Robot Framework integration tests, and ASV performance benchmarks.
|
||
- Added plan-level and project-level advisory locking with configurable timeouts, re-entrant
|
||
acquisition, conflict detection, lock renewal, graceful shutdown release, startup cleanup of
|
||
expired locks, and diagnostics check for stale lock reporting. (#327)
|
||
- Added core plan apply service with diff review output (plain, rich, JSON, YAML), artifact
|
||
summaries, apply summary persistence, merge-failure handling with sandbox rollback, and
|
||
empty ChangeSet guard. (#155)
|
||
- Added validation pipeline with rule-based checks, severity levels (required vs informational),
|
||
result aggregation, deterministic execution ordering, per-validation timeouts, and gate
|
||
enforcement that blocks apply when required validations fail. (#175)
|
||
- Added validation-gated apply pipeline that blocks the Apply phase when required Execute-phase
|
||
validations have not passed, transitions plans to constrained state with actionable CLI
|
||
hints, and runs validation attachments during apply. (#176)
|
||
- Added diff review artifact model with inline comments, approval status, per-resource grouping,
|
||
before/after content hashes, and plan apply service integration for `plan diff` and
|
||
`plan status` outputs. (#303)
|
||
- Added definition-of-done gating that evaluates DoD criteria before apply, blocks phase
|
||
transitions when required conditions are unmet, and stores pass/fail reasoning in the plan
|
||
validation summary. (#178)
|
||
- Added error recovery patterns (retry, fallback, skip, abort) with structured recovery hints
|
||
in CLI error output, plan executor integration, and error recovery service for capturing
|
||
error category, recovery action, and retry history. (#186)
|
||
- Hardened template rendering by replacing unsafe Jinja2 usage with a sandboxed renderer that
|
||
denies attribute access, function calls, and filters, allowing only `{var}` substitution
|
||
from a fixed allowlist with max template length and max output size enforcement. (#319)
|
||
- Enforced explicit exception handling by introducing a `CleverAgentsError` base class with
|
||
structured error types for configuration, provider, and file I/O failures, error code
|
||
mapping, bare-except prohibition, and secret redaction in error details. (#320)
|
||
- Added 32 BDD scenarios to boost unit test coverage from 97.0% to 97.2%.
|
||
- Added Behave BDD scenarios for six under-tested modules (container, correction
|
||
service, plan lifecycle service, plan CLI, skill CLI, database models) to exercise
|
||
uncovered lines, exception-handling paths, and partial branches. (#446)
|
||
- Fixed failing Robot Framework integration tests related to security secrets handling.
|
||
- Fixed style check violations across the codebase.
|
||
- Fixed failing unit tests.
|
||
- Added changeset persistence and diff artifact storage for tracking multi-file changes
|
||
across plan execution phases. (#163)
|
||
- Added `AsyncResourceTracker` for unified async resource lifecycle with timeout-bounded
|
||
cleanup, leak detection via finalizer, and async context manager support.
|
||
- Enhanced `LangGraphBridge` with graceful task cancellation that awaits in-flight tasks.
|
||
- Added `StateManager.close()` and `A2aEventQueue.close()` for proper resource disposal.
|
||
- Tightened read-only enforcement: write-capable tools are now blocked on read-only plans
|
||
regardless of the tool's own `read_only` flag.
|
||
- Added `ReadOnlyViolationError` to `ChangeSetCapture` to prevent write artifacts on
|
||
read-only plans.
|
||
- Added CLI fail-fast guards on `plan execute` and `plan apply` for read-only plans.
|
||
- Added `DecisionService` with record/list/tree helpers and `SnapshotStore` for hash-based
|
||
deduplication of context snapshots during plan execution.
|
||
- Expanded CONTRIBUTING.md with detailed guidance on the issue creation process, label system,
|
||
ticket lifecycle, pull request requirements, and review/merge process.
|
||
- Added commit scope, quality, and message format guidelines to CONTRIBUTING.md.
|
||
- Migrated implementation timeline from the monolithic implementation plan to `docs/timeline.md`.
|
||
- Migrated implementation notes to `docs/implementation-notes.md`.
|
||
- Relocated remaining implementation plan content to the specification and CONTRIBUTING.md,
|
||
and removed `implementation_plan.md` from the repository.
|
||
- Updated CONTRIBUTING.md to include project-specific conventions for tooling, testing,
|
||
type checking, and code style.
|
||
|
||
## v1.0.0
|
||
|
||
First release.
|