Add --project flag to config set, config get, and config list CLI
commands, enabling per-project configuration overrides stored under
[project."<name>"] TOML tables in the global config file.
Project-scoped resolution slots between environment variable and global
levels in the ConfigService resolution chain. config list --project
shows only overrides for the named project with source annotations.
Implementation:
- ConfigService: add set_project_value() and get_project_overrides()
methods for TOML-backed project-scoped persistence and retrieval
- CLI config commands: wire --project flag through set, get, and list
subcommands; project-scoped list filters to overrides only
- Database persistence: project-scoped config stored as alternative
backend for projects not using TOML
- Documentation: update docs/reference/config_resolution.md with
project-scopable key lists, CLI examples, precedence diagram, and
non-scopable key rejection behavior
Tests:
- Behave: 12 BDD scenarios in features/config_project_scope.feature
covering set/get/list, precedence over global defaults, and
non-scopable key rejection
- Robot: 5 integration smoke tests in robot/config_project_scope.robot
for end-to-end project-scoped round-trip verification
- ASV: benchmarks/config_project_scope_bench.py measuring resolution
overhead with project scope active
All nox quality gates pass: lint, typecheck, unit_tests (7522 scenarios),
integration_tests (Config Project Scope suite passed), and coverage at
98% line rate (threshold 97%).
Closes#259
Consolidate 141 BDD feature files that each complete in under 0.1 seconds
into 25 domain-grouped feature files, reducing subprocess count from 339
to ~223. Each consolidated file groups scenarios from the same domain/module
that share step definitions and fixtures. All scenarios are preserved with
clear comment headers indicating their original source file.
This reduces subprocess overhead by ~116 invocations (141 original files
replaced by 25 consolidated files), targeting the 42% of subprocess count
that contributed only 0.2% of actual test runtime.
ISSUES CLOSED: #485
Add protocol stubs for remote server communication infrastructure:
- ServerClient: Health check and version negotiation protocol
- RemoteExecutionClient: Remote plan execution and status protocol
- AuthClient: Authentication and token management protocol
- StubServerClient/StubRemoteExecutionClient/StubAuthClient: Stub
implementations raising NotImplementedError for all methods
- ServerConnectionConfig: Validated config model (server_url, namespace,
auth_token_ref, tls_verify)
- Config keys: core.server_url, core.server_namespace, core.server_tls_verify
- CLI: agents connect <url> stub command with explicit warning
- CLI: agents info shows Server Mode (disabled/stubbed)
ISSUES CLOSED: #201
Add SubplanService for building child plans from DecisionService spawn
entries and SubplanConfig. The service validates resource scopes, merge
strategies, and max_parallel bounds before spawning.
Key additions:
- SubplanService: Orchestrates child plan creation from spawn entries
- SpawnMetadata: Persisted metadata (spawn_decision_id, parent/root
plan IDs, execution mode) for status output
- Spawn validation: Checks resource scopes, merge strategy, and
parallelism bounds before spawning
- Documentation: subplan_service.md with spawn workflow and lifecycle
ISSUES CLOSED: #197
Add ASV benchmark modules to track test suite performance over time:
- bench_unit_tests.py: Feature discovery, step module loading, parallel
chunk computation, behave config parsing, and test suite metrics
- bench_coverage_report.py: Coverage collection pipeline timing and
report generation overhead
- bench_subprocess_overhead.py: Subprocess count tracking and per-feature
startup cost measurement
These benchmarks establish baselines for the post-optimization state
and enable CI-driven regression detection.
ISSUES CLOSED: #486
- Return parent_decision_id in handler result dict and add meaningful
assertion to the corresponding Behave scenario (B1)
- Add ULID format validators to SubplanPayload for plan_id, dependencies,
and parent_decision_id, consistent with Decision model validation (S1)
- Add non-empty string validator for resource_scopes items (S2)
- Narrow broad except Exception to except (ValidationError, ValueError)
so programming errors surface as real failures (S3)
- Document strategy_with_subplan.yaml in actors_examples.md and assert
it is referenced in the "Verify all examples are documented" scenario (B2)
ISSUES CLOSED: #198
Replace the subprocess-per-feature execution model (342 Python interpreter
startups) with direct use of behave's Runner API for in-process
execution.
Sequential mode (--processes 1 or BEHAVE_PARALLEL_COVERAGE=1): All
features run in a single Runner.run() call. Steps and hooks load once.
Parallel mode (--processes N, N>1): Features split into N chunks,
dispatched via multiprocessing.Pool with fork. Heavy modules shared
copy-on-write.
Proper format defaulting (mirrors behave.__main__.run_behave() logic
for -q flag). Summary extracted from runner.features status attributes
instead of regex-parsing stdout.
Simplified coverage pipeline: single slipcover invocation wraps the
entire behave-parallel process. No per-worker UUID files, no --merge
step needed. Coverage data produced in one build/coverage.json file.
Removed: behave-parallel tarball download from PyPI, tarfile and
urllib.request imports, per-worker subprocess.run() calls,
__SLIPCOVER_OUT__ placeholder mechanism, _build_base_args(),
_parse_summary(), regex-based summary parsing.
Results: nox -s unit_tests 24m21s -> 2m05s (91%); nox -s coverage_report
75m20s -> 3m00s (96%). Coverage: 98% (above 97% threshold).
ISSUES CLOSED: #481
Cap time.sleep and asyncio.sleep globally at 10ms in before_all to
eliminate retry/backoff waits (tenacity wait_fixed, retry_auto_debug
exponential backoff). Save originals as time._original_sleep and
asyncio._original_sleep for tests that need real wall-clock delays
(CircuitBreaker recovery, debounce timers, validation timeouts).
Enhance template-DB patch: add "db." prefix to _SCENARIO_DB_PREFIXES
and check db_path.stat().st_size > 0 so 0-byte auto-created SQLite
files receive the template copy instead of falling through to real
migrations.
Replace subprocess.run CLI invocations with typer.testing.CliRunner
in module_coverage, main_coverage_complete, and coverage_extras step
files, eliminating ~6s Python cold-start overhead per call.
Switch plan_persistence and action_persistence to in-memory SQLite by
default; only cross-restart scenarios use file-based via an explicit
Given step.
Replace MigrationRunner.run_migrations with Base.metadata.create_all
in plan_service_steps for freshly created in-memory engines.
Removed leftover debug comment in environment.py.
Total tier runtime: 565s -> 21s (96% reduction), 20 features all
under 5s behave-internal time.
ISSUES CLOSED: #480
Optimized the 8 features accounting for 64% of total BDD test runtime
(1,505s of 2,352s):
- features/environment.py: Added _ensure_template_db() for direct behave
invocations, expanded _SCENARIO_DB_PREFIXES to include "test_" for
step-file-created DBs, added @mock_only tag support to skip
unnecessary DB setup.
- features/plan_commands_coverage.feature: Added @mock_only tag (fully
mocked, no DB needed).
- features/plan_service.feature: 14 actor-resolution scenarios now use
lightweight in-memory plan service instead of heavyweight file-based
DB + project init.
- features/steps/plan_service_steps.py: New
step_create_lightweight_plan_service for actor-resolution tests.
- features/steps/services_coverage_steps.py: Extracted 3 helper functions
to consolidate 8 near-identical Given steps (~200 lines of duplicated
boilerplate removed).
Per-feature results: services_coverage 245s->2.8s (99%), context_service
215s->2.1s (99%), project_service 140s->0.7s (99%), plan_service
215s->4.6s (98%), core_cli_commands 113s->4.2s (96%), cli_streaming
213s->6.4s (97%), plan_commands_coverage 116s->20s (83%),
cli_plan_context_commands 248s->31.6s (87%).
ISSUES CLOSED: #479
Created scripts/create_template_db.py that builds a pre-migrated SQLite
template database using Base.metadata.create_all() + alembic stamp
(~5ms for 34 tables, vs ~0.5-3s x 25 Alembic migrations per scenario).
Nox unit_tests and coverage_report sessions generate the template before
test execution and propagate CLEVERAGENTS_TEMPLATE_DB env var to all
workers.
features/environment.py before_all() installs a monkey-patch on
MigrationRunner.init_or_upgrade that copies the template for fresh
scenario temp DBs, falling through to real migrations for :memory:,
existing files, and migration-runner unit tests.
Quick wins: sleep(0.5) -> sleep(0.05) in cli_streaming wait step;
removed redundant Background re-declaration in cli_streaming.feature
scenario 7.
ISSUES CLOSED: #483
Replace coverage.py (sys.settrace-based) with slipcover (bytecode-based
instrumentation) for significantly faster coverage collection:
- Each behave-parallel worker runs under slipcover, producing per-feature
JSON coverage files with unique UUIDs to avoid write contention.
- After all workers finish, slipcover --merge combines per-worker data
into a single build/coverage.json report.
- XML report generated via slipcover --merge --xml for CI tooling.
- Terminal report with --fail-under=97 threshold enforcement.
- Robust JSON key-fallback logic handles both slipcover and coverage.py
output formats.
- CI workflow (ci.yml, nightly-quality.yml) updated with defensive key
lookup instead of hardcoded coverage.py format.
- Documentation updated to reflect slipcover as the coverage tool.
- CHANGELOG.md updated.
ISSUES CLOSED: #482
Created ADR-042 (Resource Type Inheritance) defining single-inheritance
`inherits` field on resource type definitions with field resolution,
collection merging, handler inheritance, polymorphic tool binding,
auto-discovery, and DAG query matching. Max depth 5, single inheritance.
Created ADR-043 (Devcontainer Integration) defining devcontainer-instance
as a subtype of container-instance with lazy activation lifecycle,
devcontainer.json parsing, three container-project association patterns
(auto-detect, explicit mount, clone-into), and execution environment
routing with a 6-level precedence chain.
Specification updates across 20+ sections:
- Glossary: Resource Type Inheritance, Devcontainer, Execution Environment
- Resource type YAML schema: `inherits` field with structure reference
- Handler, sandbox strategy, and coherence tables: devcontainer-instance
- Auto-discovery: devcontainer detection subsection with WBS diagram
- Tool capability metadata: structured environment subfields
- Technology stack: devcontainer CLI row
- Project model: execution environment subsection with YAML example
- CLI agents resource add: --mount, --clone-into flags + 4 new examples
- CLI agents plan use: --execution-environment, --execution-env-priority
- CLI agents project context set: same execution environment flags
- Command synopsis block: updated for all new flags
- Execution environment routing section with precedence table + algorithm
- Resource type YAML Example 6: Devcontainer Instance (inherited type)
- End-to-end Example 16: Devcontainer-Driven Development
- End-to-end Example 17: Explicit Container with Directory Mount
- End-to-end Example 18: Container with Remote Repo Clone
Also fixed ADR index: added missing ADR-036 through ADR-040 entries,
updated next ADR number to 044.
CONTRIBUTING.md fixes backported from sister project:
- Fixed State label capitalization (State/In Progress -> State/In progress,
State/In Review -> State/In review) to match label definitions.
- Subtasks section: changed from optional to required, with exception
clause for trivially simple issues.
- Parent links: updated to use Forgejo dependency system instead of
textual references in issue descriptions.
- Fixed nox session flag typos: nox -e -> nox -s (7 occurrences).
- Replaced "Epics and Legendaries" section with comprehensive "Ticket
Type Hierarchy" defining the three-tier hierarchy (Issue -> Epic ->
Legendary) with formal criteria tables, cross-cutting rules for
hierarchy enforcement, completion semantics, promotion/demotion, and
milestone relationship rules.
- Fixed broken internal links referencing old "Epics and Legendaries"
anchor to use "Ticket Type Hierarchy" (3 occurrences).
- Traceability example: replaced project-specific code reference with
generic example.
- Removed stray horizontal rule before Project-Specific Guidelines.
ISSUES CLOSED: #491
Implement SubplanExecutionService and SubplanMergeService to enable
parent plans to decompose work into coordinated child subplans with
configurable execution modes (sequential, parallel, dependency-ordered)
and merge strategies (git_three_way, sequential_apply, fail_on_conflict,
last_wins).
- SubplanExecutionService: schedules subplan execution with retry
support via SubplanFailureHandler, max_parallel limits, fail_fast
semantics, and topological DAG ordering for dependency mode
- SubplanMergeService: wraps sandbox merge infrastructure to combine
subplan sandbox outputs using the configured merge strategy
- BDD tests: 21 scenarios covering all execution modes, merge
strategies, validation, and integration flows
- Robot tests: 11 integration test cases with helper module
- ASV benchmarks: performance benchmarks for execution and merge
- Documentation: reference guide in docs/reference/subplans.md
ISSUES CLOSED: #184
Implement the complete configuration system with multi-level resolution
chain, typed key registry, and CLI integration per specification.
ConfigService changes:
- Expand _build_catalog() to register all 102 spec-aligned config keys
across 8 groups: core (14), server (4), actor (5), plan (8),
sandbox (5), index (12), context (43), provider (11)
- Each key carries exact dotted-dash name, Python type, default value,
explicit env var name per spec, project-scopability flag, and
description
- Fix _env_name() to convert dots and dashes to underscores
- Provider keys use standard env var names (e.g., OPENAI_API_KEY)
CLI commands rewiring:
- Rewrite config set/get/list to use ConfigService instead of Settings
- Add --verbose flag to config get showing full 5-level resolution chain
- Add --project flag to config set/get/list for project-scoped overrides
- Support both glob and regex patterns in config list
- Validate keys against ConfigService registry with actionable errors
- Retain backward-compatible helper functions delegating to ConfigService
Documentation:
- Add docs/reference/config_resolution.md covering resolution chain,
all 102 config keys, CLI commands, TOML format, and provider credentials
Testing:
- Update all 4 Behave feature files and step definitions to use new
spec-aligned key names, env vars, and defaults (119 scenarios passing)
- Add robot/config_resolution.robot with 10 integration test cases
- Add benchmarks/config_resolution_bench.py with 8 time + 2 memory suites
ISSUES CLOSED: #258
Added targeted Behave BDD feature files and step definitions to improve
unit test coverage for:
- decision_service.py: Full coverage of all 7 service methods (18 scenarios)
- plan_apply_service.py: Branch coverage for handle_merge_failure (2 scenarios)
- plan_executor.py: Edge cases for rollback, checkpoint, and parse_steps (15 scenarios)
- cli/commands/plan.py: Uncovered region lines 1950-2273 (23 scenarios)
- repositories.py: Remaining missed branches and lines (14 scenarios)
- sandbox/checkpoint.py: Full coverage of CheckpointManager (26 scenarios)
- langgraph/bridge.py: Remaining uncovered lines and branches (10 scenarios)
- cli/commands/config.py: Safety net to maintain 100% coverage (42 scenarios)
Total: 150 new scenarios, 596 steps, all passing.
Also fixed a step definition collision in plan_lifecycle_coverage by renaming
"the delete result should be false" to "the plan delete result should be false".
ISSUES CLOSED: #475
Add runtime autonomy constraints (max steps, tool budget, required
confirmations) and a structured audit trail for plan execution.
New domain models:
- AutonomyGuardrails: enforces step limits, tool budgets, and
confirmation gates with validators and check methods
- GuardrailAuditEntry: records each enforcement event with timestamp,
event type, guard name, result, reason, and context
- GuardrailAuditTrail: ordered collection of audit entries persisted
to plan metadata
New service:
- AutonomyGuardrailService: high-level service for configuring
guardrails per plan, checking constraints, recording audit entries,
and serializing/restoring state via plan metadata
Tests:
- Behave: 69 scenarios covering model validation, step/budget/
confirmation checks, audit trail recording, and service operations
- Robot: 8 test cases for autonomy guardrail CLI flag smoke testing
- ASV: 6 benchmark suites measuring enforcement overhead
Documentation:
- Updated docs/reference/automation_profiles.md with guardrail fields,
enforcement behavior, audit trail schema, and event type reference
ISSUES CLOSED: #204
Implemented namespace/project/plan/skill permission model with role
bindings (owner/admin/editor/viewer) and default deny policy. Added
enforcement hooks at CLI/service boundaries that are server-only; local
mode returns permissive defaults. Includes role enums, permission check
service, role matrix documentation.
Includes Behave BDD scenarios, Robot integration tests, ASV benchmarks,
and reference documentation.
ISSUES CLOSED: #344
Patch time.monotonic alongside time.sleep in the retry timeout step
so stop_after_delay cannot elapse due to wall-clock overhead on
busy CI runners. The fake monotonic clock advances by 1µs per call,
keeping elapsed time well under the 0.01s timeout while preserving
retry semantics.
Closes#200
Add SafetyProfile as a first-class concept in the specification, composed
within AutomationProfile via a 'safety' field. This eliminates the
dual-authority problem where both AutomationProfile and a separate
SafetyProfile defined the same three safety booleans (require_sandbox,
require_checkpoints, allow_unsafe_tools) with no spec-defined resolution.
Changes:
- specification.md: Add Safety Profile glossary entry, split Automatable
Tasks into thresholds + Safety Profile sub-section, update built-in
profile matrix with safety.* prefix, update YAML examples
- ADR-041 (new): Document composition decision, field schema, relationship
to Guards, constraints, consequences, rejected alternatives (inheritance,
mixin, flat)
- ADR-017: Update profile fields table, built-in profiles, constraints,
risks, and cross-reference to ADR-041
- reference/automation_profiles.md: Rename Safety Fields to Safety Profile
sub-section, expand built-in matrix, update YAML examples
- schema/automation_profile.schema.yaml: Nest safety fields under safety
object with all SafetyProfile fields
- adr/index.md: Add ADR-041 to Tier 3 inventory
Resolves spec gap identified in issue #332.
Revert module-level import refactor and instead patch at
cleveragents.application.container.get_container — matching the
established pattern used by all existing context CLI tests.
Lazy imports inside each function re-read from the source module,
so patching there intercepts correctly.
Closes#200
Move get_container from lazy per-function imports to module-level in
context.py (10 sites) and project_context.py (3 sites). This makes
the symbol a patchable module attribute so BDD step files can
mock the DI container with unittest.mock.patch.
Closes#200
Introduce a lightweight checkpoint/rollback system for sandbox state
during plan execute and apply flows. CheckpointManager snapshots
the sandbox working directory before each phase and can restore it
on failure, giving the execution engine a reliable undo mechanism.
Key changes:
- SandboxCheckpoint model, Checkpointable protocol, and
CheckpointManager in infrastructure/sandbox/checkpoint.py
- PlanExecutor gains optional checkpoint_manager with pre/post
execute hooks and automatic rollback on failure
- PlanApplyService gains optional checkpoint_manager with pre-apply
checkpoint and rollback helper
- 12 BDD scenarios (features/sandbox_checkpoints.feature)
- 5 Robot Framework smoke tests (robot/sandbox_checkpoint_smoke.robot)
- ASV benchmarks for creation, rollback, and listing operations
- Reference documentation in docs/reference/sandbox.md
ISSUES CLOSED: #183