## Summary
Fixes the `plan correct` CLI handler which never passes the decision tree or influence edges to `CorrectionService`, producing single-node impact analysis regardless of the plan's actual decision structure.
Closes#606
## Root Cause
The `correct_decision()` handler in `plan.py` created a bare `CorrectionService()` and called `analyze_impact()` and `execute_correction()` without passing `decision_tree` or `influence_edges`. Both parameters default to `None` -> empty dicts, causing `_compute_affected_subtree()` BFS to return only the single target decision, ignoring all descendants and influence-DAG dependents.
## Changes
### Production Fix (`src/cleveragents/cli/commands/plan.py`)
- Resolve `DecisionService` via `get_container()` (following the pattern used by `plan explain` and `plan tree`)
- Build structural tree adjacency list from `decision_svc.list_decisions(plan_id)` using `parent_decision_id` relationships
- Fetch influence edges from `decision_svc.get_influence_edges(plan_id)`
- Pass both `decision_tree` and `influence_edges` to `svc.analyze_impact()` and `svc.execute_correction()`
### Existing Test Fixups
- Updated 4 step definition files that mock `CorrectionService` to also mock the new `DecisionService` resolution path
### New Tests
- **Behave BDD**: 3 scenarios verifying tree/edge forwarding (dry-run subtree, execution subtree, leaf node)
- **Robot Framework**: 3 integration smoke tests
- **ASV Benchmark**: Tree building and analyze_impact overhead benchmarks
## Quality Gates
- `nox -s lint` — PASSED
- `nox -s typecheck` — 0 errors
- `nox -s unit_tests` — 9,109 scenarios, 0 failures
- `nox -s coverage_report` — 97%
ISSUES CLOSED: #606
Reviewed-on: cleveragents/cleveragents-core#639
Co-authored-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-committed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Implement spec-mandated pre-flight guardrail checks that validate plan
readiness before entering the Strategize phase:
1. Action schema validation — verifies action exists and is well-formed.
2. Actor availability — confirms all 4 actor roles registered.
3. Skill/tool existence — transitively resolves all tools/skills.
4. Automation policy — verifies profile permits execution.
5. Rollback feasibility — ensures all tools checkpointable if required.
6. Resource accessibility — shallow connectivity check on linked resources.
7. Validation attachment resolution — pre-resolves all applicable validations.
PlanPreflightGuardrail service runs all 7 checks via run_all_checks().
On first failure, raises PreflightRejection with check name and message.
Wired into plan_lifecycle_service before Strategize phase.
Behave BDD: 18 scenarios covering all 7 checks (positive + negative).
Robot Framework: 3 integration smoke tests.
ASV benchmarks: pre-flight check execution time.
ISSUES CLOSED: #582
ToolRegistryRepository.create(), .update(), and .delete() called
session.flush() but never session.commit(). The CLI factory creates a
raw sessionmaker without a UnitOfWork wrapper, so the transaction was
never committed and SQLAlchemy performed an implicit rollback when the
session was garbage-collected.
The same bug existed in ValidationAttachmentRepository.attach() and
.detach().
Changes:
- Add session.commit() after session.flush() in all five mutating
methods across ToolRegistryRepository and
ValidationAttachmentRepository.
- Add finally: session.close() to guarantee session cleanup regardless
of success or failure.
- Update class docstrings to reflect the new commit-on-write semantics.
- Add Behave BDD feature (tool_add_persist.feature) with scenarios for
single-tool round-trip, multi-tool persistence, and duplicate
rejection, using file-based SQLite to reproduce the cross-session
issue.
- Add Robot Framework integration test (tool_add_persist.robot) with
add-then-list and fresh-list-empty scenarios.
- Add ASV benchmark (tool_add_persist_bench.py) with
track_list_after_add_count metric.
Key decisions:
- File-based SQLite (not in-memory) is used in tests because the bug
only manifests when the session/engine is fully disposed between add
and list, simulating separate CLI invocations.
- Step patterns are prefixed with "tool-persist" to avoid AmbiguousStep
collisions with existing tool_registry_steps.py.
- The commit-in-repository approach was chosen over adding a UnitOfWork
to the CLI factory because the CLI commands are simple CRUD operations
that should auto-persist without requiring callers to remember to
commit.
ISSUES CLOSED: #621
Wire SkillService to use SkillRepository for database persistence,
fixing the bug where `agents skill add` stored skills only in an
in-memory OrderedDict that was lost when the CLI process exited.
Changes:
- SkillService now accepts optional skill_repo and session_factory
parameters. When provided, add_skill() persists to the database
and the constructor pre-loads existing skills from DB rows.
- _get_skill_service() in the CLI now creates a DB-backed service
following the same engine/sessionmaker/repository pattern used by
the tool CLI (tool.py).
- _reset_skill_service() now installs a fresh in-memory SkillService
(instead of setting None) to avoid DB side-effects during unit
testing with parallel runners.
- remove_skill() also persists the deletion to the database.
Test coverage:
- Behave BDD: features/skill_add_persist.feature (4 scenarios)
- Robot Framework: robot/skill_add_persist.robot (3 smoke tests)
- ASV benchmark: benchmarks/skill_add_persist_bench.py
ISSUES CLOSED: #620
Add a call to bootstrap_builtin_types() in init_command() (project.py)
immediately after initialize_project() returns. This seeds the built-in
resource types (fs-directory, git-checkout, etc.) into the database so
that "resource add" commands succeed without "Resource type not found"
errors.
The call is idempotent — invoking it multiple times will not create
duplicate types.
Also fix the TDD robot test (resource_type_bootstrap_git.robot) to
initialize a project before running "resource add", and fix a
pre-existing parallel test failure in plan_commands_new_coverage where
unittest.mock.patch could not reliably intercept PlanApplyService under
behave-parallel fork() workers.
ISSUES CLOSED: #523, #524
Add TDD-style failing tests that verify the built-in git-checkout resource
type is available after initialization. Tests assert the correct expected
behavior: after agents init, the git-checkout type should exist in the
registry and 'agents resource add git-checkout' should succeed.
Tests are expected to fail until bug #524 is fixed, because
bootstrap_builtin_types() is never called during initialization. The fix
branch should be based on this branch so the fix commit inherits these tests.
Files added:
- features/resource_type_bootstrap_git.feature (2 Behave scenarios)
- features/steps/resource_type_bootstrap_git_steps.py (step definitions)
- robot/resource_type_bootstrap_git.robot (Robot Framework smoke test)
ISSUES CLOSED: #553
Added --yes/-y flag to the agents init CLI command to support non-interactive
initialization as specified in the documentation. When --yes is passed, all
interactive prompts are skipped and default values are applied. The command
produces the expected initialization summary output including data directory,
config file, database, and created directories.
The flag was added to both the top-level "agents init" command in main.py and
the "agents project init" subcommand in project.py. Both delegate to the shared
init_command() function which now accepts a yes parameter and produces
spec-aligned Rich panel output (Data Dir, Config, Database, Directories) with
the "Initialized (non-interactive)" status message when yes=True. When
yes=False (default), existing interactive behavior is preserved unchanged.
ISSUES CLOSED: #522
Implemented the remaining ACMS pipeline components and advanced context
strategies:
Pipeline Phase 3:
- FragmentOrdererProtocol + RelevanceCoherenceOrderer: orders fragments
by relevance while maintaining narrative coherence via UKO node prefix
grouping. Groups related fragments together, sorts groups by max
relevance, and within groups orders by relevance desc / depth asc.
- PreambleGeneratorProtocol + ProvenancePreambleGenerator: generates
provenance preamble with strategy contributions (fragment counts and
token percentages), confidence indicators (avg/min/max), tier and
depth distribution, UKO node coverage, and coverage gap detection.
Advanced Strategies:
- ArceStrategy (quality 0.95): adaptive recursive context expansion with
iterative multi-backend refinement and configurable iteration limit
(default 5) to prevent unbounded refinement. Uses composite scoring
(relevance + depth + diversity) with contextual boosting for fragments
related to the current top-ranked anchor set.
- TemporalArchaeologyStrategy (quality 0.5): historical context retrieval
from graph+cold backends. Prioritises cold-tier fragments using a
temporal scoring model (tier bonus + relevance + depth).
- PlanDecisionContextStrategy (quality 0.7): decision history retrieval
from warm/cold backends. Prioritises warm then cold tier fragments
for correction and retry scenarios.
All strategies registered in strategy registry with correct quality scores
and backend requirements. All components implement their respective
Protocol interfaces and can be injected into the ContextAssemblyPipeline
via constructor dependency injection.
Tests:
- 33 BDD scenarios in features/acms_pipeline_phase3.feature
- Robot Framework integration tests in robot/acms_pipeline_phase3.robot
- ASV performance benchmarks in benchmarks/acms_pipeline_phase3_bench.py
ISSUES CLOSED: #545
The "Resource Add Git Checkout Should Not Fail With Type Not Found"
Robot test was calling the CLI process directly, which hits the real
database file that does not exist in CI — causing
sqlite3.OperationalError. Rewrite the test to use an in-memory
SQLite database with register_resource(), matching the pattern used
by the fs-directory Robot tests.
Also fix common.resource path to use ${CURDIR}/common.resource.
Refs: #524
Day 26 (2026-03-06) comprehensive PM update:
- Schedule adherence entry with milestone forecasts and developer status
- Updated Current Status Summary: 272 issues closed (77%), 1570/1922 SP (82%)
- PR #617 (UKO Layer 1) and #611 (domain analyzers) merged late Day 26
- Open PR count: 19 -> 15 (merge rate exceeding new PR rate for first time)
- Spec gap analysis: no major coverage gaps found across all 44K-line spec
- All 15 open PRs have PM status comments with action items and deadlines
- Updated milestone roadmap, track forecasts, developer forecasts, risk summary
- Workstream and completion sections updated with Day 26 merges
ISSUES CLOSED: none (PM/docs update only)
Add TDD-style Behave BDD tests for the built-in fs-directory resource type
bootstrap (bug #523). Three Gherkin scenarios: one failing TDD test
reproducing the bug (no bootstrap called during init, tagged @wip), and
two regression tests verifying bootstrap_builtin_types() seeds correct
data and resource add fs-directory succeeds after bootstrap. Includes
Robot Framework regression tests.
Review feedback addressed:
- Removed all 21 unnecessary # type: ignore comments (hurui200320 M1)
- Fixed is not True to is False for clarity (Aditya F2)
- Fixed Robot common.resource path to ${CURDIR}/common.resource (hurui200320 L1)
- Squashed all commits into one and rebased onto master (C1, C2)
- Added CHANGELOG entry with correct scenario count
Closes#537
Move the (#588) CHANGELOG entry from after (#203) to the top of the
## Unreleased section, matching the convention that newest entries are
prepended first. All pre-existing entries (#473, #495, #203, #494)
remain intact and unmodified.
Addresses reviewer finding F8 from PR #611 second-pass review.
Refs: #588
Added minimal LSP server entrypoint supporting initialize/shutdown/exit
handshake over JSON-RPC stdin/stdout transport with Content-Length
framing. Unsupported methods return MethodNotFound error with descriptive
message. Wired LSP requests through ACP facade in local mode. Added
agents lsp serve CLI command with --log-level flag, PID output, and
startup banner. Created reference documentation for the stub server.
Includes Behave BDD tests for protocol handshake, Robot smoke test, and
ASV startup latency benchmark.
ISSUES CLOSED: #203
Core domain types (FragmentProvenance, ContextFragment, ContextBudget,
ContextPayload) now extend their CRP counterparts via Pydantic v2
inheritance, ensuring isinstance compatibility across the model
hierarchy.
Key changes:
- CRP base types made frozen=True (no consumer mutates them)
- CRP AssembledContext fields changed from list to tuple (frozen consistency)
- Core types extend CRP bases: FragmentProvenance(CRPFragmentProvenance),
ContextFragment(CRPContextFragment), ContextBudget(CRPContextBudget),
ContextPayload(CRPAssembledContext)
- Removed duplicate ContextFragment dataclass from skeleton_compressor
- Updated project_context.py to pass tuples to frozen AssembledContext
- Added Behave tests (10 scenarios), Robot integration tests (3 cases),
and ASV benchmarks for the unified hierarchy
- Updated Known Limitations table in docs/reference/acms.md
ISSUES CLOSED: #569
- SPEC-1: Added genuine TDD failing Scenario 1 (@wip) that creates registry
WITHOUT bootstrap and asserts git-checkout exists — reproduces bug #524.
Existing scenarios retained as regression tests (no @wip since they pass).
Added NOTE FOR FIX AUTHOR comment documenting fix-path expectations.
- BUG-1: Removed colliding @when('I run "agents resource add..."') step.
Replaced with uniquely-prefixed bootstrap-git step pattern that invokes
resource_add() directly with mocked DI, avoiding AmbiguousStep collision
with wildcard @when('I run "{command}"') in cli_plan_context_commands_steps.
- BUG-2: Removed duplicate @then('the CLI exit code should be {code:d}').
Replaced with prefixed bootstrap-git assertion steps.
- BUG-3: Removed duplicate @then('the CLI output should not contain...").
Replaced with prefixed bootstrap-git assertion steps.
- TEST-1: Replaced bare MagicMock() with direct service patching via
_PATCH_SERVICE, consistent with PR #567 pattern.
- TEST-2: Updated Robot docs from 'expected to FAIL' to 'regression tests'
since both Robot tests call bootstrap explicitly and pass.
- CODE-1: Simplified hasattr guards on enum fields — removed redundant
hasattr checks, using .value directly since ResourceKind and
SandboxStrategy are always enums.
- TEST-3: Added assertion on bootstrap_builtin_types() return value via
new Then step 'the bootstrap-git registered types should include'.
- Updated CHANGELOG from 'Two scenarios' to 'Three scenarios'.
Refs: #553
Implemented the analyzer plugin framework with AnalyzerProtocol,
AnalyzerRegistry for registration/discovery by file extension,
PythonAnalyzer (AST-based extraction of modules, classes, functions,
imports, docstrings), and MarkdownAnalyzer (section, code block, and
link extraction). Both analyzers produce well-formed UKO triples with
proper URI schemes.
ISSUES CLOSED: #551