# Work Remaining This section describes what remains to be done to bring the current CleverAgents codebase up to the specification. **Last Updated:** February 6, 2026 ### Current State Assessment #### What's Implemented Based on analysis of the current codebase: 1. **Plan Lifecycle Foundation** - The 4-phase lifecycle (Action → Strategize → Execute → Apply), where Action serves as a reusable template stage, is partially implemented in `plan_lifecycle_service.py` 2. **Database Models** - Models exist for projects, plans, contexts, changes, and actors 3. **LangGraph Integration** - Graph-based workflow support exists but is not fully connected to the plan workflow 4. **Reactive System** - A reactive system with stream routing is present 5. **Actor System** - Actor models and services exist but lack full behavioral definitions 6. **Basic Context Analysis** - Simple context loading and analysis capabilities 7. **Change Tracking** - Basic `Change` and `ChangeSet` models exist #### What's Missing Critical gaps between the specification and current implementation: 1. **No Resource Abstraction** - The unified resource layer for files, databases, APIs is entirely missing 2. **No Sandboxing** - Execute phase writes directly to files without isolation 3. **No Decision Tree** - No decision tracking, storage, or correction mechanism 4. **No Skills/MCP Integration** - No skill registry, skill YAML parsing, tool adapters, or MCP protocol support 5. **Single-File Limitation** - Hard-coded to generate exactly one file per plan 6. **Text-Based Code Generation** - Still parsing LLM output instead of tool-based approach 7. **No Checkpointing** - No rollback or checkpoint capabilities 8. **No Code Intelligence / ACMS** - Missing ACMS implementation: UKO ontology analyzers, Context Assembly Pipeline (the 10-component pluggable pipeline for strategy orchestration, fragment fusion, and context finalization), CRP skill, context budget calculator, and backend indices (text, vector, graph). The full design is specified in the "Advanced Context Management System (ACMS)" section. 9. **No Context Tiers** - No hot/warm/cold memory architecture, no warm-to-cold demotion, no temporal versioning of UKO nodes 10. **Limited Validation** - Basic stub validation instead of semantic checks ### Work Items by Priority ### 1) Implement Tool-Based Resource Modification (Critical Foundation) #### Problem The current system generates code as a single text blob that gets written to one file. The specification requires a modern tool-based approach where LLMs invoke discrete operations on resources. #### Implementation Steps 1. **Create Resource Abstraction Layer**

   # New modules needed:
   src/cleveragents/domain/models/resources.py
   src/cleveragents/domain/resources/handlers.py
   src/cleveragents/domain/resources/sandbox.py
   
2. **Implement Built-in Tools (via built-in skills)**

   # Core file operations (provided by built-in skill groups)
   read_file(path: str) -> str
   write_file(path: str, content: str) -> None
   edit_file(path: str, changes: list[Edit]) -> None
   delete_file(path: str) -> None
   move_file(src: str, dst: str) -> None
   create_directory(path: str) -> None
   list_files(pattern: str) -> list[str]
   search_files(pattern: str, content_pattern: str) -> list[Match]
   
3. **Connect Tools to ChangeSet** - Each tool invocation that modifies resources creates a `Change` record - ChangeSet accumulates these changes during execution - No more parsing LLM text output for code #### Estimated Effort: 2-3 weeks ### 2) Implement Sandbox Infrastructure (Critical for Safety) #### Problem Execute phase currently writes directly to the project. The specification requires all changes happen in an isolated sandbox that can be reviewed before applying. #### Implementation Steps 1. **Define Sandbox Interface**

   class Sandbox(Protocol):
       def create() -> SandboxRef
       def read(path: str) -> Content
       def write(path: str, content: Content) -> Change
       def diff() -> DiffView
       def commit() -> None
       def rollback() -> None
   
2. **Implement Sandbox Strategies** - `GitWorktreeSandbox` - For git repositories (preferred) - `FilesystemCopySandbox` - For non-git projects - `TransactionSandbox` - For databases - `NoOpSandbox` - For non-sandboxable resources 3. **Lazy Sandbox Creation** - Only create sandboxes when resources are accessed - Each plan gets its own sandbox namespace #### Estimated Effort: 2 weeks ### 3) Build Decision Tree System (Core Innovation) #### Problem No decision tracking exists. The specification's key innovation is recording every decision with full context, enabling correction without full re-execution. #### Implementation Steps 1. **Create Decision Models**

   -- New tables needed
   CREATE TABLE decisions (
       decision_id TEXT PRIMARY KEY,  -- ULID
       plan_id TEXT NOT NULL,
       parent_decision_id TEXT,
       decision_type TEXT NOT NULL,
       question TEXT NOT NULL,
       chosen_option TEXT NOT NULL,
       alternatives_considered TEXT,  -- JSON array
       context_snapshot TEXT NOT NULL,  -- JSON
       created_at TEXT NOT NULL
   );
   
2. **Implement Decision Recording** - Every choice during Strategize creates a Decision record - Capture complete context snapshot with each decision - Track downstream dependencies 3. **Build Correction Mechanism**

   agents plan correct <decision_id> --mode=revert --guidance "<new decision>"
   
- Mark decision as superseded - Recompute only affected subtree - Preserve unaffected work #### Estimated Effort: 3 weeks ### 4) Implement Advanced Context Management System (ACMS) (Scalability Enabler) #### Problem Current context is limited to a few hundred characters from a few files. Large codebases require intelligent, pluggable context discovery with multiple retrieval strategies, a universal knowledge representation, and demand-driven context assembly. The full ACMS design is specified in the "Advanced Context Management System (ACMS)" section of this document. #### Implementation Steps 1. **Universal Knowledge Ontology (UKO) Foundation** - Implement the four-layer UKO ontology hierarchy (Layer 0: universal foundation, Layer 1: domain specializations, Layer 2: paradigm/format specializations, Layer 3: technology-specific) - Implement the DetailDepth / DetailLevelMap system with inheritance-based resolution (Layer 0 universal integer semantics, per-domain named level maps at Layer 1, paradigm refinements at Layer 2, language-specific insertions at Layer 3) - Build source code analyzers for Python, TypeScript, Rust, Java (extract UKO nodes from source code via `uko-code:`, `uko-oo:`, `uko-func:`, `uko-py:`, etc.) - Implement document analyzers for Markdown, reStructuredText, HTML (`uko-doc:`) with semantic awareness (topic extraction, implicit cross-reference inference via `uko-doc:discussesTopic`, citation resolution) - Implement data schema analyzers for PostgreSQL, MySQL, SQLite (`uko-data:`) with schema introspection, foreign key graph construction, view dependency analysis, and stored procedure parsing - Implement infrastructure analyzers for Docker Compose, Kubernetes (`uko-infra:`) with service dependency graph construction and config reference resolution - Implement temporal versioning (`validFrom`, `validUntil`, `isCurrent`, `isRevisionOf`) - Integrate eager indexing on resource add via the Resource Registry lifecycle hooks 2. **Three-Index Backend Architecture** - **Text Index**: Tantivy/SQLite FTS for exact keyword matches - **Vector Index**: FAISS/Qdrant for semantic embedding search - **Graph Store**: RDF store for UKO relationship traversal - Implement `ScopedBackendView` to auto-filter all queries to plan project resources 3. **Context Request Protocol (CRP)** - Implement the `builtin/context` skill with tools: `request_context`, `query_history`, `get_context_budget` - Implement `ContextRequest` and `ContextResponse` data structures - Integrate CRP with actor execution loop 4. **Context Assembly Pipeline (10-component pluggable pipeline)** - Implement `ContextStrategy` base class and strategy registration - Build six built-in strategies: `simple-keyword` (0.3), `semantic-embedding` (0.6), `breadth-depth-navigator` (0.85), `arce` (0.95), `temporal-archaeology` (0.5), `plan-decision-context` (0.7) - **Phase 1 — Strategy Orchestration**: Implement `StrategySelectorProtocol` / `ConfidenceWeightedSelector`, `BudgetAllocatorProtocol` / `ProportionalBudgetAllocator`, `StrategyExecutorProtocol` / `ParallelStrategyExecutor` (with circuit breaker) - **Phase 2 — Fragment Fusion**: Implement `FragmentDeduplicatorProtocol` / `ContentHashDeduplicator`, `DetailDepthResolverProtocol` / `MaxDepthResolver`, `FragmentScorerProtocol` / `WeightedCompositeScorer`, `BudgetPackerProtocol` / `GreedyKnapsackPacker`, `FragmentOrdererProtocol` / `RelevanceCoherenceOrderer` - **Phase 3 — Context Finalization**: Implement `PreambleGeneratorProtocol` / `ProvenancePreambleGenerator`, `SkeletonCompressorProtocol` / `DepthReductionCompressor` - Implement `ContextAssemblyPipeline` orchestrator and `PipelineConfigResolver` (scope chain: plan > project > global > built-in) 6. **Context Budget Calculator & Refresh Manager** - Dynamic budget computation from model context window, response reserve, tool definitions, system prompt - Skeleton compression and plan context inheritance - Automatic re-assembly on budget change exceeding refresh threshold 7. **Tiered Context System** - **Hot**: Current working set assembled by ACMS (in LLM context) - **Warm**: Recent decisions and search results (warm retention window) - **Cold**: Historical data, patterns, and archived UKO snapshots (queryable via retrieval) #### Estimated Effort: 8-10 weeks ### 5) Complete Actor System with Behavioral Definitions #### Problem Actors exist but don't define behavior beyond model selection. The specification requires actors to be composable graphs with tools, memory, and context policies. #### Implementation Steps 1. **Extend Actor Configuration**

   actors:
     my_strategist:
       type: graph
       config:
         provider: anthropic
         model: claude-3-opus
         memory_policy: per_plan
         context_view: architect # High-level view
       skills:
         - local/file-ops             # provides read_file
         - local/code-intelligence    # provides search_code, analyze_dependencies
       routes:
         strategize:
           entry_point: analyze
           nodes:
             - name: analyze
               type: llm
             - name: plan
               type: llm
   
2. **Implement Tool Access Policies** - Strategy actors: read-only tools - Execute actors: read/write within sandbox - Apply actors: commit tools 3. **Actor-Specific Context Views** - Strategist: Architecture, dependencies, patterns - Executor: Implementation details, specific files - Reviewer: Diffs, tests, risk analysis #### Estimated Effort: 2 weeks ### 6) Resource System, Resource Types, and Resource Registry #### Problem Resources are currently defined inline within projects. The specification requires resources to be independently registered first-class entities with a type system, DAG relationships (physical/virtual), auto-discovery, and content-identity tracking. Tools need resource bindings to declare and resolve their resource dependencies. #### Implementation Steps 1. **Resource Type Registry and CLI** - Implement 24 built-in resource types: 15 physical (`git-checkout`, `git`, `git-remote`, `git-branch`, `git-tag`, `git-commit`, `git-tree`, `git-tree-entry`, `git-stash`, `git-submodule`, `fs-mount`, `fs-directory`, `fs-file`, `fs-symlink`, `fs-hardlink`) and 9 virtual (`file`, `directory`, `symlink`, `commit`, `branch`, `tag`, `remote`, `submodule`, `tree`) - Implement resource type YAML parsing and validation for custom types - Implement `agents resource type add [--update]/remove/list/show` commands - Implement dynamic CLI subcommand registration (custom types create new `agents resource add` subcommands) 2. **Resource Registry and CLI** - Implement `agents resource add /remove/list/show/tree` commands - Implement Resource Registry persistence in database - Implement DAG parent/child relationships with type constraints and cycle detection - Implement `agents resource link-child`/`unlink-child` for manual DAG management 3. **Auto-Discovery System** - Implement handler-driven child resource discovery (git-checkout discovers git + fs-directory worktree root; git discovers remotes, branches, tags, commits, stashes, submodules; git-commit discovers root git-tree; git-tree discovers tree entries + subtrees; fs-mount discovers root fs-directory; fs-directory discovers subdirectories, files, symlinks, hardlinks) - Implement resource reuse during discovery (link existing resources instead of duplicating) - Implement refresh mechanism for keeping discovered children up to date 4. **Physical/Virtual Resource Model** - Implement content hashing for identity tracking - Implement virtual resource linking (shared virtual parents for identical physical resources) - Implement divergence detection (unlink when content changes) 5. **Project-Resource Linking** - Replace `agents project add-resource`/`remove-resource` with `link-resource`/`unlink-resource` - Implement project-level overrides (read-only, alias) - Migrate any existing inline resource data to Resource Registry 6. **Resource Handlers** - Implement `GitCheckoutHandler` with auto-discovery, sandbox creation, checkpoint support - Implement `GitHandler` for git repository structure discovery (remotes, branches, tags, commits, stashes, submodules) - Implement `GitObjectHandler` for git-commit, git-tree, git-tree-entry (read-only access to git objects) - Implement `GitRefHandler` for git-branch, git-tag, git-stash (read/write ref manipulation) - Implement `GitConfigHandler` for git-remote, git-submodule (read-only config access) - Implement `FilesystemHandler` with auto-discovery for fs-mount, fs-directory, fs-file, fs-symlink, fs-hardlink - Implement handler plugin system for custom resource types #### Estimated Effort: 5 weeks ### 7) MCP Integration, Tool System, Skill System, and Tool-Resource Bindings #### Problem No tool or skill abstraction exists. The specification defines tools as independently registered operations managed via `agents tool` CLI commands, and skills as namespaced collections of tools managed via `agents skill` CLI commands. Tools can be sourced from MCP servers, Agent Skills folders, built-ins, and custom code. Tools also need resource bindings to declare and resolve their dependencies on resources. #### Implementation Steps 1. **Tool Registry and CLI** - Implement tool YAML parsing and validation (including `resources` section for resource slots) - Implement `agents tool add [--update]/remove/list/show` commands - Implement Tool Registry persistence in database - Implement anonymous tool support (inline definitions without registration) 2. **Skill Registry and CLI** - Implement skill YAML parsing and validation - Implement `agents skill add [--update]/remove/list/show/tools` commands - Implement hierarchical skill composition (includes) with per-tool metadata overrides - Implement Skill Registry persistence in database - Implement named tool references (resolving from Tool Registry) and anonymous inline tools 3. **Tool-Resource Binding System** - Implement resource slot parsing and validation in tool YAML - Implement three binding modes: contextual, static, parameter - Implement binding resolution at activation time (contextual/static) and invocation time (parameter) - Implement resource type compatibility validation - Implement built-in tool implicit resource bindings - Inject bound resources into `ToolExecutionContext.resources` 4. **MCP Tool Adapter**

   class MCPToolAdapter:
       def wrap_mcp_tool(self, tool) -> Tool:
           # Add sandbox interception
           # Add change tracking
           # Add capability metadata
           # Add resource binding support
   
5. **Tool Capability Metadata**

   class ToolCapability:
       read_only: bool
       write_scope: list[str]       # References resource slot names
       checkpointable: bool
       idempotent: bool
       side_effects: list[str]
   
6. **External MCP Server Support**

   # In tool or skill YAML configuration
   mcp_servers:
     - name: github
       command: "npx @anthropic/mcp-github"
       env: {GITHUB_TOKEN: "${GITHUB_TOKEN}"}
   
7. **Agent Skills Adapter** - Implement SKILL.md frontmatter parsing for discovery - Implement progressive disclosure (metadata → instructions → resources) - Implement sandboxed script execution 8. **Metadata Override System** - Implement shallow-merge override logic for tool capability metadata - Support overrides at skill-level (tool refs), include-level (tool_overrides), and actor graph node-level - Built-in tool metadata is not overridable #### Estimated Effort: 5 weeks ### 8) Implement Validation and Semantic Error Prevention #### Problem Current validation is a stub. The specification requires multi-layer semantic validation. Validations are now a specialized subtype of Tool — they extend the Tool class with a `mode` (required/informational) and a structured JSON return format (`passed`, `message`, `data`). They can be attached to resources directly, through projects, or through plans. #### Implementation Steps 1. **Decision-Time Validation** - Validate choices during Strategize - Check alternatives for feasibility - Record validation in decision metadata 2. **Execution-Time Guards**

   # Actor references a skill containing validation tools
   skills:
     - local/semantic-validators   # validate_api_compatibility,
                                   # check_invariants, verify_test_coverage
   
3. **Validation as Tool Subtype**

   # Register validations (tool subtypes)
   agents validation add --config ./validations/run-tests.yaml --required local/run-tests
   agents validation add --config ./validations/lint-check.yaml --required local/lint-check

   # Attach to resource (directly or scoped to a project)
   agents validation attach --project my-api my-api-repo local/run-tests
   agents validation attach my-api-repo local/lint-check
   
4. **Validation Class Inheritance** - Implement Validation as a class extending Tool - Reuse Tool Registry, binding resolution, and execution flow - Add only: mode field, return format enforcement, attachment scoping 5. **Tool Wrapping** - Implement `wraps` field resolution: validate wrapped Tool exists at registration time, store reference - Implement `transform` function: sandboxed Python execution that converts wrapped Tool output to `{ "passed", "message", "data" }` - Implement property inheritance from wrapped Tool (`input_schema`, `resource_slots`, `timeout`) with override semantics - Enforce read-only semantics on wrapped Tool invocation regardless of wrapped Tool's `writes` flag - Emit registration-time warning when wrapping a Tool with `writes: true` - Implement `argument_mapping` field: map Validation input arguments to wrapped Tool input arguments with support for renaming and fixed values - (Deferred) Wrapped Tool deduplication: when multiple Validations wrap the same Tool with identical forwarded arguments, invoke it once and fan out output to each `transform`. Deferred due to complexity around argument equivalence and cache invalidation during fix-then-revalidate loops - Implement default transform (pass-through for output already containing `passed` field) #### Estimated Effort: 2 weeks ### 9) Connect LangGraph to Plan Lifecycle #### Problem The reactive/LangGraph infrastructure exists but isn't connected to the main plan workflow. #### Implementation Steps 1. **Create Unified Plan Graph**

   class PlanLifecycleGraph:
       def strategize_subgraph(self) -> StateGraph
       def execute_subgraph(self) -> StateGraph
       def apply_subgraph(self) -> StateGraph
   
2. **Wire Phase Transitions** - `use` command triggers strategize graph - `execute` command triggers execute graph - `apply` command triggers apply graph 3. **Remove Linear Pipeline** - Replace tell/build/apply with graph execution - Maintain backward compatibility at CLI level #### Estimated Effort: 1 week ### Implementation Roadmap #### Phase 1: Foundation (8-10 weeks) 1. Tool-based resource modification (critical foundation) 2. Sandbox infrastructure 3. Decision tree system #### Phase 2: Core Systems (10-12 weeks) 4. Advanced Context Management System (ACMS) — UKO, CRP, strategies, fusion, budget 5. Actor system with behavioral definitions 6. Resource system, resource types, and resource registry 7. MCP integration, tool system, skill system, and tool-resource bindings #### Phase 3: Validation and Integration (3 weeks) 8. Validation and semantic error prevention 9. Connect LangGraph to plan lifecycle #### Phase 4: Production Hardening (4 weeks) 10. Checkpointing and rollback 11. Cost controls and rate limiting 12. Security fixes (remove eval, fix async) ### Total Estimated Timeline: 6-7 months ### Key Success Metrics 1. **Multi-file generation**: Can generate a REST API with routes/, models/, tests/ 2. **Safe execution**: All changes happen in sandbox, reviewed before apply 3. **Decision correction**: Can correct a decision and recompute only affected work 4. **Scale to large codebases**: ACMS with UKO, pluggable strategies, and fusion can work with 10K+ file projects efficiently 5. **Semantic safety**: Catches breaking changes before they're applied ### Migration Strategy The implementation can proceed incrementally: 1. Start with resource abstraction (enables everything else) 2. Add sandboxing to existing execute phase 3. Gradually replace text parsing with tool invocations 4. Build decision tree alongside existing flow 5. Enhance context as indexing comes online This allows the system to remain functional during development while progressively adding the architectural improvements described in the specification.