Files
cleveragents-core/specification.md
T

167 KiB
Raw Blame History

CleverAgents Documentation (Detailed Spec)

Source Material

Key themes include:

  • A four-phase plan lifecycle (Action → Strategize → Execute → Apply).
  • Actors as a unifying abstraction (an LLM/agent or a whole graph).
  • A sandbox + diff review workflow and CLI-first interaction model.
  • A scalable context/memory architecture (hot/warm/cold tiers, per-actor views).
  • A future-facing correction model where the user can "edit the decision tree" and only recompute affected subtrees.

Big Picture: What CleverAgents is

CleverAgents is your command center for AI agents—a unified platform for orchestrating any task you want agents to accomplish, from developing large software projects to writing comprehensive technical papers, administering databases, managing cloud infrastructure, or any complex multi-step workflow. The core value proposition is enabling long-running, complex, large-scale tasks to execute autonomously with minimal human intervention, making it ideal for building entire software systems, producing extensive documentation, or managing sophisticated operations largely hands-off.

In server mode, CleverAgents becomes a collaborative hub where teams can share resources—prompts, actors, actions, and projects—while executing plans in the cloud. This enables a consistent experience across all your devices: start a complex task on your laptop, check progress from your phone, and review results from any machine.

While CleverAgents leverages LangGraph and LangChain for the underlying LLM runtime primitives (tool calling, graphs, routing), its value lies in what it builds on top:

  • CleverAgents provides:

    • A first-class plan lifecycle (Action/Strategize/Execute/Apply) for breaking down and tracking complex work,
    • A project + resource model for grounding tasks in real codebases, databases, documents, and infrastructure,
    • A consistent actor abstraction for defining and composing intelligent agents,
    • A consistent skill abstraction for anything an agent can execute,
    • A sandbox + checkpoint safety model for safe, reversible execution,
    • A CLI/TUI/Web UX for controlling and monitoring large multi-step autonomous work.

Glossary (Terms Used Precisely)

  • Plan: A tracked lifecycle for a single unit-of-work (which may spawn subplans). Plans follow the same namespace rules as actors.
  • Action: A reusable plan template not tied to any project yet. Created via CLI commands (not YAML files). Actions follow the same namespace rules as actors.
  • Strategize: Read-only planning phase that produces a strategy and subplan blueprint. All decisions are made during this phase.
  • Execute: Phase that performs work in a sandbox; spawns subplans (based on decisions made in Strategize); produces artifacts/diffs.
  • Apply: Phase that commits sandbox results into the real project (and records an "applied" plan state).
  • Project: A collection of resources + configuration that define "where work happens" and "what can be touched." Created via CLI commands. Can be local (contains local-only and remote resources) or remote (all resources remotely accessible).
  • Resource: Anything that can be read/written/queried (files, repo trees, DB endpoints, cloud clusters, documents). Extends the MCP resource concept to support both read and write operations. Each resource defines its own sandbox strategy.
  • Skill: A callable capability defined inline in actor YAML configuration as tool nodes. Extends the MCP standard and Agent Skills standard. Skills follow the same naming scheme as plans, actions, etc.
  • Actor: Anything conversational; may be a single agent/LLM or an entire graph of actors/tools. Defined via YAML configuration files (LangGraph definitions). Always named using <namespace>/<name> format.
  • Session: A user interaction context and conversation thread that can span multiple plans.
  • Server: Optional shared service for multi-user storage, permissions, and orchestration. Plans on remote projects can execute on the server.
  • Namespace: Scoping mechanism for actors, actions, plans, skills, etc. local/ is reserved for local-only items. User namespaces (<username>/) and organization namespaces (<orgname>/) are stored on the server. Built-in LLM actors use provider namespaces (e.g., openai/, anthropic/).
  • Decision: A recorded choice point made during Strategize that affects downstream work. Decisions form a tree structure that enables correction and replay.
  • ULID: Universally Unique Lexicographically Sortable Identifier. Preferred over UUID for plan and decision IDs due to time-sortability.

Components

Plan

A plan is the fundamental unit of orchestration and traceability.

Plan Lifecycle Phases

A plan always moves through the following phases, in order:

Action → Strategize → Execute → Apply → Applied (terminal)

In this spec:

  • Strategize is the phase name (the output is a strategy).
  • Execute is the phase name (the output is a changeset).
  • Apply is the phase name (the output is an applied change).
  • Applied is the resulting terminal state after Apply succeeds.

This four-stage model is explicitly called out as the new architecture replacing a prior linear pipeline.

Phase Transition Verbs (CLI / UX Contract)

Verbs that trigger phase transitions. CleverAgents should standardize these verbs as the public API (CLI, TUI, web):

Current Phase Command Verb Next Phase
(none) create Action
Action use Strategize
Strategize execute Execute
Execute apply Applied

Important behavioral rule: CleverAgents must support multiple "automation levels" that can automatically progress through these verbs without the user explicitly issuing them, but the verbs remain the conceptual contract.

Plan States (Per Phase)

A plan's phase indicates "what step of the lifecycle it is in." Separately, the plan has a processing state indicating "what is happening right now."

Recommended state model:

  • Action phase states

    • available (action exists and can be used)
    • draft (action is being authored/edited)
    • archived (soft-deleted or hidden, optional)
  • Strategize / Execute / Apply phase states

    • queued (waiting for compute/worker)
    • processing (currently running)
    • errored (failed; includes error metadata)
    • complete (finished successfully)
    • cancelled (user/system cancelled; safe terminal for that phase)

Plan Identity and Traceability

Every plan should have:

  • plan_id: Unique, immutable ID (UUID or ULID).
  • parent_plan_id: Nullable; present for subplans.
  • root_plan_id: The top-most plan in the tree.
  • attempt: An integer attempt counter that increments when re-running a phase (e.g., re-executing after a fix).
  • created_at / updated_at / completed_at timestamps.
  • created_by (user identity / session identity).

Plan Hierarchy (Subplans) and Parallelism

A single plan should usually represent the smallest "complete" unit of work (similar to what would fit in one git commit). However:

  • Plans are hierarchical.
  • Decisions about subplans are made during Strategize (as subplan_spawn decision types).
  • Subplans are actually spawned during Execute (based on those decisions).
  • Subplans can run in parallel or sequentially.
  • The parent plan is responsible for merging results.

This is core to the long-term objective: tackling large tasks while only recomputing parts of the decision tree when corrected.

Hierarchical Decomposition for Scale

When handling massive tasks (e.g., converting Firefox to Rust), the system uses hierarchical decomposition:

  1. Root Level: High-level architectural decisions

    • "Convert Firefox Renderer to Rust"
    • Decision: "Start with leaf modules, work inward"
    • Context: Module dependency graph (2,847 modules)
  2. Subsystem Level: Major component decisions

    • "Phase 1: Convert utility libraries (no external deps)"
    • Each subsystem gets its own bounded context
  3. Module Level: Individual module conversions

    • "Convert string_utils module"
    • Context: Only the 47 functions and 12 dependent files
    • Decision: "Use Rust's String type"
  4. File Level: Specific file changes

    • Actual code transformations
    • Minimal context needed

At each level, only the relevant context is loaded. The persistent decision graph means we can always reconstruct why we're converting a particular module and what constraints apply from higher-level decisions.

Subplan Spawning Mechanism

In the actor definition for the execution actor, there are nodes that act as skills whose purpose is to generate subplans. The execution actor can call these skills to trigger subplans.

# Example: Execution actor with subplan spawning capability
actors:
  code_executor:
    type: graph
    routes:
      execute_workflow:
        nodes:
          - name: spawn_test_subplan
            type: tool
            config:
              tools:
                - name: create_subplan
                  code: |
                    # This skill creates a subplan
                    subplan = context.spawn_subplan(
                        action="local/write-tests",
                        target_files=input_data.files_to_test
                    )
                    result = subplan.id

Subplan Execution Modes

  • Sequential: Subplans execute one after another. If one fails, subsequent subplans are not started.
  • Parallel: Subplans execute concurrently. If one fails, others can continue.

Subplan Failure Handling

  • Parallel execution: Other parallel subplans continue even if one fails.
  • Sequential execution: Subsequent subplans are not called if a prior one fails.
  • Note: An "error" only occurs if an exception is thrown by the application (a bug). Plan failures (e.g., tests don't pass) are handled within the plan's logic, not as application errors.

Result Merging

The way results are merged depends on the resource type:

  • Git-compatible resources (source code, text files): Git-style merge
  • Databases: Transaction coordination or sequential application
  • Other resources: Pluggable merge strategies based on resource type
  • Non-mergeable resources: May require sequential execution only

The Plan "Decision Tree" and Visualization

CleverAgents intends to record enough information to render:

  • an ASCII tree in the TUI, and
  • optionally a GUI tree via visualization tools (D3/Cytoscape) once the data exists.

This implies each plan should persist:

  • decisions made,
  • the rationale (or at least the prompt/context snapshot that produced it),
  • dependencies ("this decision influenced these child plans").

This is required for "correcting plans" (see Behavior section).

Decision Data Model

Relationship Between Plan Description and Decisions

Each plan has a description field (inherited from the action's description, potentially with argument substitutions). This description acts as the primary component of the prompt fed to the strategy actor during the Strategize phase.

Decisions are choices that are NOT explicitly defined by the plan description. They represent the gaps, ambiguities, or implementation details that must be resolved to execute the plan.

For example:

  • Plan description: "Increase test coverage to 85%"
  • Decisions that emerge:
    • "Which modules should be prioritized?" (not specified in description)
    • "Should we use mocks or integration tests for the database layer?" (not specified)
    • "Should we refactor the auth module to make it more testable, or write tests around it as-is?" (not specified)

Decision Making Based on Autonomy Level

Who makes decisions depends on the plan's automation level:

Automation Level Who Makes Decisions
Manual User is prompted for each decision point
Review-before-apply Actor makes decisions automatically during Strategize, user reviews before Apply
Full automation Actor makes all decisions automatically

When automation allows automatic decisions, the strategy actor uses its best judgment based on context, and records its reasoning in the decision's rationale field.

When user input is required, the system pauses and prompts the user:

Decision required: Which modules should be prioritized for test coverage?

Options identified by the strategy actor:
  1. auth module (currently 45% coverage, high risk)
  2. payment module (currently 52% coverage, high risk)  
  3. user module (currently 71% coverage, medium risk)

Your choice (or provide custom guidance): _

The Prompt as the Root Decision

The prompt passed to the strategize actor is itself a decision node in the decision tree—specifically, it's the root decision of type prompt_definition.

This is important because:

  1. Every plan has its own prompt: The root plan's prompt comes from the action description + user arguments. Subplan prompts are created by parent plans during their execution.

  2. Parent plans create subplan prompts: When a parent plan spawns a subplan, it decides what prompt to give that subplan. This is recorded as a prompt_definition decision in the parent's tree, and becomes the root decision of the child plan's tree.

  3. Unified correction mechanism: Since the prompt is just another decision, correcting it uses the same agents plan correct command as any other decision.

Plan Tree Example:
├── [prompt_definition] "Increase test coverage to 85%"          <- Root decision (correctable)
│   ├── [strategy_choice] "Prioritize auth and payment modules"
│   ├── [subplan_spawn] "Write tests for auth module"
│   │   └── Subplan: auth-tests
│   │       ├── [prompt_definition] "Write unit tests for auth module using mocks"  <- Created by parent
│   │       ├── [implementation_choice] "Test login flow first"
│   │       └── ...
│   └── [subplan_spawn] "Write tests for payment module"
│       └── Subplan: payment-tests
│           ├── [prompt_definition] "Write unit tests for payment module"  <- Created by parent
│           └── ...

Correcting Decisions (Including Prompts)

All corrections use the same unified command:

agents plan correct <decision_id> --mode=<mode> --guidance "<corrected decision text>"

Parameters:

  • <decision_id>: The ULID of the decision to correct
  • --mode: Either revert (rollback and re-run) or append (add fix at end)
  • --guidance: Free-form text specifying what the correct decision should be

Examples:

# Correct a strategy choice
agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=revert \
  --guidance "Prioritize the payment module first, not auth, due to upcoming deadline"

# Correct the root prompt to be more specific
agents plan tree <plan_id>
# Shows: [prompt_definition] id=01ARZ3NDEKTSV4RRFFQ69G5FAV "Increase test coverage to 85%"

agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=revert \
  --guidance "Increase test coverage to 85%, prioritizing auth and payment modules. Use mocks for database tests, not integration tests."

# Correct a subplan's prompt (originally created by parent plan)
agents plan correct 01BRZ4PDFLUTW5SSGR70H6GBW --mode=revert \
  --guidance "Write unit tests for auth module, focusing on edge cases for token expiration"

# Append a fix rather than rewriting history
agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=append \
  --guidance "The previous approach missed error handling tests - add comprehensive error path coverage"

Note: CLI commands should not require interactive input. The --guidance parameter provides the correction inline. For very long guidance text, use a file:

agents plan correct <decision_id> --mode=revert --guidance-file ./correction.txt

When to correct the prompt vs. a specific decision:

Situation Correction Approach
Original request was too vague Correct the prompt_definition decision
Strategy actor made a bad choice on a specific question Correct that specific decision
Parent plan gave a subplan a bad prompt Correct the subplan's prompt_definition

Because the prompt is part of the decision tree, the system automatically knows that correcting it invalidates all downstream decisions in that plan (and its subplans).

Decisions are only created during the Strategize phase. The decision tree captures what choices were made and why, enabling correction and replay.

Decision Record Structure

Decision:
  # Identity
  decision_id: ULID                    # Unique identifier
  plan_id: ULID                        # Parent plan this decision belongs to
  parent_decision_id: ULID | null      # Parent decision (for tree structure)
  sequence_number: int                 # Order within the plan's decisions
  
  # Classification
  decision_type: enum
    - prompt_definition      # The prompt/description for this plan (root decision)
    - strategy_choice        # High-level approach decision during Strategize
    - implementation_choice  # How to implement a specific task
    - resource_selection     # Which resources to read/modify
    - subplan_spawn          # Decision to create a subplan (spawned later in Execute)
    - tool_invocation        # Which skill/tool to use
    - error_recovery         # How to handle a failure
    - validation_response    # Response to validation failure
    - user_intervention      # User provided guidance/correction
  
  # The Decision Itself
  question: str                        # What question was being answered
  chosen_option: str                   # What was decided
  alternatives_considered: list[str]   # Other options that were evaluated
  confidence_score: float | null       # 0.0-1.0 if the actor provided confidence
  
  # Context Snapshot (for replay)
  context_snapshot:
    hot_context_hash: str              # Cryptographic hash of the exact context
    hot_context_ref: str               # Pointer to the full stored snapshot
    relevant_resources: list[ResourceRef]  # Every file/symbol that influenced this decision
    actor_state_ref: str               # Complete LangGraph checkpoint
    
  # When the system decides "refactor the authentication module to use async patterns," 
  # it permanently records:
  # - Which files were examined to make that decision
  # - What symbols and dependencies were traced
  # - The exact code state that was analyzed
  # - The reasoning chain that led to this choice
  # - Alternative approaches that were considered but rejected
  
  # Rationale
  rationale: str                       # Why this option was chosen
  actor_reasoning: str | null          # Raw LLM reasoning if available
  
  # Downstream Impact (populated during Execute phase)
  downstream_decision_ids: list[ULID]  # Decisions that depend on this one
  downstream_plan_ids: list[ULID]      # Subplans spawned because of this decision
  artifacts_produced: list[ArtifactRef] # Files/outputs created under this decision
  
  # Timestamps
  created_at: datetime
  
  # Correction Metadata
  is_correction: bool                  # Was this decision a correction of another?
  corrects_decision_id: ULID | null    # If correction, which decision was replaced
  correction_reason: str | null        # Why the correction was made
  superseded_by: ULID | null           # If this decision was later corrected

Decision Timing

Phase Decision Activity
Strategize Decisions are created. downstream_plan_ids is empty.
Execute Subplans are spawned. downstream_plan_ids is populated when subplans are created based on subplan_spawn decisions.
Apply No new decisions. History can be flagged for cleanup after successful apply.

Decision Tree Storage Schema

-- Core decision table
CREATE TABLE decisions (
    decision_id TEXT PRIMARY KEY,      -- ULID
    plan_id TEXT NOT NULL,
    parent_decision_id TEXT,
    sequence_number INTEGER NOT NULL,
    decision_type TEXT NOT NULL,
    question TEXT,
    chosen_option TEXT NOT NULL,
    alternatives_considered TEXT,      -- JSON array
    confidence_score REAL,
    rationale TEXT,
    actor_reasoning TEXT,
    context_snapshot TEXT NOT NULL,    -- JSON blob
    is_correction BOOLEAN DEFAULT FALSE,
    corrects_decision_id TEXT,
    correction_reason TEXT,
    superseded_by TEXT,
    created_at TEXT NOT NULL,
    
    FOREIGN KEY (plan_id) REFERENCES plans(plan_id),
    FOREIGN KEY (parent_decision_id) REFERENCES decisions(decision_id),
    FOREIGN KEY (corrects_decision_id) REFERENCES decisions(decision_id),
    FOREIGN KEY (superseded_by) REFERENCES decisions(decision_id)
);

-- Downstream relationships (many-to-many for DAG)
CREATE TABLE decision_dependencies (
    upstream_decision_id TEXT NOT NULL,
    downstream_decision_id TEXT NOT NULL,
    dependency_type TEXT NOT NULL,     -- 'decision', 'plan', 'artifact'
    downstream_ref TEXT NOT NULL,      -- The actual ID of decision/plan/artifact
    
    PRIMARY KEY (upstream_decision_id, downstream_decision_id, downstream_ref),
    FOREIGN KEY (upstream_decision_id) REFERENCES decisions(decision_id)
);

-- Correction history
CREATE TABLE correction_attempts (
    attempt_id TEXT PRIMARY KEY,       -- ULID
    plan_id TEXT NOT NULL,
    original_decision_id TEXT NOT NULL,
    new_decision_id TEXT,
    original_subtree_snapshot TEXT,    -- Reference to archived state
    correction_reason TEXT,
    status TEXT NOT NULL,              -- 'pending', 'executing', 'completed', 'failed'
    created_at TEXT NOT NULL,
    completed_at TEXT,
    
    FOREIGN KEY (plan_id) REFERENCES plans(plan_id),
    FOREIGN KEY (original_decision_id) REFERENCES decisions(decision_id),
    FOREIGN KEY (new_decision_id) REFERENCES decisions(decision_id)
);

Action

What an Action Is

An action is a reusable plan template that is not associated with any projects yet.

Important: Actions are created via CLI commands, NOT YAML configuration files. YAML configuration files are only used to define actors.

Examples:

  • "Increase test coverage to 80%"
  • "Refactor module X to be async-safe"
  • "Write an RFC for feature Y"
  • "Provision an infra cluster and validate access" (non-code)

Actions are intentionally project-agnostic so they can be reused across projects.

Action Creation (CLI)

Actions are created using the CLI:

agents action create \
  --name "local/code-coverage" \
  --description "Increase test coverage to target percentage" \
  --strategy-actor "local/coverage-strategist" \
  --execution-actor "local/coverage-executor" \
  --definition-of-done "Coverage reaches target percentage; All new tests pass" \
  --arg "target_coverage_percent:int:required:Target coverage percentage (0-100)" \
  --arg "test_framework:str:optional:Test framework to use (pytest, unittest, etc.)"

Required parameters:

  • --name: Namespaced name (e.g., local/code-coverage, myorg/deploy-action)
  • --strategy-actor: Name of the actor to use for Strategize phase
  • --execution-actor: Name of the actor to use for Execute phase
  • --definition-of-done: Free-form text describing completion criteria

Optional parameters:

  • --description: Human-readable description
  • --arg: Argument definitions (can be repeated). Format: name:type:required|optional:description
  • --reusable: Whether action remains available after use (default: true)
  • --read-only: Whether action only performs read operations (default: false)

Arguments defined with --arg are values that will be:

  • Injected into the description and/or definition of done (via templating)
  • Passed into the context of the actors
  • Required when using the action on projects

Action Data Model (Expanded)

A plan in the Action phase has:

1) name (namespaced)

Format:

  • [server:][namespace/]<name>

Rules:

  • If server is omitted, default server is assumed unless namespace is local.
  • If namespace is omitted, default is local.
  • Names should be stable identifiers (kebab-case recommended).

Examples:

  • local/code-coverage
  • myusername/code-coverage
  • myorgname/code-coverage
  • prod:myorgname/code-coverage (server-qualified)

2) short_description

Optional at creation; auto-filled if blank.

3) long_description

Optional but recommended for reusable actions.

4) definition_of_done (DoD)

Required. Must be explicit and testable.

5) actors

Two actors minimum:

  • strategy_actor (planner/architect)
  • execution_actor (builder/implementer)

You can optionally add:

  • review_actor (code review / QA)
  • apply_actor (release/merge specialist) …but these can also be handled as subplans.

Actors can be:

  • an LLM agent,
  • a graph,
  • or another actor reference.

Actor abstraction is central: an actor may be a single agent or an entire graph.

6) reusable (boolean)

  • Default: true.
  • If true: using the action creates a new plan in Strategize while leaving the action available.
  • If false: action self-deletes (or auto-archives) after first use.

7) read_only (boolean)

  • Default: false.
  • If true: the plan must only use read-only skills and must never modify resources (even in sandbox).
  • Read-only actions are still useful for "investigation reports," architecture reviews, or dry-run planning.

To make actions genuinely reusable, actions should declare their inputs:

  • required args (e.g., target coverage percent),
  • optional args (e.g., test framework),
  • validation rules (types, bounds).

Example:

  • target_coverage_percent: integer 0100

A policy bundle that can be applied to enforce safe execution:

  • allowed skill categories,
  • require checkpoints,
  • require sandbox,
  • require human approval at Apply.

This relates to the "checkpointable skills + sandbox" approach for safe writing.

Strategy (Strategize Phase)

Using an Action (Transition to Strategize)

The use command transitions an Action into the Strategize phase by applying it to one or more projects:

# Basic usage
agents plan use local/code-coverage --project my-api-service

# Multiple projects
agents plan use local/schema-update \
  --project api-service \
  --project web-frontend \
  --project mobile-app

# With action arguments
agents plan use local/code-coverage \
  --project my-api-service \
  --arg target_coverage_percent=85 \
  --arg test_framework=pytest

# With explicit automation level
agents plan use local/deploy-action \
  --project staging-env \
  --automation-level manual

Parameters:

  • --project: Project to apply the action to (can be repeated for multi-project plans)
  • --arg: Action argument values (format: name=value)
  • --automation-level: Override automation level for this plan

When the action is used:

  1. A new plan is created with a unique ULID
  2. The plan's automation level is determined (plan > session > global)
  3. The plan enters the Strategize phase
  4. The strategy_actor begins analyzing the project(s)

What Strategize Does

When an action is used on projects, it becomes a plan in Strategize.

Strategize is:

  • read-only, producing a plan of attack,
  • responsible for gathering context from project resources,
  • responsible for generating a strategy and subplan blueprint,
  • not allowed to execute subplans or modify resources.

This "architect vs coder" separation is explicitly described as a core motivation.

Resource-aware dependency analysis: During the Strategize phase, the strategy actor employs specialized mechanisms to compute precise dependency closures:

# Pseudocode of what happens inside a strategy actor
def compute_closure_for_refactoring(target_module):
    closure = ResourceClosure()
    
    # Direct file dependencies
    closure.add_files(find_imports(target_module))
    closure.add_files(find_includes(target_module))
    
    # Symbol dependencies
    for symbol in extract_exported_symbols(target_module):
        closure.add_files(find_symbol_usage(symbol, scope='project'))
    
    # Test dependencies
    closure.add_files(find_tests_for_module(target_module))
    
    # Build system dependencies
    closure.add_files(find_build_references(target_module))
    
    return closure

The system leverages several key insights:

  • Modular boundaries exist: Even in legacy codebases, there are natural boundaries
  • Changes are incremental: We don't convert 50,000 files atomically
  • Dependencies are sparse: Most modules depend on a small fraction of the codebase
  • Interfaces are narrow: Public APIs are much smaller than implementations

Strategize Data Model

A plan in Strategize contains all Action fields plus:

1) projects

A list of projects the plan is used on.

Important: A strategy plan may target multiple projects. Multi-project work in one "window" is considered a major usability advantage over tools that require being run from a single directory.

2) strategy_context

A structured object describing:

  • what resources were considered,
  • how they were retrieved,
  • what filtering/limits were applied,
  • what the actor saw.

This matters because a plan must be debuggable and correctable later.

Recommended fields:

  • resource_refs: IDs of resources used
  • queries: search queries performed
  • selected_chunks: chunk IDs + sources + reasons
  • constraints: context window limits, file ignore patterns
  • generated_summaries: if summarization occurred

3) strategy

The output plan:

  • steps (ordered and/or DAG),
  • conditions/branches ("if tests fail, do X"),
  • subplans to spawn (including which action templates to use),
  • evaluation criteria (how to know success),
  • risk assessment.

Strategize should output not only narrative text but also a machine-usable blueprint:

  • list of tasks,
  • required skills,
  • expected outputs,
  • dependencies between tasks.

This blueprint becomes the input to Execute.

5) cost_estimate and risk_estimate (optional)

Cost and risk estimation is optional but recommended for production use.

When enabled, a specialized estimation actor analyzes:

  • The initial prompt/request
  • The strategy produced by the Strategize phase
  • Historical data from similar plans (if available)

And produces estimates for:

  • LLM tokens/cost range
  • Number of steps/subplans expected
  • Expected risk of rollbacks
  • Estimated execution time

Implementation: Similar to how there's a strategy_actor and execution_actor for each action, there can be an optional estimation_actor whose entire job is cost/risk estimation. This actor runs after Strategize completes (before Execute) and its output is informational only.

# Example: Action with estimation actor
agents action create \
  --name "local/expensive-refactor" \
  --strategy-actor "local/refactor-planner" \
  --execution-actor "local/refactor-executor" \
  --estimation-actor "local/cost-estimator" \
  --definition-of-done "All code refactored according to plan"

This becomes critical in server/multi-user usage and cost controls.

Execution (Execute Phase)

What Execute Does

Execute is where the plan actually performs work, but in a sandboxed environment that can later be reviewed and applied.

Key properties:

  1. Work happens in a sandbox All file modifications, generated artifacts, and intermediate outputs live in an isolated "execution workspace" until Apply.

  2. Execute may spawn subplans Subplans are a first-class behavior of Execute: a parent plan can distribute work to child plans and merge results.

  3. Execute must support checkpointing / rollback (when enabled) Checkpointable skills allow rolling back to a checkpoint ID to recover from partial failure or wrong turns.

  4. Execute produces a "reviewable diff" Diff review sandbox is described as a differentiating feature: users can inspect changes before applying.

Execution Workspace / Sandbox Model (Detailed)

A sandbox isolates plan execution from the real project resources until Apply.

Key Sandbox Principles

  1. Lazy Sandboxing: Resources are sandboxed only when accessed, not upfront.

    • A project may have many resources (git repo + 10 databases + cloud accounts)
    • A plan may only modify one resource
    • Only accessed resources are sandboxed
    • Efficient for large projects
  2. Per-Plan Sandboxes: Each plan and subplan has its own sandbox containing only the resources it edits.

  3. Resource-Defined Strategy: The sandbox strategy is defined on each resource, not by skills or globally.

  4. Cleanup Behavior:

    • Sandboxes are cleaned up before application exit when possible
    • Abandoned sandboxes (from crashes, etc.) are cleaned up on next application run
    • Completed plan sandboxes are cleaned or archived based on retention policy

Sandbox Implementation Strategies

Different resource types require different sandbox strategies:

Resource Type Strategy Rollback Mechanism
Git Repository git_worktree Git reset/checkout
Filesystem copy_on_write or overlay Restore from snapshot
Database transaction_rollback Transaction rollback
Cloud Infra terraform_state Terraform plan (reversed)
API Endpoint none Often not sandboxable

1. Git worktree / branch sandbox (preferred for code)

  • Create a worktree or temporary branch
  • All modifications are commits or staged changes
  • Apply merges/cherry-picks

Pros: natural rollback, diff support, efficient Cons: requires git

2. Filesystem copy sandbox

  • Copy project directory to a sandbox directory
  • Execute modifies sandbox copy
  • Apply syncs diff back

Pros: simple Cons: expensive for huge repos

3. Overlay filesystem sandbox

  • Use overlayfs-style "copy-on-write" to avoid full copies

Pros: efficient Cons: more complex, OS-dependent

4. Transaction-based sandbox (for databases)

  • Begin transaction at sandbox creation
  • All operations within transaction
  • Rollback on failure, commit on apply

Pros: native to databases Cons: long-running transactions can cause issues

5. No sandbox (for non-sandboxable resources)

  • Some resources cannot be sandboxed (certain APIs, cloud services)
  • User proceeds at their own risk
  • Plan should warn about non-sandboxable resources

Multi-Resource Sandboxing

When a plan accesses multiple resources:

  • Each resource gets its own sandbox (based on its defined strategy)
  • Sandboxes are independent
  • Apply commits each sandbox separately
  • If any sandbox Apply fails, others may still succeed (partial apply)

Complete isolation during execution prevents compound errors: Each plan executes in its own sandbox, which means:

Plan A (refactoring auth module):
- Sandbox A1: Contains only auth/*.cpp, auth_tests/*.cpp
- Cannot see Plan B's intermediate states
- Cannot accidentally depend on Plan B's half-done work

Plan B (updating API endpoints):
- Sandbox B1: Contains only api/*.cpp, api_tests/*.cpp  
- Makes changes assuming current auth interface
- Protected from Plan A's intermediate refactoring

Hierarchical merge resolution: When subplans complete, the parent plan performs intelligent merging:

def merge_subplan_results(subplan_results):
    # Group by resource type
    by_resource = group_by_resource_type(subplan_results)
    
    # Apply resource-specific merge strategies
    for resource_type, changes in by_resource:
        if resource_type == 'git_repo':
            merge_git_changes(changes)  # Three-way merge
        elif resource_type == 'database':
            merge_db_changes(changes)   # Sequential application
        elif resource_type == 'config_files':
            merge_config_changes(changes)  # Smart JSON/YAML merge
    
    # Validate merged state
    run_integration_tests()

Execution Data Model

A plan in Execute contains:

1) execution_context

The context used for execution (often smaller/more tactical than strategy context).

2) execution_log

Structured timeline of:

  • skill calls,
  • actor calls,
  • outputs,
  • errors and retries,
  • checkpoints created.

This log is essential for debugging.

3) artifacts

Outputs produced:

  • changed files,
  • generated files,
  • reports,
  • diagrams,
  • test outputs,
  • diffs.

4) sandbox_ref

Pointer to the sandbox location/state:

  • path, branch name, workspace ID, container ID, etc.

5) checkpoint_graph (if enabled)

A record of checkpoints:

  • checkpoint ID
  • timestamp
  • skill responsible
  • resources affected
  • rollback instructions / metadata

Checkpointing in Execute (Core Safety Mechanism)

The meeting notes repeatedly emphasize the need for checkpoints tied to skills and rollbacks. The intended user-level behavior is:

  • "Give me a checkpoint ID."
  • Perform additional operations.
  • "Roll back to checkpoint X."

Not all skills can support this; checkpointing must be declared per skill.

Skill-level checkpointability

Each skill declares:

  • checkpointable: true|false
  • rollback_mechanism: how rollback occurs
  • scope: what resources it can revert

Examples:

  • File skill: snapshot file states pre-modification
  • Git skill: create commit or stash; rollback is reset/checkout
  • CLI skill inside a container: rollback by restoring filesystem snapshot or reloading base image state

There's explicit discussion that checkpointing is easier when skill scope is constrained (e.g., "only files within a docker image + git").

Plan-level rollback policy

Plans should have an option:

  • rollback_enabled: true|false

If disabled, the plan may use more generic/unsafe skills with fewer restrictions (useful for low-stakes tasks).

Execution should be treated like a transactional pipeline:

  • Each step either:

    • commits a checkpoint on success, or
    • rolls back to the previous checkpoint on failure.

This is explicitly motivated by "partial failure leaves codebase inconsistent" and the need for transaction rollback.

Tool-Based Resource Modification (Modern Architecture)

IMPORTANT: CleverAgents does NOT parse LLM output to extract code. Instead, it uses the modern tool-based approach pioneered by Claude Code, Cursor, and Aider where:

  1. LLMs call tools/skills directly (edit_file(), write_file(), delete_file(), etc.)
  2. Tools operate on the sandbox - each tool invocation modifies sandbox state directly
  3. ChangeSet is built from tool invocations - not by parsing LLM text output
  4. Validation runs on sandbox state - after tools execute, not on parsed output

This architecture provides:

  • Atomic operations: Each tool call is a discrete, trackable change
  • No parsing ambiguity: Tools have structured parameters (path, content, etc.)
  • Resource-agnostic: Same pattern works for files, databases, APIs, any resource type
  • Safety by design: Tools run in sandbox with defined capabilities and restrictions
  • MCP compatibility: Skills map directly to MCP tools for external integrations

How It Works

LLM Response (with tool calls)
    ↓
┌─────────────────────────────────────┐
│ Skill/Tool Router                   │
│ - Routes each tool call to handler  │
│ - Validates parameters              │
│ - Enforces capability restrictions  │
└─────────────────────────────────────┘
    ↓
┌─────────────────────────────────────┐
│ Sandbox Execution                   │
│ - Tool operates on sandboxed state  │
│ - Each invocation recorded          │
│ - Checkpoint created if needed      │
└─────────────────────────────────────┘
    ↓
┌─────────────────────────────────────┐
│ ChangeSet Accumulation              │
│ - Each resource-modifying call →    │
│   becomes a Change record           │
│ - ChangeSet = history of changes    │
└─────────────────────────────────────┘
    ↓
┌─────────────────────────────────────┐
│ Validation & Review                 │
│ - Run validators on sandbox state   │
│ - Generate diff from ChangeSet      │
│ - Present for review before Apply   │
└─────────────────────────────────────┘

Built-in Resource Skills

CleverAgents provides these core skills for resource manipulation:

Skill Description Creates Change?
read_file(path) Read file contents No
write_file(path, content) Create/overwrite file Yes
edit_file(path, changes) Apply targeted edits Yes
delete_file(path) Remove file Yes
move_file(src, dst) Rename/move file Yes
create_directory(path) Create directory Yes
list_files(pattern) List files matching glob No
search_files(pattern, content) Search file contents No
get_file_info(path) Get file metadata No

Each skill automatically:

  • Operates within sandbox boundaries
  • Records changes to the ChangeSet
  • Validates parameters against project configuration
  • Enforces deny-list patterns (.git/, node_modules/, etc.)

Why Not Parse LLM Output?

The obsolete approach of parsing markdown code fences has fundamental problems:

  1. Ambiguity: Is text explanation or code? Where does one file end and another begin?
  2. Fragility: Models output varying formats; regex parsing is brittle
  3. Loss of semantics: You lose the intent (create vs modify vs delete)
  4. No atomicity: Can't rollback individual operations
  5. Resource-limited: Only works for files, not databases or other resources

The tool-based approach solves all of these by making each operation explicit, typed, and trackable.

Semantic Error Prevention

CleverAgents provides multiple layers of proactive error prevention that catch semantic errors before they can propagate through the system.

Layer 1: Decision-time Validation During Strategize

Every decision includes semantic validation:

Decision: Refactor payment module to async
alternatives_considered:
  - "Convert to async/await patterns" (chosen)
  - "Use thread pool with channels" (rejected: doesn't integrate with async ecosystem)
  - "Keep synchronous with timeout" (rejected: doesn't solve core latency issue)
confidence_score: 0.85
validation_performed:
  - Checked all payment API consumers can handle async
  - Verified database driver supports async operations
  - Confirmed no regulatory requirement for sync processing

Layer 2: Execution-time Semantic Guards

The execution actor configuration includes validation nodes that understand semantics:

actors:
  code_executor:
    type: graph
    nodes:
      - name: semantic_validator
        type: tool
        config:
          tools:
            - name: validate_api_compatibility
              code: |
                # Not just syntax checking - semantic validation
                old_api = extract_api_signature(previous_version)
                new_api = extract_api_signature(current_version)
                
                breaking_changes = find_breaking_changes(old_api, new_api)
                if breaking_changes:
                    # Don't just fail - understand the impact
                    affected_consumers = find_api_consumers(breaking_changes)
                    migration_plan = generate_migration(breaking_changes)
                    
                    if can_auto_migrate(affected_consumers, migration_plan):
                        apply_migration(migration_plan)
                    else:
                        raise SemanticError(
                            "Breaking API changes require manual review",
                            changes=breaking_changes,
                            affected=affected_consumers
                        )

Layer 3: Invariant Enforcement

The system maintains semantic invariants:

class RefactoringInvariants:
    # User-defined invariants for the codebase
    invariants = [
        "All public APIs must maintain backward compatibility",
        "Database transactions must complete within 5 seconds",
        "Authentication must always use OAuth2",
        "Payment processing must be idempotent"
    ]
    
    def check_invariant_preservation(self, changes):
        for invariant in self.invariants:
            if not self.verify_invariant(invariant, changes):
                return InvariantViolation(invariant, changes)
        return Success()

Layer 4: Predictive Error Prevention

The system learns from past failures:

Error Pattern Database:
  - pattern: "Async conversion in payment module"
    historical_failures:
      - "Race condition in payment confirmation"
      - "Timeout handling breaks idempotency"
    preventive_checks:
      - "Add explicit transaction boundaries"
      - "Verify idempotency keys are preserved"
      - "Check distributed lock acquisition"

Applied (Apply Phase)

What Apply Does

Apply takes the sandboxed work product and makes it "real" in the project.

Core properties:

  1. Apply is a controlled commit step Apply exists specifically to separate "generated work" from "committed work," enabling review and safer automation.

  2. Apply is often the highest-risk step It changes real systems. This is where permissions, approvals, and checks matter most.

  3. Apply produces a terminal 'applied' plan After successful apply, the plan becomes Applied.

Apply should perform (configurable) validations before committing:

  • Diff review gate

    • If automation level requires review, show:

      • changed files summary,
      • full diff,
      • risk warnings.
  • Pre-apply tests

    • Run validation as defined by the actor and/or project (see below).
  • Conflict resolution

    • If applying to a git repo, handle rebase/merge conflicts safely.
  • Audit log

    • Record who applied, what changed, when, and why.

Validation Configuration

Validation is defined by the actor configuration and project settings, not hardcoded.

Actor-defined validation: The execution actor's workflow should include validation nodes that:

  1. Read project-defined test/validation commands from context
  2. Execute appropriate tests based on resource type
  3. Handle failures by iterating (attempting fixes) or escalating to user

Project-defined validation: Each project should define how its resources are validated:

# Example: Add validation configuration to a project
agents project set-validation \
  --project my-api-service \
  --resource api-repo \
  --test-command "pytest" \
  --lint-command "ruff check ." \
  --type-check-command "pyright"

This information is passed into the actor's context, allowing generic actors (not specific to any project) to execute appropriate validation.

Validation Failure Handling

When validation fails:

  1. Iteration: Actor attempts to fix the issue (e.g., fix failing tests)
  2. Retry limit: After several failed attempts, iteration stops
  3. User intervention: User can provide additional instructions
  4. Resume: Plan continues with new guidance

The user can prompt the plan with additional instructions when stuck:

agents plan prompt <plan_id> "Try using mock objects for the database tests"

Apply Data Model

A plan in Apply includes:

  • apply_summary
  • applied_artifacts (final commit hash, merged PR link, file list)
  • final_validation_results (test outputs, lint outputs)
  • approval_record (if human approvals are required)
  • deployment_record (optional, if apply triggers deploy)

"Applied" Terminal State

When Apply succeeds:

  • plan.phase = applied
  • plan.state = complete
  • the sandbox may be cleaned up or archived depending on retention policy

When Apply fails:

  • plan.phase remains Apply
  • plan.state = errored
  • sandbox remains intact for inspection/retry

Project

A project is the boundary that answers:

  • "Where is the work happening?"
  • "What can this plan read and write?"
  • "What tools can this plan use?"
  • "What context is available?"

A project is a collection of resources and configuration.

Important: Projects are created via CLI commands, NOT YAML configuration files.

Project Types: Local vs Remote

Projects are classified based on their resources:

Type Definition Where Plans Can Execute
Local Contains at least one local-only resource Client only
Remote All resources are remotely accessible Client or Server

This distinction matters for server mode: the server can only execute plans on remote projects because it needs network access to all resources.

Project Creation (CLI)

# Create a new project
agents project create \
  --name "my-api-service" \
  --tag "python" \
  --tag "backend"

# Add resources to the project
agents project add-resource \
  --project "my-api-service" \
  --name "api-repo" \
  --type "git_repository" \
  --location "git@github.com:org/api-service.git" \
  --sandbox-strategy "git_worktree"

agents project add-resource \
  --project "my-api-service" \
  --name "staging-db" \
  --type "database" \
  --location "postgresql://staging.example.com/mydb" \
  --sandbox-strategy "transaction_rollback" \
  --read-only

Project Data Model

A project includes:

1) Identity

  • project_id (ULID)
  • name
  • namespace (follows same rules as actors: local/, <username>/, <orgname>/)
  • tags (e.g., "python", "infra", "paper", "prod")
  • is_remote (boolean, derived from resources)

2) Resources

Resources are the "things you can act on." Each resource defines its own sandbox strategy.

Resource types include:

  • filesystem root(s)
  • git repository
  • database endpoints
  • cloud accounts
  • document corpora (papers, PDFs)
  • API schemas

Each resource has:

  • resource_id (ULID)
  • name
  • type (e.g., git_repository, database, filesystem, api_endpoint)
  • location (path/url/connection string)
  • is_remote (boolean - can it be accessed over network?)
  • sandbox_strategy (defined on the resource - see Sandboxing section)
  • read_only (boolean - whether writes are allowed)
  • metadata (language, repo size, etc.)

3) Context configuration

Project-level defaults:

  • ignore patterns (like .gitignore semantics)
  • max file size
  • indexing strategy
  • preferred chunking/summarization policy (even if evolving)
  • context retention policy

4) Security / permissions defaults

  • who can run write plans
  • which skills are restricted
  • whether apply requires approvals

Multi-Project Operations

A single plan may target multiple projects (e.g., updating shared schemas across services). This is considered a key UX advantage over "run in one directory" systems.

In multi-project execution:

  • Strategize must clarify which steps affect which projects.
  • Execution must isolate sandboxes per project OR define a composite sandbox.
  • Apply must commit changes to each project separately, with separate approval records if necessary.

Namespaces

Namespaces define ownership, scoping, and discoverability of actions, projects, actors, and plans.

All named entities use the format <namespace>/<name>.

Namespace Types

Namespace Scope Storage Examples
local/ Current machine only Local database local/my-reviewer, local/test-action
<username>/ Personal server namespace Server database freemo/code-analyzer, jsmith/deploy-script
<orgname>/ Organization namespace Server database cleverthis/standard-review, acme/deploy-action
openai/, anthropic/, etc. Built-in LLM actors N/A (built-in) openai/gpt-4, anthropic/claude-3-opus

Namespace Rules

  • local/

    • Reserved namespace for local-only items
    • Exists only on the current machine
    • Stored in local database
    • Fast iteration, no sharing
    • Default namespace when none specified
  • <username>/ (e.g., freemo/, jsmith/)

    • Personal namespace on the server
    • Created when user registers an account
    • Stored on server, synced when connected
    • Used for reusable actions/actors a user wants across machines
    • Only the owning user can create/modify items
  • <orgname>/ (e.g., cleverthis/, acme/)

    • Organization namespace on the server
    • Created when organization is registered
    • Shared across team members
    • Permissions and approvals managed at org level
    • Actions/actors/projects can be centrally managed
  • Built-in Provider Namespaces (openai/, anthropic/, google/, etc.)

    • Reserved for built-in LLM actors
    • Automatically available when API keys are configured
    • In server mode: available if logged in and server has keys
    • In local mode: requires environment variables or app configuration
    • Cannot be used for custom actors

Server-qualified Names

To disambiguate between servers (when connected to multiple):

  • dev:freemo/code-coverage (personal namespace on dev server)
  • prod:cleverthis/deploy-action (org namespace on prod server)

This enables a pattern where:

  • local machine runs a lightweight client
  • server stores canonical definitions
  • multiple servers can coexist

Actor

What an Actor Is

An actor is the abstraction that generalizes "agent" into "anything conversational."

  • It can be as small as a single LLM agent.
  • It can also be an entire graph that itself calls other actors/tools.
  • Actors can be nested/hierarchical, enabling "orchestrator of orchestrators."

Every custom actor IS a graph (a LangGraph defined via YAML configuration). Even a simple actor wrapping a single LLM is technically a graph with one node.

Actor Naming

Actors are always named using <namespace>/<name> format:

  • local/my-reviewer - Local actor
  • freemo/code-analyzer - Personal server actor
  • cleverthis/deploy-specialist - Organization actor
  • openai/gpt-4 - Built-in LLM actor

Actor Definition (YAML Configuration)

Actors are defined via YAML configuration files. This is the ONLY place YAML configuration is used (not for actions or projects).

Example actor configuration (see examples/ directory for full examples):

cleveragents:
  version: "3.0"
  default_actor: workflow_controller

actors:
  # Simple LLM actor
  my_assistant:
    type: llm
    config:
      actor: openai/gpt-4           # Reference to built-in actor
      temperature: 0.7
      system_prompt: |
        You are a helpful assistant.
        Current task: {{ context.task_description }}

  # Tool actor with inline Python
  data_processor:
    type: tool
    config:
      tools:
        - name: process_data
          code: |
            # Inline Python code
            result = do_something(input_data, context)

  # Actor referencing another actor
  reviewer:
    type: llm
    config:
      actor: local/code-reviewer    # Reference to another custom actor
      memory_enabled: true
      max_history: 20

routes:
  main_workflow:
    type: graph
    entry_point: start
    nodes:
      - name: analyze
        type: agent
        agent: my_assistant
      - name: process
        type: agent
        agent: data_processor
    edges:
      - source: start
        target: analyze
      - source: analyze
        target: process
      - source: process
        target: end

context:
  global:
    task_description: "Default task"

Actor Arguments

All actors can receive arguments when invoked, including built-in actors. Arguments are passed when:

  1. An action is used on projects (arguments flow to strategy/execution actors)
  2. An actor is directly invoked

Arguments are injected into the actor's context and can be used in Jinja2 templates within prompts.

For built-in actors (like openai/gpt-4), common arguments include:

  • temperature
  • max_tokens
  • system_prompt

Actor Composition (Hierarchical References)

Actors can reference other actors by name:

actors:
  complex_workflow:
    type: llm
    config:
      actor: local/base-analyzer    # References another actor

Load order matters: Referenced actors must be loaded/defined before actors that depend on them.

This enables hierarchical composition where:

  • Actor A's graph can include nodes that call Actor B (by name)
  • Actor B itself is a graph that might call Actor C
  • And so on...

Actor vs Agent (Relationship)

  • Agent: an actor that is specifically an LLM with tools and reasoning behaviors.

  • Actor: may be an agent, but may also be:

    • a composite workflow,
    • a multi-step graph,
    • a wrapper around a third-party system (as long as it's "text in → text out" conversationally).

Actor Definition Fields (From Notes + Extended)

The notes describe an actor as a configuration profile with fields like name, provider, model, configuration, blob, and graph description.

A robust actor schema should include:

  • name (namespaced)
  • provider (LLM provider or runtime target)
  • model
  • system_prompt (or prompt template)
  • tool_access_policy
  • graph_descriptor (for composite actors)
  • memory_policy (per-plan/per-actor—see memory section)
  • context_view_policy (what context this actor sees)
  • limits (token limits, tool call limits, retries)
  • cost_policy (caps, budgets)
  • metadata (tags, use cases, version)

Actor Composition and Graphs

Actors can reference:

  • other actors
  • skills (MCP/custom tools)
  • subgraphs

This is central to enabling both:

  • multi-agent orchestration, and
  • modular reuse of workflows.

Nodes in the Graph: Actor, Skill, or Custom Tool

The transcript explicitly frames the graph nodes as being any of:

  • an actor,
  • an MCP skill,
  • a custom tool node (arbitrary python code).

This is a powerful simplification: "everything is a node."

Agent

Agent Definition

In CleverAgents, an agent is a specialized actor with:

  • a conversational interface,
  • tool-calling capability,
  • potentially memory, planning heuristics, and role identity.

Examples of agent roles:

  • planner/architect (strategy actor)
  • coder/implementer (execution actor)
  • reviewer/qa agent
  • release/apply agent

The transcript explicitly discusses role separation like planner/coder/reviewer in context views/memory proposals.

Agent Behavior Configuration

Agents should be configurable without code changes:

  • prompt templates
  • tool sets
  • safety constraints
  • style constraints (verbosity, code style)
  • reliability controls (self-checks, validations)

A design goal is user empowerment: "users customize LLM behavior without modifying core code."

Skills

What a Skill Is

A skill is an executable capability exposed to the system. Skills are defined inline in actor YAML configuration files as tool nodes.

Examples:

  • file read/write
  • git operations
  • shell command execution
  • database query execution
  • API calls
  • vector search / RAG query
  • cloud administration actions

Skill Definition

Skills do not have separate names. They are defined as part of actor configurations:

actors:
  my_processor:
    type: tool
    config:
      tools:
        - name: process_data        # Tool name within this actor
          code: |
            # Inline Python code
            result = process(input_data, context)
            
        - name: validate_output
          code: |
            # Another tool in the same actor
            if not is_valid(input_data):
                raise ValueError("Invalid output")
            result = input_data

To create a "named skill": Define an actor with a single tool node. The actor name then effectively becomes the skill name.

Skill Standards

Skills should extend two complementary standards:

  1. MCP (Model Context Protocol): https://modelcontextprotocol.io/

    • Tools: Schema-defined operations LLMs can invoke
    • Resources: Read-only data sources
    • Prompts: User-controlled templates
  2. Agent Skills Standard: https://agentskills.io/

    • Folders of instructions, scripts, and resources
    • Portable procedural knowledge across agent products

CleverAgents extends the MCP resource concept to support both read and write operations.

Skill Sources

Skills come from multiple places:

  1. Built-in skills (first-party, when implemented)

    • files, git, search, indexing, context mgmt, sandbox ops
  2. MCP skills (external MCP servers)

  3. Imported skills from other ecosystems

    • wrap tools from other agent frameworks as long as interface is compatible
  4. Custom tool nodes (most common)

    • arbitrary python code embedded in actor configuration
    • behaves as a graph node

Skill Capability Metadata (Critical for Safety)

MCP's metadata is not sufficient (read-only/idempotent isn't enough; write scope is unclear). CleverAgents needs extended metadata for each skill.

Required skill metadata:

  • read_only: bool - Whether skill only performs read operations
  • writes: bool - Whether skill can modify resources
  • write_scope:
    • file paths allowed
    • resource IDs allowed
    • environment boundaries (container only vs host)
  • idempotent: bool - Whether repeated calls produce same result
  • checkpointable: bool - Whether skill supports checkpoint/rollback
  • checkpoint_scope - What can be rolled back
  • side_effects - install packages, mutate infra, etc.
  • required_permissions - What permissions needed to use
  • rate_limits / cost_profile - Usage constraints
  • human_approval_required: bool - Optional approval gate

Read-Only Actions

When an action is marked read_only: true, it can only use skills that have read_only: true in their metadata. This is enforced at runtime.

Skill Registry / Catalog

To scale, CleverAgents should maintain a catalog of skills with metadata:

  • auto-extracted from MCP descriptors (where possible)
  • refined manually via annotations
  • enhanced by CleverAgents-specific extensions

This registry supports:

  • plan validation ("this plan requires checkpointable write skills; do we have them?")
  • safe automation ("don't ask permission for every tiny command—use sandbox/checkpoints instead")

MCP Integration Architecture

CleverAgents fully integrates with the Model Context Protocol (MCP) while extending it for agentic workflows:

MCP Concepts Mapping

MCP Concept CleverAgents Equivalent Extension
Tool Skill Extended metadata (write_scope, checkpointable)
Resource Resource Read AND write operations
Prompt Action template Full plan lifecycle
Server MCP Server (external) Integrated via skill adapters

Using External MCP Servers

CleverAgents can connect to any MCP server and expose its tools as skills:

actors:
  github_ops:
    type: tool
    config:
      mcp_servers:
        - name: github
          command: "npx @anthropic/mcp-github"
          env:
            GITHUB_TOKEN: "${GITHUB_TOKEN}"
        - name: filesystem
          command: "npx @anthropic/mcp-filesystem"
          args: ["--root", "/workspace"]

When an actor specifies mcp_servers, all tools from those servers become available as skills within that actor's execution context.

MCP Tool → Skill Adapter

External MCP tools are automatically wrapped with CleverAgents skill semantics:

┌─────────────────────────────────────┐
│ MCP Server                          │
│ - Exposes tools via JSON-RPC        │
│ - Has MCP metadata (read-only, etc) │
└─────────────────────────────────────┘
              ↓
┌─────────────────────────────────────┐
│ MCPSkillAdapter                     │
│ - Wraps MCP tool as Skill           │
│ - Infers extended metadata          │
│ - Intercepts calls for:             │
│   - Sandbox path rewriting          │
│   - Change tracking                 │
│   - Permission enforcement          │
└─────────────────────────────────────┘
              ↓
┌─────────────────────────────────────┐
│ Skill Execution                     │
│ - Runs in plan's sandbox context    │
│ - Changes recorded to ChangeSet     │
│ - Checkpoints created as needed     │
└─────────────────────────────────────┘

Skill Execution Flow (Tool-Based Architecture)

When an LLM decides to use a skill, the following flow occurs:

1. LLM generates tool call: edit_file(path="src/main.py", changes=[...])
                                ↓
2. Tool Router receives call
   - Validates parameters against schema
   - Checks skill capability metadata
   - Enforces permission restrictions
                                ↓
3. Sandbox Context Resolution
   - Maps logical path to sandbox path
   - Ensures sandbox exists for resource
   - Creates checkpoint if skill is checkpointable
                                ↓
4. Skill Execution
   - Runs skill code (built-in or MCP)
   - Operations occur on sandboxed state
   - Result captured
                                ↓
5. Change Recording
   - If skill modifies resources, create Change record
   - Append Change to plan's ChangeSet
   - Update sandbox state
                                ↓
6. Return to LLM
   - Return skill result
   - LLM continues with next action

Built-in Skills (Core Resource Operations)

CleverAgents provides these built-in skills that work with any resource through the unified abstraction layer:

File Operations:

read_file(path: str) -> str
write_file(path: str, content: str) -> None
edit_file(path: str, edits: list[Edit]) -> None
delete_file(path: str) -> None
move_file(source: str, destination: str) -> None
copy_file(source: str, destination: str) -> None

Directory Operations:

create_directory(path: str) -> None
list_directory(path: str, pattern: str = "*") -> list[str]
delete_directory(path: str, recursive: bool = False) -> None

Search Operations:

search_files(pattern: str, content_pattern: str = None) -> list[Match]
find_definition(symbol: str) -> list[Location]
find_references(symbol: str) -> list[Location]

Git Operations (when resource is git repository):

git_status() -> GitStatus
git_diff(path: str = None) -> str
git_log(count: int = 10) -> list[Commit]
git_blame(path: str) -> list[BlameLine]

Each built-in skill:

  • Has fully defined capability metadata
  • Operates through the resource abstraction layer
  • Automatically tracks changes to the ChangeSet
  • Respects sandbox boundaries and deny-lists

Change Tracking from Tool Invocations

Critical Architecture Point: The ChangeSet is NOT built by parsing LLM output. It is built by recording the effects of tool/skill invocations:

class SkillExecutionContext:
    """Context provided to skill execution."""
    
    def __init__(self, plan: Plan, sandbox: Sandbox):
        self.plan = plan
        self.sandbox = sandbox
        self.changes: list[Change] = []
    
    def record_change(self, change: Change) -> None:
        """Record a change made by a skill."""
        self.changes.append(change)
        self.plan.changeset.add_change(change)


class WriteFileSkill:
    """Built-in skill for writing files."""
    
    def execute(self, path: str, content: str, ctx: SkillExecutionContext) -> None:
        # Get the resource handler for this path
        handler = ctx.sandbox.get_handler(path)
        
        # Perform the write (returns Change record)
        change = handler.write(path, content, ctx.sandbox)
        
        # Record the change
        ctx.record_change(change)

This approach means:

  • Every resource modification is explicit and tracked
  • The ChangeSet accurately reflects what was done, not what was said
  • Rollback is precise (replay inverse of recorded changes)
  • Audit logs show exactly what each skill invocation did

Session

What a Session Is

A session is a user's interactive thread with CleverAgents across time.

A session should:

  • maintain conversational continuity,
  • store plan references,
  • persist memory (if enabled),
  • provide a UI anchor (CLI invocation, TUI workspace, web session).

Session and Memory Persistence

The notes include a known issue: conversation history can be lost between CLI invocations depending on connection string configuration, implying the system needs a stable memory service backend.

Therefore, CleverAgents should specify:

  • sessions have stable IDs
  • sessions can be resumed
  • session storage backend is configured explicitly
  • if session persistence is disabled, the UX should be explicit about it (no silent history loss)

Server

What a Server Is

A server is an optional mode that enables:

  • multi-user access
  • shared org namespaces (<username>/ and <orgname>/)
  • persistent plan records
  • remote plan execution
  • permissioning and governance

Single-user local mode is the default (so setup is easy), but the architecture anticipates server mode for shared skills and org-level actions.

Client-Only vs Server Mode

Mode Description Plan Execution
Client-only No server connection. All data in local database. Always local
Server mode Connected to a CleverAgents server. Namespaced items sync. Local or server

It is possible to run a client with no server at all. Server is optional.

Plan Execution Location

Where a plan executes depends on the project type:

Project Type Client Execution Server Execution
Local (has local-only resources) Yes No (server can't access local resources)
Remote (all resources remotely accessible) Yes Yes

When acting on local projects, the client must be running because only the client can access local resources.

Remote projects can execute on either:

  • The client (if user prefers local execution)
  • The server (for long-running plans, since client may be transient)

Server Execution Benefits

Server execution is useful when:

  • Plans take a long time to execute
  • Client may disconnect (laptop closes, network issues)
  • Multiple team members need to monitor plan progress
  • Centralized logging and auditing required

No Plan Queuing

Plans are not queued. When a plan is used on projects and executed, it runs immediately. There is no worker queue or delayed execution model.

Multi-user Risks and Prompt Injection

Prompt injection isn't critical in single-user mode but becomes important for multi-user server environments.

Server mode must include:

  • permission boundaries
  • prompt sanitization / safe templating
  • resource access controls
  • auditing

Permissions

Permissions exist at multiple layers:

1) Namespace-level permissions

  • Who can create/edit org actions/actors/projects?
  • Who can run them?

2) Project-level permissions

  • Who can modify project resources?
  • Who can apply changes?

3) Plan-level permissions

  • Can this plan write?
  • Does it require approvals?
  • Can it access restricted skills?

4) Skill-level permissions

  • Some skills should require:

    • explicit user approval per call, or
    • elevated role membership.

A simple and powerful governance model:

  • Strategize: generally safe, read-only → minimal restrictions
  • Execute: writes occur but sandboxed → moderate restrictions
  • Apply: writes are real → strict restrictions + optional mandatory review

This aligns with the four-phase model's safety rationale.

Resources

Resources are any objects a plan can reason about or manipulate.

CleverAgents extends the MCP resource concept to support both read AND write operations (MCP resources are read-only).

Resource Types (Examples)

  • FilesystemResource: directories, files
  • GitRepoResource: git repos, branches
  • DatabaseResource: SQL/NoSQL endpoints
  • CloudResource: clusters, accounts, infra configs
  • DocumentCorpusResource: PDFs, markdown docs, wikis
  • APISpecResource: OpenAPI/Swagger, Postman collections
  • IssueTrackerResource: tickets, bugs, tasks (optional)

Resource Sandbox Strategy

Each resource defines its own sandbox strategy. This is critical because:

  1. The same resource may be accessed through different skills
  2. Different resource types require different sandboxing approaches
  3. Some resources cannot be sandboxed at all
Resource Type Sandbox Strategy Rollback Mechanism
Git Repository git_worktree Git reset/checkout
Filesystem copy_on_write or overlay Restore from snapshot
Database transaction_rollback Transaction rollback
Cloud Infra terraform_state Terraform apply (reversed)
API Endpoint none (often not sandboxable) N/A

Example resource with sandbox strategy:

agents project add-resource \
  --project "my-api" \
  --name "main-repo" \
  --type "git_repository" \
  --location "git@github.com:org/api.git" \
  --sandbox-strategy "git_worktree"

Lazy Sandboxing

Resources are sandboxed lazily when accessed, not upfront. Note that this is different from indexing - resources are indexed immediately when added to a project, but sandboxes are only created when execution needs to modify a resource:

  1. A project may contain many resources (e.g., git repo + 10 databases + cloud accounts)
  2. A plan may only need to modify one resource
  3. Only the accessed resources are sandboxed
  4. Each plan/subplan has its own sandbox containing only edited resources

This is efficient for large projects where most resources remain untouched.

Resource Access Tracking

The system needs to know which skills touch which resources to reason about safety and checkpointing.

Every skill call should log:

  • resource IDs accessed
  • read/write actions
  • file paths or object IDs touched

This enables:

  • better context assembly
  • better rollback feasibility analysis
  • better auditing
  • accurate sandbox scoping

Unified Resource Abstraction Layer

CleverAgents provides a unified abstraction that allows skills to work with any resource type through a consistent interface. This enables:

  1. Resource-agnostic skills: A skill like read_content(path) works whether the path refers to a file, database record, or API endpoint
  2. Consistent sandbox semantics: All resources support the same sandbox lifecycle (create, read, write, checkpoint, rollback)
  3. Pluggable resource handlers: New resource types can be added without modifying existing skills
  4. Unified change tracking: All resource modifications flow into the same ChangeSet model

Resource Handler Interface

Every resource type implements this interface:

class ResourceHandler(Protocol):
    """Handler for a specific resource type."""
    
    def read(self, path: str, sandbox: Sandbox) -> Content:
        """Read content from the sandboxed resource."""
        ...
    
    def write(self, path: str, content: Content, sandbox: Sandbox) -> Change:
        """Write content and return the Change record."""
        ...
    
    def delete(self, path: str, sandbox: Sandbox) -> Change:
        """Delete resource and return the Change record."""
        ...
    
    def list(self, pattern: str, sandbox: Sandbox) -> list[str]:
        """List paths matching pattern."""
        ...
    
    def diff(self, path: str, sandbox: Sandbox) -> str:
        """Generate diff between sandbox and original state."""
        ...
    
    def supports_operation(self, operation: OperationType) -> bool:
        """Check if this resource supports the given operation."""
        ...

Built-in Resource Handlers

Resource Type Handler Read Write Delete Sandbox Strategy
Filesystem FilesystemHandler copy_on_write
Git Repository GitHandler git_worktree
PostgreSQL PostgresHandler transaction
SQLite SQLiteHandler copy_on_write
HTTP API HTTPHandler ✓* ✓* none
S3 Bucket S3Handler versioning

*HTTP writes may not be sandboxable depending on the API

Resource Path Resolution

Paths in skills are resolved through a resource routing system:

path://resource-name/relative/path
  ↓
┌─────────────────────────────────────┐
│ Resource Router                     │
│ - Parses path scheme                │
│ - Looks up resource by name         │
│ - Routes to appropriate handler     │
└─────────────────────────────────────┘
  ↓
┌─────────────────────────────────────┐
│ Handler (e.g., GitHandler)          │
│ - Resolves relative path            │
│ - Operates on sandboxed state       │
│ - Returns Change record             │
└─────────────────────────────────────┘

For convenience, paths without a scheme default to the project's primary filesystem resource.

Code Intelligence & Context Discovery

Overview

CleverAgents employs a sophisticated multi-layered indexing and discovery system that enables agents to efficiently navigate and understand codebases of any scale. This system goes far beyond simple text search, providing semantic understanding of code structure, dependencies, and relationships through a combination of indexed embeddings, vector search, and an RDF-based graph store.

Critical Design Decision: All indexing happens immediately when resources are added to projects or when code changes. There is no "on-demand" indexing during agent execution. This ensures that agents always have instant access to search capabilities without any indexing delays. The computational cost is paid once upfront, not repeatedly during agent operations.

Key Design Principles:

  1. Pluggable Architecture: Every component can be extended or replaced
  2. Progressive Enhancement: System works with basic text search, enhances with advanced features
  3. Eager Indexing: Indices are built immediately when resources are added and kept continuously up-to-date
  4. Agent Awareness: Agents understand available indices through skills
  5. Real-time Synchronization: Indices update immediately as code changes

Architecture Components

1. Multi-Modal Indexing Engine

The indexing engine operates across three complementary modalities:

IndexingEngine:
  modalities:
    # Traditional text-based indexing
    text_index:
      type: "full_text_search"
      backend: "tantivy" | "elasticsearch" | "sqlite_fts"
      features:
        - Token-based search
        - Regex patterns
        - Language-aware tokenization
        - File path indexing
    
    # Semantic understanding via embeddings
    vector_index:
      type: "embedding_search"
      backend: "faiss" | "qdrant" | "weaviate" | "pgvector"
      models:
        - code: "codegen-6B-multi"
        - docs: "instructor-xl"
        - cross-modal: "clip-code"
      features:
        - Function-level embeddings
        - Class-level embeddings
        - Module-level embeddings
        - Documentation embeddings
        - Cross-language similarity
    
    # Structural understanding via graph
    graph_index:
      type: "rdf_knowledge_graph"
      backend: "blazegraph" | "stardog" | "apache_jena" | "neo4j"
      ontology: "CodeOntology"
      features:
        - AST-based relationships
        - Dependency graphs
        - Call graphs
        - Inheritance hierarchies
        - Data flow analysis

2. RDF-Based Code Knowledge Graph

The graph store represents code as a rich semantic network using RDF (Resource Description Framework) triples. This enables sophisticated queries about code structure and relationships.

Core Ontology Design:

# CodeOntology - Core vocabulary for code representation
@prefix code: <https://cleveragents.ai/ontology/code#> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .
@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .

# Core Classes
code:Module a rdfs:Class ;
    rdfs:comment "A code module (file, package, namespace)" .

code:Class a rdfs:Class ;
    rdfs:comment "A class or similar construct" .

code:Function a rdfs:Class ;
    rdfs:comment "A function, method, or procedure" .

code:Variable a rdfs:Class ;
    rdfs:comment "A variable, constant, or field" .

code:Type a rdfs:Class ;
    rdfs:comment "A type definition" .

# Core Properties
code:contains a rdf:Property ;
    rdfs:domain code:Module ;
    rdfs:range code:Entity ;
    rdfs:comment "Module contains entity" .

code:imports a rdf:Property ;
    rdfs:domain code:Module ;
    rdfs:range code:Module ;
    rdfs:comment "Module imports another module" .

code:extends a rdf:Property ;
    rdfs:domain code:Class ;
    rdfs:range code:Class ;
    rdfs:comment "Class inheritance relationship" .

code:calls a rdf:Property ;
    rdfs:domain code:Function ;
    rdfs:range code:Function ;
    rdfs:comment "Function calls another function" .

code:references a rdf:Property ;
    rdfs:domain code:Entity ;
    rdfs:range code:Entity ;
    rdfs:comment "Entity references another entity" .

code:hasParameter a rdf:Property ;
    rdfs:domain code:Function ;
    rdfs:range code:Parameter ;
    rdfs:comment "Function has parameter" .

code:returns a rdf:Property ;
    rdfs:domain code:Function ;
    rdfs:range code:Type ;
    rdfs:comment "Function return type" .

# Annotations
code:hasDocstring a rdf:Property ;
    rdfs:domain code:Entity ;
    rdfs:range xsd:string .

code:hasComplexity a rdf:Property ;
    rdfs:domain code:Function ;
    rdfs:range xsd:integer ;
    rdfs:comment "Cyclomatic complexity" .

code:hasTestCoverage a rdf:Property ;
    rdfs:domain code:Entity ;
    rdfs:range xsd:decimal ;
    rdfs:comment "Test coverage percentage" .

Example Knowledge Graph Fragment:

# Concrete example: Authentication module
<file:///src/auth/auth_manager.py> a code:Module ;
    code:imports <file:///src/core/user.py> ;
    code:imports <file:///src/utils/crypto.py> ;
    code:contains <class://AuthManager> .

<class://AuthManager> a code:Class ;
    code:hasMethod <method://AuthManager.authenticate> ;
    code:hasMethod <method://AuthManager.validate_token> ;
    code:extends <class://BaseManager> ;
    code:hasDocstring "Manages user authentication and session tokens" .

<method://AuthManager.authenticate> a code:Function ;
    code:hasParameter <param://username> ;
    code:hasParameter <param://password> ;
    code:returns <type://AuthToken> ;
    code:calls <method://CryptoUtils.hash_password> ;
    code:calls <method://UserDB.find_user> ;
    code:hasComplexity 8 ;
    code:hasTestCoverage 0.95 .

Advanced Graph Queries:

# Find all functions that manipulate user authentication
PREFIX code: <https://cleveragents.ai/ontology/code#>
SELECT ?function ?module
WHERE {
    ?function a code:Function ;
              code:calls*/code:references ?entity .
    ?entity rdfs:label ?label .
    FILTER(CONTAINS(LCASE(?label), "auth") || CONTAINS(LCASE(?label), "user"))
    ?module code:contains ?function .
}

# Find circular dependencies
SELECT ?module1 ?module2
WHERE {
    ?module1 code:imports+ ?module2 .
    ?module2 code:imports+ ?module1 .
    FILTER(?module1 != ?module2)
}

# Find most complex untested functions
SELECT ?function ?complexity
WHERE {
    ?function a code:Function ;
              code:hasComplexity ?complexity ;
              code:hasTestCoverage ?coverage .
    FILTER(?complexity > 10 && ?coverage < 0.5)
}
ORDER BY DESC(?complexity)
LIMIT 10

3. Intelligent Context Assembly Pipeline

The context assembly pipeline leverages all three indices to build optimal context for each agent:

class ContextAssemblyPipeline:
    def assemble_context(self, 
                        query: str, 
                        actor_type: str,
                        resource_scope: List[Resource],
                        max_tokens: int) -> Context:
        
        # Stage 1: Query Understanding
        intent = self.analyze_query_intent(query)
        entities = self.extract_entities(query)  # Classes, functions, concepts
        
        # Stage 2: Multi-Modal Search
        results = SearchResults()
        
        # Text search for exact matches
        if intent.needs_exact_match:
            text_results = self.text_index.search(
                query=query,
                filters={"resources": resource_scope},
                limit=100
            )
            results.add(text_results)
        
        # Vector search for semantic similarity
        if intent.needs_semantic_match:
            query_embedding = self.embed_query(query, actor_type)
            vector_results = self.vector_index.search(
                embedding=query_embedding,
                filters={"resources": resource_scope},
                limit=50
            )
            results.add(vector_results)
        
        # Graph traversal for structural relationships
        if entities:
            graph_results = self.graph_index.traverse(
                start_nodes=entities,
                patterns=self.get_patterns_for_actor(actor_type),
                max_depth=3,
                limit=50
            )
            results.add(graph_results)
        
        # Stage 3: Relevance Ranking
        ranked_results = self.rank_by_relevance(
            results=results,
            actor_type=actor_type,
            query_intent=intent
        )
        
        # Stage 4: Context Optimization
        context = self.optimize_context(
            ranked_results=ranked_results,
            max_tokens=max_tokens,
            strategy=self.get_strategy_for_actor(actor_type)
        )
        
        return context
    
    def get_patterns_for_actor(self, actor_type: str) -> List[GraphPattern]:
        """Different actors need different traversal patterns"""
        patterns = {
            "strategist": [
                "module_dependencies",      # Understand architecture
                "interface_boundaries",     # Find API surfaces
                "test_coverage_gaps"        # Identify risks
            ],
            "executor": [
                "implementation_details",   # Get full function bodies
                "local_dependencies",       # Find what to change
                "usage_patterns"           # Understand call sites
            ],
            "reviewer": [
                "change_impact_analysis",   # What could break
                "similar_patterns",        # Consistency checks
                "test_relationships"       # Verification paths
            ]
        }
        return patterns.get(actor_type, ["general_traversal"])

4. Plugin Architecture for Extensibility

The system is designed for extensibility at every level:

PluginSystem:
  # Language-specific analyzers
  analyzers:
    python:
      class: "PythonAnalyzer"
      features:
        - AST parsing via ast module
        - Type inference via mypy
        - Import resolution
        - Docstring extraction
    
    typescript:
      class: "TypeScriptAnalyzer"
      features:
        - TSC-based parsing
        - Type extraction
        - Module resolution
        - JSDoc parsing
    
    rust:
      class: "RustAnalyzer"
      features:
        - rust-analyzer integration
        - Lifetime analysis
        - Trait resolution
        - Macro expansion
    
    # Custom analyzer example
    custom_dsl:
      class: "CustomDSLAnalyzer"
      config:
        grammar: "path/to/grammar.peg"
        semantic_rules: "path/to/rules.yaml"
  
  # Index backend providers
  backends:
    graph:
      - name: "blazegraph"
        class: "BlazegraphBackend"
        scalability: "billions of triples"
        features: ["SPARQL", "reasoning", "geospatial"]
      
      - name: "neo4j"
        class: "Neo4jBackend"
        scalability: "enterprise"
        features: ["Cypher", "APOC", "GDS"]
      
      - name: "custom_graph"
        class: "MyCustomGraphDB"
        config:
          connection: "custom://localhost:7687"
    
    vector:
      - name: "faiss"
        class: "FaissBackend"
        scalability: "100M vectors"
        features: ["GPU acceleration", "HNSW"]
      
      - name: "qdrant"
        class: "QdrantBackend"
        scalability: "distributed"
        features: ["filtering", "payloads", "snapshots"]
  
  # Embedding model providers
  embedders:
    - name: "openai"
      class: "OpenAIEmbedder"
      models: ["text-embedding-3-large", "text-embedding-3-small"]
    
    - name: "local"
      class: "LocalEmbedder"
      models: ["all-MiniLM-L6-v2", "instructor-xl"]
    
    - name: "custom"
      class: "MyFineTunedEmbedder"
      model_path: "path/to/model"

5. Agent Skills for Code Intelligence

Agents interact with the code intelligence system through specialized skills:

actors:
  code_explorer:
    type: tool
    config:
      tools:
        # Semantic code search
        - name: search_code_semantically
          code: |
            # Find code similar to a concept
            results = context.code_intelligence.vector_search(
                query="validate user input against schema",
                scope=input_data.get("scope", "all"),
                limit=input_data.get("limit", 10)
            )
            
            # Enrich with graph context
            for result in results:
                dependencies = context.code_intelligence.get_dependencies(
                    entity=result.entity_id,
                    depth=2
                )
                result.context = dependencies
            
            return results
        
        # Structural analysis
        - name: analyze_dependencies
          code: |
            # Use graph store for dependency analysis
            module = input_data["module_path"]
            
            # SPARQL query for full dependency closure
            query = f'''
            PREFIX code: <https://cleveragents.ai/ontology/code#>
            SELECT ?dep ?type
            WHERE {{
                <{module}> code:imports* ?dep .
                ?dep a ?type .
            }}
            '''
            
            deps = context.code_intelligence.graph_query(query)
            
            # Compute metrics
            return {
                "direct_deps": len([d for d in deps if d.distance == 1]),
                "transitive_deps": len(deps),
                "circular_deps": context.code_intelligence.find_circular_deps(module),
                "dependency_graph": deps
            }
        
        # Intelligent refactoring assistant
        - name: suggest_refactoring_targets
          code: |
            # Combine all three indices for comprehensive analysis
            
            # 1. Text search for TODO/FIXME/HACK comments
            todos = context.code_intelligence.text_search(
                pattern="(TODO|FIXME|HACK):",
                scope=input_data["scope"]
            )
            
            # 2. Graph analysis for high complexity
            complex_functions = context.code_intelligence.graph_query('''
                SELECT ?func ?complexity ?coverage
                WHERE {
                    ?func a code:Function ;
                          code:hasComplexity ?complexity ;
                          code:hasTestCoverage ?coverage .
                    FILTER(?complexity > 15 || ?coverage < 0.3)
                }
                ORDER BY DESC(?complexity)
            ''')
            
            # 3. Vector search for code smells
            smells = []
            for pattern in ["duplicate code", "long method", "large class"]:
                similar = context.code_intelligence.find_similar_code(
                    pattern=pattern,
                    threshold=0.8
                )
                smells.extend(similar)
            
            # Synthesize recommendations
            return {
                "high_priority": complex_functions[:5],
                "technical_debt": todos,
                "code_smells": smells,
                "suggested_order": context.code_intelligence.rank_by_impact(
                    complex_functions + todos + smells
                )
            }

6. Real-time Index Synchronization

The system maintains index freshness through immediate, proactive updates:

class IndexSynchronizer:
    def __init__(self):
        self.file_watcher = FileSystemWatcher()
        self.git_monitor = GitChangeMonitor()
        self.incremental_indexer = IncrementalIndexer()
    
    def on_resource_added(self, resource: Resource, project: Project):
        """When a resource is added to a project, index it immediately"""
        # Full initial indexing - happens once when resource is added
        with self.progress_reporter(f"Indexing {resource.name}") as progress:
            files = self.scan_resource(resource)
            total = len(files)
            
            # Parallel indexing for performance
            with ThreadPoolExecutor(max_workers=cpu_count()) as executor:
                futures = []
                
                for i, file in enumerate(files):
                    future = executor.submit(self.index_file_complete, file)
                    futures.append(future)
                    progress.update(i / total)
                
                # Wait for all indexing to complete
                for future in futures:
                    future.result()
        
        # Now set up watchers for incremental updates
        self.setup_watchers(resource)
        
        # Mark resource as indexed and ready
        resource.indexing_status = "ready"
        self.notify_agents_index_ready(resource)
        
        # CRITICAL: Agents can now immediately search this resource
        # No "warming up" period - indices are complete and ready
    
    def index_file_complete(self, file_path: str):
        """Comprehensive initial indexing of a file"""
        # Parse file once
        ast = self.parse_file(file_path)
        
        # Update all indices immediately
        self.update_text_index(file_path, ast)
        self.update_vector_embeddings(file_path, ast)
        self.update_graph_triples(file_path, ast)
        
        # Extract and index all metadata
        self.index_symbols(file_path, ast)
        self.index_dependencies(file_path, ast)
        self.index_complexity_metrics(file_path, ast)
    
    def setup_watchers(self, project: Project):
        # File system watching for immediate updates
        self.file_watcher.watch(
            path=project.root_path,
            events=["create", "modify", "delete"],
            callback=self.on_file_change
        )
        
        # Git monitoring for batch updates
        self.git_monitor.watch(
            repo=project.git_repo,
            events=["commit", "merge", "rebase"],
            callback=self.on_git_change
        )
    
    def on_file_change(self, event: FileEvent):
        # Quick incremental update
        if event.type in ["create", "modify"]:
            # Parse changed file
            ast = self.parse_file(event.path)
            
            # Update indices
            self.update_text_index(event.path, ast)
            self.update_vector_embeddings(event.path, ast)
            self.update_graph_triples(event.path, ast)
        
        elif event.type == "delete":
            self.remove_from_indices(event.path)
    
    def on_git_change(self, event: GitEvent):
        # Batch update for git operations
        changed_files = event.get_changed_files()
        
        # Optimize batch processing
        with self.batch_updater() as updater:
            for file in changed_files:
                updater.queue_update(file)
            
            # Process in parallel
            updater.execute(parallel=True)
    
    def update_graph_triples(self, file_path: str, ast: AST):
        # Generate RDF triples from AST
        triples = []
        
        # Module-level triples
        module_uri = self.uri_for_file(file_path)
        for import_stmt in ast.imports:
            imported_uri = self.resolve_import(import_stmt)
            triples.append((module_uri, "code:imports", imported_uri))
        
        # Function-level triples
        for func in ast.functions:
            func_uri = self.uri_for_function(func)
            triples.append((module_uri, "code:contains", func_uri))
            triples.append((func_uri, "a", "code:Function"))
            triples.append((func_uri, "code:hasComplexity", func.complexity))
            
            # Call relationships
            for call in func.calls:
                called_uri = self.resolve_call(call)
                triples.append((func_uri, "code:calls", called_uri))
        
        # Update graph store
        self.graph_store.update_triples(triples)

When advanced features are unavailable, the system gracefully degrades:

class FallbackSearchProvider:
    def search(self, query: str, resources: List[Resource]) -> SearchResults:
        # Try advanced search first
        try:
            if self.vector_index.is_available():
                return self.vector_search(query, resources)
        except ServiceUnavailable:
            pass
        
        # Fallback to graph search
        try:
            if self.graph_index.is_available():
                return self.graph_search(query, resources)
        except ServiceUnavailable:
            pass
        
        # Ultimate fallback: grep-like text search
        return self.basic_text_search(query, resources)
    
    def basic_text_search(self, query: str, resources: List[Resource]):
        # Use ripgrep or similar for fast text search
        results = []
        
        for resource in resources:
            matches = ripgrep.search(
                pattern=query,
                path=resource.path,
                context_lines=3
            )
            
            for match in matches:
                results.append(SearchResult(
                    file=match.file,
                    line=match.line,
                    content=match.content,
                    score=1.0  # Basic scoring
                ))
        
        return results

Index Lifecycle

The system follows a clear lifecycle for index management:

Index Lifecycle:
  1_resource_added:
    trigger: "agents project add-resource"
    action: "Immediate full indexing"
    duration: "Depends on size (10K files ~1 minute)"
    result: "All indices ready for instant search"
    
  2_code_changed:
    trigger: "File modification detected"
    action: "Immediate incremental update"
    duration: "Milliseconds per file"
    result: "Indices stay synchronized"
    
  3_resource_removed:
    trigger: "agents project remove-resource"
    action: "Immediate index cleanup"
    duration: "Seconds"
    result: "No stale data in indices"
    
  4_maintenance:
    trigger: "Scheduled or manual"
    action: "Reindex for consistency"
    duration: "Background process"
    result: "Indices optimized and verified"

Key Guarantees:
  - "No search happens on stale data"
  - "No 'index building' delays during agent execution"
  - "Changes visible in search immediately"
  - "Initial indexing is a one-time cost per resource"

Integration with Context Tiers

The Code Intelligence system directly feeds into the three-tier context architecture:

Context Tier Integration:
  hot_tier:
    source: "Real-time results from code intelligence"
    content:
      - Currently edited files
      - Direct dependencies
      - Immediately relevant functions
    
  warm_tier:
    source: "Indexed embeddings and graph queries"
    content:
      - Recent search results
      - Cached graph traversals
      - Vector similarity matches
      - Active decision contexts
    
  cold_tier:
    source: "Historical indices and compressed data"
    content:
      - Previous plan analyses
      - Archived dependency graphs
      - Historical refactoring patterns
      - Learned codebase conventions

Performance Characteristics

The system maintains pre-computed indices for instant search performance:

Performance Metrics:
  initial_indexing_speed:
    text_index: "10,000 files/minute"
    vector_index: "1,000 files/minute (with GPU)"
    graph_index: "5,000 files/minute"
  
  query_performance:
    text_search: "< 100ms for 1M files"
    vector_search: "< 200ms for 10M embeddings"
    graph_traversal: "< 500ms for 3-hop queries"
    
  storage_requirements:
    text_index: "~10% of source size"
    vector_index: "~1GB per 100K functions"
    graph_store: "~100MB per 10K files"
    
  scalability:
    max_files: "No hard limit (tested to 10M files)"
    max_graph_size: "1B+ triples"
    max_vectors: "100M+ embeddings"

Progressive Enhancement Path

Organizations can adopt Code Intelligence features progressively. At each stage, existing resources are reindexed to take advantage of new capabilities:

adoption_stages:
  stage_1_basic:
    features: ["text search", "file watching"]
    requirements: ["ripgrep", "sqlite"]
    initial_setup: "Index all text content on resource add"
    benefit: "Instant exact-match search"
    
  stage_2_semantic:
    features: ["vector embeddings", "similarity search"]
    requirements: ["embedding model", "vector DB"]
    initial_setup: "Generate embeddings for all code (one-time cost)"
    benefit: "Instant semantic similarity search"
    
  stage_3_structural:
    features: ["RDF graph", "relationship queries"]
    requirements: ["graph database", "language analyzers"]
    initial_setup: "Parse and build complete knowledge graph"
    benefit: "Instant relationship queries"
    
  stage_4_intelligent:
    features: ["ML-driven ranking", "automated analysis"]
    requirements: ["GPU", "training data"]
    initial_setup: "Pre-compute ML features and rankings"
    benefit: "Instant intelligent suggestions"
    
  stage_5_custom:
    features: ["Domain-specific ontologies", "Custom analyzers"]
    requirements: ["Domain expertise", "Custom development"]
    initial_setup: "Build domain-specific indices"
    benefit: "Instant domain-aware intelligence"

This Code Intelligence & Context Discovery system ensures that CleverAgents can efficiently work with codebases of any size, providing agents with the contextual understanding they need to make intelligent decisions about code changes, refactoring, and feature development.

Summary of Timing:

  • Indexing: Eager (happens immediately when resources are added/changed)
  • Searching: Instant (because indices are pre-computed and ready)
  • Sandboxing: Lazy (only when execution needs to modify a resource)
  • Context Assembly: Real-time (but fast because it queries ready indices)

This design ensures agents never wait for index building during execution, providing a responsive and predictable experience even on massive codebases.

Context

Note: This section describes the high-level context management system. For details on how context is discovered and indexed, see the Code Intelligence & Context Discovery section above.

Context in CleverAgents is not "dump all files into an LLM." It is a system that:

  • finds relevant information from resources,
  • injects appropriate subsets into each actor/node,
  • and scales to large repositories.

Current Reality and Planned Improvements

The transcript implies:

  • There is a global context concept,
  • But nodes only see what is injected into their prompts,
  • It's functional but not yet elegant,
  • A more advanced automated system is planned.

Tiered Context Architecture (Hot/Warm/Cold)

The system uses a sophisticated three-tier memory architecture that enables working with massive codebases without holding everything in memory:

  • Hot context (hot cache) The small set of immediately relevant chunks injected into the current actor prompt. When working on a 50,000 file codebase, hot context focuses on the immediate task (e.g., 10-20 files for a specific refactoring).

  • Warm context Recent decisions and their contexts from this plan tree - quickly accessible. Includes indexed embeddings, vector search results, and graph store representations. Maintains the decision chain that led to the current work.

  • Cold storage Historical decisions from past plans on this codebase - queryable but not in active memory. Long-term storage in SQLite or caching systems containing prior summaries, older plan artifacts, and historical patterns (e.g., "last time we refactored auth, we also had to update these services").

Promotion/demotion behavior:

  • System analyzes current query
  • Promotes relevant data upward (cold → warm → hot)
  • Demotes stale data out of hot to keep prompts tight
  • Preserves complete context snapshots for every decision

This architecture leverages the key insight that software development is inherently local - even in huge codebases, individual changes typically touch a bounded set of files. The Decision Tree captures these localities.

Actor-Specific Context Views

A key missing feature identified in the notes is per-actor context views, filtering, relevance, and actor-aware context limits.

The intended direction:

  • global context exists at plan level,
  • each actor gets a "view" of that context tuned to their role,
  • memory may be shared or per-plan depending on design choices.

Actor Context View Service (Proposed)

A dedicated module/service that:

  • maintains actor-specific "context views,"
  • tracks actor memory and relevance,
  • enforces actor-specific limits (tokens, file types, etc.).

Initial Context vs Deep Context

A practical approach mentioned:

  • Initial context is high-level (repo tree, language, overview),
  • Then the system searches for relevant details iteratively via RAG.

This strongly suggests CleverAgents should define:

  • an "initial context recipe" per project type (codebase vs documents vs infra),
  • iterative context refinement loops during strategize/execute.

Behavior

Automation Levels

Automation levels determine which phase transitions happen automatically.

Automation Level Modes

Mode Behavior
Manual User explicitly triggers: create → use → execute → apply. Every phase transition requires a command. Every decision point pauses for human input. User sees context, alternatives, and recommendation. User provides explicit choice or custom guidance.
Review-before-apply Strategize + Execute happen automatically. AI makes all decisions autonomously. System pauses before Apply to show diff and ask for approval. User can approve, reject, or correct specific decisions.
Full automation System runs all phases automatically end-to-end. AI makes all decisions. Execution proceeds through apply. Human notified of completion. Rollback available if issues detected.

Progressive Trust Building

New users typically follow this progression:

  1. Start with manual mode to understand system behavior
  2. Move to review-before-apply as confidence builds
  3. Enable full automation for specific task types
  4. Gradually expand full automation scope

Semantic Escalation

Even in full automation mode, the system understands when it needs help:

class AutonomyController:
    def assess_decision_confidence(self, decision, context):
        factors = {
            'past_success_rate': self.get_historical_success(decision.type),
            'codebase_familiarity': self.get_familiarity_score(context.project),
            'risk_assessment': self.evaluate_risk(decision),
            'invariant_complexity': self.analyze_invariants(decision)
        }
        
        confidence = self.compute_confidence(factors)
        
        if confidence < self.threshold:
            if self.automation_level == 'full':
                # Even in full automation, critical decisions escalate
                return RequestHumanGuidance(decision, factors)
        
        return ProceedAutonomously(decision)

Automation Level Hierarchy

Automation levels are determined using this precedence (highest to lowest):

  1. Plan-level: Explicitly set when using an action on projects
  2. Session-level: Set for the current session
  3. Global-level: Persisted application configuration
# Set global automation level (persists across sessions)
agents config set automation-level review-before-apply

# Set session automation level (overrides global for this session)
agents session set automation-level full-automation

# Use action with explicit automation level (overrides session and global)
agents plan use local/my-action --project my-proj --automation-level manual

Automation Level Persistence Rules

  1. Global level: Persisted in application configuration. Default is manual on fresh install.

  2. Session level: Lives for the duration of the session. Not persisted.

  3. Plan level: Once a plan's automation level is determined (at the moment of use), it is locked to that plan. Even if the session or global level changes later, the plan retains its original automation level.

  4. Explicit change: A plan's automation level can be changed explicitly after creation:

agents plan set-automation-level <plan_id> full-automation

Subplan Automation Levels

Subplans inherit the parent plan's automation level.

However, if the parent plan's automation level is changed explicitly mid-execution, new subplans will use the new level while already-completed subplans retain their original level.

Granular Automation Flags

For fine-grained control, additional flags can modify behavior:

  • auto_strategize - Automatically proceed from Action to Strategize
  • auto_execute - Automatically proceed from Strategize to Execute
  • auto_apply - Automatically proceed from Execute to Apply
  • auto_retry_on_failure - Automatically retry failed phases

This allows combinations like:

  • Auto strategize, manual execute, manual apply (for risky infra tasks)
  • Auto strategize + execute, manual apply (the review-before-apply pattern)

Validation and Guardrails

Plan generation validation

There is a note that validation logic is stubbed and must be implemented. The spec should require:

  • validate action schema
  • validate actor availability
  • validate required skills exist
  • validate permission policy
  • validate rollback feasibility (if enabled)
  • validate project resource accessibility

This prevents "plan runs with fake providers" and other surprises.

Cost / rate limits

The notes mention future concerns:

  • API call limits
  • cost caps So CleverAgents should define:
  • per-plan budgets
  • per-session budgets
  • per-org budgets
  • per-actor max tool calls / max retries

Correcting Plans (Core Feature)

Correcting plans is where CleverAgents becomes more than "a fancy prompt runner."

The Goal

When a plan makes a wrong decision early, we want to:

  • correct the decision,
  • recompute only the affected subtree,
  • preserve unaffected work.

This is explicitly described: "redo everything below that decision, not the entire code base."

Decision Tree Representation

Every plan records (see Decision Data Model section):

  • decisions (choice points) - created during Strategize
  • dependencies (which later work depended on that decision)
  • child plans spawned because of that decision - populated during Execute
  • artifacts generated under that branch

This makes plan runs auditable and correctable.

Two Correction Modes

  1. Revert-from-history correction (--mode=revert)
  • Find the decision point in the tree
  • Roll back all changes (code and non-code) to that point
  • Re-run from that decision point forward
  • Keep old execution artifacts for comparison
  • Potentially expensive if high up in the tree
  1. Add-at-end correction (--mode=append)
  • Leave history intact
  • Append a new plan at the end that fixes the outcome
  • Cheaper and safer sometimes
  • Does not rewrite history

Correction Flow (Revert Mode)

When user requests correction at Decision B:

  1. Mark for Correction

    Decision B.superseded_by = new_decision_id
    
  2. Identify Downstream Impact

    • Recursively collect all decisions that depend on Decision B
    • Collect all subplans spawned from those decisions
    • These form the "affected subtree"
  3. Rollback Resources

    • For each affected decision's artifacts_produced:
      • Rollback to the checkpoint before that artifact was created
    • For affected subplans:
      • Rollback their sandboxes entirely
  4. Preserve for Comparison

    • Archive the original subtree's artifacts
    • Create a CorrectionAttempt record linking old and new
  5. Re-execute from Decision Point

    • Restore context to Decision B.context_snapshot
    • User provides new guidance/correction
    • Re-run execution from that point
    • New decisions get is_correction: true, corrects_decision_id: B
  6. Apply as Normal

    • The corrected plan goes through normal Apply gating
    • Diff shows changes from the correction

History Cleanup

History can only be flagged for cleanup after a plan is Applied.

Once a plan is applied:

  • It can no longer be rolled back
  • Old correction artifacts can be archived or deleted based on retention policy
  • The decision tree is preserved for audit purposes

CLI Commands for Correction

# View decision tree
agents plan tree <plan_id>
agents plan tree <plan_id> --format=json  # For visualization tools

# Inspect a specific decision
agents plan explain <decision_id>
# Shows: question, chosen option, alternatives, rationale, downstream impact

# Correct via revert-and-replay
agents plan correct <decision_id> --mode=revert --guidance "<what the decision should be>"
# Re-executes from that point with the new guidance

# Correct via append (add fix at end)  
agents plan correct <decision_id> --mode=append --guidance "<description of the fix>"
# Creates a new subplan to fix the outcome without rewriting history

# For long guidance text, use a file
agents plan correct <decision_id> --mode=revert --guidance-file ./correction.txt

# Compare old vs new after correction
agents plan diff <correction_attempt_id>

Correction Safety

Corrections always:

  • Create a new attempt revision (increment plan.attempt)
  • Preserve old artifacts for diff/compare
  • Run execute in sandbox again
  • Require apply gating again
  • Never modify already-applied changes

This keeps history reproducible and prevents accidental destructive edits.

Human-in-the-Loop Collaboration

Even though the direction is "more autonomous," the transcript explicitly recognizes that real workflows require engineers to collaborate with the system, editing code while it works, and using better UX integration (TUI/web/IDE).

So ClverAgents should aim for:

  • visibility: what is it doing now?
  • interruptibility: pause/cancel/retry
  • editability: allow user to modify strategy before execute
  • reconciliation: detect if user changed sandbox files mid-run and handle it

UI / Interaction Model

CLI-first + TUI + Web App + IDE

The system intends to be CLI-first, with:

  • a TUI built using Textual,
  • which can generate a web app "for free,"
  • and later an IDE plugin that embeds the TUI in the IDE.

This implies a "single UI codebase" model:

  • same underlying view logic,
  • multiple frontends.

Plan Tree Visualization

The TUI should show:

  • plan list
  • plan details
  • plan tree (ASCII)
  • diff view
  • approvals

And should later allow exporting the tree as image (PNG) or JSON for other visualization tools.

Additional Recommended Sections (Missing from Draft but Important)

Storage and Persistence

CleverAgents should define where each concept lives:

  • Actions: stored in a registry (local files or server DB)
  • Actors: stored similarly (config files + DB indexing)
  • Projects: stored locally or on server
  • Plans: stored in plan DB with full logs
  • Context indexes: vector store / graph store / SQLite
  • Artifacts: filesystem or object storage

Observability

To debug large plans:

  • every phase should emit events
  • every actor call should log prompt/context references
  • every skill call should log resource access
  • every checkpoint should be recorded

Security Model

  • sandbox isolation
  • resource-level ACLs
  • prompt injection mitigations (server mode)
  • secret management (API keys, DB credentials)
  • audit logs for apply

Extensibility

  • plugin system for skills
  • custom node types
  • action templates
  • actor templates (noted as missing currently)

Summary of Key Intended Behaviors (If You Only Read One Section)

  • Plans follow Action → Strategize → Execute → Apply with automation levels controlling transitions.
  • Strategize is read-only and produces a strategy + blueprint.
  • Execute happens in a sandbox, can spawn subplans, and should support checkpoints/rollback when enabled.
  • Apply commits changes from sandbox to real project after review/validation.
  • Actors are hierarchical: an actor can be a single agent or an entire graph.
  • Graph nodes can be actors, MCP skills, or custom tool nodes.
  • Context should evolve toward hot/warm/cold tiers and actor-specific context views.
  • The system is designed for large tasks where the user can correct a decision and only recompute downstream work, visualizable as a plan decision tree.

The system can handle Firefox-scale projects not through magic, but through:

  • Hierarchical decomposition breaking massive tasks into bounded work
  • Persistent decision graphs maintaining context across any scale
  • Isolated execution preventing cascading failures
  • Semantic validation catching errors before propagation
  • Progressive automation building trust through incremental success

If you want, I can also produce:

  • a canonical JSON/YAML schema for Actions, Actors, Projects, Plans, Skills, and Context Views,
  • a CLI command reference (every command, flags, examples),
  • and a set of end-to-end example workflows (single project, multi-project, infra task, paper-writing task) consistent with this spec.

Work Remaining to Make CleverAgents Fully Functional

This section describes—exhaustively and in implementation terms—what remains to be done to bring the current CleverAgents codebase up to the intended CleverAgents behavior described in the transcript/meeting notes (Action → Strategize → Execute → Apply; actors as composable graphs; sandbox + diff review; context scaling; checkpointable skills; etc.).

It is based primarily on the attached "what's missing / code analysis" document (last updated Jan 29, 2026) which outlines the current master branch's functional gaps, architectural limitations, and regressions introduced by new reactive/langgraph code.


Current State vs Intended: The Gap in One Sentence

Right now, the codebase still behaves like a linear, single-file "LLM dump to generated.py" pipeline (with tiny context and stubbed validation), while intended CleverAgents requires a plan lifecycle engine that can reliably generate, validate, sandbox, diff-review, and apply multi-file/multi-project changes—backed by checkpointable skills, rich context, and composable actor graphs.

Everything below is the concrete work required to close that gap.


1) Restore the Core Capability: MultiFile Change Generation (Hard Blocker)

Problem (today)

Per the code analysis, the current plan generation workflow is hard-coded to emit exactly one Change and write one file containing raw LLM output (often including markdown + explanations). This alone prevents "repo writing," multi-service changes, scaffolding, and most realistic coding tasks.

What "done" looks like

The system must be able to produce N changes across N files, where each change is one of:

  • create (new file)
  • modify (update an existing file)
  • delete
  • rename/move (path change + optional content change)

…and do so reliably, with predictable structure and minimal ambiguity.

High-level implementation plan

A. Upgrade the internal "change set" model (do this first)

Create a canonical internal representation for generated output, e.g.:

  • ChangeSet

    • changes: list[Change]
    • warnings: list[str]
    • generated_by_actor: ActorRef
    • validation: ValidationResult
    • source_prompt_snapshot: PromptSnapshot (for audit/debug)

Where Change includes:

  • operation (enum)
  • path (required except delete/move cases)
  • new_path (for move)
  • content (full content for create, optional for modify if using patches)
  • patch (optional unified diff)
  • language (optional but helpful for validation)

Key requirement: the pipeline must stop treating a plan as "one output string"; it must treat the plan as "a structured set of resource ops."

B. Implement Tool-Based Resource Modification (Modern Architecture)

CRITICAL: Do NOT parse LLM output to extract code. Instead, use the modern tool-based approach:

LLM → calls tools/skills directly → tools modify sandbox → ChangeSet built from tool invocations

Implementation details:

  1. Provide built-in resource skills that LLMs call directly:

    • read_file(path) - Read file contents
    • write_file(path, content) - Create/overwrite file
    • edit_file(path, edits) - Apply targeted edits (search/replace or line-based)
    • delete_file(path) - Remove file
    • move_file(src, dst) - Rename/move file
    • create_directory(path) - Create directory
    • list_files(pattern) - List files matching glob
    • search_files(pattern) - Search file contents
  2. Each skill invocation that modifies resources creates a Change record:

    # When write_file("src/main.py", content) is called:
    change = Change(
        operation=OperationType.CREATE,
        path="src/main.py",
        content=content,
        language="python"
    )
    changeset.add_change(change)
    
  3. The ChangeSet is the accumulated history of resource-modifying skill calls, NOT parsed LLM text.

Why tool-based is superior to parsing:

Aspect Parsing Approach Tool-Based Approach
Ambiguity "Is this code or explanation?" Each operation is explicit
Atomicity Parse entire response Each tool call is discrete
Rollback Reconstruct from diff Replay inverse of recorded changes
Resource types Files only Any resource (files, DBs, APIs)
Audit trail What model said What actually happened

C. Support MCP Tools and Custom Skills

Skills can come from multiple sources, all producing the same Change records:

  1. Built-in skills: Core file operations, git, search
  2. MCP servers: External tools wrapped with sandbox interception
  3. Custom inline skills: Python code in actor YAML

All skills operate through the unified resource abstraction layer:

  • Paths are resolved to sandboxed locations
  • Write operations create Change records
  • Capability metadata enforces safety restrictions

D. Implement directory creation and scaffolding rules

Skills automatically handle directory operations:

  • Auto-create parent directories for new file paths (safe, idempotent)
  • Enforce project root boundaries (never escape the sandbox)
  • Enforce deny-list patterns (.git/, node_modules/, etc.)
  • Respect project .gitignore and custom exclusion rules

This directly addresses the "create REST API" expectation (routes/, models/, etc.) that current code cannot meet.


2) Connect the New Reactive/LangGraph Infrastructure to the Actual Plan Workflow (Hard Blocker)

Problem (today)

The Jan 29 update introduced substantial new code under src/cleveragents/langgraph/ and src/cleveragents/reactive/ (graphs, nodes, state, RxPy routing), but the analysis states it is not connected to the main plan generation workflow and therefore does not fix the fundamental functional issues.

What "done" looks like

  • A plan's Strategize and Execute phases run through a real graph engine (LangGraph-like execution).

  • The graph actually drives:

    • context selection,
    • strategy creation,
    • code generation,
    • validation,
    • sandbox writes,
    • diff production,
    • apply actions.

High-level implementation plan

A. Create a single "PlanLifecycleGraph" entrypoint

Define a top-level orchestrator graph per plan instance:

Strategize subgraph

  • gather context (read-only)
  • produce strategy (machine-usable blueprint + narrative)
  • produce an "execution plan" (tasks/subplans)

Execute subgraph

  • allocate sandbox
  • execute tasks (spawn subplans where needed)
  • produce ChangeSet + artifacts + logs
  • run validation gates
  • produce diff view summary

Apply subgraph

  • apply changes transactionally
  • record audit log + final state

B. Replace (or wrap) the current "tell/build/apply/file" pipeline

Maintain backward compatibility at the CLI layer (for now), but internally:

  • tell creates/updates an Action (or a plan request)
  • build maps to Strategize + Execute
  • apply maps to Apply

The key is to stop having "build" be "one LLM call that returns one file."

C. Make the new reactive routing optional—not mandatory

The reactive router (RxPy streams) is powerful, but don't make it the required execution substrate until it's stable and secure (see security section below). Use it as:

  • an experimental routing layer,
  • or a message bus,
  • or a future optimization for streaming UX.

But the "plan must work" path should be stable with deterministic orchestration.


3) Implement Real Sandbox + Diff Review + Transactional Apply (Hard Blocker)

Problem (today)

The analysis notes that apply currently writes files directly and can leave the codebase inconsistent on partial failures; diff preview is missing or insufficient; rollback is not robust.

What "done" looks like (expectation)

  • Execute never mutates the real project directly.
  • All changes happen in a sandbox.
  • User (or automation policy) reviews diffs.
  • Apply commits changes atomically (all-or-nothing) to real project resources.

High-level implementation plan

A. Standardize sandbox backends (choose one as default per project type)

For code repos, default should usually be git-based:

  • git worktree or temporary branch
  • sandbox has its own working directory
  • diff is native git diff

For non-git projects:

  • filesystem copy or overlay sandbox

Each project declares:

  • preferred sandbox backend
  • whether sandbox is mandatory

B. Apply must be transactional

Transactional apply means:

  • if any step fails (permission error, merge conflict, write failure), the system:

    • aborts apply,
    • reverts partial writes,
    • preserves sandbox for inspection,
    • provides a resumable/retry path.

Implementation details:

  • In git mode: apply is a merge/cherry-pick of sandbox commits (atomic-ish).
  • In filesystem mode: write to temp files, then rename swap, then finalize (atomic within filesystem constraints).

C. Diff review must become a first-class artifact

A plan should store:

  • file list changed
  • per-file diff
  • summary (what changed + why)
  • "risk markers" (e.g., touched auth code, migrations, infra)

This is a core differentiator and must be polished because it is the human safety boundary.


4) Build the "Checkpointable Skills" Layer (Hard Blocker for Reliable Autonomy)

Problem (today)

The system lacks robust rollback, and the analysis emphasizes the need for checkpoints to avoid partial failures and inconsistent states. The meeting notes also highlight the intent to build checkpointable skills because MCP metadata is insufficient for write scope and rollback guarantees.

What "done" looks like

  • Skills declare whether they are checkpointable.

  • The execution engine can:

    • create a checkpoint ID,
    • perform actions,
    • roll back to checkpoint.

High-level implementation plan

A. Define a checkpoint contract at the skill interface level

Every skill should optionally support:

  • checkpoint() -> CheckpointId
  • rollback(checkpoint_id) -> None
  • describe_effects(call) -> ResourceEffects (what it read/wrote)

Not every skill can implement rollback. But for "write" skills you rely on, you need either:

  • skill-level rollback, or
  • containment (sandbox) + external rollback (git reset, snapshot restore).

B. Implement checkpointable "core skills" first

Start with:

  • file skill (write with pre-write snapshots)
  • git skill (commit/stash checkpoints)
  • shell skill inside sandbox container (filesystem snapshot or git-backed)

This provides enough reliability to support multi-step Execute.

C. Add a skill metadata registry (even if MCP is used)

The system needs to know:

  • read-only vs write
  • write scope (paths/resources)
  • idempotency
  • checkpointability
  • side effects

The meeting notes explicitly call out that MCP doesn't provide enough writing-scope metadata for safe automation; CleverAgents needs an internal extension/registry.


5) Expand Context Beyond "300 chars × 5 files" and Make It Actor-Aware (Hard Blocker for Real Code Understanding)

Problem (today)

The analysis states the context system is effectively a placeholder:

  • only a small preview of a few files,
  • no robust repo understanding,
  • inadequate for realistic tasks.

What "done" looks like

  • The system can ingest and reason over large repos via indexing and retrieval.
  • Context is not a single global blob; it is assembled per actor/node based on role and task.
  • Users can pin/force include critical resources.

High-level implementation plan

A. Implement repo indexing + retrieval

Minimum viable:

  • file tree + language detection
  • full-text search
  • embedding index (optional but recommended)
  • ignore patterns applied consistently

B. Adopt the tiered context architecture (hot/warm/cold)

As described in the meeting notes:

  • hot cache = prompt-injected snippets
  • warm = indexes/embeddings/graph storage
  • cold = historical artifacts, summaries, prior plan outputs

Then implement automatic promote/demote.

C. Add "Actor Context Views"

Actors require tailored views:

  • strategist sees architecture docs, READMEs, module boundaries, dependency graphs
  • executor sees precise code sections relevant to edits
  • reviewer sees diffs + tests + risk zones

This actor-aware context view system is explicitly described as missing today and a planned improvement.


6) Fix Provider/LLM Plumbing So It Cannot Fail Silently (Hard Blocker)

Problem (today)

The analysis identifies multiple provider-level failures:

  • defaulting to a FakeListLLM that returns hardcoded strings,
  • returning None provider silently,
  • hardcoding auto-debug to OpenAI GPT-4 regardless of configuration,
  • OpenRouter listed but not implemented.

What "done" looks like

  • If no provider is configured, the user gets a clear, actionable error.
  • Auto-debug uses the configured actor/provider.
  • Providers are discoverable, consistent, and testable.

High-level implementation plan

A. Remove FakeListLLM as default behavior

  • Keep FakeListLLM only for tests.
  • In production, enforce: "no provider configured → fail fast with clear message."

B. Make provider selection actor-driven end-to-end

Actor config should determine:

  • provider
  • model
  • tool access
  • safety settings

Auto-debug should be:

  • a plan/action using an actor, not a hard-coded special-case.

C. Implement "provider auto-support" strategy

To avoid manual churn:

  • derive supported providers from LangChain/LangGraph provider availability
  • or implement a provider adapter registry that can load provider plugins dynamically

7) Replace Stubbed Validation With Real Validation Gates (Hard Blocker for Quality)

Problem (today)

Validation is currently a stub ("PASS" or output length > 10). This guarantees that broken outputs will be considered valid.

What "done" looks like

Validation in Execute must include:

  • syntax checks
  • lint/format checks (optional)
  • unit tests (when available)
  • build checks (typecheck, compile)
  • security sanity checks (optional)
  • schema validation for structured outputs

High-level implementation plan

Create a ValidationPipeline that can run per project type:

Python project

  • python -m py_compile for changed files
  • ruff / black (optional)
  • pytest (optional)

Node project

  • npm test, tsc, etc.

Infra

  • terraform validate, etc.

Then enforce:

  • Execute phase cannot complete if validation fails (unless user overrides explicitly).

Also fix the "invalid python docstring wrap" behavior:

  • invalid output should trigger a repair loop or be quarantined as an artifact,
  • not silently converted into useless code.

8) Make Apply/Execute Safe Under Failure, Concurrency, and Resume (Stability Blockers)

Problems (today)

The analysis calls out:

  • partial failure leaves repo inconsistent,
  • no concurrency protection (two builds can race),
  • no resume/progress persistence,
  • missing cleanup for abandoned plans.

What "done" looks like

  • Plans can be resumed after interruption.
  • Plan execution is locked per project/sandbox to avoid races.
  • Abandoned sandboxes are cleaned up safely.
  • The system can report progress reliably.

High-level implementation plan

  • Add a plan lock (DB row lock or filesystem lock).

  • Persist step-level progress events.

  • Implement resumable execution by checkpoint:

    • "resume from checkpoint X" or "resume from step Y"
  • Add lifecycle cleanup jobs:

    • garbage collect old sandboxes
    • expire old checkpoint files (see also new LangGraph checkpoint file leak)

9) Address Critical Security Issues Introduced by New Reactive Code (Security Blocker)

Problems (today)

The analysis reports a critical eval() vulnerability in stream routing configuration (remote code execution vector), plus additional concerns:

  • silent exception swallowing,
  • event loop leaks,
  • thread safety issues,
  • template injection risks.

What "done" looks like

  • No configuration-driven arbitrary code execution (no eval).
  • Errors are surfaced, not swallowed.
  • Long-running systems do not leak tasks/subscriptions/checkpoint files.
  • Template rendering is safe.

High-level implementation plan

A. Remove eval() from any config parsing path

Replace with one of:

  • a whitelist of allowed transform operators
  • a small safe expression language (parsed, not executed)
  • "transform" must reference a named function from a registry, not arbitrary text

B. Stop swallowing exceptions

In routing/execution:

  • capture exception
  • attach it to message metadata
  • fail the stream/plan with clear error state

C. Fix async lifecycle correctness

  • do not create new event loops without closing
  • prefer asyncio.run or properly managed loop policy in one place
  • ensure Rx subscriptions are disposed on shutdown and on stream reconfiguration

D. Replace str.format templating with a safe template engine

If templating is needed, use a sandboxed template engine or restrict tokens severely.


10) Finish the "Plan System" (Actions, Strategize, Execute, Apply) in Storage + CLI + UI

Problem (today)

The analysis describes the current master as still conceptually the linear tell/build/apply pipeline. Meanwhile, intended CleverAgents introduces reusable Actions decoupled from Projects, and separate Strategize/Execute/Apply phases with automation levels.

What "done" looks like

  • Actions exist as reusable templates.
  • Using an Action creates a plan in Strategize.
  • Execute operates on a sandbox.
  • Apply commits.
  • Plan tree / decision tree is recorded.

High-level implementation plan

A. Data model changes

Add/ensure tables for:

  • actions (templates)
  • plans (instances with phase + state)
  • plan_steps or plan_nodes (for decision tree/subplans)
  • plan_artifacts
  • plan_events (structured log)

B. CLI changes

Implement commands aligned to the lifecycle:

  • agents create action ...
  • agents use <action> --project <p>
  • agents execute <plan>
  • agents apply <plan>

Keep aliases for old commands for usability, but map them internally.

C. TUI changes (once data exists)

  • show plan tree (ASCII first)
  • show phase transitions and states
  • show diffs in Execute output
  • show approvals for Apply

This is explicitly a near-term intended UX: visual decision tree + correction ability.


11) Actor System Completion: Make Actors Actually Define Behavior (Not Just Model Selection)

Problem (today)

Meeting notes and analysis agree that actor CRUD exists or is being ported, but key pieces are missing:

  • graph descriptor not used,
  • no per-actor tool access policies,
  • no per-actor memory,
  • no actor-defined orchestration behavior in practice.

What "done" looks like

An actor should be able to define:

  • model/provider,
  • system prompt,
  • tool/skill access policy,
  • memory policy,
  • graph descriptor (if actor is a composite workflow),
  • context view policy.

High-level implementation plan

  • Treat actor config as the single source of truth for behavior.

  • If graph_descriptor exists:

    • instantiate a LangGraph runtime graph using it,
    • nodes may reference other actors or skills.
  • Add a policy layer:

    • this actor can only call read-only skills in Strategize,
    • this actor can call write skills in Execute only within sandbox.

12) Persistence + DB Hygiene (Stability Blocker)

Problems (today)

The analysis lists:

  • database directory creation issues (first run failures),
  • missing/messy migrations,
  • global mutable provider registry state causing test/config/thread issues,
  • memory service defaults to in-memory and loses history.

High-level implementation plan

  • Ensure .cleveragents/ created before DB use.
  • Require migrations to exist and be applied (Alembic).
  • Remove global registry singletons; make DI container own state.
  • Make session memory persistent by default (SQLite is acceptable).
  • Store session IDs in DB; never "generate and forget."

13) Cost Controls, Rate Limits, and Provider Fallback (Production Readiness)

Problem (today)

The analysis notes there are no cost/rate controls, and the meeting notes acknowledge this is required for practical full automation.

High-level implementation plan

  • Track token usage per:

    • plan,
    • phase,
    • actor,
    • session,
    • org (future server mode)
  • Add configurable budgets and caps:

    • max cost per plan
    • max calls per minute
    • max retries per node
  • Implement provider fallback policy:

    • if provider A fails transiently, retry or fail over to provider B
    • preserve determinism where required (e.g., pin for Apply reviews)

14) Server / Multi-User Readiness (Not Required for Local MVP, But Must Be Designed In)

Problem (today)

The meeting notes say multi-user prompt injection and governance are less relevant for single-user local mode, but become critical for server mode.

High-level implementation plan

  • Add permissions at:

    • namespace level (org actions)
    • project level (apply rights)
    • skill level (dangerous operations)
  • Add audit logs for Apply.

  • Add prompt-safety hardening:

    • treat user input as data, not instruction overrides
    • strict templating and role separation
  • Separate "execution happens locally" from "coordination happens on server" (optional architecture).


Suggested Build Order (So You Don't Get Stuck)

If you want a pragmatic sequencing that preserves momentum and prevents rewrites:

  1. Multi-file ChangeSet + structured output parsing (unblocks everything)
  2. Sandbox + diff + transactional apply
  3. Wire new graph engine into plan workflow
  4. Checkpointable skills + rollback
  5. Real validation pipeline
  6. Context indexing + actor context views
  7. Provider + memory persistence fixes
  8. Security hardening of reactive/router code
  9. TUI plan tree + correction UX
  10. Cost controls + server-mode foundations

This ordering matches the reality that without multi-file output and safe apply, higher-level orchestration and context improvements won't matter—the system still can't "do the work" safely.


Acceptance Criteria for "Fully Functional" (Concrete)

A realistic "fully functional" bar (aligned to intended behavior) is:

  • Can strategize read-only using real context across many files.
  • Can execute in sandbox, producing a multi-file ChangeSet for non-trivial features.
  • Can run validation and fail safely without corrupting project state.
  • Can show a diff review and apply atomically.
  • Can roll back to a checkpoint during Execute (at least for file/git/shell-in-sandbox).
  • Can run multi-project plans without mixing resources.
  • Does not silently fall back to fake providers or swallow errors.
  • Has no config-driven RCE paths (no eval).
  • Has persistent session/memory so iterative CLI use works.

Everything listed earlier maps directly to closing the gaps documented in the attached analysis and to achieving the behaviors described in the transcript.

CleverAgents Architecture FAQ

Q: How does CleverAgents handle persistent repository knowledge beyond ephemeral context windows?

What exists today architecturally: The specification defines a sophisticated multi-tier memory system that goes far beyond ephemeral context windows. At its core is the Decision Tree structure which provides a durable, queryable record of every choice made during planning, along with the complete context that informed those choices.

How the persistent model works in practice:

When a strategy actor analyzes a codebase during the Strategize phase, it doesn't just make decisions in isolation. Each decision creates a comprehensive Decision record that includes:

context_snapshot:
  hot_context_hash: str              # Cryptographic hash of the exact context
  hot_context_ref: str               # Pointer to the full stored snapshot
  relevant_resources: list[ResourceRef]  # Every file/symbol that influenced this decision
  actor_state_ref: str               # Complete LangGraph checkpoint

This means when the system decides "refactor the authentication module to use async patterns," it permanently records:

  • Which files were examined to make that decision
  • What symbols and dependencies were traced
  • The exact code state that was analyzed
  • The reasoning chain that led to this choice
  • Alternative approaches that were considered but rejected

The three-tier memory architecture enables scale:

  1. Hot tier: Immediate working context (what's in the current LLM context window)
  2. Warm tier: Recent decisions and their contexts from this plan tree - quickly accessible
  3. Cold tier: Historical decisions from past plans on this codebase - queryable but not in active memory

When working on a 50,000 file codebase, the system doesn't need to hold all files in memory. Instead:

  • Hot context focuses on the immediate task (e.g., 10-20 files for a specific refactoring)
  • Warm context maintains the decision chain that got us here
  • Cold context provides historical patterns ("last time we refactored auth, we also had to update these services")

Why this scales to massive codebases:

The key insight is that software development is inherently local - even in huge codebases, individual changes typically touch a bounded set of files. The Decision Tree captures these localities. When converting Firefox to Rust (your example), the system would:

  1. Make high-level architectural decisions (captured as root decision nodes)
  2. Decompose into major subsystem conversions (each a decision spawning subplans)
  3. Each subsystem plan makes decisions about its modules
  4. Module plans make decisions about individual files

At each level, only the relevant context is loaded. The persistent decision graph means we can always reconstruct why we're converting a particular module and what constraints apply from higher-level decisions.

Concrete example of persistence in action:

Plan: Convert Firefox Renderer to Rust
├── [Decision] Architecture approach: Start with leaf modules, work inward
│   Context: Analyzed module dependency graph, 2,847 modules total
│   Resources: module_graph.json, architecture_docs.md
│   
├── [Decision] Phase 1: Convert utility libraries (no external deps)
│   └── [Subplan] Convert string_utils module
│       ├── [Decision] Use Rust's String type, not custom implementation
│       │   Context: Analyzed 47 string_utils.cpp functions
│       │   Resources: string_utils.cpp, string_utils.h, 12 dependent files
│       │   Rationale: Rust's String provides same guarantees with better ergonomics

Even months later, we can query: "Why did we use Rust's String type?" and get the exact context and reasoning, without reprocessing the entire codebase.

Q: How does the system compute task-specific dependency closures for large-scale operations?

What exists today architecturally: The specification defines multiple mechanisms for computing and maintaining minimal dependency closures. The execution blueprint produced during the Strategize phase doesn't just list steps - it includes a complete dependency graph with explicit scoping for each operation.

How dependency closure computation works:

During the Strategize phase, the strategy actor employs several mechanisms to compute precise dependency closures:

  1. Resource-aware analysis: The actor uses specialized skills to trace dependencies:

    # Pseudocode of what happens inside a strategy actor
    def compute_closure_for_refactoring(target_module):
        closure = ResourceClosure()
    
        # Direct file dependencies
        closure.add_files(find_imports(target_module))
        closure.add_files(find_includes(target_module))
    
        # Symbol dependencies
        for symbol in extract_exported_symbols(target_module):
            closure.add_files(find_symbol_usage(symbol, scope='project'))
    
        # Test dependencies
        closure.add_files(find_tests_for_module(target_module))
    
        # Build system dependencies
        closure.add_files(find_build_references(target_module))
    
        return closure
    
  2. Hierarchical scoping: When spawning subplans, each subplan receives:

    • An explicit relevant_resources list
    • A sandbox_strategy appropriate for those resources
    • Clear boundaries of what it can and cannot modify
  3. Decision-based tracking: Each subplan_spawn decision records:

    decision_type: subplan_spawn
    chosen_option: "Refactor authentication module"
    downstream_plan_ids: ["plan-auth-refactor-123"]
    artifacts_produced: 
      - auth_module_files: ["auth.rs", "auth_test.rs", "auth_types.rs"]
      - api_updates: ["api/v2/login.rs", "api/v2/logout.rs"]
    

Concrete example - Converting a subsystem to Rust:

Let's trace how the system handles "Convert Firefox's Network Stack to Rust":

STRATEGIZE PHASE:
1. Analyze network stack structure
   - Identifies 847 C++ files in netwerk/ directory
   - Traces public API surface (237 exported functions)
   - Maps internal dependencies (1,432 internal calls)
   
2. Compute minimal closure for Phase 1 (DNS resolver):
   - Core files: dns_resolver.cpp, dns_cache.cpp, dns_config.cpp (3 files)
   - Direct dependencies: 12 files in netwerk/base/
   - Test files: 8 test files specific to DNS
   - Build files: 2 moz.build files
   - Total closure: 25 files (not 847!)
   
3. Generate execution blueprint with subplans:
   - convert-dns-types: Closure of 5 files (type definitions)
   - convert-dns-cache: Closure of 8 files (cache + tests)  
   - convert-dns-resolver: Closure of 12 files (resolver + integration)

Why this is tractable even for massive codebases:

The system leverages several key insights about real software:

  1. Modular boundaries exist: Even in legacy codebases, there are natural boundaries
  2. Changes are incremental: We don't convert 50,000 files atomically
  3. Dependencies are sparse: Most modules depend on a small fraction of the codebase
  4. Interfaces are narrow: Public APIs are much smaller than implementations

The Firefox example would decompose into ~1,000 bounded subplans, each touching 10-100 files. The parent plan tracks the overall architecture, while each subplan maintains its focused closure.

How we prevent closure explosion:

  • Lazy expansion: Dependencies are traced only as deep as needed for correctness
  • Interface-based boundaries: When possible, work against stable interfaces
  • Incremental validation: Each subplan validates its changes don't break dependents
  • Hierarchical merge strategies: Parent plans resolve conflicts between subplan changes

Q: What mechanisms enforce global consistency during parallel execution across many files?

What exists today architecturally: The sandbox model combined with hierarchical plan execution provides strong guarantees about consistency during parallel execution. This isn't just process isolation - it's semantic isolation with intelligent merge strategies.

How the coordination mechanism prevents compound errors:

  1. Complete isolation during execution: Each plan executes in its own sandbox, which means:

    Plan A (refactoring auth module):
    - Sandbox A1: Contains only auth/*.cpp, auth_tests/*.cpp
    - Cannot see Plan B's intermediate states
    - Cannot accidentally depend on Plan B's half-done work
    
    Plan B (updating API endpoints):
    - Sandbox B1: Contains only api/*.cpp, api_tests/*.cpp  
    - Makes changes assuming current auth interface
    - Protected from Plan A's intermediate refactoring
    
  2. Resource-specific sandbox strategies provide natural coordination:

    Git repositories:
      - Strategy: git worktrees
      - Coordination: Git's three-way merge algorithm
      - Conflict detection: Built into Git
      - Rollback: git reset/checkout
    
    Databases:
      - Strategy: Transaction isolation  
      - Coordination: MVCC (multi-version concurrency control)
      - Conflict detection: Serialization failures
      - Rollback: Transaction abort
    
    Cloud Infrastructure:
      - Strategy: Terraform workspaces
      - Coordination: State locking
      - Conflict detection: Resource conflicts in plan
      - Rollback: Previous state restoration
    
  3. Hierarchical merge resolution: When subplans complete, the parent plan performs intelligent merging:

    def merge_subplan_results(subplan_results):
        # Group by resource type
        by_resource = group_by_resource_type(subplan_results)
    
        # Apply resource-specific merge strategies
        for resource_type, changes in by_resource:
            if resource_type == 'git_repo':
                merge_git_changes(changes)  # Three-way merge
            elif resource_type == 'database':
                merge_db_changes(changes)   # Sequential application
            elif resource_type == 'config_files':
                merge_config_changes(changes)  # Smart JSON/YAML merge
    
        # Validate merged state
        run_integration_tests()
    

Concrete example - Preventing cascading failures:

Consider refactoring a shared authentication library used by 15 services:

PARALLEL EXECUTION WITHOUT COORDINATION (what we prevent):
- Service A refactors to async auth → breaks Service B
- Service B compensates with workaround → breaks Service C  
- Service C changes error handling → breaks Services D, E, F
- Cascade of failures!

CLEVERAGENTS COORDINATED EXECUTION:
Parent Plan: Refactor auth library
├── Subplan 1: Update auth library interface
│   Sandbox: Only auth library files
│   Output: New interface definition
│   
├── Barrier: Wait for Subplan 1 completion
│   
├── Parallel Subplans 2-16: Update each service
│   Each sandbox: Only that service's files
│   Each uses: New interface from Subplan 1
│   No inter-service dependencies during execution
│   
└── Merge Phase:
    - Collect all service updates
    - Apply to main branch in order
    - Run integration tests
    - If conflicts: Parent plan resolves using semantic understanding

Advanced coordination patterns:

  1. Optimistic concurrency with semantic conflict resolution:

    Two subplans both modify api/user.rs:
    - Plan A: Adds async fn get_user_profile()
    - Plan B: Adds fn validate_user_permissions()
    
    Merge strategy:
    - Git merge succeeds (different functions)
    - Semantic validation ensures both functions work together
    - Parent plan adds integration glue if needed
    
  2. Checkpoint-based coordination:

    Execution timeline:
    T1: Subplan A creates checkpoint before major refactor
    T2: Subplan B creates checkpoint before API changes
    T3: Subplan A encounters error, rolls back to T1
    T4: Subplan B completes successfully
    T5: Subplan A retries with knowledge of B's success
    
  3. Resource locking for critical sections:

    When modifying shared schema files:
    - Acquire exclusive lock on schema resources
    - Make changes atomically
    - Release lock with new version
    - Other plans rebase on new schema
    

Q: How does the system proactively prevent semantic errors before they propagate?

What exists today architecturally: The specification defines multiple layers of proactive error prevention that go far beyond traditional testing. This is a comprehensive defense-in-depth approach that catches semantic errors before they can propagate.

Layer 1: Decision-time validation during Strategize:

Every decision includes semantic validation:

Decision: Refactor payment module to async
alternatives_considered:
  - "Convert to async/await patterns" (chosen)
  - "Use thread pool with channels" (rejected: doesn't integrate with async ecosystem)
  - "Keep synchronous with timeout" (rejected: doesn't solve core latency issue)
confidence_score: 0.85
validation_performed:
  - Checked all payment API consumers can handle async
  - Verified database driver supports async operations
  - Confirmed no regulatory requirement for sync processing

Layer 2: Execution-time semantic guards:

The execution actor configuration includes validation nodes that understand semantics:

actors:
  code_executor:
    type: graph
    nodes:
      - name: semantic_validator
        type: tool
        config:
          tools:
            - name: validate_api_compatibility
              code: |
                # Not just syntax checking - semantic validation
                old_api = extract_api_signature(previous_version)
                new_api = extract_api_signature(current_version)
                
                breaking_changes = find_breaking_changes(old_api, new_api)
                if breaking_changes:
                    # Don't just fail - understand the impact
                    affected_consumers = find_api_consumers(breaking_changes)
                    migration_plan = generate_migration(breaking_changes)
                    
                    if can_auto_migrate(affected_consumers, migration_plan):
                        apply_migration(migration_plan)
                    else:
                        raise SemanticError(
                            "Breaking API changes require manual review",
                            changes=breaking_changes,
                            affected=affected_consumers
                        )

Layer 3: Invariant enforcement through the type system:

# The system maintains semantic invariants
class RefactoringInvariants:
    # User-defined invariants for the codebase
    invariants = [
        "All public APIs must maintain backward compatibility",
        "Database transactions must complete within 5 seconds",
        "Authentication must always use OAuth2",
        "Payment processing must be idempotent"
    ]
    
    def check_invariant_preservation(self, changes):
        for invariant in self.invariants:
            if not self.verify_invariant(invariant, changes):
                return InvariantViolation(invariant, changes)
        return Success()

Layer 4: Predictive error prevention through pattern matching:

The system learns from past failures:

Error Pattern Database:
  - pattern: "Async conversion in payment module"
    historical_failures:
      - "Race condition in payment confirmation"
      - "Timeout handling breaks idempotency"
    preventive_checks:
      - "Add explicit transaction boundaries"
      - "Verify idempotency keys are preserved"
      - "Check distributed lock acquisition"

Concrete example - Preventing a subtle distributed systems bug:

Scenario: Refactoring a service to use event sourcing:

PROACTIVE CONTAINMENT IN ACTION:

1. Strategy Phase Semantic Analysis:
   - Decision: "Convert order service to event sourcing"
   - Semantic check: "Event sourcing requires eventual consistency"
   - Identifies: 3 services assume immediate consistency
   - Adds decision: "Update dependent services for eventual consistency"

2. Execution Phase Invariant Checking:
   - Detects: PaymentService.chargeCard() called after OrderCreated event
   - Semantic issue: Payment before order confirmation violates business rules
   - Automatic fix: Insert OrderConfirmed event requirement
   
3. Validation Node Catches Edge Case:
   - Discovers: Audit service expects synchronous order numbers
   - Impact: Async events break compliance reporting
   - Resolution: Add audit event buffer with guaranteed ordering
   
4. Pre-Apply Semantic Verification:
   - Simulates production event flow
   - Detects: Under high load, events can arrive out of order
   - Adds: Event ordering guarantees via vector clocks

Why this prevents issues that traditional testing misses:

Traditional tests check "does the code work?" Our semantic containment asks:

  • Does it preserve business invariants?
  • Does it maintain architectural patterns?
  • Does it respect distributed systems principles?
  • Does it handle the edge cases we've seen before?

Integration with Definition of Done (DoD):

Each plan's DoD includes semantic requirements:

definition_of_done:
  must:
    - "All API changes maintain backward compatibility"
    - "No increase in p99 latency"
    - "Audit trail remains complete"
  should:
    - "Improve code coverage by 10%"
    - "Reduce cyclomatic complexity"
  may:
    - "Optimize for memory usage"

The validation nodes enforce these semantics, not just test passage.

Q: How does the system balance human supervision with autonomous operation?

What exists today architecturally: The specification defines a sophisticated gradation of automation levels that precisely controls when human intervention is needed. This isn't a binary human/AI split - it's a spectrum that can be adjusted per task, per project, or per organization.

How the automation levels work in practice:

Manual Mode:
  - Every decision point pauses for human input
  - User sees: Context, alternatives, recommendation
  - User provides: Explicit choice or custom guidance
  - Use case: Critical production changes, learning new codebases

Review-before-apply Mode:
  - AI makes all decisions autonomously
  - Execution completes in sandbox
  - Human reviews complete diff before apply
  - User can: Approve, reject, or correct specific decisions
  - Use case: Normal feature development, refactoring

Full Automation Mode:
  - AI makes all decisions
  - Execution proceeds through apply
  - Human notified of completion
  - Rollback available if issues detected
  - Use case: Routine updates, test generation, documentation

The decision correction mechanism enables progressive automation:

The agents plan correct command is crucial for building trust:

# User observes AI made suboptimal choice
agents plan tree <plan_id>
# Sees: [Decision] "Use REST API for service communication"

# User knows gRPC would be better for this use case
agents plan correct <decision_id> --mode=revert \
  --guidance "Use gRPC instead of REST. This service requires streaming 
             updates and binary protocol efficiency. Set up protocol 
             buffer definitions and generate client/server stubs."

# System:
# 1. Marks original decision as superseded
# 2. Creates new decision with user guidance
# 3. Recomputes ONLY affected downstream decisions
# 4. Preserves all unrelated work

Progressive trust building through automation levels:

New users typically follow this progression:

  1. Start with manual mode to understand system behavior
  2. Move to review-before-apply as confidence builds
  3. Enable full automation for specific task types
  4. Gradually expand full automation scope

Real autonomy through semantic understanding:

True autonomy isn't about removing humans - it's about the system understanding when it needs help:

class AutonomyController:
    def assess_decision_confidence(self, decision, context):
        factors = {
            'past_success_rate': self.get_historical_success(decision.type),
            'codebase_familiarity': self.get_familiarity_score(context.project),
            'risk_assessment': self.evaluate_risk(decision),
            'invariant_complexity': self.analyze_invariants(decision)
        }
        
        confidence = self.compute_confidence(factors)
        
        if confidence < self.threshold:
            if self.automation_level == 'full':
                # Even in full automation, critical decisions escalate
                return RequestHumanGuidance(decision, factors)
        
        return ProceedAutonomously(decision)

Concrete example - Autonomous handling of a complex refactoring:

Scenario: "Modernize legacy e-commerce system"

INITIAL PLAN (Full Automation Mode):
1. System analyzes 50,000 line codebase
2. Identifies modernization opportunities
3. Creates plan with 47 subplans

AUTONOMOUS EXECUTION WITH SMART ESCALATION:

Subplan 1-15: Update utility functions (executes autonomously)
- Confidence: 0.95 (straightforward transformations)
- Result: Success

Subplan 16: Refactor payment processing
- Confidence: 0.4 (critical business logic)
- Action: ESCALATES to human
- Human provides: "Preserve exact penny rounding behavior"
- Continues autonomously with constraint

Subplan 17-30: UI component updates (executes autonomously)
- Confidence: 0.9 (isolated changes)
- Result: Success

Subplan 31: Database schema migration
- Detects: Would require 6-hour downtime
- Action: ESCALATES to human
- Human provides: "Use online migration with feature flags"
- Re-plans with zero-downtime approach

Subplan 32-47: Complete autonomously

The path to greater autonomy:

The system becomes more autonomous through:

  1. Learning from corrections:

    • Every correction teaches the system about user preferences
    • Patterns emerge: "This team always prefers gRPC for microservices"
    • Future decisions incorporate these learnings
  2. Building project-specific context:

    • Each successful plan adds to project knowledge
    • System learns codebase patterns, team conventions, business rules
    • Confidence increases with familiarity
  3. Hierarchical delegation:

    • Proven subplan patterns become fully autonomous
    • Human focuses on high-level decisions
    • System handles implementation details
  4. Semantic safety nets:

    • Comprehensive invariant checking reduces risk
    • Rollback capabilities provide recovery path
    • Humans can trust system won't cause catastrophic failures

Q: What's actually implemented today versus planned for the future?

Concrete implementations in the architecture:

  1. Decision Tree with Complete Context Capture

    • Full schema defined
    • Storage model specified
    • Correction mechanism detailed
    • Query patterns established
  2. Hierarchical Plan/Subplan System

    • Spawning mechanism defined
    • Execution semantics specified
    • Merge strategies documented
    • Failure handling described
  3. Resource-Aware Sandbox Isolation

    • Multiple strategies defined (git worktree, filesystem overlay, transactions)
    • Lazy sandboxing for efficiency
    • Resource-specific merge algorithms
    • Cleanup behavior specified
  4. Multi-Layer Error Prevention

    • Decision validation during planning
    • Semantic validation nodes
    • Invariant enforcement
    • Definition of Done checking
  5. Graduated Automation Controls

    • Three levels clearly defined
    • Decision correction without full re-execution
    • Confidence-based escalation
    • Progressive trust building

Near-term implementations (architecture complete, engineering straightforward):

  1. Memory Tier Management

    • Hot/warm/cold distinction clear
    • Context loading patterns defined
    • Just needs LRU cache and storage backend
  2. Cross-Plan Learning

    • Decision history provides training data
    • Pattern extraction is standard ML
    • Confidence scoring is well-understood
  3. Cost/Risk Estimation

    • Dedicated estimation actor role defined
    • Historical data provides baselines
    • Standard prediction problem
  4. Extended Validation Patterns

    • Pluggable validation architecture
    • Project-specific rules as configuration
    • Industry patterns can be packaged

Research territory (requires innovation but architecture supports):

  1. Optimal Context Selection for 100K+ file codebases

    • Current: Heuristic-based selection
    • Research: ML-driven relevance ranking
    • Architecture supports: Any selection algorithm can plug in
  2. Automated Invariant Discovery

    • Current: User-defined invariants
    • Research: Mining invariants from code patterns
    • Architecture supports: Invariants are just validation rules
  3. Cross-Project Knowledge Transfer

    • Current: Project-specific learning
    • Research: Generalized pattern recognition
    • Architecture supports: Cold tier can span projects
  4. Fully Autonomous Recovery Strategies

    • Current: Rollback and retry with guidance
    • Research: Automatic error understanding and fixing
    • Architecture supports: Recovery is just another plan type

Why we can confidently handle Firefox-scale projects:

The architecture doesn't require magical AI breakthroughs. It requires:

  • Hierarchical decomposition: ✓ Fully specified
  • Bounded context operations: ✓ Dependency closure computation defined
  • Parallel execution with isolation: ✓ Sandbox model complete
  • Semantic validation: ✓ Multi-layer approach specified
  • Progressive automation: ✓ Automation levels and correction defined

The difference between handling a 1,000 file project and a 100,000 file project is:

  • More subplans (hierarchical decomposition handles this)
  • Larger cold storage (standard database scaling)
  • Better context selection (improves with use but works with heuristics)
  • More validation patterns (accumulate over time)

This isn't speculative architecture astronautics - it's applying proven distributed systems principles to AI agent coordination. The innovation is in the integration, not in requiring fundamental breakthroughs.