# CleverAgents Documentation (Detailed Spec) ### Source Material Key themes include: * A **four-phase plan lifecycle** (Action → Strategize → Execute → Apply). * **Actors** as a unifying abstraction (an LLM/agent *or* a whole graph). * A **sandbox + diff review** workflow and **CLI-first** interaction model. * A scalable **context/memory architecture** (hot/warm/cold tiers, per-actor views). * A future-facing correction model where the user can "edit the decision tree" and only recompute affected subtrees. ## Big Picture: What CleverAgents *is* CleverAgents is your **command center for AI agents**—a unified platform for orchestrating any task you want agents to accomplish, from developing large software projects to writing comprehensive technical papers, administering databases, managing cloud infrastructure, or any complex multi-step workflow. The core value proposition is enabling **long-running, complex, large-scale tasks to execute autonomously with minimal human intervention**, making it ideal for building entire software systems, producing extensive documentation, or managing sophisticated operations largely hands-off. In **server mode**, CleverAgents becomes a collaborative hub where teams can share resources—prompts, actors, actions, and projects—while executing plans in the cloud. This enables a consistent experience across all your devices: start a complex task on your laptop, check progress from your phone, and review results from any machine. While CleverAgents leverages LangGraph and LangChain for the underlying LLM runtime primitives (tool calling, graphs, routing), its value lies in what it builds on top: * **CleverAgents** provides: * A **first-class plan lifecycle** (Action/Strategize/Execute/Apply) for breaking down and tracking complex work, * A **project + resource model** for grounding tasks in real codebases, databases, documents, and infrastructure, * A consistent **actor abstraction** for defining and composing intelligent agents, * A consistent **skill abstraction** for anything an agent can execute, * A **sandbox + checkpoint** safety model for safe, reversible execution, * A **CLI/TUI/Web UX** for controlling and monitoring large multi-step autonomous work. ## Glossary (Terms Used Precisely) * **Plan**: A tracked lifecycle for a single unit-of-work (which may spawn subplans). Plans follow the same namespace rules as actors. * **Action**: A reusable plan template not tied to any project yet. Created via CLI commands (not YAML files). Actions follow the same namespace rules as actors. * **Strategize**: Read-only planning phase that produces a strategy and subplan blueprint. **All decisions are made during this phase.** * **Execute**: Phase that performs work in a sandbox; **spawns subplans** (based on decisions made in Strategize); produces artifacts/diffs. * **Apply**: Phase that commits sandbox results into the real project (and records an "applied" plan state). * **Project**: A collection of resources + configuration that define "where work happens" and "what can be touched." Created via CLI commands. Can be **local** (contains local-only and remote resources) or **remote** (all resources remotely accessible). * **Resource**: Anything that can be read/written/queried (files, repo trees, DB endpoints, cloud clusters, documents). **Extends the MCP resource concept** to support both read and write operations. Each resource defines its own **sandbox strategy**. * **Skill**: A callable capability defined **inline in actor YAML configuration** as tool nodes. Extends the **MCP standard** and **Agent Skills standard**. Skills follow the same naming scheme as plans, actions, etc. * **Actor**: Anything conversational; may be a single agent/LLM or an entire graph of actors/tools. Defined via **YAML configuration files** (LangGraph definitions). Always named using `/` format. * **Session**: A user interaction context and conversation thread that can span multiple plans. * **Server**: Optional shared service for multi-user storage, permissions, and orchestration. Plans on remote projects can execute on the server. * **Namespace**: Scoping mechanism for actors, actions, plans, skills, etc. `local/` is reserved for local-only items. User namespaces (`/`) and organization namespaces (`/`) are stored on the server. Built-in LLM actors use provider namespaces (e.g., `openai/`, `anthropic/`). * **Decision**: A recorded choice point made during Strategize that affects downstream work. Decisions form a tree structure that enables correction and replay. * **ULID**: Universally Unique Lexicographically Sortable Identifier. Preferred over UUID for plan and decision IDs due to time-sortability. ## CLI Commands ### Command Synopsis ```text agents|cleveragents [--data-dir PATH] [--config-path PATH] [--help] [--version] [--install-completion [SHELL]] [--show-completion [SHELL]] [args] agents version agents info agents diagnostics agents init [--yes] agents session create [--name NAME] [--actor ACTOR] [--metadata KEY=VALUE ...] agents session list [--format table|json] agents session show agents session delete [--yes] agents session export --output FILE agents session import --input FILE [--name NAME] agents session tell "" [--session SESSION_ID|NAME] [--actor ACTOR] [--stream] agents project create --name/-n NAME [--description/-d TEXT] [--tag/-t TAG ...] agents project add-resource --project/-p PROJECT --name/-n NAME --type/-t TYPE --location/-l LOCATION --sandbox-strategy/-s STRATEGY [--read-only] [--metadata/-m KEY=VALUE ...] agents project remove-resource --project/-p PROJECT --name/-n NAME [--force/-f] [--yes] agents project list [--namespace/-n NS] [--tag/-t TAG] [--format table|json] agents project show [--format rich|json] agents project set-validation --project/-p PROJECT [--resource NAME] [--test-command CMD] [--lint-command CMD] [--type-check-command CMD] [--build-command CMD] [--timeout SECONDS] [--clear] agents project delete [--force/-f] [--yes] agents project context set --project PROJECT [--view strategize|execute|apply|default] [--include-resource NAME ...] [--exclude-resource NAME ...] [--include-path GLOB ...] [--exclude-path GLOB ...] [--hot-max-tokens N] [--warm-max-decisions N] [--cold-max-decisions N] [--query-limit N] [--max-file-size BYTES] [--max-total-size BYTES] [--summarize/--no-summarize] [--summary-max-tokens N] [--clear] agents project context show --project PROJECT [--view strategize|execute|apply|default] [--format table|json] agents actor run --config/-c FILE... --prompt/-p TEXT [--output/-o FILE] [--verbose/-v] [--unsafe/-u] [--context NAME] [--context-dir PATH] [--load-context FILE] [--temperature/-t FLOAT] [--allow-rxpy-in-run-mode] agents actor add --config/-c FILE [--unsafe] [--set-default] [--option/-o key=value ...] agents actor update [--config/-c FILE] [--unsafe|--safe] [--set-default] [--option/-o key=value ...] agents actor remove agents actor list agents actor show agents actor set-default agents actor context add --name NAME [-r/--recursive] agents actor context load --name NAME [-r/--recursive] agents actor context rm --name NAME agents actor context list [NAME] [--format table|json] agents actor context show --name NAME [PATH] agents actor context export --name NAME --output FILE agents actor context import --name NAME --input FILE agents actor context delete --name NAME [--yes] agents actor context clear --name NAME [--yes] agents plan list [--phase PHASE] [--state STATE] [--project PROJECT] [--action ACTION] [--format table|json] agents plan use --project/-p PROJECT [--project/-p PROJECT ...] [--arg/-a name=value ...] [--automation-level manual|review|auto] [--safety-profile NAME] [--require-sandbox/--no-require-sandbox] [--require-checkpoints/--no-require-checkpoints] [--require-apply-approval/--no-require-apply-approval] [--allow-skill-category NAME ...] [--deny-skill-category NAME ...] agents plan execute [PLAN_ID] agents plan apply [--yes] agents plan status [PLAN_ID] agents plan cancel [--reason/-r TEXT] agents plan tree [PLAN_ID] [--format tree|json|flat] [--show-superseded] agents plan explain [--show-context] [--show-reasoning] agents plan correct --mode revert|append --guidance/-g TEXT [--guidance-file/-f PATH|-] [--dry-run] [--yes] agents plan diff --plan agents plan diff --correction agents plan artifacts agents plan prompt "" agents plan resume agents plan checkpoint [--label TEXT] agents plan rollback [--yes] agents action create --strategy-actor/-s ACTOR --execution-actor/-e ACTOR --definition-of-done/-d TEXT [--description TEXT] [--long-description TEXT] [--arg/-a spec ...] [--reusable/--no-reusable] [--read-only] [--tag/-t TAG ...] [--available] [--estimation-actor ACTOR] [--safety-profile NAME] [--require-sandbox/--no-require-sandbox] [--require-checkpoints/--no-require-checkpoints] [--require-apply-approval/--no-require-apply-approval] [--allow-skill-category NAME ...] [--deny-skill-category NAME ...] agents action list [--namespace/-n NS] [--state/-s STATE] [--available] agents action show agents action available agents action archive agents config set agents config get agents config list agents providers list ``` ### Command Reference #### Global Options **Purpose** Configure global state locations and shell integration for every command. **Arguments** - `--data-dir PATH`: Overrides the global data directory (database, caches, sessions, logs). When omitted, the default data location is used. - `--config-path PATH`: Overrides the global configuration file path. When omitted, the default config path is used. - `--help`, `-h`: Print help for the current command. - `--version`, `-V`: Print the version and exit. - `--install-completion [SHELL]`: Install shell completion for the given shell. - `--show-completion [SHELL]`: Show the completion script for the given shell. **Examples**

$ agents --data-dir /srv/cleveragents --config-path /srv/cleveragents/config.toml info

╭─ System Snapshot ─────────────────╮
│ CleverAgents 1.0.0                │
│ Mode: local                       │
│ Automation: review                │
│ Default Actor: local/orchestrator │
│ Status: ready                     │
╰───────────────────────────────────╯

╭─ Paths ───────────────────────────────╮
│ Data Dir: /srv/cleveragents           │
│ Config: /srv/cleveragents/config.toml │
│ Logs: /srv/cleveragents/logs          │
│ Cache: /srv/cleveragents/cache        │
│ Database: /srv/cleveragents/agents.db │
╰───────────────────────────────────────╯

╭─ Runtime ──────────╮
│ PID: 4127          │
│ Uptime: 00:14:32   │
│ Python: 3.13.1     │
│ Host: devbox.local │
│ Platform: linux    │
╰────────────────────╯

╭─ Projects & Sessions ─╮
│ Projects: 2          │
│ Sessions: 1 active  │
│ Active Plans: 0     │
│ Actors: 3           │
╰───────────────────────╯

✓ OK Environment loaded
#### agents version **Purpose** Print the current CLI version. **Arguments** None. **Examples**

$ agents version

╭─ CLI Version ────╮
│ CleverAgents CLI │
│ Version: 1.0.0   │
│ Channel: stable  │
│ Python: 3.13     │
╰──────────────────╯

╭─ Build ────────────────╮
│ Build Date: 2026-02-08 │
│ Commit: a17c3f9        │
│ Schema: v3             │
│ Platform: linux-x86_64 │
╰────────────────────────╯

╭─ Dependencies ─────────────────╮
│ LangGraph: 0.2.60           │
│ LangChain: 0.3.18          │
│ MCP SDK: 1.4.0             │
│ Pydantic: 2.10.4           │
╰────────────────────────────────╯

✓ OK Version reported
#### agents info **Purpose** Show configuration and runtime information useful for debugging or support. **Arguments** None. **Examples**

$ agents info

╭─ Environment ───────────────────────────────────────────────╮
│ Data Dir: /home/alex/.cleveragents                          │
│ Config: /home/alex/.cleveragents/config.toml                │
│ Database: sqlite:///home/alex/.cleveragents/cleveragents.db │
│ Server Mode: disabled                                       │
│ Platform: Linux 6.8.0 (x86_64)                              │
╰─────────────────────────────────────────────────────────────╯

╭─ Runtime ─────────────────────────╮
│ Default Actor: local/orchestrator │
│ Automation: review                │
│ Providers: 3 configured           │
│ Sessions: 2 active                │
│ Active Plans: 1                   │
╰───────────────────────────────────╯

╭─ Storage ─────╮
│ Cache: 118 MB │
│ Logs: 42 MB   │
│ Backups: 3    │
│ DB Size: 8 MB │
╰───────────────╯

╭─ Indexing ─────────────╮
│ Text Index: ready     │
│ Vector Index: ready   │
│ Graph Store: disabled │
│ Indexed Files: 1,247  │
╰────────────────────────╯

✓ OK Environment details ready
#### agents diagnostics **Purpose** Run health checks for configuration, providers, and filesystem permissions. **Arguments** None. **Examples**

$ agents diagnostics

╭─ Checks ────────────────────────────────╮
│ Check            Status  Details        │
│ ───────────────  ──────  ────────────── │
│ Config file      OK      readable       │
│ Database         OK      writable       │
│ OPENAI_API_KEY   WARN    missing        │
│ Anthropic key    OK      configured     │
│ Disk space       OK      2.1 GB free    │
│ Text index       OK      tantivy 0.22   │
│ Vector index     OK      faiss (CPU)    │
│ Graph store      WARN    not configured │
│ File permissions OK      data dir r/w   │
│ Git              OK      git 2.43.0     │
╰────────────────────────────────────────╯

╭─ Summary ────────╮
│ Checks: 10 total │
│ Warnings: 2      │
│ Errors: 0         │
│ Duration: 0.6s   │
╰──────────────────╯

╭─ Recommendations ─────────────────────────────────────────────╮
│ - Set OPENAI_API_KEY to enable OpenAI models                  │
│ - Configure a graph store backend for structural code queries │
│ - Run agents providers list to verify credentials             │
╰───────────────────────────────────────────────────────────────╯

⚠ WARN 2 warnings require attention
#### agents init **Purpose** Initialize or reset the global CleverAgents environment. This wipes any existing data and re-creates the global config and database. **Arguments** - `--yes`: Skip the confirmation prompt and proceed with the wipe. **Examples**

$ agents init

Warning: This will remove all data in /home/alex/.cleveragents
Continue? [y/N]: y

╭─ Environment Reset ────────────────────────────────╮
│ Config: /home/alex/.cleveragents/config.toml       │
│ Database: /home/alex/.cleveragents/cleveragents.db │
│ Backup: /home/alex/.cleveragents.backup-2026-02-08 │
│ Status: ready                                      │
╰────────────────────────────────────────────────────╯

╭─ Defaults ────────────────────────╮
│ Default Actor: local/orchestrator │
│ Automation: review                │
│ Sandbox: required                 │
│ Checkpoints: enabled              │
│ Apply Approval: required          │
╰───────────────────────────────────╯

╭─ Created ─────────────────╮
│ Config: config.toml       │
│ Database: cleveragents.db │
│ Logs: logs/               │
│ Cache: cache/             │
│ Backups: backups/         │
╰───────────────────────────╯

╭─ Schema ───────────────────────╮
│ Version: v3                   │
│ Tables: 12 created            │
│ Migrations: up to date        │
╰────────────────────────────────╯

✓ OK Environment initialized
#### agents session **Purpose** Manage interactive sessions that hold a conversation history and orchestrator state. ##### agents session create **Purpose** Create a new session for interactive work. **Arguments** - `--name NAME`: Human-friendly session name. - `--actor ACTOR`: Orchestrator actor to use for `session tell` (defaults to configured default). - `--metadata KEY=VALUE`: Attach metadata to the session (repeatable). **Examples**

$ agents session create --name weekly-planning --actor local/orchestrator --metadata team=platform

╭─ Session ──────────────────────╮
│ Name: weekly-planning          │
│ ID: 01HXM2A6K1P2E9Q9D4GQ7J4S7Z │
│ Actor: local/orchestrator      │
│ Created: 2026-02-08 12:44      │
│ Namespace: local                │
╰────────────────────────────────╯

╭─ Settings ─────────────╮
│ Automation: review     │
│ Streaming: off         │
│ Context: default       │
│ Memory: enabled        │
│ Max History: 50 turns │
╰────────────────────────╯

╭─ Actor Details ───────────────────╮
│ Provider: anthropic              │
│ Model: claude-3.5                │
│ Temperature: 0.7                 │
│ Context Window: 200K tokens      │
╰───────────────────────────────────╯

╭─ Metadata ──────╮
│ - team=platform │
│ - quarter=Q1    │
╰─────────────────╯

✓ OK Session created
##### agents session list **Purpose** List sessions available on the local machine. **Arguments** - `--format table|json`: Output format. **Examples**

$ agents session list --format table

╭─ Sessions ───────────────────────────────────────────────────────────────────╮
│ ID        Name             Actor               Messages  Updated          │
│ ────────  ───────────────  ──────────────────  ────────  ──────────────── │
│ 01HXM2A6  weekly-planning  local/orchestrator  6         2026-02-08 12:44 │
│ 01HXM1F2  refactor-sprint  local/orchestrator  14        2026-02-07 18:11 │
╰──────────────────────────────────────────────────────────────────────────────╯

╭─ Summary ────────────────────╮
│ Total: 2                     │
│ Most Recent: weekly-planning │
│ Oldest: refactor-sprint      │
│ Total Messages: 20           │
│ Storage: 42 KB               │
╰──────────────────────────────╯

✓ OK 2 sessions listed
##### agents session show **Purpose** Show details and recent messages for a session. **Arguments** - ``: The session identifier or name. **Examples**

$ agents session show weekly-planning

╭─ Session Summary ──────────────╮
│ Name: weekly-planning          │
│ ID: 01HXM2A6K1P2E9Q9D4GQ7J4S7Z │
│ Actor: local/orchestrator      │
│ Messages: 6                    │
│ Created: 2026-02-08 12:30      │
│ Updated: 2026-02-08 12:44      │
│ Automation: review              │
╰────────────────────────────────╯

╭─ Recent Messages ──────────────────────────────────╮
│ user  Create an action to refresh dependency locks │
│ assistant  Plan created, running commands...       │
│ assistant  Completed 2 commands                    │
╰────────────────────────────────────────────────────╯

╭─ Linked Plans ──────────────────────────────╮
│ Plan ID                     Phase   State    │
│ ──────────────────────────  ──────  ──────── │
│ 01HXM8C2ZK4Q7C2B3F2R4VYV6J  execute  complete │
╰─────────────────────────────────────────────╯

╭─ Token Usage ──────────────╮
│ Input Tokens: 3,420       │
│ Output Tokens: 1,185     │
│ Estimated Cost: $0.0184  │
╰────────────────────────────╯

✓ OK Session details loaded
##### agents session delete **Purpose** Delete a session and its stored conversation history. **Arguments** - ``: The session identifier or name. - `--yes`: Skip the confirmation prompt. **Examples**

$ agents session delete weekly-planning

Delete session weekly-planning? [y/N]: y

╭─ Deletion Summary ─────────────╮
│ Session: weekly-planning       │
│ ID: 01HXM2A6K1P2E9Q9D4GQ7J4S7Z │
│ Messages: 6 removed            │
│ Storage: 18 KB freed           │
│ Plans Orphaned: 0              │
╰────────────────────────────────╯

╭─ Cleanup ───────────╮
│ Backups: none       │
│ Logs: preserved     │
│ Context: cleared    │
│ Checkpoints: none  │
╰─────────────────────╯

✓ OK Session deleted
##### agents session export **Purpose** Export a session as a portable JSON file. **Arguments** - ``: The session identifier or name. - `--output FILE`: Output file path. **Examples**

$ agents session export weekly-planning --output /tmp/weekly-planning.json

╭─ Session Export ──────────────────╮
│ Session: weekly-planning          │
│ Output: /tmp/weekly-planning.json │
│ Messages: 6                       │
│ Size: 24 KB                       │
│ Format: JSON                      │
╰───────────────────────────────────╯

╭─ Contents ─────────────────╮
│ Messages: 6               │
│ Plan References: 1       │
│ Metadata Keys: 2         │
│ Actor Config: included   │
│ Schema Version: v3       │
╰────────────────────────────╯

╭─ Integrity ──────────────────╮
│ Checksum: sha256:7a9b...42c1 │
│ Encrypted: no                │
╰──────────────────────────────╯

✓ OK Export completed
##### agents session import **Purpose** Import a session JSON file. **Arguments** - `--input FILE`: Input JSON file. - `--name NAME`: Optional override for the session name. **Examples**

$ agents session import --input /tmp/weekly-planning.json --name weekly-planning-restored

╭─ Session Import ───────────────────────╮
│ Input: /tmp/weekly-planning.json       │
│ Name: weekly-planning-restored         │
│ Session ID: 01HXM3D3B2W4CQYQ3P4ZB8A5T1 │
│ Messages: 6                            │
│ Schema: v3                              │
╰────────────────────────────────────────╯

╭─ Validation ────────────╮
│ Checksum: verified     │
│ Schema: compatible     │
│ Actor Ref: resolved    │
╰─────────────────────────╯

╭─ Merge ──────────────╮
│ Existing: none       │
│ Strategy: create new │
╰──────────────────────╯

✓ OK Import completed
##### agents session tell **Purpose** Send a natural-language request to the orchestrator. The orchestrator can create actions, plans, or project changes by issuing the necessary CleverAgents commands under the hood. **Arguments** - `""`: Instruction text. - `--session SESSION_ID|NAME`: Session to use (defaults to the most recent session). - `--actor ACTOR`: Override the session actor for this request. - `--stream`: Stream progress as the orchestrator works. **Examples**

$ agents session tell "Create an action to refresh dependency locks and add it to the platform project" \
  --session weekly-planning

╭─ Plan Request ──────────────────────────────────────────╮
│ Actor: local/orchestrator                               │
│ Session: weekly-planning                                │
│ Automation: review                                      │
│ Prompt: Create an action to refresh dependency locks... │
╰─────────────────────────────────────────────────────────╯

╭─ Commands Executed ────────────────────────────────────────────────────────────────────────────────────────╮
│ - agents action create local/refresh-locks --strategy-actor local/planner --execution-actor local/executor │
│ - agents project add-resource --project local/platform --name repo --type git_repository                   │
╰────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭─ Result ────────────────────────────╮
│ Action: local/refresh-locks (draft) │
│ Project: local/platform             │
│ Resource: repo                      │
╰─────────────────────────────────────╯

╭─ Usage ─────────────────────╮
│ Input Tokens: 1,842       │
│ Output Tokens: 624       │
│ Cost: $0.0094             │
│ Duration: 3.2s             │
│ Tool Calls: 2             │
╰─────────────────────────────╯

✓ OK Orchestrator completed 2 commands
#### agents project **Purpose** Manage projects and their resources. ##### agents project create **Purpose** Create a new project record. **Arguments** - `--name/-n NAME`: Namespaced project name. - `--description/-d TEXT`: Optional description. - `--tag/-t TAG`: Project tag (repeatable). **Examples**

$ agents project create --name local/api-service --description "Backend API" --tag python --tag backend

╭─ Project ──────────────────────╮
│ Name: local/api-service        │
│ ID: 01HXM4T08Y0N5R9VZ4QX4BPTZ1 │
│ Description: Backend API       │
│ Tags: python, backend          │
│ Type: local                    │
│ Created: 2026-02-08 12:46      │
╰────────────────────────────────╯

╭─ Paths ────────────────────────────────────╮
│ Root: /repos/api-service                   │
│ Data Dir: /repos/api-service/.cleveragents │
╰────────────────────────────────────────────╯

╭─ Defaults ────────────────────╮
│ Sandbox: git_worktree         │
│ Validation: unset             │
│ Context Filters: none         │
│ Automation: (inherits global) │
│ Apply Approval: required      │
╰───────────────────────────────╯

╭─ Resources ──────╮
│ Total: 0        │
│ Indexed: 0      │
│ Sandboxable: 0 │
╰──────────────────╯

✓ OK Project created
##### agents project add-resource **Purpose** Attach a resource to a project with a specific sandbox strategy. **Arguments** - `--project/-p PROJECT`: Project name. - `--name/-n NAME`: Resource name. - `--type/-t TYPE`: Resource type (git_repository, filesystem, database, api_endpoint, cloud_infra, document, repo_tree). - `--location/-l LOCATION`: Path, URL, or connection string. - `--sandbox-strategy/-s STRATEGY`: Sandbox strategy (git_worktree, copy_on_write, overlay, transaction_rollback, terraform_state, none). - `--read-only`: Marks the resource as read-only. - `--metadata/-m KEY=VALUE`: Resource metadata (repeatable). **Examples**

$ agents project add-resource --project local/api-service --name repo --type git_repository \
  --location /repos/api --sandbox-strategy git_worktree --metadata branch=main

╭─ Resource ─────────────────╮
│ Project: local/api-service │
│ Name: repo                 │
│ Type: git_repository       │
│ Location: /repos/api       │
│ Sandbox: git_worktree      │
│ Remote: no                 │
╰────────────────────────────╯

╭─ Permissions ────────────╮
│ Read: allowed            │
│ Write: allowed           │
│ Apply: requires approval │
╰──────────────────────────╯

╭─ Indexing ─────────────────────╮
│ Status: indexing...          │
│ Files Found: 347             │
│ Language: Python (primary)   │
│ Estimated Time: ~20 seconds  │
╰────────────────────────────────╯

╭─ Metadata ──────────────────────────╮
│ - branch=main                       │
│ - origin=git@github.com:org/api.git │
╰─────────────────────────────────────╯

✓ OK Resource added
##### agents project remove-resource **Purpose** Remove a resource from a project. **Arguments** - `--project/-p PROJECT`: Project name. - `--name/-n NAME`: Resource name. - `--force/-f`: Skip additional checks. - `--yes`: Skip confirmation prompt. **Examples**

$ agents project remove-resource --project local/api-service --name repo

Remove resource repo from local/api-service? [y/N]: y

╭─ Resource Removed ─────────╮
│ Project: local/api-service │
│ Name: repo                 │
│ Type: git_repository       │
│ Sandbox: git_worktree      │
╰────────────────────────────╯

╭─ Index Cleanup ──────────╮
│ Text Index: 347 removed │
│ Vectors: 892 removed    │
│ Graph Triples: cleared  │
│ Duration: 0.3s          │
╰──────────────────────────╯

╭─ Project Summary ──────────────╮
│ Resources: 1 remaining         │
│ Last Updated: 2026-02-08 12:50 │
│ Active Plans: 0                │
╰────────────────────────────────╯

✓ OK Resource removed
##### agents project list **Purpose** List projects with optional filters. **Arguments** - `--namespace/-n NS`: Filter by namespace. - `--tag/-t TAG`: Filter by tag. - `--format table|json`: Output format. **Examples**

$ agents project list --format table

╭─ Projects ──────────────────────────────────────────────────────────────────────╮
│ ID        Name               Resources  Tags             Remote  Active Plans │
│ ────────  ─────────────────  ─────────  ───────────────  ──────  ──────────── │
│ 01HXM4T0  local/api-service  2          python, backend  No      1            │
│ 01HXM4B9  local/docs         1          docs             No      0            │
╰────────────────────────────────────────────────────────────────────────────────╯

╭─ Summary ──────────────────╮
│ Total: 2                   │
│ With Resources: 2          │
│ Remote: 0                  │
│ Total Resources: 3         │
│ Indexed Files: 1,247       │
│ Active Plans: 1            │
╰────────────────────────────╯

✓ OK 2 projects listed
##### agents project show **Purpose** Show full project details. **Arguments** - ``: Project name. - `--format rich|json`: Output format. **Examples**

$ agents project show local/api-service

╭─ Project Details ──────────────╮
│ Name: local/api-service        │
│ ID: 01HXM4T08Y0N5R9VZ4QX4BPTZ1 │
│ Description: Backend API       │
│ Tags: python, backend          │
│ Resources: 2                   │
│ Remote: no                     │
│ Created: 2026-02-08 12:46      │
╰────────────────────────────────╯

╭─ Resources ─────────────────────────────────────────────────────────────╮
│ Name  Type            Location       Sandbox              Read-Only │
│ ────  ──────────────  ─────────────  ────────────────────  ───────── │
│ repo  git_repository  /repos/api     git_worktree         no        │
│ db    database        postgresql://  transaction_rollback yes       │
╰─────────────────────────────────────────────────────────────────────────╯

╭─ Validation ────────╮
│ Test: pytest        │
│ Lint: ruff check .  │
│ Type Check: pyright │
│ Build: (none)       │
│ Timeout: 300s       │
╰─────────────────────╯

╭─ Context ───────────────────╮
│ Include: repo               │
│ Exclude: **/node_modules/** │
│ Max File Size: 1 MB        │
╰─────────────────────────────╯

╭─ Indexing Status ──────────╮
│ Text Index: ready         │
│ Vector Index: ready       │
│ Graph Store: disabled     │
│ Indexed Files: 347        │
│ Last Indexed: 12:48       │
╰────────────────────────────╯

╭─ Active Plans ──────────────────────────╮
│ Plan ID   Action               Phase    │
│ ────────  ───────────────────  ─────── │
│ 01HXM7A9  local/code-coverage  execute │
╰─────────────────────────────────────────╯

✓ OK Project loaded
##### agents project set-validation **Purpose** Set or update validation commands for a project or a specific resource. **Arguments** - `--project/-p PROJECT`: Project name. - `--resource NAME`: Optional resource name. - `--test-command CMD`: Test command. - `--lint-command CMD`: Lint command. - `--type-check-command CMD`: Type-check command. - `--build-command CMD`: Build command. - `--timeout SECONDS`: Timeout per command. - `--clear`: Clear validation config. **Examples**

$ agents project set-validation --project local/api-service --resource repo \
  --test-command "pytest" --lint-command "ruff check ." --type-check-command "pyright"

╭─ Validation Commands ╮
│ Test: pytest         │
│ Lint: ruff check .   │
│ Type Check: pyright  │
│ Build: (none)        │
│ Timeout: 300s        │
╰──────────────────────╯

╭─ Scope ─────────────────────╮
│ Project: local/api-service  │
│ Resource: repo              │
│ Applies To: sandbox + apply │
╰─────────────────────────────╯

╭─ Execution Order ─────────────────╮
│ 1. Type Check → pyright          │
│ 2. Lint → ruff check .            │
│ 3. Test → pytest                  │
│ On Failure: iterate up to 3 tries │
│ Then: escalate to user            │
╰───────────────────────────────────╯

╭─ Notes ──────────────────────╮
│ - Commands run after Execute │
│ - Failures block Apply phase │
│ - Actor can auto-fix issues  │
╰──────────────────────────────╯

✓ OK Validation updated
##### agents project delete **Purpose** Delete a project and all associated resources. **Arguments** - ``: Project name. - `--force/-f`: Delete even if active plans exist. - `--yes`: Skip confirmation prompt. **Examples**

$ agents project delete local/docs

Delete project local/docs? This cannot be undone. [y/N]: y

╭─ Deletion Summary ──────────────────╮
│ Project: local/docs                 │
│ ID: 01HXM4B9F2C1V8X2N6Q7K9L0M1      │
│ Resources: 1 removed                │
│ Data Dir: /repos/docs/.cleveragents │
╰─────────────────────────────────────╯

╭─ Index Cleanup ────────╮
│ Text Index: cleared   │
│ Vectors: 240 removed  │
│ Graph Triples: none  │
│ Storage Freed: 12 MB │
╰────────────────────────╯

╭─ Backups ────────────────────────────────────╮
│ Snapshot: /backups/local-docs-2026-02-08.tgz │
│ Retention: 7 days                            │
╰──────────────────────────────────────────────╯

✓ OK Project deleted
##### agents project context **Purpose** Manage context policies for the hot/warm/cold tiers and per-view context selection. ###### agents project context set **Purpose** Set the context policy for a project and (optionally) a specific view. **Arguments** - `--project PROJECT`: Project name. - `--view strategize|execute|apply|default`: Which view this policy applies to. - `--include-resource NAME`: Resource allowlist (repeatable). - `--exclude-resource NAME`: Resource denylist (repeatable). - `--include-path GLOB`: Path allowlist (repeatable). - `--exclude-path GLOB`: Path denylist (repeatable). - `--hot-max-tokens N`: Maximum token budget for hot context. This is a soft cap and may be null. The actor/LLM hard limit can be lower; the effective hot context is the lesser of the two. - `--warm-max-decisions N`: Maximum number of decisions kept in warm context. - `--cold-max-decisions N`: Maximum number of decisions kept in cold context. - `--query-limit N`: Max number of retrieval results per query. - `--max-file-size BYTES`: Max file size included in context. - `--max-total-size BYTES`: Max total size across included files. - `--summarize/--no-summarize`: Enable or disable summarization for large context segments. - `--summary-max-tokens N`: Token limit for generated summaries. - `--clear`: Clear the policy for the selected view. **Examples**

$ agents project context set --project local/api-service --view strategize \
  --include-resource repo --exclude-path "**/node_modules/**" \
  --hot-max-tokens 12000 --warm-max-decisions 50 --cold-max-decisions 200 \
  --summarize --summary-max-tokens 800

╭─ Context Policy ────────────╮
│ Project: local/api-service  │
│ View: strategize            │
│ Include: repo               │
│ Exclude: **/node_modules/** │
╰─────────────────────────────╯

╭─ Limits ─────────────────────╮
│ Hot Tokens: 12000 (soft cap) │
│ Warm Decisions: 50           │
│ Cold Decisions: 200          │
│ Query Limit: 20              │
│ Max File Size: 1 MB          │
│ Max Total Size: 50 MB        │
╰──────────────────────────────╯

╭─ Summarization ─╮
│ Enabled: yes    │
│ Max Tokens: 800 │
╰─────────────────╯

╭─ Other Views ────────╮
│ execute: (default) │
│ apply: (default)   │
│ default: (unset)   │
╰──────────────────────╯

✓ OK Context policy updated
###### agents project context show **Purpose** Show the active context policy for a project. **Arguments** - `--project PROJECT`: Project name. - `--view strategize|execute|apply|default`: View to display. - `--format table|json`: Output format. **Examples**

$ agents project context show --project local/api-service --view strategize

╭─ Context Policy ────────────╮
│ Project: local/api-service  │
│ View: strategize            │
│ Include: repo               │
│ Exclude: **/node_modules/** │
╰─────────────────────────────╯

╭─ Limits ─────────────────────╮
│ Hot Tokens: 12000 (soft cap) │
│ Warm Decisions: 50           │
│ Cold Decisions: 200          │
│ Query Limit: 20              │
│ Max File Size: 1 MB          │
│ Max Total Size: 50 MB        │
╰──────────────────────────────╯

╭─ Summarization ─╮
│ Enabled: yes    │
│ Max Tokens: 800 │
╰─────────────────╯

╭─ Current Usage ──────────────╮
│ Hot Context: 8,420 / 12,000  │
│ Warm Entries: 12 / 50        │
│ Cold Entries: 47 / 200       │
│ Indexed Resources: 1        │
╰──────────────────────────────╯

✓ OK Context policy loaded
#### agents actor **Purpose** Manage actors and run actor configurations directly. ##### agents actor run **Purpose** Run an actor configuration in isolation with simple, manual context. **Arguments** - `--config/-c FILE...`: YAML or JSON config files. - `--prompt/-p TEXT`: Prompt to send. - `--output/-o FILE`: Output file path. - `--verbose/-v`: Increase verbosity (repeatable). - `--unsafe/-u`: Allow unsafe configs. - `--context NAME`: Named actor context to attach. - `--context-dir PATH`: Context storage location. - `--load-context FILE`: Load context from JSON. - `--temperature/-t FLOAT`: Override temperature. - `--allow-rxpy-in-run-mode`: Allow RxPy routes in run mode. **Examples**

$ agents actor run -c ./actors/code_reader.yaml -p "Summarize the README" --context docs

╭─ Run Summary ─────────────────────╮
│ Actor: local/code_reader          │
│ Context: docs                     │
│ Config: ./actors/code_reader.yaml │
│ Temperature: 0.2                  │
│ Provider: anthropic               │
│ Model: claude-3.5                  │
╰───────────────────────────────────╯

╭─ Inputs ─────────────────────╮
│ Prompt: Summarize the README │
│ Context Files: 3             │
│ Context Size: 12.4 KB       │
╰──────────────────────────────╯

╭─ Result Metrics ──────╮
│ Output: stdout         │
│ Input Tokens: 1,524   │
│ Output Tokens: 842    │
│ Duration: 1.8s         │
│ Cost: $0.0021          │
│ Tool Calls: 0          │
╰───────────────────────╯

✓ OK Summary generated
##### agents actor add **Purpose** Add a new actor configuration. **Arguments** - ``: Actor name. - `--config/-c FILE`: Actor config file. - `--unsafe`: Mark actor as unsafe. - `--set-default`: Set as default actor. - `--option/-o key=value`: Option override (repeatable). **Examples**

$ agents actor add local/reviewer --config ./actors/reviewer.yaml --set-default

╭─ Actor Added ────────╮
│ Name: local/reviewer │
│ Provider: openai     │
│ Model: gpt-4         │
│ Default: yes         │
│ Unsafe: no           │
│ Type: graph          │
╰──────────────────────╯

╭─ Config ─────────────────────╮
│ Path: ./actors/reviewer.yaml │
│ Hash: 8b3f3d2                │
│ Options: 4                   │
│ Nodes: 3                     │
│ Edges: 4                     │
╰──────────────────────────────╯

╭─ Capabilities ───────╮
│ - code review        │
│ - diff summarization │
│ - lint guidance      │
╰──────────────────────╯

╭─ Tools ────────────────────────╮
│ Tool          Read-Only  Safe │
│ ────────────  ─────────  ──── │
│ read_file     yes        yes  │
│ search_files  yes        yes  │
│ git_diff      yes        yes  │
╰────────────────────────────────╯

✓ OK Actor added
##### agents actor update **Purpose** Update an existing actor configuration. **Arguments** - ``: Actor name. - `--config/-c FILE`: Updated config file. - `--unsafe`: Mark actor as unsafe. - `--safe`: Mark actor as safe. - `--set-default`: Set as default actor. - `--option/-o key=value`: Option override (repeatable). **Examples**

$ agents actor update local/reviewer --option temperature=0.2

╭─ Actor Updated ───────────╮
│ Name: local/reviewer      │
│ Change: temperature → 0.2 │
│ Updated: 2026-02-08 13:02 │
│ Config Hash: 9c4e2a1     │
╰───────────────────────────╯

╭─ Effective Options ╮
│ - temperature: 0.2 │
│ - max_tokens: 2048 │
│ - top_p: 1.0       │
╰────────────────────╯

╭─ Previous Value ───────╮
│ temperature: 0.7 → 0.2 │
╰────────────────────────╯

✓ OK Actor updated
##### agents actor remove **Purpose** Remove a custom actor. **Arguments** - ``: Actor name. **Examples**

$ agents actor remove local/reviewer

╭─ Actor Removed ──────╮
│ Name: local/reviewer │
│ Provider: openai     │
│ Model: gpt-4         │
╰──────────────────────╯

╭─ Impact ───────────────────────────────────╮
│ Default Actor: reset to local/orchestrator │
│ Sessions: 0 affected                       │
│ Active Plans: 0 affected                   │
│ Actions Referencing: 0                     │
╰────────────────────────────────────────────╯

╭─ Cleanup ──────────────╮
│ Config: kept on disk  │
│ Contexts: 1 orphaned  │
╰────────────────────────╯

✓ OK Actor removed
##### agents actor list **Purpose** List all actors. **Arguments** None. **Examples**

$ agents actor list

╭─ Actors ──────────────────────────────────────────────────────────────╮
│ Name            Provider   Model       Default  Built-in  Unsafe │
│ ──────────────  ─────────  ──────────  ───────  ────────  ────── │
│ local/reviewer  openai     gpt-4       ✓                  no     │
│ openai/gpt-4    openai     gpt-4                ✓         no     │
│ anthropic/3.5   anthropic  claude-3.5           ✓         no     │
╰──────────────────────────────────────────────────────────────────────╯

╭─ Summary ──────────────╮
│ Total: 3               │
│ Built-in: 2            │
│ Custom: 1              │
│ Unsafe: 0              │
│ Providers Used: 2     │
╰────────────────────────╯

✓ OK 3 actors listed
##### agents actor show **Purpose** Show details for a single actor. **Arguments** - ``: Actor name. **Examples**

$ agents actor show local/reviewer

╭─ Actor Details ────────────────────╮
│ Name: local/reviewer               │
│ Provider: openai                   │
│ Model: gpt-4                       │
│ Default: yes                       │
│ Built-in: no                       │
│ Unsafe: no                         │
│ Type: graph                        │
│ Created: 2026-02-08 12:35          │
│ Updated: 2026-02-08 12:40          │
│ Config: ./actors/reviewer.yaml     │
│ Config Hash: 9c4e2a1              │
╰────────────────────────────────────╯

╭─ Options ──────────╮
│ - temperature: 0.2 │
│ - max_tokens: 2048 │
│ - top_p: 1.0       │
╰────────────────────╯

╭─ Graph Structure ─╮
│ Nodes: 3          │
│ Edges: 4          │
│ Entry: analyze    │
│ Exit: report      │
╰───────────────────╯

╭─ Tools ────────────────────────╮
│ Tool          Read-Only  Safe │
│ ────────────  ─────────  ──── │
│ read_file     yes        yes  │
│ search_files  yes        yes  │
│ git_diff      yes        yes  │
╰────────────────────────────────╯

╭─ Permissions ───────╮
│ Unsafe: no          │
│ Filesystem: allowed │
│ Network: restricted │
╰─────────────────────╯

╭─ Usage ─────────────────────────────────────────╮
│ Referenced by Actions: 1 (local/code-coverage) │
│ Active in Sessions: 0                          │
│ Total Runs: 14                                 │
│ Avg Cost/Run: $0.0032                          │
╰─────────────────────────────────────────────────╯

✓ OK Actor loaded
##### agents actor set-default **Purpose** Set the default actor. **Arguments** - ``: Actor name. **Examples**

$ agents actor set-default local/reviewer

╭─ Default Actor ──────────────╮
│ Current: local/reviewer      │
│ Previous: local/orchestrator │
│ Provider: openai             │
│ Model: gpt-4                 │
╰──────────────────────────────╯

╭─ Impact ──────────────────────────────────╮
│ Sessions: new sessions use local/reviewer │
│ Existing Sessions: unchanged              │
│ Active Plans: unchanged                   │
╰───────────────────────────────────────────╯

╭─ Saved To ──────────────────────────╮
│ File: ~/.cleveragents/config.toml │
│ Key: default-actor                │
╰─────────────────────────────────────╯

✓ OK Default actor updated
##### agents actor context **Purpose** Manage manual context for actor runs. These commands mirror the legacy context behavior but are scoped to an actor context name. ###### agents actor context add **Purpose** Add files or directories to an actor context. **Arguments** - `--name NAME`: Context name. - ``: Files or directories to add. - `-r/--recursive`: Add directories recursively. **Examples**

$ agents actor context add --name docs README.md docs/

╭─ Context Files Added ╮
│ Context: docs        │
│ Added: 12 items      │
│ Total: 12 items      │
│ Total Size: 52 KB   │
╰──────────────────────╯

╭─ Added Items ───────────╮
│ Name       Type  Size   │
│ ─────────  ────  ────── │
│ README.md  file  4.2 KB │
│ docs/      dir   48 KB  │
╰─────────────────────────╯

╭─ Limits ────────╮
│ Max Files: 500  │
│ Max Size: 25 MB │
│ Used: 0.2%      │
╰─────────────────╯

╭─ Skipped ─────────────────────────╮
│ Binary Files: 2 (not text)       │
│ Over Size Limit: 0              │
╰───────────────────────────────────╯

✓ OK Context updated
###### agents actor context load **Purpose** Alias for `context add`. **Arguments** Same as `context add`. **Examples**

$ agents actor context load --name docs README.md

╭─ Context Files Added ╮
│ Context: docs        │
│ Added: 1 item        │
│ Total: 13 items      │
╰──────────────────────╯

╭─ Recent Item ───╮
│ Name: README.md │
│ Size: 4.2 KB    │
╰─────────────────╯

✓ OK Context updated
###### agents actor context rm **Purpose** Remove files or directories from an actor context. **Arguments** - `--name NAME`: Context name. - ``: Paths to remove. **Examples**

$ agents actor context rm --name docs README.md

╭─ Context Files Removed ╮
│ Context: docs          │
│ Removed: README.md     │
│ Total: 12 items        │
╰────────────────────────╯

╭─ Stats ───────────────────╮
│ Remaining Size: 48 KB     │
│ Updated: 2026-02-08 13:06 │
╰───────────────────────────╯

✓ OK Context updated
###### agents actor context list **Purpose** List files stored in an actor context. **Arguments** - `[NAME]`: Context name (optional to list all contexts). - `--format table|json`: Output format. **Examples**

$ agents actor context list docs

╭─ Context Files ──────────────────────────────╮
│ Name              Type  Size     Added      │
│ ────────────────  ────  ───────  ────────── │
│ README.md         file  4.2 KB   02-08 12:10 │
│ docs/overview.md  file  12.8 KB  02-08 12:10 │
│ docs/cli.md       file  9.5 KB   02-08 12:10 │
╰──────────────────────────────────────────────╯

╭─ Stats ───────────────────────╮
│ Total Files: 3               │
│ Total Size: 26.5 KB          │
│ Estimated Tokens: ~6,600    │
│ Languages: Markdown          │
╰───────────────────────────────╯

✓ OK 3 files listed
###### agents actor context show **Purpose** Show content of a file in an actor context. **Arguments** - `--name NAME`: Context name. - `[PATH]`: File path (optional for summary). **Examples**

$ agents actor context show --name docs README.md

╭─ Metadata ──────────────────╮
│ Path: README.md             │
│ Size: 4.2 KB                │
│ Added: 2026-02-08 12:10     │
│ Language: Markdown           │
│ Estimated Tokens: ~1,050   │
│ Lines: 87                    │
╰────────────────────────────╯

╭─ Preview ─────────────────────────────────────────╮
│ # CleverAgents                                    │
│                                                   │
│ CleverAgents is your command center for agents... │
│ ...                                               │
╰───────────────────────────────────────────────────╯

✓ OK Content displayed
###### agents actor context export **Purpose** Export a context as JSON. **Arguments** - `--name NAME`: Context name. - `--output FILE`: Output file path. **Examples**

$ agents actor context export --name docs --output /tmp/docs-context.json

╭─ Context Export ───────────────╮
│ Context: docs                  │
│ Output: /tmp/docs-context.json │
│ Items: 12                      │
│ Size: 48 KB                    │
╰────────────────────────────────╯

╭─ Integrity ──────────────────╮
│ Checksum: sha256:19b2...a7d0 │
│ Compressed: no               │
╰──────────────────────────────╯

✓ OK Export completed
###### agents actor context import **Purpose** Import a context from JSON. **Arguments** - `--name NAME`: Context name. - `--input FILE`: Input JSON file. **Examples**

$ agents actor context import --name docs --input /tmp/docs-context.json

╭─ Context Import ──────────────╮
│ Context: docs                 │
│ Input: /tmp/docs-context.json │
│ Items: 12                     │
╰───────────────────────────────╯

╭─ Merge ───────────╮
│ Strategy: replace │
│ Conflicts: 0      │
╰───────────────────╯

✓ OK Import completed
###### agents actor context delete **Purpose** Delete an entire context. **Arguments** - `--name NAME`: Context name. - `--yes`: Skip confirmation. **Examples**

$ agents actor context delete --name docs

Delete context docs? [y/N]: y

╭─ Context Deleted ────╮
│ Context: docs        │
│ Items: 12 removed    │
│ Storage: 48 KB freed │
╰──────────────────────╯

╭─ Cleanup ───────╮
│ Backups: none   │
│ Logs: preserved │
╰─────────────────╯

✓ OK Context deleted
###### agents actor context clear **Purpose** Clear all files from a context but keep the context itself. **Arguments** - `--name NAME`: Context name. - `--yes`: Skip confirmation. **Examples**

$ agents actor context clear --name docs

Clear context docs? [y/N]: y

╭─ Context Cleared ────╮
│ Context: docs        │
│ Items: 12 removed    │
│ Storage: 48 KB freed │
╰──────────────────────╯

╭─ Retention ────────╮
│ Context: preserved │
│ Files: removed     │
╰────────────────────╯

✓ OK Context cleared
#### agents plan **Purpose** Manage plans through the Action -> Strategize -> Execute -> Apply lifecycle. ##### agents plan list **Purpose** List plans with optional filtering. **Arguments** - `--phase PHASE`: Filter by phase. - `--state STATE`: Filter by processing state. - `--project PROJECT`: Filter by project. - `--action ACTION`: Filter by action name. - `--format table|json`: Output format. **Examples**

$ agents plan list --phase execute --format table

╭─ Plans ──────────────────────────────────────────────────────────────────────────────╮
│ ID        Phase    State       Action               Project            Elapsed  │
│ ────────  ───────  ──────────  ───────────────────  ─────────────────  ───────── │
│ 01HXM7A9  execute  processing  local/code-coverage  local/api-service  00:01:12  │
╰──────────────────────────────────────────────────────────────────────────────────────╯

╭─ Filters ──────╮
│ Phase: execute │
│ State: (any)   │
│ Project: (any) │
│ Action: (any)  │
╰────────────────╯

╭─ Summary ─────────╮
│ Total: 1          │
│ Processing: 1    │
│ Completed: 0     │
│ Errored: 0        │
╰───────────────────╯

✓ OK 1 plan listed
##### agents plan use **Purpose** Apply an action to one or more projects and start the Strategize phase. **Arguments** - ``: Action name. - `--project/-p PROJECT`: Project name (repeatable). - `--arg/-a name=value`: Action argument (repeatable). - `--automation-level manual|review|auto`: Automation level. - `--safety-profile NAME`: Named safety profile. - `--require-sandbox/--no-require-sandbox`: Enforce sandboxing. - `--require-checkpoints/--no-require-checkpoints`: Enforce checkpoints. - `--require-apply-approval/--no-require-apply-approval`: Require human approval before apply. - `--allow-skill-category NAME`: Allow skill categories (repeatable). - `--deny-skill-category NAME`: Deny skill categories (repeatable). **Examples**

$ agents plan use local/code-coverage --project local/api-service \
  --arg target_coverage_percent=85 --automation-level review

╭─ Plan Created ──────────────────────╮
│ Plan ID: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Phase: strategize                   │
│ Action: local/code-coverage         │
│ Project: local/api-service          │
│ Automation: review                  │
│ Attempt: 1                          │
╰─────────────────────────────────────╯

╭─ Inputs ─────────────────────╮
│ - target_coverage_percent=85 │
│ - safety_profile=default     │
╰──────────────────────────────╯

╭─ Actors ────────────────────────╮
│ Strategy: local/strategist      │
│ Execution: local/executor       │
│ Estimation: (none)              │
╰─────────────────────────────────╯

╭─ Safety ──────────────────╮
│ Sandbox: required         │
│ Apply Approval: required  │
│ Checkpoints: enabled      │
│ Read-Only: no             │
╰───────────────────────────╯

╭─ Context ───────────────────────╮
│ Resources: 2 (repo, db)        │
│ Indexed Files: 347             │
│ View: strategize               │
│ Hot Token Budget: 12,000       │
╰─────────────────────────────────╯

╭─ Next Steps ─────────────────────────────────────╮
│ - agents plan execute 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ - agents plan status 01HXM8C2ZK4Q7C2B3F2R4VYV6J  │
│ - agents plan tree 01HXM8C2ZK4Q7C2B3F2R4VYV6J    │
╰──────────────────────────────────────────────────╯

✓ OK Plan created
##### agents plan execute **Purpose** Start or resume execution for a plan. **Arguments** - `[PLAN_ID]`: Plan ID (optional if the most recent plan is unambiguous). **Examples**

$ agents plan execute 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Execution ──────────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Phase: execute                   │
│ Sandbox: git_worktree            │
│ Worker: local/executor           │
│ Started: 12:58:10                │
│ Attempt: 1                       │
╰──────────────────────────────────╯

╭─ Sandbox ──────────────────────────────────╮
│ Strategy: git_worktree                   │
│ Path: /repos/api/.worktrees/plan-01HXM8 │
│ Branch: cleveragents/plan-01HXM8C2      │
│ Status: active                           │
╰────────────────────────────────────────────╯

╭─ Strategy Summary ─────────────────────╮
│ Decisions: 4                          │
│ Planned Subplans: 2                  │
│ Estimated Files: ~12                  │
│ Risk: low                              │
╰────────────────────────────────────────╯

╭─ Progress ────────╮
│  Collect context │
│  Run tools       │
│  Build changeset │
│  Validate        │
╰───────────────────╯

✓ OK Execution started
##### agents plan apply **Purpose** Apply sandboxed changes to real resources. **Arguments** - ``: Plan ID. - `--yes`: Skip confirmation. **Examples**

$ agents plan apply 01HXM8C2ZK4Q7C2B3F2R4VYV6J

Apply changes for plan 01HXM8C2ZK4Q7C2B3F2R4VYV6J? [y/N]: y

╭─ Apply Summary ─────────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J    │
│ Artifacts: 6 files updated          │
│ Changes: 42 insertions, 9 deletions │
│ Project: local/api-service          │
│ Applied At: 2026-02-08 13:04        │
╰─────────────────────────────────────╯

╭─ Validation ───────────────────╮
│ Tests: passed (24/24)          │
│ Lint: passed (0 warnings)      │
│ Type Check: passed (0 errors)  │
│ Duration: 12.4s                │
╰────────────────────────────────╯

╭─ Sandbox Cleanup ─────────╮
│ Worktree: removed        │
│ Branch: merged to main   │
│ Checkpoint: archived     │
╰───────────────────────────╯

╭─ Plan Lifecycle ────────────────────────╮
│ Phase: applied                         │
│ State: complete                        │
│ Total Duration: 00:06:14              │
│ Total Cost: $0.0847                    │
│ Decisions Made: 4                     │
│ Subplans: 2 (completed)                │
╰─────────────────────────────────────────╯

╭─ Next Steps ──────╮
│ - Review git diff │
│ - Commit changes  │
╰───────────────────╯

✓ OK Changes applied
##### agents plan status **Purpose** Show detailed status for a plan. **Arguments** - `[PLAN_ID]`: Plan ID (optional if the most recent plan is unambiguous). **Examples**

$ agents plan status 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Plan Status ────────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Phase: execute                   │
│ State: processing                │
│ Action: local/code-coverage      │
│ Project: local/api-service       │
│ Automation: review               │
│ Attempt: 1                       │
╰──────────────────────────────────╯

╭─ Progress ───────╮
│  Strategize     │
│  Execute        │
│  Apply (queued) │
╰──────────────────╯

╭─ Timing ──────────╮
│ Started: 12:57:01 │
│ Elapsed: 00:01:12 │
│ ETA: 00:03:45     │
╰───────────────────╯

╭─ Execution Detail ──────────╮
│ Sandbox: git_worktree       │
│ Tool Calls: 8              │
│ Files Modified: 3          │
│ Subplans: 1/2 complete     │
│ Checkpoints: 2 created     │
╰─────────────────────────────╯

╭─ Cost ───────────────╮
│ Tokens Used: 12,420 │
│ Cost So Far: $0.041 │
│ Estimated: $0.085   │
╰──────────────────────╯

✓ OK Status refreshed
##### agents plan cancel **Purpose** Cancel a plan that is not terminal. **Arguments** - ``: Plan ID. - `--reason/-r TEXT`: Optional reason. **Examples**

$ agents plan cancel 01HXM8C2ZK4Q7C2B3F2R4VYV6J --reason "blocked on credentials"

╭─ Plan Cancelled ─────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Phase: execute                   │
│ Reason: blocked on credentials   │
│ State: cancelled                 │
│ Cancelled At: 13:02:15          │
╰──────────────────────────────────╯

╭─ Sandbox ────────────────╮
│ Status: preserved        │
│ Files Modified: 3       │
│ Checkpoints: 2          │
╰──────────────────────────╯

╭─ Subplans ────────────────╮
│ Completed: 1             │
│ Cancelled: 1             │
│ Artifacts Preserved: yes │
╰───────────────────────────╯

╭─ Recovery ────────────────────────────────────────╮
│ - Resolve credentials                             │
│ - Run agents plan resume 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
╰───────────────────────────────────────────────────╯

✓ OK Plan cancelled
##### agents plan tree **Purpose** Render the decision tree for a plan. **Arguments** - `[PLAN_ID]`: Plan ID (optional if a current plan exists). - `--format tree|json|flat`: Output format. - `--show-superseded`: Include superseded decisions. **Examples**

$ agents plan tree 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Decision Tree ──────────────────────────────────────────────────────────────╮
│ - [prompt_definition] "Increase test coverage to 85%"                      │
│   ├─ [strategy_choice] "Prioritize auth and payments" (confidence: 0.82) │
│   │  ├─ [subplan_spawn] "Write auth tests" → 01HXM9F1A                    │
│   │  └─ [subplan_spawn] "Write payment tests" → 01HXM9F2B                 │
│   └─ [implementation_choice] "Use mocks for database layer"                │
╰──────────────────────────────────────────────────────────────────────────────╯

╭─ Tree Summary ─────────╮
│ Nodes: 5              │
│ Depth: 3              │
│ Subplans: 2           │
│ Superseded: 0 (hidden) │
╰────────────────────────╯

╭─ Subplans ─────────────────────────────────────────╮
│ ID          Name           Phase    State     │
│ ──────────  ─────────────  ───────  ───────── │
│ 01HXM9F1A   auth-tests     execute  processing │
│ 01HXM9F2B   payment-tests  execute  queued     │
╰────────────────────────────────────────────────────╯

╭─ Decision IDs (for correction) ──────────────╮
│ Root: 01HXM9A0B1Q2W3R5G8Z0P4Q1X8         │
│ Strategy: 01HXM9A1C2Q7W3R5G8Z0P4Q1X9     │
│ Spawn Auth: 01HXM9A2D3Q8W4R6H9Z1P5Q2X0   │
│ Spawn Payment: 01HXM9A3E4Q9W5R7I0Z2P6Q3X1 │
│ Implementation: 01HXM9A4F5Q0W6R8J1Z3P7Q4X2 │
╰──────────────────────────────────────────────╯

✓ OK Decision tree rendered
##### agents plan explain **Purpose** Show a detailed explanation for a decision. **Arguments** - ``: Decision ID. - `--show-context`: Include the context snapshot. - `--show-reasoning`: Include raw model reasoning if available. **Examples**

$ agents plan explain 01HXM9A1C2Q7W3R5G8Z0P4Q1X9

╭─ Decision ─────────────────────────────────────╮
│ ID: 01HXM9A1C2Q7W3R5G8Z0P4Q1X9                 │
│ Type: strategy_choice                          │
│ Question: Which modules should be prioritized? │
│ Chosen: Auth and payments                      │
│ Confidence: 0.82                               │
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J              │
│ Sequence: 2 of 5                               │
│ Created: 2026-02-08 12:58                     │
╰────────────────────────────────────────────────╯

╭─ Alternatives Considered ──────────────────────╮
│ 1. Auth and payments (chosen)                 │
│ 2. User module first (coverage 71%, med risk) │
│ 3. All modules equally (spread thin)          │
╰────────────────────────────────────────────────╯

╭─ Impact ─────────────────────╮
│ Downstream Decisions: 3    │
│ Downstream Subplans: 2    │
│ Artifacts Produced: 5     │
│ Correction Impact: medium │
╰──────────────────────────────╯

╭─ Context Snapshot ───────────────╮
│ - Coverage < 70% in auth         │
│ - Payments failures last release │
│ - Auth: 12 files, 45% coverage   │
│ - Payments: 8 files, 52% cover.  │
│ Hot Context Hash: sha256:4b2e... │
╰──────────────────────────────────╯

╭─ Rationale ───────────────────────────────────────╮
│ Auth and payment modules have the lowest coverage │
│ and highest business risk. Auth handles security  │
│ tokens, payments handles money. Both had bugs in  │
│ the last release traceable to missing tests.      │
╰───────────────────────────────────────────────────╯

╭─ Correction ──────────────────────────────────────────────╮
│ agents plan correct 01HXM9A1C2Q7W3R5G8Z0P4Q1X9           │
│   --mode revert --guidance "Prioritize payments first..." │
╰───────────────────────────────────────────────────────────╯

✓ OK Decision explained
##### agents plan correct **Purpose** Correct a decision either by reverting and re-executing or by appending a fix. **Arguments** - ``: Decision ID. - `--mode revert|append`: Correction mode. - `--guidance/-g TEXT`: Guidance text. - `--guidance-file/-f PATH|-`: Guidance file (use `-` for stdin). - `--dry-run`: Show impact without executing. - `--yes`: Skip confirmation for revert mode. **Examples**

$ agents plan correct 01HXM9A1C2Q7W3R5G8Z0P4Q1X9 --mode revert \
  --guidance "Prioritize payments first" --yes

╭─ Correction ─────────────────────────────────────╮
│ Mode: revert                                     │
│ Impact: 3 decisions, 2 subplans, 5 artifacts     │
│ New Decision: 01HXM9B7Z3Q1Q8K2E9H7K3W2M8         │
│ Corrects: 01HXM9A1C2Q7W3R5G8Z0P4Q1X9            │
│ Attempt: 2                                       │
╰──────────────────────────────────────────────────╯

╭─ Affected Subtree ───────────╮
│ Decisions Invalidated: 3   │
│ Subplans Rolled Back: 2   │
│ Artifacts Archived: 5     │
│ Unaffected Decisions: 2   │
╰──────────────────────────────╯

╭─ Sandbox Rollback ─────────────╮
│ Checkpoint: cp_01HXM8C2       │
│ Files Reverted: 5             │
│ Status: restored              │
╰────────────────────────────────╯

╭─ Recompute ────────╮
│ Queued: 2 subplans │
│ ETA: 4m            │
╰────────────────────╯

╭─ History ──────────────────────────────────────────╮
│ - Original decision superseded                      │
│ - Prior artifacts archived for comparison           │
│ - agents plan diff --correction 01HXM9B7Z3Q1Q8K2.. │
╰─────────────────────────────────────────────────────╯

✓ OK Correction applied
##### agents plan diff **Purpose** Show diffs for a plan or a correction attempt. **Arguments** - `--plan `: Show diff for a plan. - `--correction `: Compare correction outputs. **Examples**

$ agents plan diff --plan 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Diff Summary ─────────────────────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J            │
│ Project: local/api-service                  │
│ Files Changed: 2                             │
│ Insertions: 12                               │
│ Deletions: 4                                 │
│ Net Change: +8 lines                         │
╰────────────────────────────────────────────────╯

╭─ Files ───────────────────────────────╮
│ Path                 Change  Status   │
│ ───────────────────  ──────  ──────── │
│ src/auth/session.py  +8 -2   modified │
│ src/auth/tokens.py   +4 -2   modified │
╰───────────────────────────────────────╯

╭─ Patch Preview ─────────────────────────╮
│ --- a/src/auth/session.py              │
│ +++ b/src/auth/session.py              │
│ @@ -12,4 +12,10 @@                      │
│ - import jwt                            │
│ + import sessionlib                     │
│ - def validate_token(...)               │
│ + def validate_session(...)             │
│ --- a/src/auth/tokens.py               │
│ +++ b/src/auth/tokens.py               │
│ @@ -5,3 +5,7 @@                         │
│ - TOKEN_EXPIRY = 3600                    │
│ + TOKEN_EXPIRY = 7200                    │
╰─────────────────────────────────────────╯

╭─ Risk Assessment ────────────────╮
│ API Compatibility: preserved    │
│ Test Coverage: maintained       │
│ Breaking Changes: none detected │
╰──────────────────────────────────╯

✓ OK Diff generated
##### agents plan artifacts **Purpose** List artifacts produced by a plan. **Arguments** - ``: Plan ID. **Examples**

$ agents plan artifacts 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Artifacts ──────────────────────────────────────────────╮
│ Path                   Type   Size    Change     Subplan │
│ ─────────────────────  ─────  ──────  ─────────  ─────── │
│ src/auth/session.py    write  2.1 KB  +8 -2      root    │
│ tests/test_session.py  write  4.7 KB  +47 -0     root    │
│ src/auth/tokens.py     edit   1.8 KB  +4 -2      root    │
│ tests/test_tokens.py   write  3.2 KB  +32 -0     auth    │
╰──────────────────────────────────────────────────────────╯

╭─ Summary ──────────╮
│ Total: 4           │
│ Writes: 2 (new)    │
│ Edits: 2 (modified) │
│ Deletes: 0         │
│ Total Size: 11.8 KB │
╰────────────────────╯

╭─ By Subplan ──────────────╮
│ Root Plan: 3 artifacts    │
│ auth-tests: 1 artifact   │
│ payment-tests: (pending) │
╰───────────────────────────╯

✓ OK 4 artifacts listed
##### agents plan prompt **Purpose** Provide additional guidance to a plan, typically when it is errored or awaiting input. **Arguments** - ``: Plan ID. - `""`: Guidance text. **Examples**

$ agents plan prompt 01HXM8C2ZK4Q7C2B3F2R4VYV6J "Use mocks for database tests"

╭─ Guidance Added ───────────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J       │
│ Guidance: Use mocks for database tests │
│ Scope: next execution step             │
│ Phase: execute                          │
│ State: errored → processing            │
╰────────────────────────────────────────╯

╭─ Decision Created ────────────────────────╮
│ Type: user_intervention                 │
│ ID: 01HXM9C5G7R2X8S3K4Z5Q8R6Y3        │
│ Parent: 01HXM9A1C2Q7W3R5G8Z0P4Q1X9     │
╰───────────────────────────────────────────╯

╭─ Queue ────╮
│ Pending: 1 │
│ Applied: 0 │
╰────────────╯

✓ OK Guidance queued
##### agents plan resume **Purpose** Resume an interrupted plan from its last checkpoint. **Arguments** - ``: Plan ID. **Examples**

$ agents plan resume 01HXM8C2ZK4Q7C2B3F2R4VYV6J

╭─ Plan Resumed ───────────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Checkpoint: cp_01HXM8C2          │
│ Phase: execute                   │
│ State: processing                │
│ Attempt: 2                       │
╰──────────────────────────────────╯

╭─ Checkpoint Details ──────────────╮
│ Label: before auth refactor     │
│ Created: 2026-02-08 13:04       │
│ Files at Checkpoint: 3          │
│ Decisions at Checkpoint: 2      │
╰───────────────────────────────────╯

╭─ Pending Guidance ─────────────────╮
│ 1. Use mocks for database tests │
╰────────────────────────────────────╯

✓ OK Plan resumed
##### agents plan checkpoint **Purpose** Create an explicit checkpoint in the plan sandbox. **Arguments** - ``: Plan ID. - `--label TEXT`: Optional label. **Examples**

$ agents plan checkpoint 01HXM8C2ZK4Q7C2B3F2R4VYV6J --label "before auth refactor"

╭─ Checkpoint ─────────────────────╮
│ ID: cp_01HXM8C2                  │
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Label: before auth refactor      │
│ Created: 2026-02-08 13:04        │
│ Sequence: 3 of 10               │
╰──────────────────────────────────╯

╭─ Captured State ──────────╮
│ Files Tracked: 3         │
│ Decisions Recorded: 4   │
│ Tool Calls Logged: 8    │
│ Sandbox Branch: cp-3    │
╰───────────────────────────╯

╭─ Retention ───────────╮
│ Policy: keep last 10  │
│ Auto-Cleanup: enabled │
╰───────────────────────╯

╭─ Storage ──────────╮
│ Delta Size: 640 KB │
│ Total: 3.1 MB      │
╰────────────────────╯

✓ OK Checkpoint created
##### agents plan rollback **Purpose** Rollback a plan sandbox to a checkpoint. **Arguments** - ``: Plan ID. - ``: Checkpoint ID. - `--yes`: Skip confirmation. **Examples**

$ agents plan rollback 01HXM8C2ZK4Q7C2B3F2R4VYV6J cp_01HXM8C2

Rollback plan 01HXM8C2ZK4Q7C2B3F2R4VYV6J to cp_01HXM8C2? [y/N]: y

╭─ Rollback Summary ───────────────╮
│ Plan: 01HXM8C2ZK4Q7C2B3F2R4VYV6J │
│ Checkpoint: cp_01HXM8C2          │
│ Label: before auth refactor      │
│ Files: 6 reverted                │
╰──────────────────────────────────╯

╭─ Changes Reverted ──────────────────╮
│ File                    Action     │
│ ──────────────────────  ────────── │
│ src/auth/session.py     restored   │
│ src/auth/tokens.py      restored   │
│ tests/test_session.py   removed    │
│ tests/test_tokens.py    removed    │
│ src/auth/fixtures.py    restored   │
│ src/auth/__init__.py    restored   │
╰─────────────────────────────────────╯

╭─ Impact ──────────────────────────────╮
│ Subplans Invalidated: 2             │
│ Sandbox: restored to cp_01HXM8C2   │
│ Decisions After CP: 2 discarded     │
│ Tool Calls After CP: 5 undone      │
╰───────────────────────────────────────╯

╭─ Post-Rollback State ─────────╮
│ Phase: execute                │
│ State: queued (awaiting input) │
│ Checkpoints Remaining: 2     │
╰───────────────────────────────╯

✓ OK Rollback complete
#### agents action **Purpose** Manage reusable actions. ##### agents action create **Purpose** Create a new action template. **Arguments** - ``: Namespaced action name. - `--strategy-actor/-s ACTOR`: Strategize actor. - `--execution-actor/-e ACTOR`: Execution actor. - `--definition-of-done/-d TEXT`: Completion criteria. - `--description TEXT`: Short description. - `--long-description TEXT`: Long description. - `--arg/-a spec`: Argument definition (repeatable). - `--reusable/--no-reusable`: Keep action after use. - `--read-only`: Read-only action. - `--tag/-t TAG`: Tags (repeatable). - `--available`: Make action available immediately. - `--estimation-actor ACTOR`: Optional estimation actor. - `--safety-profile NAME`: Named safety profile. - `--require-sandbox/--no-require-sandbox`: Enforce sandboxing. - `--require-checkpoints/--no-require-checkpoints`: Enforce checkpoints. - `--require-apply-approval/--no-require-apply-approval`: Require approval before apply. - `--allow-skill-category NAME`: Allow skill categories. - `--deny-skill-category NAME`: Deny skill categories. **Examples**

$ agents action create local/code-coverage \
  --strategy-actor local/strategist \
  --execution-actor local/executor \
  --definition-of-done "Coverage reaches 85%" \
  --arg "target_coverage_percent:int:required:Target coverage percentage" \
  --available

╭─ Action Created ──────────────────────╮
│ Name: local/code-coverage             │
│ ID: 01HXMAY3D1JQ0C3G1H0Q7B2W7M        │
│ State: available                      │
│ Strategy Actor: local/strategist      │
│ Execution Actor: local/executor       │
│ Reusable: yes                         │
│ Read Only: no                         │
│ Created: 2026-02-08 12:20             │
╰───────────────────────────────────────╯

╭─ Definition of Done ─╮
│ Coverage reaches 85% │
╰──────────────────────╯

╭─ Arguments ──────────────────────────────────────────────────────╮
│ Name                     Type    Required  Description          │
│ ───────────────────────  ──────  ────────  ───────────────────── │
│ target_coverage_percent  int     yes       Target coverage %    │
│ test_command             string  no        Test framework to use │
╰──────────────────────────────────────────────────────────────────╯

╭─ Safety ──────────────────╮
│ Sandbox: required         │
│ Apply Approval: required  │
│ Checkpoints: enabled      │
│ Skill Allow: (all)        │
│ Skill Deny: (none)        │
╰───────────────────────────╯

╭─ Usage ─────────────────────────────────────────────────────────────────╮
│ agents plan use local/code-coverage --project local/api-service         │
│   --arg target_coverage_percent=85                                      │
╰─────────────────────────────────────────────────────────────────────────╯

✓ OK Action created
##### agents action list **Purpose** List actions with optional filters. **Arguments** - `--namespace/-n NS`: Filter by namespace. - `--state/-s STATE`: Filter by state. - `--available`: Show only available actions. **Examples**

$ agents action list --available

╭─ Actions ──────────────────────────────────────────────────────────────────────────────────╮
│ Name                 State      Strategy Actor    Execution Actor  Reusable  Plans │
│ ───────────────────  ─────────  ────────────────  ───────────────  ────────  ───── │
│ local/code-coverage  available  local/strategist  local/executor   ✓         3     │
╰──────────────────────────────────────────────────────────────────────────────────────────╯

╭─ Filters ────────╮
│ State: available │
│ Namespace: (any) │
╰──────────────────╯

╭─ Summary ────────────╮
│ Total: 1             │
│ Available: 1         │
│ Draft: 0              │
│ Archived: 0          │
│ Total Plans Created: 3 │
╰──────────────────────╯

✓ OK 1 action listed
##### agents action show **Purpose** Show details for an action. **Arguments** - ``: Action ID or name. **Examples**

$ agents action show local/code-coverage

╭─ Action Details ──────────────────────╮
│ Name: local/code-coverage             │
│ ID: 01HXMAY3D1JQ0C3G1H0Q7B2W7M        │
│ State: available                      │
│ Strategy Actor: local/strategist      │
│ Execution Actor: local/executor       │
│ Reusable: yes                         │
│ Read Only: no                         │
│ Created: 2026-02-08 12:20             │
╰───────────────────────────────────────╯

╭─ Definition of Done ─╮
│ Coverage reaches 85% │
╰──────────────────────╯

╭─ Arguments ──────────────────────────────────────────────────────╮
│ Name                     Type    Required  Description          │
│ ───────────────────────  ──────  ────────  ───────────────────── │
│ target_coverage_percent  int     yes       Target coverage %    │
│ test_command             string  no        Test framework to use │
╰──────────────────────────────────────────────────────────────────╯

╭─ Safety ──────────────────╮
│ Sandbox: required         │
│ Apply Approval: required  │
│ Checkpoints: enabled      │
╰───────────────────────────╯

╭─ History ────────────────────╮
│ Plans Created: 3           │
│ Plans Completed: 2        │
│ Plans Failed: 0            │
│ Avg Duration: 00:04:30    │
│ Avg Cost: $0.072           │
╰──────────────────────────────╯

╭─ Usage ───────────────────────────────────────────────────────────╮
│ - agents plan use local/code-coverage --project local/api-service │
│     --arg target_coverage_percent=85                              │
╰───────────────────────────────────────────────────────────────────╯

✓ OK Action loaded
##### agents action available **Purpose** Mark a draft action as available. **Arguments** - ``: Action ID or name. **Examples**

$ agents action available 01HXMAY3D1JQ0C3G1H0Q7B2W7M

╭─ Action Available ─────────────╮
│ ID: 01HXMAY3D1JQ0C3G1H0Q7B2W7M │
│ State: draft → available       │
│ Name: local/code-coverage      │
╰────────────────────────────────╯

╭─ Visibility ─────╮
│ Namespace: local │
│ Listed: yes      │
│ Usable: yes      │
╰──────────────────╯

╭─ Validation ────────────────────────╮
│ Strategy Actor: resolved           │
│ Execution Actor: resolved          │
│ Definition of Done: present        │
│ Arguments: valid schema            │
╰─────────────────────────────────────╯

✓ OK Action marked available
##### agents action archive **Purpose** Archive an action. **Arguments** - ``: Action ID or name. **Examples**

$ agents action archive local/old-action

╭─ Action Archived ──────────╮
│ Name: local/old-action     │
│ State: available → archived │
│ Archived: 2026-02-08 12:22 │
╰────────────────────────────╯

╭─ Impact ───────────────────────╮
│ Availability: hidden from list │
│ Existing Plans: unchanged      │
│ Active Plans: 0 affected       │
╰────────────────────────────────╯

╭─ History ─────────────────╮
│ Total Plans: 5          │
│ Completed: 4            │
│ Failed: 1               │
│ Last Used: 2026-02-06  │
╰───────────────────────────╯

✓ OK Action archived
#### agents config **Purpose** Manage global configuration values. ##### agents config set **Purpose** Set a configuration key. **Arguments** - ``: automation-level, default-actor, log-level. - ``: Value to set. **Examples**

$ agents config set automation-level review

╭─ Config Updated ──────────────╮
│ Key: automation-level         │
│ Value: review                 │
│ Previous: manual              │
│ Source: config                │
│ Scope: global                 │
╰──────────────────────────────╯

╭─ Effective ────────────────────────╮
│ Sessions: new sessions           │
│ Plans: future plans (unless set) │
│ Existing: unchanged              │
╰────────────────────────────────────╯

╭─ Saved To ──────────────────────────╮
│ File: ~/.cleveragents/config.toml │
│ Line: 8                           │
╰─────────────────────────────────────╯

✓ OK Config updated
##### agents config get **Purpose** Get a configuration value. **Arguments** - ``: Key to read. **Examples**

$ agents config get default-actor

╭─ Config ──────────────────╮
│ Key: default-actor        │
│ Value: local/orchestrator │
│ Source: config            │
│ Overridden: no            │
│ Type: string              │
╰───────────────────────────╯

╭─ Origin ──────────────────────────╮
│ File: ~/.cleveragents/config.toml │
│ Line: 12                          │
│ Default: local/orchestrator      │
╰───────────────────────────────────╯

╭─ Resolution Chain ──────────────╮
│ 1. CLI flag: (not set)          │
│ 2. Env var: (not set)           │
│ 3. Config file: local/orchestra │
│ 4. Default: local/orchestrator  │
│ Winner: config file (level 3)   │
╰─────────────────────────────────╯

✓ OK Config read
##### agents config list **Purpose** List all configuration values. **Arguments** None. **Examples**

$ agents config list

╭─ Config ───────────────────────────────────────────────╮
│ Key               Value               Source   Modified  │
│ ────────────────  ──────────────────  ───────  ──────── │
│ automation-level  review              config   yes       │
│ default-actor     local/orchestrator  config   yes       │
│ log-level         INFO                default  no        │
│ sandbox-required  true                default  no        │
│ apply-approval    true                default  no        │
╰────────────────────────────────────────────────────────╯

╭─ Overrides ─────╮
│ Env: none       │
│ CLI Flags: none │
╰─────────────────╯

╭─ Config File ───────────────────────╮
│ Path: ~/.cleveragents/config.toml │
│ Size: 284 bytes                   │
│ Valid: yes                        │
╰─────────────────────────────────────╯

✓ OK 5 settings listed
#### agents providers **Purpose** Inspect provider availability. ##### agents providers list **Purpose** List available providers and whether credentials are configured. **Arguments** None. **Examples**

$ agents providers list

╭─ Providers ───────────────────────────────────────────────────────╮
│ Provider    Status       Default Model  Models Available │
│ ──────────  ───────────  ─────────────  ──────────────── │
│ openai      missing key  gpt-4          4                │
│ anthropic   configured   claude-3.5     3                │
│ openrouter  configured   gpt-4o         12               │
│ google      missing key  gemini-2.0     2                │
╰──────────────────────────────────────────────────────────────────╯

╭─ Credentials ─╮
│ Configured: 2 │
│ Missing: 2    │
╰───────────────╯

╭─ Routing ──────────────────────╮
│ Default: anthropic/claude-3.5  │
│ Fallback: openrouter/gpt-4o   │
╰────────────────────────────────╯

╭─ Key Sources ─────────────────────╮
│ anthropic: ANTHROPIC_API_KEY     │
│ openrouter: OPENROUTER_API_KEY   │
│ openai: OPENAI_API_KEY (missing) │
│ google: GOOGLE_API_KEY (missing) │
╰───────────────────────────────────╯

╭─ Rate Limits ───────────────╮
│ anthropic: 1M tokens/min  │
│ openrouter: 500K tok/min │
╰─────────────────────────────╯

✓ OK Providers listed
## Components ### Plan A **plan** is the fundamental unit of orchestration and traceability. #### Plan Lifecycle Phases A plan always moves through the following phases, in order: **Action → Strategize → Execute → Apply → Applied (terminal)** In this spec: * **Strategize** is the phase name (the output is a **strategy**). * **Execute** is the phase name (the output is a **changeset**). * **Apply** is the phase name (the output is an **applied change**). * **Applied** is the resulting terminal state after Apply succeeds. This four-stage model is explicitly called out as the new architecture replacing a prior linear pipeline. #### Phase Transition Verbs (CLI / UX Contract) Verbs that trigger phase transitions. CleverAgents should standardize these verbs as the **public API** (CLI, TUI, web): | Current Phase | Command Verb | Next Phase | | ------------- | ------------ | ---------- | | (none) | `create` | Action | | Action | `use` | Strategize | | Strategize | `execute` | Execute | | Execute | `apply` | Applied | **Important behavioral rule:** CleverAgents must support multiple "automation levels" that can automatically progress through these verbs without the user explicitly issuing them, but the verbs remain the conceptual contract. #### Plan States (Per Phase) A plan's phase indicates "what step of the lifecycle it is in." Separately, the plan has a **processing state** indicating "what is happening right now." Recommended state model: * **Action phase states** * `available` (action exists and can be used) * `draft` (action is being authored/edited) * `archived` (soft-deleted or hidden, optional) * **Strategize / Execute / Apply phase states** * `queued` (waiting for compute/worker) * `processing` (currently running) * `errored` (failed; includes error metadata) * `complete` (finished successfully) * `cancelled` (user/system cancelled; safe terminal for that phase) #### Plan Identity and Traceability Every plan should have: * **plan_id**: Unique, immutable ID (UUID or ULID). * **parent_plan_id**: Nullable; present for subplans. * **root_plan_id**: The top-most plan in the tree. * **attempt**: An integer attempt counter that increments when re-running a phase (e.g., re-executing after a fix). * **created_at / updated_at / completed_at** timestamps. * **created_by** (user identity / session identity). #### Plan Hierarchy (Subplans) and Parallelism A single plan should usually represent the smallest "complete" unit of work (similar to what would fit in one git commit). However: * **Plans are hierarchical.** * **Decisions about subplans are made during Strategize** (as `subplan_spawn` decision types). * **Subplans are actually spawned during Execute** (based on those decisions). * Subplans can run **in parallel** or **sequentially**. * The parent plan is responsible for **merging results**. This is core to the long-term objective: tackling large tasks while only recomputing parts of the decision tree when corrected. ##### Hierarchical Decomposition for Scale When handling massive tasks (e.g., converting Firefox to Rust), the system uses hierarchical decomposition: 1. **Root Level**: High-level architectural decisions - "Convert Firefox Renderer to Rust" - Decision: "Start with leaf modules, work inward" - Context: Module dependency graph (2,847 modules) 2. **Subsystem Level**: Major component decisions - "Phase 1: Convert utility libraries (no external deps)" - Each subsystem gets its own bounded context 3. **Module Level**: Individual module conversions - "Convert string_utils module" - Context: Only the 47 functions and 12 dependent files - Decision: "Use Rust's String type" 4. **File Level**: Specific file changes - Actual code transformations - Minimal context needed At each level, only the relevant context is loaded. The persistent decision graph means we can always reconstruct why we're converting a particular module and what constraints apply from higher-level decisions. ##### Subplan Spawning Mechanism In the actor definition for the execution actor, there are **nodes that act as skills** whose purpose is to generate subplans. The execution actor can call these skills to trigger subplans. ```yaml # Example: Execution actor with subplan spawning capability actors: code_executor: type: graph routes: execute_workflow: nodes: - name: spawn_test_subplan type: tool config: tools: - name: create_subplan code: | # This skill creates a subplan subplan = context.spawn_subplan( action="local/write-tests", target_files=input_data.files_to_test ) result = subplan.id ``` ##### Subplan Execution Modes * **Sequential**: Subplans execute one after another. If one fails, subsequent subplans are not started. * **Parallel**: Subplans execute concurrently. If one fails, others can continue. ##### Subplan Failure Handling * **Parallel execution**: Other parallel subplans continue even if one fails. * **Sequential execution**: Subsequent subplans are not called if a prior one fails. * **Note**: An "error" only occurs if an exception is thrown by the application (a bug). Plan failures (e.g., tests don't pass) are handled within the plan's logic, not as application errors. ##### Result Merging The way results are merged depends on the resource type: * **Git-compatible resources** (source code, text files): Git-style merge * **Databases**: Transaction coordination or sequential application * **Other resources**: Pluggable merge strategies based on resource type * **Non-mergeable resources**: May require sequential execution only #### The Plan "Decision Tree" and Visualization CleverAgents intends to record enough information to render: * an **ASCII tree** in the TUI, and * optionally a GUI tree via visualization tools (D3/Cytoscape) once the data exists. This implies each plan should persist: * decisions made, * the rationale (or at least the prompt/context snapshot that produced it), * dependencies ("this decision influenced these child plans"). This is required for "correcting plans" (see Behavior section). #### Decision Data Model ##### Relationship Between Plan Description and Decisions Each plan has a **description field** (inherited from the action's description, potentially with argument substitutions). This description acts as the **primary component of the prompt** fed to the strategy actor during the Strategize phase. **Decisions are choices that are NOT explicitly defined by the plan description.** They represent the gaps, ambiguities, or implementation details that must be resolved to execute the plan. For example: - **Plan description**: "Increase test coverage to 85%" - **Decisions that emerge**: - "Which modules should be prioritized?" (not specified in description) - "Should we use mocks or integration tests for the database layer?" (not specified) - "Should we refactor the auth module to make it more testable, or write tests around it as-is?" (not specified) ##### Decision Making Based on Autonomy Level Who makes decisions depends on the plan's **automation level**: | Automation Level | Who Makes Decisions | |-----------------|---------------------| | **Manual** | User is prompted for each decision point | | **Review-before-apply** | Actor makes decisions automatically during Strategize, user reviews before Apply | | **Full automation** | Actor makes all decisions automatically | When **automation allows automatic decisions**, the strategy actor uses its best judgment based on context, and records its reasoning in the decision's `rationale` field. When **user input is required**, the system pauses and prompts the user: ``` Decision required: Which modules should be prioritized for test coverage? Options identified by the strategy actor: 1. auth module (currently 45% coverage, high risk) 2. payment module (currently 52% coverage, high risk) 3. user module (currently 71% coverage, medium risk) Your choice (or provide custom guidance): _ ``` ##### The Prompt as the Root Decision **The prompt passed to the strategize actor is itself a decision node** in the decision tree—specifically, it's the **root decision** of type `prompt_definition`. This is important because: 1. **Every plan has its own prompt**: The root plan's prompt comes from the action description + user arguments. Subplan prompts are created by parent plans during their execution. 2. **Parent plans create subplan prompts**: When a parent plan spawns a subplan, it decides what prompt to give that subplan. This is recorded as a `prompt_definition` decision in the parent's tree, and becomes the root decision of the child plan's tree. 3. **Unified correction mechanism**: Since the prompt is just another decision, correcting it uses the same `agents plan correct` command as any other decision. ``` Plan Tree Example: ├── [prompt_definition] "Increase test coverage to 85%" <- Root decision (correctable) │ ├── [strategy_choice] "Prioritize auth and payment modules" │ ├── [subplan_spawn] "Write tests for auth module" │ │ └── Subplan: auth-tests │ │ ├── [prompt_definition] "Write unit tests for auth module using mocks" <- Created by parent │ │ ├── [implementation_choice] "Test login flow first" │ │ └── ... │ └── [subplan_spawn] "Write tests for payment module" │ └── Subplan: payment-tests │ ├── [prompt_definition] "Write unit tests for payment module" <- Created by parent │ └── ... ``` ##### Correcting Decisions (Including Prompts) All corrections use the same unified command: ```bash agents plan correct --mode= --guidance "" ``` **Parameters:** * ``: The ULID of the decision to correct * `--mode`: Either `revert` (rollback and re-run) or `append` (add fix at end) * `--guidance`: Free-form text specifying what the correct decision should be **Examples:** ```bash # Correct a strategy choice agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=revert \ --guidance "Prioritize the payment module first, not auth, due to upcoming deadline" # Correct the root prompt to be more specific agents plan tree # Shows: [prompt_definition] id=01ARZ3NDEKTSV4RRFFQ69G5FAV "Increase test coverage to 85%" agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=revert \ --guidance "Increase test coverage to 85%, prioritizing auth and payment modules. Use mocks for database tests, not integration tests." # Correct a subplan's prompt (originally created by parent plan) agents plan correct 01BRZ4PDFLUTW5SSGR70H6GBW --mode=revert \ --guidance "Write unit tests for auth module, focusing on edge cases for token expiration" # Append a fix rather than rewriting history agents plan correct 01ARZ3NDEKTSV4RRFFQ69G5FAV --mode=append \ --guidance "The previous approach missed error handling tests - add comprehensive error path coverage" ``` **Note:** CLI commands should not require interactive input. The `--guidance` parameter provides the correction inline. For very long guidance text, use a file: ```bash agents plan correct --mode=revert --guidance-file ./correction.txt ``` **When to correct the prompt vs. a specific decision:** | Situation | Correction Approach | |-----------|---------------------| | Original request was too vague | Correct the `prompt_definition` decision | | Strategy actor made a bad choice on a specific question | Correct that specific decision | | Parent plan gave a subplan a bad prompt | Correct the subplan's `prompt_definition` | Because the prompt is part of the decision tree, the system automatically knows that correcting it invalidates all downstream decisions in that plan (and its subplans). **Decisions are only created during the Strategize phase.** The decision tree captures what choices were made and why, enabling correction and replay. ##### Decision Record Structure ```yaml Decision: # Identity decision_id: ULID # Unique identifier plan_id: ULID # Parent plan this decision belongs to parent_decision_id: ULID | null # Parent decision (for tree structure) sequence_number: int # Order within the plan's decisions # Classification decision_type: enum - prompt_definition # The prompt/description for this plan (root decision) - strategy_choice # High-level approach decision during Strategize - implementation_choice # How to implement a specific task - resource_selection # Which resources to read/modify - subplan_spawn # Decision to create a subplan (spawned later in Execute) - tool_invocation # Which skill/tool to use - error_recovery # How to handle a failure - validation_response # Response to validation failure - user_intervention # User provided guidance/correction # The Decision Itself question: str # What question was being answered chosen_option: str # What was decided alternatives_considered: list[str] # Other options that were evaluated confidence_score: float | null # 0.0-1.0 if the actor provided confidence # Context Snapshot (for replay) context_snapshot: hot_context_hash: str # Cryptographic hash of the exact context hot_context_ref: str # Pointer to the full stored snapshot relevant_resources: list[ResourceRef] # Every file/symbol that influenced this decision actor_state_ref: str # Complete LangGraph checkpoint # When the system decides "refactor the authentication module to use async patterns," # it permanently records: # - Which files were examined to make that decision # - What symbols and dependencies were traced # - The exact code state that was analyzed # - The reasoning chain that led to this choice # - Alternative approaches that were considered but rejected # Rationale rationale: str # Why this option was chosen actor_reasoning: str | null # Raw LLM reasoning if available # Downstream Impact (populated during Execute phase) downstream_decision_ids: list[ULID] # Decisions that depend on this one downstream_plan_ids: list[ULID] # Subplans spawned because of this decision artifacts_produced: list[ArtifactRef] # Files/outputs created under this decision # Timestamps created_at: datetime # Correction Metadata is_correction: bool # Was this decision a correction of another? corrects_decision_id: ULID | null # If correction, which decision was replaced correction_reason: str | null # Why the correction was made superseded_by: ULID | null # If this decision was later corrected ``` ##### Decision Timing | Phase | Decision Activity | |-------|-------------------| | **Strategize** | Decisions are **created**. `downstream_plan_ids` is empty. | | **Execute** | Subplans are **spawned**. `downstream_plan_ids` is **populated** when subplans are created based on `subplan_spawn` decisions. | | **Apply** | No new decisions. History can be flagged for cleanup after successful apply. | ##### Decision Tree Storage Schema ```sql -- Core decision table CREATE TABLE decisions ( decision_id TEXT PRIMARY KEY, -- ULID plan_id TEXT NOT NULL, parent_decision_id TEXT, sequence_number INTEGER NOT NULL, decision_type TEXT NOT NULL, question TEXT, chosen_option TEXT NOT NULL, alternatives_considered TEXT, -- JSON array confidence_score REAL, rationale TEXT, actor_reasoning TEXT, context_snapshot TEXT NOT NULL, -- JSON blob is_correction BOOLEAN DEFAULT FALSE, corrects_decision_id TEXT, correction_reason TEXT, superseded_by TEXT, created_at TEXT NOT NULL, FOREIGN KEY (plan_id) REFERENCES plans(plan_id), FOREIGN KEY (parent_decision_id) REFERENCES decisions(decision_id), FOREIGN KEY (corrects_decision_id) REFERENCES decisions(decision_id), FOREIGN KEY (superseded_by) REFERENCES decisions(decision_id) ); -- Downstream relationships (many-to-many for DAG) CREATE TABLE decision_dependencies ( upstream_decision_id TEXT NOT NULL, downstream_decision_id TEXT NOT NULL, dependency_type TEXT NOT NULL, -- 'decision', 'plan', 'artifact' downstream_ref TEXT NOT NULL, -- The actual ID of decision/plan/artifact PRIMARY KEY (upstream_decision_id, downstream_decision_id, downstream_ref), FOREIGN KEY (upstream_decision_id) REFERENCES decisions(decision_id) ); -- Correction history CREATE TABLE correction_attempts ( attempt_id TEXT PRIMARY KEY, -- ULID plan_id TEXT NOT NULL, original_decision_id TEXT NOT NULL, new_decision_id TEXT, original_subtree_snapshot TEXT, -- Reference to archived state correction_reason TEXT, status TEXT NOT NULL, -- 'pending', 'executing', 'completed', 'failed' created_at TEXT NOT NULL, completed_at TEXT, FOREIGN KEY (plan_id) REFERENCES plans(plan_id), FOREIGN KEY (original_decision_id) REFERENCES decisions(decision_id), FOREIGN KEY (new_decision_id) REFERENCES decisions(decision_id) ); ``` ### Action #### What an Action Is An **action** is a reusable plan template that is not associated with any projects yet. **Important: Actions are created via CLI commands, NOT YAML configuration files.** YAML configuration files are only used to define actors. Examples: * "Increase test coverage to 80%" * "Refactor module X to be async-safe" * "Write an RFC for feature Y" * "Provision an infra cluster and validate access" (non-code) Actions are **intentionally project-agnostic** so they can be reused across projects. #### Action Creation (CLI) Actions are created using the CLI: ```bash agents action create \ --name "local/code-coverage" \ --description "Increase test coverage to target percentage" \ --strategy-actor "local/coverage-strategist" \ --execution-actor "local/coverage-executor" \ --definition-of-done "Coverage reaches target percentage; All new tests pass" \ --arg "target_coverage_percent:int:required:Target coverage percentage (0-100)" \ --arg "test_framework:str:optional:Test framework to use (pytest, unittest, etc.)" ``` **Required parameters:** * `--name`: Namespaced name (e.g., `local/code-coverage`, `myorg/deploy-action`) * `--strategy-actor`: Name of the actor to use for Strategize phase * `--execution-actor`: Name of the actor to use for Execute phase * `--definition-of-done`: Free-form text describing completion criteria **Optional parameters:** * `--description`: Human-readable description * `--arg`: Argument definitions (can be repeated). Format: `name:type:required|optional:description` * `--reusable`: Whether action remains available after use (default: true) * `--read-only`: Whether action only performs read operations (default: false) Arguments defined with `--arg` are values that will be: * Injected into the description and/or definition of done (via templating) * Passed into the context of the actors * Required when using the action on projects #### Action Data Model (Expanded) A plan in the Action phase has: ##### 1) `name` (namespaced) Format: * `[server:][namespace/]` Rules: * If server is omitted, default server is assumed unless namespace is `local`. * If namespace is omitted, default is `local`. * Names should be stable identifiers (kebab-case recommended). Examples: * `local/code-coverage` * `myusername/code-coverage` * `myorgname/code-coverage` * `prod:myorgname/code-coverage` (server-qualified) ##### 2) `short_description` Optional at creation; auto-filled if blank. ##### 3) `long_description` Optional but recommended for reusable actions. ##### 4) `definition_of_done` (DoD) Required. Must be explicit and testable. ##### 5) `actors` Two actors minimum: * **strategy_actor** (planner/architect) * **execution_actor** (builder/implementer) Actors can be: * an LLM agent (built-in), * a graph (custom yaml, or built-in), Note: graphs are hierarchical allowing them to reference other actors as nodes. Actor abstraction is central: an actor may be a single agent or an entire graph. ##### 6) `reusable` (boolean) * Default: `true`. * If `true`: using the action creates a *new* plan in Strategize while leaving the action available. * If `false`: action self-deletes (or auto-archives) after first use. ##### 7) `read_only` (boolean) * Default: `false`. * If `true`: the plan must only use read-only skills and must never modify resources (even in sandbox). * Read-only actions are still useful for "investigation reports," architecture reviews, or dry-run planning. ##### 8) `inputs_schema` (recommended addition) To make actions genuinely reusable, actions should declare their inputs: * required args (e.g., target coverage percent), * optional args (e.g., test framework), * validation rules (types, bounds). Example: * `target_coverage_percent`: integer 0–100 ##### 9) `safety_profile` (recommended addition) A policy bundle that can be applied to enforce safe execution: * allowed skill categories, * require checkpoints, * require sandbox, * require human approval at Apply. This relates to the "checkpointable skills + sandbox" approach for safe writing. ### Strategy (Strategize Phase) #### Using an Action (Transition to Strategize) The `use` command transitions an Action into the Strategize phase by applying it to one or more projects: ```bash # Basic usage agents plan use local/code-coverage --project my-api-service # Multiple projects agents plan use local/schema-update \ --project api-service \ --project web-frontend \ --project mobile-app # With action arguments agents plan use local/code-coverage \ --project my-api-service \ --arg target_coverage_percent=85 \ --arg test_framework=pytest # With explicit automation level agents plan use local/deploy-action \ --project staging-env \ --automation-level manual ``` **Parameters:** * `--project`: Project to apply the action to (can be repeated for multi-project plans) * `--arg`: Action argument values (format: `name=value`) * `--automation-level`: Override automation level for this plan When the action is used: 1. A new plan is created with a unique ULID 2. The plan's automation level is determined (plan > session > global) 3. The plan enters the **Strategize** phase 4. The `strategy_actor` begins analyzing the project(s) #### What Strategize Does When an action is **used** on projects, it becomes a **plan in Strategize**. Strategize is: * **read-only**, producing a plan of attack, * responsible for gathering context from project resources, * responsible for generating a strategy and subplan blueprint, * not allowed to execute subplans or modify resources. This "architect vs coder" separation is explicitly described as a core motivation. **Resource-aware dependency analysis**: During the Strategize phase, the strategy actor employs specialized mechanisms to compute precise dependency closures: ```python # Pseudocode of what happens inside a strategy actor def compute_closure_for_refactoring(target_module): closure = ResourceClosure() # Direct file dependencies closure.add_files(find_imports(target_module)) closure.add_files(find_includes(target_module)) # Symbol dependencies for symbol in extract_exported_symbols(target_module): closure.add_files(find_symbol_usage(symbol, scope='project')) # Test dependencies closure.add_files(find_tests_for_module(target_module)) # Build system dependencies closure.add_files(find_build_references(target_module)) return closure ``` The system leverages several key insights: - **Modular boundaries exist**: Even in legacy codebases, there are natural boundaries - **Changes are incremental**: We don't convert 50,000 files atomically - **Dependencies are sparse**: Most modules depend on a small fraction of the codebase - **Interfaces are narrow**: Public APIs are much smaller than implementations #### Strategize Data Model A plan in Strategize contains all Action fields plus: ##### 1) `projects` A list of projects the plan is used on. Important: A strategy plan may target **multiple projects**. Multi-project work in one "window" is considered a major usability advantage over tools that require being run from a single directory. ##### 2) `strategy_context` A structured object describing: * what resources were considered, * how they were retrieved, * what filtering/limits were applied, * what the actor saw. This matters because a plan must be debuggable and correctable later. Recommended fields: * `resource_refs`: IDs of resources used * `queries`: search queries performed * `selected_chunks`: chunk IDs + sources + reasons * `constraints`: context window limits, file ignore patterns * `generated_summaries`: if summarization occurred ##### 3) `strategy` The output plan: * steps (ordered and/or DAG), * conditions/branches ("if tests fail, do X"), * subplans to spawn (including which action templates to use), * evaluation criteria (how to know success), * risk assessment. ##### 4) `execution_blueprint` (recommended addition) Strategize should output not only narrative text but also a machine-usable blueprint: * list of tasks, * required skills, * expected outputs, * dependencies between tasks. This blueprint becomes the input to Execute. ##### 5) `cost_estimate` and `risk_estimate` (optional) **Cost and risk estimation is optional** but recommended for production use. When enabled, a specialized **estimation actor** analyzes: * The initial prompt/request * The strategy produced by the Strategize phase * Historical data from similar plans (if available) And produces estimates for: * LLM tokens/cost range * Number of steps/subplans expected * Expected risk of rollbacks * Estimated execution time **Implementation**: Similar to how there's a `strategy_actor` and `execution_actor` for each action, there can be an optional `estimation_actor` whose entire job is cost/risk estimation. This actor runs after Strategize completes (before Execute) and its output is informational only. ```bash # Example: Action with estimation actor agents action create \ --name "local/expensive-refactor" \ --strategy-actor "local/refactor-planner" \ --execution-actor "local/refactor-executor" \ --estimation-actor "local/cost-estimator" \ --definition-of-done "All code refactored according to plan" ``` This becomes critical in server/multi-user usage and cost controls. ### Execution (Execute Phase) #### What Execute Does Execute is where the plan actually performs work, but **in a sandboxed environment** that can later be reviewed and applied. Key properties: 1. **Work happens in a sandbox** All file modifications, generated artifacts, and intermediate outputs live in an isolated "execution workspace" until Apply. 2. **Execute may spawn subplans** Subplans are a first-class behavior of Execute: a parent plan can distribute work to child plans and merge results. 3. **Execute must support checkpointing / rollback (when enabled)** Checkpointable skills allow rolling back to a checkpoint ID to recover from partial failure or wrong turns. 4. **Execute produces a "reviewable diff"** Diff review sandbox is described as a differentiating feature: users can inspect changes before applying. #### Execution Workspace / Sandbox Model (Detailed) A sandbox isolates plan execution from the real project resources until Apply. ##### Key Sandbox Principles 1. **Lazy Sandboxing**: Resources are sandboxed **only when accessed**, not upfront. - A project may have many resources (git repo + 10 databases + cloud accounts) - A plan may only modify one resource - Only accessed resources are sandboxed - Efficient for large projects 2. **Per-Plan Sandboxes**: Each plan and subplan has its own sandbox containing only the resources it edits. 3. **Resource-Defined Strategy**: The sandbox strategy is defined on each **resource**, not by skills or globally. 4. **Cleanup Behavior**: - Sandboxes are cleaned up before application exit when possible - Abandoned sandboxes (from crashes, etc.) are cleaned up on next application run - Completed plan sandboxes are cleaned or archived based on retention policy ##### Sandbox Implementation Strategies Different resource types require different sandbox strategies: | Resource Type | Strategy | Rollback Mechanism | |--------------|----------|-------------------| | Git Repository | `git_worktree` | Git reset/checkout | | Filesystem | `copy_on_write` or `overlay` | Restore from snapshot | | Database | `transaction_rollback` | Transaction rollback | | Cloud Infra | `terraform_state` | Terraform plan (reversed) | | API Endpoint | `none` | Often not sandboxable | **1. Git worktree / branch sandbox (preferred for code)** * Create a worktree or temporary branch * All modifications are commits or staged changes * Apply merges/cherry-picks Pros: natural rollback, diff support, efficient Cons: requires git **2. Filesystem copy sandbox** * Copy project directory to a sandbox directory * Execute modifies sandbox copy * Apply syncs diff back Pros: simple Cons: expensive for huge repos **3. Overlay filesystem sandbox** * Use overlayfs-style "copy-on-write" to avoid full copies Pros: efficient Cons: more complex, OS-dependent **4. Transaction-based sandbox (for databases)** * Begin transaction at sandbox creation * All operations within transaction * Rollback on failure, commit on apply Pros: native to databases Cons: long-running transactions can cause issues **5. No sandbox (for non-sandboxable resources)** * Some resources cannot be sandboxed (certain APIs, cloud services) * User proceeds at their own risk * Plan should warn about non-sandboxable resources ##### Multi-Resource Sandboxing When a plan accesses multiple resources: * Each resource gets its own sandbox (based on its defined strategy) * Sandboxes are independent * Apply commits each sandbox separately * If any sandbox Apply fails, others may still succeed (partial apply) **Complete isolation during execution prevents compound errors**: Each plan executes in its own sandbox, which means: ``` Plan A (refactoring auth module): - Sandbox A1: Contains only auth/*.cpp, auth_tests/*.cpp - Cannot see Plan B's intermediate states - Cannot accidentally depend on Plan B's half-done work Plan B (updating API endpoints): - Sandbox B1: Contains only api/*.cpp, api_tests/*.cpp - Makes changes assuming current auth interface - Protected from Plan A's intermediate refactoring ``` **Hierarchical merge resolution**: When subplans complete, the parent plan performs intelligent merging: ```python def merge_subplan_results(subplan_results): # Group by resource type by_resource = group_by_resource_type(subplan_results) # Apply resource-specific merge strategies for resource_type, changes in by_resource: if resource_type == 'git_repo': merge_git_changes(changes) # Three-way merge elif resource_type == 'database': merge_db_changes(changes) # Sequential application elif resource_type == 'config_files': merge_config_changes(changes) # Smart JSON/YAML merge # Validate merged state run_integration_tests() ``` #### Execution Data Model A plan in Execute contains: ##### 1) `execution_context` The context used for execution (often smaller/more tactical than strategy context). ##### 2) `execution_log` Structured timeline of: * skill calls, * actor calls, * outputs, * errors and retries, * checkpoints created. This log is essential for debugging. ##### 3) `artifacts` Outputs produced: * changed files, * generated files, * reports, * diagrams, * test outputs, * diffs. ##### 4) `sandbox_ref` Pointer to the sandbox location/state: * path, branch name, workspace ID, container ID, etc. ##### 5) `checkpoint_graph` (if enabled) A record of checkpoints: * checkpoint ID * timestamp * skill responsible * resources affected * rollback instructions / metadata #### Checkpointing in Execute (Core Safety Mechanism) The intended user-level behavior is: * "Give me a checkpoint ID." * Perform additional operations. * "Roll back to checkpoint X." Not all skills can support this; checkpointing must be declared per skill. ##### Skill-level checkpointability Each skill declares: * `checkpointable: true|false` * `rollback_mechanism`: how rollback occurs * `scope`: what resources it can revert Examples: * File skill: snapshot file states pre-modification * Git skill: create commit or stash; rollback is reset/checkout * CLI skill inside a container: rollback by restoring filesystem snapshot or reloading base image state It should be noted that checkpointing is easier when skill scope is constrained (e.g., "only files within a docker image + git"). ##### Plan-level rollback policy Plans should have an option: * `rollback_enabled: true|false` If disabled, the plan may use more generic/unsafe skills with fewer restrictions (useful for low-stakes tasks). #### Execution as Transactions (Recommended) Execution should be treated like a transactional pipeline: * Each step either: * commits a checkpoint on success, or * rolls back to the previous checkpoint on failure. This is explicitly motivated by "partial failure leaves codebase inconsistent" and the need for transaction rollback. #### Tool-Based Resource Modification (Modern Architecture) **IMPORTANT: CleverAgents does NOT parse LLM output to extract code.** Instead, it uses the modern tool-based approach pioneered by Claude Code, Cursor, and Aider where: 1. **LLMs call tools/skills directly** (`edit_file()`, `write_file()`, `delete_file()`, etc.) 2. **Tools operate on the sandbox** - each tool invocation modifies sandbox state directly 3. **ChangeSet is built from tool invocations** - not by parsing LLM text output 4. **Validation runs on sandbox state** - after tools execute, not on parsed output This architecture provides: * **Atomic operations**: Each tool call is a discrete, trackable change * **No parsing ambiguity**: Tools have structured parameters (path, content, etc.) * **Resource-agnostic**: Same pattern works for files, databases, APIs, any resource type * **Safety by design**: Tools run in sandbox with defined capabilities and restrictions * **MCP compatibility**: Skills map directly to MCP tools for external integrations ##### How It Works ``` LLM Response (with tool calls) ↓ ┌─────────────────────────────────────┐ │ Skill/Tool Router │ │ - Routes each tool call to handler │ │ - Validates parameters │ │ - Enforces capability restrictions │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ Sandbox Execution │ │ - Tool operates on sandboxed state │ │ - Each invocation recorded │ │ - Checkpoint created if needed │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ ChangeSet Accumulation │ │ - Each resource-modifying call → │ │ becomes a Change record │ │ - ChangeSet = history of changes │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ Validation & Review │ │ - Run validators on sandbox state │ │ - Generate diff from ChangeSet │ │ - Present for review before Apply │ └─────────────────────────────────────┘ ``` ##### Built-in Resource Skills CleverAgents provides these core skills for resource manipulation: | Skill | Description | Creates Change? | |-------|-------------|-----------------| | `read_file(path)` | Read file contents | No | | `write_file(path, content)` | Create/overwrite file | Yes | | `edit_file(path, changes)` | Apply targeted edits | Yes | | `delete_file(path)` | Remove file | Yes | | `move_file(src, dst)` | Rename/move file | Yes | | `create_directory(path)` | Create directory | Yes | | `list_files(pattern)` | List files matching glob | No | | `search_files(pattern, content)` | Search file contents | No | | `get_file_info(path)` | Get file metadata | No | Each skill automatically: * Operates within sandbox boundaries * Records changes to the ChangeSet * Validates parameters against project configuration * Enforces deny-list patterns (`.git/`, `node_modules/`, etc.) ##### Why Not Parse LLM Output? The obsolete approach of parsing markdown code fences has fundamental problems: 1. **Ambiguity**: Is text explanation or code? Where does one file end and another begin? 2. **Fragility**: Models output varying formats; regex parsing is brittle 3. **Loss of semantics**: You lose the intent (create vs modify vs delete) 4. **No atomicity**: Can't rollback individual operations 5. **Resource-limited**: Only works for files, not databases or other resources The tool-based approach solves all of these by making each operation explicit, typed, and trackable. ### Semantic Error Prevention CleverAgents provides multiple layers of proactive error prevention that catch semantic errors before they can propagate through the system. #### Layer 1: Decision-time Validation During Strategize Every decision includes semantic validation: ```yaml Decision: Refactor payment module to async alternatives_considered: - "Convert to async/await patterns" (chosen) - "Use thread pool with channels" (rejected: doesn't integrate with async ecosystem) - "Keep synchronous with timeout" (rejected: doesn't solve core latency issue) confidence_score: 0.85 validation_performed: - Checked all payment API consumers can handle async - Verified database driver supports async operations - Confirmed no regulatory requirement for sync processing ``` #### Layer 2: Execution-time Semantic Guards The execution actor configuration includes validation nodes that understand semantics: ```yaml actors: code_executor: type: graph nodes: - name: semantic_validator type: tool config: tools: - name: validate_api_compatibility code: | # Not just syntax checking - semantic validation old_api = extract_api_signature(previous_version) new_api = extract_api_signature(current_version) breaking_changes = find_breaking_changes(old_api, new_api) if breaking_changes: # Don't just fail - understand the impact affected_consumers = find_api_consumers(breaking_changes) migration_plan = generate_migration(breaking_changes) if can_auto_migrate(affected_consumers, migration_plan): apply_migration(migration_plan) else: raise SemanticError( "Breaking API changes require manual review", changes=breaking_changes, affected=affected_consumers ) ``` #### Layer 3: Invariant Enforcement The system maintains semantic invariants: ```python class RefactoringInvariants: # User-defined invariants for the codebase invariants = [ "All public APIs must maintain backward compatibility", "Database transactions must complete within 5 seconds", "Authentication must always use OAuth2", "Payment processing must be idempotent" ] def check_invariant_preservation(self, changes): for invariant in self.invariants: if not self.verify_invariant(invariant, changes): return InvariantViolation(invariant, changes) return Success() ``` #### Layer 4: Predictive Error Prevention The system learns from past failures: ```yaml Error Pattern Database: - pattern: "Async conversion in payment module" historical_failures: - "Race condition in payment confirmation" - "Timeout handling breaks idempotency" preventive_checks: - "Add explicit transaction boundaries" - "Verify idempotency keys are preserved" - "Check distributed lock acquisition" ``` ### Applied (Apply Phase) #### What Apply Does Apply takes the sandboxed work product and makes it "real" in the project. Core properties: 1. **Apply is a controlled commit step** Apply exists specifically to separate "generated work" from "committed work," enabling review and safer automation. 2. **Apply is often the highest-risk step** It changes real systems. This is where permissions, approvals, and checks matter most. 3. **Apply produces a terminal 'applied' plan** After successful apply, the plan becomes Applied. #### Apply Responsibilities (Recommended Checklist) Apply should perform (configurable) validations before committing: * **Diff review gate** * If automation level requires review, show: * changed files summary, * full diff, * risk warnings. * **Pre-apply tests** * Run validation as defined by the actor and/or project (see below). * **Conflict resolution** * If applying to a git repo, handle rebase/merge conflicts safely. * **Audit log** * Record who applied, what changed, when, and why. #### Validation Configuration Validation is **defined by the actor configuration and project settings**, not hardcoded. **Actor-defined validation**: The execution actor's workflow should include validation nodes that: 1. Read project-defined test/validation commands from context 2. Execute appropriate tests based on resource type 3. Handle failures by iterating (attempting fixes) or escalating to user **Project-defined validation**: Each project should define how its resources are validated: ```bash # Example: Add validation configuration to a project agents project set-validation \ --project my-api-service \ --resource api-repo \ --test-command "pytest" \ --lint-command "ruff check ." \ --type-check-command "pyright" ``` This information is passed into the actor's context, allowing generic actors (not specific to any project) to execute appropriate validation. #### Validation Failure Handling When validation fails: 1. **Iteration**: Actor attempts to fix the issue (e.g., fix failing tests) 2. **Retry limit**: After several failed attempts, iteration stops 3. **User intervention**: User can provide additional instructions 4. **Resume**: Plan continues with new guidance The user can prompt the plan with additional instructions when stuck: ```bash agents plan prompt "Try using mock objects for the database tests" ``` #### Apply Data Model A plan in Apply includes: * `apply_summary` * `applied_artifacts` (final commit hash, merged PR link, file list) * `final_validation_results` (test outputs, lint outputs) * `approval_record` (if human approvals are required) * `deployment_record` (optional, if apply triggers deploy) #### "Applied" Terminal State When Apply succeeds: * plan.phase = `applied` * plan.state = `complete` * the sandbox may be cleaned up or archived depending on retention policy When Apply fails: * plan.phase remains Apply * plan.state = `errored` * sandbox remains intact for inspection/retry ### Project A **project** is the boundary that answers: * "Where is the work happening?" * "What can this plan read and write?" * "What tools can this plan use?" * "What context is available?" A project is a **collection of resources and configuration**. **Important: Projects are created via CLI commands, NOT YAML configuration files.** #### Project Types: Local vs Remote Projects are classified based on their resources: | Type | Definition | Where Plans Can Execute | |------|------------|------------------------| | **Local** | Contains at least one local-only resource | Client only | | **Remote** | All resources are remotely accessible | Client or Server | This distinction matters for server mode: the server can only execute plans on **remote** projects because it needs network access to all resources. #### Project Creation (CLI) ```bash # Create a new project agents project create \ --name "my-api-service" \ --tag "python" \ --tag "backend" # Add resources to the project agents project add-resource \ --project "my-api-service" \ --name "api-repo" \ --type "git_repository" \ --location "git@github.com:org/api-service.git" \ --sandbox-strategy "git_worktree" agents project add-resource \ --project "my-api-service" \ --name "staging-db" \ --type "database" \ --location "postgresql://staging.example.com/mydb" \ --sandbox-strategy "transaction_rollback" \ --read-only ``` #### Project Data Model A project includes: ##### 1) Identity * `project_id` (ULID) * `name` * `namespace` (follows same rules as actors: `local/`, `/`, `/`) * `tags` (e.g., "python", "infra", "paper", "prod") * `is_remote` (boolean, derived from resources) ##### 2) Resources Resources are the "things you can act on." **Each resource defines its own sandbox strategy.** Resource types include: * filesystem root(s) * git repository * database endpoints * cloud accounts * document corpora (papers, PDFs) * API schemas Each resource has: * `resource_id` (ULID) * `name` * `type` (e.g., `git_repository`, `database`, `filesystem`, `api_endpoint`) * `location` (path/url/connection string) * `is_remote` (boolean - can it be accessed over network?) * `sandbox_strategy` (defined on the resource - see Sandboxing section) * `read_only` (boolean - whether writes are allowed) * `metadata` (language, repo size, etc.) ##### 3) Context configuration Project-level defaults: * ignore patterns (like `.gitignore` semantics) * max file size * indexing strategy * preferred chunking/summarization policy (even if evolving) * context retention policy ##### 4) Security / permissions defaults * who can run write plans * which skills are restricted * whether apply requires approvals #### Multi-Project Operations A single plan may target multiple projects (e.g., updating shared schemas across services). This is considered a key UX advantage over "run in one directory" systems. In multi-project execution: * Strategize must clarify which steps affect which projects. * Execution must isolate sandboxes per project OR define a composite sandbox. * Apply must commit changes to each project separately, with separate approval records if necessary. ### Namespaces Namespaces define **ownership, scoping, and discoverability** of actions, projects, actors, and plans. **All named entities use the format `/`.** #### Namespace Types | Namespace | Scope | Storage | Examples | |-----------|-------|---------|----------| | `local/` | Current machine only | Local database | `local/my-reviewer`, `local/test-action` | | `/` | Personal server namespace | Server database | `freemo/code-analyzer`, `jsmith/deploy-script` | | `/` | Organization namespace | Server database | `cleverthis/standard-review`, `acme/deploy-action` | | `openai/`, `anthropic/`, etc. | Built-in LLM actors | N/A (built-in) | `openai/gpt-4`, `anthropic/claude-3-opus` | #### Namespace Rules * **`local/`** * Reserved namespace for local-only items * Exists only on the current machine * Stored in local database * Fast iteration, no sharing * Default namespace when none specified * **`/`** (e.g., `freemo/`, `jsmith/`) * Personal namespace on the server * Created when user registers an account * Stored on server, synced when connected * Used for reusable actions/actors a user wants across machines * Only the owning user can create/modify items * **`/`** (e.g., `cleverthis/`, `acme/`) * Organization namespace on the server * Created when organization is registered * Shared across team members * Permissions and approvals managed at org level * Actions/actors/projects can be centrally managed * **Built-in Provider Namespaces** (`openai/`, `anthropic/`, `google/`, etc.) * Reserved for built-in LLM actors * Automatically available when API keys are configured * In server mode: available if logged in and server has keys * In local mode: requires environment variables or app configuration * Cannot be used for custom actors #### Server-qualified Names To disambiguate between servers (when connected to multiple): * `dev:freemo/code-coverage` (personal namespace on dev server) * `prod:cleverthis/deploy-action` (org namespace on prod server) This enables a pattern where: * local machine runs a lightweight client * server stores canonical definitions * multiple servers can coexist ### Actor #### What an Actor Is An **actor** is the abstraction that generalizes "agent" into "anything conversational." * It can be as small as a single LLM agent. * It can also be an entire graph that itself calls other actors/tools. * Actors can be nested/hierarchical, enabling "orchestrator of orchestrators." **Every custom actor IS a graph** (a LangGraph defined via YAML configuration). Even a simple actor wrapping a single LLM is technically a graph with one node. #### Actor Naming Actors are **always named using `/` format**: * `local/my-reviewer` - Local actor * `freemo/code-analyzer` - Personal server actor * `cleverthis/deploy-specialist` - Organization actor * `openai/gpt-4` - Built-in LLM actor #### Actor Definition (YAML Configuration) Actors are defined via **YAML configuration files**. This is the ONLY place YAML configuration is used (not for actions or projects). Example actor configuration (see `examples/` directory for full examples): ```yaml cleveragents: version: "3.0" default_actor: workflow_controller actors: # Simple LLM actor my_assistant: type: llm config: actor: openai/gpt-4 # Reference to built-in actor temperature: 0.7 system_prompt: | You are a helpful assistant. Current task: {{ context.task_description }} # Tool actor with inline Python data_processor: type: tool config: tools: - name: process_data code: | # Inline Python code result = do_something(input_data, context) # Actor referencing another actor reviewer: type: llm config: actor: local/code-reviewer # Reference to another custom actor memory_enabled: true max_history: 20 routes: main_workflow: type: graph entry_point: start nodes: - name: analyze type: agent agent: my_assistant - name: process type: agent agent: data_processor edges: - source: start target: analyze - source: analyze target: process - source: process target: end context: global: task_description: "Default task" ``` #### Actor Arguments **All actors can receive arguments when invoked**, including built-in actors. Arguments are passed when: 1. An action is used on projects (arguments flow to strategy/execution actors) 2. An actor is directly invoked Arguments are injected into the actor's context and can be used in Jinja2 templates within prompts. For built-in actors (like `openai/gpt-4`), common arguments include: * `temperature` * `max_tokens` * `system_prompt` #### Actor Composition (Hierarchical References) Actors can reference other actors **by name**: ```yaml actors: complex_workflow: type: llm config: actor: local/base-analyzer # References another actor ``` **Load order matters**: Referenced actors must be loaded/defined before actors that depend on them. This enables hierarchical composition where: * Actor A's graph can include nodes that call Actor B (by name) * Actor B itself is a graph that might call Actor C * And so on... #### Actor vs Agent (Relationship) * **Agent**: an actor that is specifically an LLM with tools and reasoning behaviors. * **Actor**: may be an agent, but may also be: * a composite workflow, * a multi-step graph, * a wrapper around a third-party system (as long as it's "text in → text out" conversationally). #### Actor Definition Fields (From Notes + Extended) A robust actor schema should include: * `name` (namespaced) * `provider` (LLM provider or runtime target) * `model` * `system_prompt` (or prompt template) * `tool_access_policy` * `graph_descriptor` (for composite actors) * `memory_policy` (per-plan/per-actor—see memory section) * `context_view_policy` (what context this actor sees) * `limits` (token limits, tool call limits, retries) * `cost_policy` (caps, budgets) * `metadata` (tags, use cases, version) #### Actor Composition and Graphs Actors can reference: * other actors * skills (MCP/custom tools) * subgraphs This is central to enabling both: * multi-agent orchestration, and * modular reuse of workflows. #### Nodes in the Graph: Actor, Skill, or Custom Tool The transcript explicitly frames the graph nodes as being any of: * an actor, * an MCP skill, * a custom tool node (arbitrary python code). This is a powerful simplification: "everything is a node." ### Agent #### Agent Definition In CleverAgents, an **agent** is a specialized actor with: * a conversational interface, * tool-calling capability, * potentially memory, planning heuristics, and role identity. Examples of agent roles: * planner/architect (strategy actor) * coder/implementer (execution actor) * reviewer/qa agent * release/apply agent The transcript explicitly discusses role separation like planner/coder/reviewer in context views/memory proposals. #### Agent Behavior Configuration Agents should be configurable without code changes: * prompt templates * tool sets * safety constraints * style constraints (verbosity, code style) * reliability controls (self-checks, validations) A design goal is user empowerment: "users customize LLM behavior without modifying core code." ### Skills #### What a Skill Is A **skill** is an executable capability exposed to the system. **Skills are defined inline in actor YAML configuration files** as tool nodes. Examples: * file read/write * git operations * shell command execution * database query execution * API calls * vector search / RAG query * cloud administration actions #### Skill Definition **Skills do not have separate names.** They are defined as part of actor configurations: ```yaml actors: my_processor: type: tool config: tools: - name: process_data # Tool name within this actor code: | # Inline Python code result = process(input_data, context) - name: validate_output code: | # Another tool in the same actor if not is_valid(input_data): raise ValueError("Invalid output") result = input_data ``` **To create a "named skill"**: Define an actor with a single tool node. The actor name then effectively becomes the skill name. #### Skill Standards Skills should extend two complementary standards: 1. **MCP (Model Context Protocol)**: https://modelcontextprotocol.io/ - Tools: Schema-defined operations LLMs can invoke - Resources: Read-only data sources - Prompts: User-controlled templates 2. **Agent Skills Standard**: https://agentskills.io/ - Folders of instructions, scripts, and resources - Portable procedural knowledge across agent products CleverAgents extends the MCP resource concept to support **both read and write operations**. #### Skill Sources Skills come from multiple places: 1. **Built-in skills** (first-party, when implemented) * files, git, search, indexing, context mgmt, sandbox ops 2. **MCP skills** (external MCP servers) 3. **Imported skills** from other ecosystems * wrap tools from other agent frameworks as long as interface is compatible 4. **Custom tool nodes** (most common) * arbitrary python code embedded in actor configuration * behaves as a graph node #### Skill Capability Metadata (Critical for Safety) MCP's metadata is not sufficient (read-only/idempotent isn't enough; write scope is unclear). CleverAgents needs extended metadata for each skill. Required skill metadata: * `read_only: bool` - Whether skill only performs read operations * `writes: bool` - Whether skill can modify resources * `write_scope`: * file paths allowed * resource IDs allowed * environment boundaries (container only vs host) * `idempotent: bool` - Whether repeated calls produce same result * `checkpointable: bool` - Whether skill supports checkpoint/rollback * `checkpoint_scope` - What can be rolled back * `side_effects` - install packages, mutate infra, etc. * `required_permissions` - What permissions needed to use * `rate_limits` / `cost_profile` - Usage constraints * `human_approval_required: bool` - Optional approval gate #### Read-Only Actions When an action is marked `read_only: true`, it can **only use skills that have `read_only: true`** in their metadata. This is enforced at runtime. #### Skill Registry / Catalog To scale, CleverAgents should maintain a catalog of skills with metadata: * auto-extracted from MCP descriptors (where possible) * refined manually via annotations * enhanced by CleverAgents-specific extensions This registry supports: * plan validation ("this plan requires checkpointable write skills; do we have them?") * safe automation ("don't ask permission for every tiny command—use sandbox/checkpoints instead") #### MCP Integration Architecture CleverAgents fully integrates with the **Model Context Protocol (MCP)** while extending it for agentic workflows: ##### MCP Concepts Mapping | MCP Concept | CleverAgents Equivalent | Extension | |------------|------------------------|-----------| | Tool | Skill | Extended metadata (write_scope, checkpointable) | | Resource | Resource | Read AND write operations | | Prompt | Action template | Full plan lifecycle | | Server | MCP Server (external) | Integrated via skill adapters | ##### Using External MCP Servers CleverAgents can connect to any MCP server and expose its tools as skills: ```yaml actors: github_ops: type: tool config: mcp_servers: - name: github command: "npx @anthropic/mcp-github" env: GITHUB_TOKEN: "${GITHUB_TOKEN}" - name: filesystem command: "npx @anthropic/mcp-filesystem" args: ["--root", "/workspace"] ``` When an actor specifies `mcp_servers`, all tools from those servers become available as skills within that actor's execution context. ##### MCP Tool → Skill Adapter External MCP tools are automatically wrapped with CleverAgents skill semantics: ``` ┌─────────────────────────────────────┐ │ MCP Server │ │ - Exposes tools via JSON-RPC │ │ - Has MCP metadata (read-only, etc) │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ MCPSkillAdapter │ │ - Wraps MCP tool as Skill │ │ - Infers extended metadata │ │ - Intercepts calls for: │ │ - Sandbox path rewriting │ │ - Change tracking │ │ - Permission enforcement │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ Skill Execution │ │ - Runs in plan's sandbox context │ │ - Changes recorded to ChangeSet │ │ - Checkpoints created as needed │ └─────────────────────────────────────┘ ``` ##### Skill Execution Flow (Tool-Based Architecture) When an LLM decides to use a skill, the following flow occurs: ``` 1. LLM generates tool call: edit_file(path="src/main.py", changes=[...]) ↓ 2. Tool Router receives call - Validates parameters against schema - Checks skill capability metadata - Enforces permission restrictions ↓ 3. Sandbox Context Resolution - Maps logical path to sandbox path - Ensures sandbox exists for resource - Creates checkpoint if skill is checkpointable ↓ 4. Skill Execution - Runs skill code (built-in or MCP) - Operations occur on sandboxed state - Result captured ↓ 5. Change Recording - If skill modifies resources, create Change record - Append Change to plan's ChangeSet - Update sandbox state ↓ 6. Return to LLM - Return skill result - LLM continues with next action ``` ##### Built-in Skills (Core Resource Operations) CleverAgents provides these built-in skills that work with any resource through the unified abstraction layer: **File Operations:** ```python read_file(path: str) -> str write_file(path: str, content: str) -> None edit_file(path: str, edits: list[Edit]) -> None delete_file(path: str) -> None move_file(source: str, destination: str) -> None copy_file(source: str, destination: str) -> None ``` **Directory Operations:** ```python create_directory(path: str) -> None list_directory(path: str, pattern: str = "*") -> list[str] delete_directory(path: str, recursive: bool = False) -> None ``` **Search Operations:** ```python search_files(pattern: str, content_pattern: str = None) -> list[Match] find_definition(symbol: str) -> list[Location] find_references(symbol: str) -> list[Location] ``` **Git Operations (when resource is git repository):** ```python git_status() -> GitStatus git_diff(path: str = None) -> str git_log(count: int = 10) -> list[Commit] git_blame(path: str) -> list[BlameLine] ``` Each built-in skill: * Has fully defined capability metadata * Operates through the resource abstraction layer * Automatically tracks changes to the ChangeSet * Respects sandbox boundaries and deny-lists ##### Change Tracking from Tool Invocations **Critical Architecture Point:** The ChangeSet is NOT built by parsing LLM output. It is built by recording the effects of tool/skill invocations: ```python class SkillExecutionContext: """Context provided to skill execution.""" def __init__(self, plan: Plan, sandbox: Sandbox): self.plan = plan self.sandbox = sandbox self.changes: list[Change] = [] def record_change(self, change: Change) -> None: """Record a change made by a skill.""" self.changes.append(change) self.plan.changeset.add_change(change) class WriteFileSkill: """Built-in skill for writing files.""" def execute(self, path: str, content: str, ctx: SkillExecutionContext) -> None: # Get the resource handler for this path handler = ctx.sandbox.get_handler(path) # Perform the write (returns Change record) change = handler.write(path, content, ctx.sandbox) # Record the change ctx.record_change(change) ``` This approach means: * Every resource modification is explicit and tracked * The ChangeSet accurately reflects what was done, not what was said * Rollback is precise (replay inverse of recorded changes) * Audit logs show exactly what each skill invocation did ### Session #### What a Session Is A **session** is a user's interactive thread with CleverAgents across time. A session should: * maintain conversational continuity, * store plan references, * persist memory (if enabled), * provide a UI anchor (CLI invocation, TUI workspace, web session). #### Session and Memory Persistence The notes include a known issue: conversation history can be lost between CLI invocations depending on connection string configuration, implying the system needs a stable memory service backend. Therefore, CleverAgents should specify: * sessions have stable IDs * sessions can be resumed * session storage backend is configured explicitly * if session persistence is disabled, the UX should be explicit about it (no silent history loss) ### Server #### What a Server Is A server is an optional mode that enables: * multi-user access * shared org namespaces (`/` and `/`) * persistent plan records * remote plan execution * permissioning and governance **Single-user local mode is the default** (so setup is easy), but the architecture anticipates server mode for shared skills and org-level actions. #### Client-Only vs Server Mode | Mode | Description | Plan Execution | |------|-------------|----------------| | **Client-only** | No server connection. All data in local database. | Always local | | **Server mode** | Connected to a CleverAgents server. Namespaced items sync. | Local or server | **It is possible to run a client with no server at all.** Server is optional. #### Plan Execution Location Where a plan executes depends on the project type: | Project Type | Client Execution | Server Execution | |--------------|-----------------|------------------| | **Local** (has local-only resources) | ✅ Yes | ❌ No (server can't access local resources) | | **Remote** (all resources remotely accessible) | ✅ Yes | ✅ Yes | When acting on **local projects**, the client must be running because only the client can access local resources. **Remote projects** can execute on either: - The client (if user prefers local execution) - The server (for long-running plans, since client may be transient) #### Server Execution Benefits Server execution is useful when: * Plans take a long time to execute * Client may disconnect (laptop closes, network issues) * Multiple team members need to monitor plan progress * Centralized logging and auditing required #### No Plan Queuing **Plans are not queued.** When a plan is used on projects and executed, it runs immediately. There is no worker queue or delayed execution model. #### Multi-user Risks and Prompt Injection Prompt injection isn't critical in single-user mode but becomes important for multi-user server environments. Server mode must include: * permission boundaries * prompt sanitization / safe templating * resource access controls * auditing ### Permissions Permissions exist at multiple layers: #### 1) Namespace-level permissions * Who can create/edit org actions/actors/projects? * Who can run them? #### 2) Project-level permissions * Who can modify project resources? * Who can apply changes? #### 3) Plan-level permissions * Can this plan write? * Does it require approvals? * Can it access restricted skills? #### 4) Skill-level permissions * Some skills should require: * explicit user approval per call, or * elevated role membership. #### Approval Gates by Phase (Recommended) A simple and powerful governance model: * Strategize: generally safe, read-only → minimal restrictions * Execute: writes occur but sandboxed → moderate restrictions * Apply: writes are real → strict restrictions + optional mandatory review This aligns with the four-phase model's safety rationale. ### Resources Resources are any objects a plan can reason about or manipulate. **CleverAgents extends the MCP resource concept** to support both read AND write operations (MCP resources are read-only). #### Resource Types (Examples) * **FilesystemResource**: directories, files * **GitRepoResource**: git repos, branches * **DatabaseResource**: SQL/NoSQL endpoints * **CloudResource**: clusters, accounts, infra configs * **DocumentCorpusResource**: PDFs, markdown docs, wikis * **APISpecResource**: OpenAPI/Swagger, Postman collections * **IssueTrackerResource**: tickets, bugs, tasks (optional) #### Resource Sandbox Strategy **Each resource defines its own sandbox strategy.** This is critical because: 1. The same resource may be accessed through different skills 2. Different resource types require different sandboxing approaches 3. Some resources cannot be sandboxed at all | Resource Type | Sandbox Strategy | Rollback Mechanism | |--------------|------------------|-------------------| | Git Repository | `git_worktree` | Git reset/checkout | | Filesystem | `copy_on_write` or `overlay` | Restore from snapshot | | Database | `transaction_rollback` | Transaction rollback | | Cloud Infra | `terraform_state` | Terraform apply (reversed) | | API Endpoint | `none` (often not sandboxable) | N/A | Example resource with sandbox strategy: ```bash agents project add-resource \ --project "my-api" \ --name "main-repo" \ --type "git_repository" \ --location "git@github.com:org/api.git" \ --sandbox-strategy "git_worktree" ``` #### Lazy Sandboxing Resources are **sandboxed lazily when accessed**, not upfront. Note that this is different from indexing - resources are indexed immediately when added to a project, but sandboxes are only created when execution needs to modify a resource: 1. A project may contain many resources (e.g., git repo + 10 databases + cloud accounts) 2. A plan may only need to modify one resource 3. Only the accessed resources are sandboxed 4. Each plan/subplan has its own sandbox containing only edited resources This is efficient for large projects where most resources remain untouched. #### Resource Access Tracking The system needs to know **which skills touch which resources** to reason about safety and checkpointing. Every skill call should log: * resource IDs accessed * read/write actions * file paths or object IDs touched This enables: * better context assembly * better rollback feasibility analysis * better auditing * accurate sandbox scoping #### Unified Resource Abstraction Layer CleverAgents provides a unified abstraction that allows skills to work with any resource type through a consistent interface. This enables: 1. **Resource-agnostic skills**: A skill like `read_content(path)` works whether the path refers to a file, database record, or API endpoint 2. **Consistent sandbox semantics**: All resources support the same sandbox lifecycle (create, read, write, checkpoint, rollback) 3. **Pluggable resource handlers**: New resource types can be added without modifying existing skills 4. **Unified change tracking**: All resource modifications flow into the same ChangeSet model ##### Resource Handler Interface Every resource type implements this interface: ```python class ResourceHandler(Protocol): """Handler for a specific resource type.""" def read(self, path: str, sandbox: Sandbox) -> Content: """Read content from the sandboxed resource.""" ... def write(self, path: str, content: Content, sandbox: Sandbox) -> Change: """Write content and return the Change record.""" ... def delete(self, path: str, sandbox: Sandbox) -> Change: """Delete resource and return the Change record.""" ... def list(self, pattern: str, sandbox: Sandbox) -> list[str]: """List paths matching pattern.""" ... def diff(self, path: str, sandbox: Sandbox) -> str: """Generate diff between sandbox and original state.""" ... def supports_operation(self, operation: OperationType) -> bool: """Check if this resource supports the given operation.""" ... ``` ##### Built-in Resource Handlers | Resource Type | Handler | Read | Write | Delete | Sandbox Strategy | |--------------|---------|------|-------|--------|------------------| | Filesystem | `FilesystemHandler` | ✓ | ✓ | ✓ | copy_on_write | | Git Repository | `GitHandler` | ✓ | ✓ | ✓ | git_worktree | | PostgreSQL | `PostgresHandler` | ✓ | ✓ | ✓ | transaction | | SQLite | `SQLiteHandler` | ✓ | ✓ | ✓ | copy_on_write | | HTTP API | `HTTPHandler` | ✓ | ✓* | ✓* | none | | S3 Bucket | `S3Handler` | ✓ | ✓ | ✓ | versioning | *HTTP writes may not be sandboxable depending on the API ##### Resource Path Resolution Paths in skills are resolved through a resource routing system: ``` path://resource-name/relative/path ↓ ┌─────────────────────────────────────┐ │ Resource Router │ │ - Parses path scheme │ │ - Looks up resource by name │ │ - Routes to appropriate handler │ └─────────────────────────────────────┘ ↓ ┌─────────────────────────────────────┐ │ Handler (e.g., GitHandler) │ │ - Resolves relative path │ │ - Operates on sandboxed state │ │ - Returns Change record │ └─────────────────────────────────────┘ ``` For convenience, paths without a scheme default to the project's primary filesystem resource. ### Code Intelligence & Context Discovery #### Overview CleverAgents employs a sophisticated multi-layered indexing and discovery system that enables agents to efficiently navigate and understand codebases of any scale. This system goes far beyond simple text search, providing semantic understanding of code structure, dependencies, and relationships through a combination of indexed embeddings, vector search, and an RDF-based graph store. **Critical Design Decision**: All indexing happens immediately when resources are added to projects or when code changes. There is no "on-demand" indexing during agent execution. This ensures that agents always have instant access to search capabilities without any indexing delays. The computational cost is paid once upfront, not repeatedly during agent operations. **Key Design Principles:** 1. **Pluggable Architecture**: Every component can be extended or replaced 2. **Progressive Enhancement**: System works with basic text search, enhances with advanced features 3. **Eager Indexing**: Indices are built immediately when resources are added and kept continuously up-to-date 4. **Agent Awareness**: Agents understand available indices through skills 5. **Real-time Synchronization**: Indices update immediately as code changes #### Architecture Components ##### 1. Multi-Modal Indexing Engine The indexing engine operates across three complementary modalities: ```yaml IndexingEngine: modalities: # Traditional text-based indexing text_index: type: "full_text_search" backend: "tantivy" | "elasticsearch" | "sqlite_fts" features: - Token-based search - Regex patterns - Language-aware tokenization - File path indexing # Semantic understanding via embeddings vector_index: type: "embedding_search" backend: "faiss" | "qdrant" | "weaviate" | "pgvector" models: - code: "codegen-6B-multi" - docs: "instructor-xl" - cross-modal: "clip-code" features: - Function-level embeddings - Class-level embeddings - Module-level embeddings - Documentation embeddings - Cross-language similarity # Structural understanding via graph graph_index: type: "rdf_knowledge_graph" backend: "blazegraph" | "stardog" | "apache_jena" | "neo4j" ontology: "CodeOntology" features: - AST-based relationships - Dependency graphs - Call graphs - Inheritance hierarchies - Data flow analysis ``` ##### 2. RDF-Based Code Knowledge Graph The graph store represents code as a rich semantic network using RDF (Resource Description Framework) triples. This enables sophisticated queries about code structure and relationships. **Core Ontology Design:** ```turtle # CodeOntology - Core vocabulary for code representation @prefix code: . @prefix rdfs: . @prefix xsd: . # Core Classes code:Module a rdfs:Class ; rdfs:comment "A code module (file, package, namespace)" . code:Class a rdfs:Class ; rdfs:comment "A class or similar construct" . code:Function a rdfs:Class ; rdfs:comment "A function, method, or procedure" . code:Variable a rdfs:Class ; rdfs:comment "A variable, constant, or field" . code:Type a rdfs:Class ; rdfs:comment "A type definition" . # Core Properties code:contains a rdf:Property ; rdfs:domain code:Module ; rdfs:range code:Entity ; rdfs:comment "Module contains entity" . code:imports a rdf:Property ; rdfs:domain code:Module ; rdfs:range code:Module ; rdfs:comment "Module imports another module" . code:extends a rdf:Property ; rdfs:domain code:Class ; rdfs:range code:Class ; rdfs:comment "Class inheritance relationship" . code:calls a rdf:Property ; rdfs:domain code:Function ; rdfs:range code:Function ; rdfs:comment "Function calls another function" . code:references a rdf:Property ; rdfs:domain code:Entity ; rdfs:range code:Entity ; rdfs:comment "Entity references another entity" . code:hasParameter a rdf:Property ; rdfs:domain code:Function ; rdfs:range code:Parameter ; rdfs:comment "Function has parameter" . code:returns a rdf:Property ; rdfs:domain code:Function ; rdfs:range code:Type ; rdfs:comment "Function return type" . # Annotations code:hasDocstring a rdf:Property ; rdfs:domain code:Entity ; rdfs:range xsd:string . code:hasComplexity a rdf:Property ; rdfs:domain code:Function ; rdfs:range xsd:integer ; rdfs:comment "Cyclomatic complexity" . code:hasTestCoverage a rdf:Property ; rdfs:domain code:Entity ; rdfs:range xsd:decimal ; rdfs:comment "Test coverage percentage" . ``` **Example Knowledge Graph Fragment:** ```turtle # Concrete example: Authentication module a code:Module ; code:imports ; code:imports ; code:contains . a code:Class ; code:hasMethod ; code:hasMethod ; code:extends ; code:hasDocstring "Manages user authentication and session tokens" . a code:Function ; code:hasParameter ; code:hasParameter ; code:returns ; code:calls ; code:calls ; code:hasComplexity 8 ; code:hasTestCoverage 0.95 . ``` **Advanced Graph Queries:** ```sparql # Find all functions that manipulate user authentication PREFIX code: SELECT ?function ?module WHERE { ?function a code:Function ; code:calls*/code:references ?entity . ?entity rdfs:label ?label . FILTER(CONTAINS(LCASE(?label), "auth") || CONTAINS(LCASE(?label), "user")) ?module code:contains ?function . } # Find circular dependencies SELECT ?module1 ?module2 WHERE { ?module1 code:imports+ ?module2 . ?module2 code:imports+ ?module1 . FILTER(?module1 != ?module2) } # Find most complex untested functions SELECT ?function ?complexity WHERE { ?function a code:Function ; code:hasComplexity ?complexity ; code:hasTestCoverage ?coverage . FILTER(?complexity > 10 && ?coverage < 0.5) } ORDER BY DESC(?complexity) LIMIT 10 ``` ##### 3. Intelligent Context Assembly Pipeline The context assembly pipeline leverages all three indices to build optimal context for each agent: ```python class ContextAssemblyPipeline: def assemble_context(self, query: str, actor_type: str, resource_scope: List[Resource], max_tokens: int) -> Context: # Stage 1: Query Understanding intent = self.analyze_query_intent(query) entities = self.extract_entities(query) # Classes, functions, concepts # Stage 2: Multi-Modal Search results = SearchResults() # Text search for exact matches if intent.needs_exact_match: text_results = self.text_index.search( query=query, filters={"resources": resource_scope}, limit=100 ) results.add(text_results) # Vector search for semantic similarity if intent.needs_semantic_match: query_embedding = self.embed_query(query, actor_type) vector_results = self.vector_index.search( embedding=query_embedding, filters={"resources": resource_scope}, limit=50 ) results.add(vector_results) # Graph traversal for structural relationships if entities: graph_results = self.graph_index.traverse( start_nodes=entities, patterns=self.get_patterns_for_actor(actor_type), max_depth=3, limit=50 ) results.add(graph_results) # Stage 3: Relevance Ranking ranked_results = self.rank_by_relevance( results=results, actor_type=actor_type, query_intent=intent ) # Stage 4: Context Optimization context = self.optimize_context( ranked_results=ranked_results, max_tokens=max_tokens, strategy=self.get_strategy_for_actor(actor_type) ) return context def get_patterns_for_actor(self, actor_type: str) -> List[GraphPattern]: """Different actors need different traversal patterns""" patterns = { "strategist": [ "module_dependencies", # Understand architecture "interface_boundaries", # Find API surfaces "test_coverage_gaps" # Identify risks ], "executor": [ "implementation_details", # Get full function bodies "local_dependencies", # Find what to change "usage_patterns" # Understand call sites ], "reviewer": [ "change_impact_analysis", # What could break "similar_patterns", # Consistency checks "test_relationships" # Verification paths ] } return patterns.get(actor_type, ["general_traversal"]) ``` ##### 4. Plugin Architecture for Extensibility The system is designed for extensibility at every level: ```yaml PluginSystem: # Language-specific analyzers analyzers: python: class: "PythonAnalyzer" features: - AST parsing via ast module - Type inference via mypy - Import resolution - Docstring extraction typescript: class: "TypeScriptAnalyzer" features: - TSC-based parsing - Type extraction - Module resolution - JSDoc parsing rust: class: "RustAnalyzer" features: - rust-analyzer integration - Lifetime analysis - Trait resolution - Macro expansion # Custom analyzer example custom_dsl: class: "CustomDSLAnalyzer" config: grammar: "path/to/grammar.peg" semantic_rules: "path/to/rules.yaml" # Index backend providers backends: graph: - name: "blazegraph" class: "BlazegraphBackend" scalability: "billions of triples" features: ["SPARQL", "reasoning", "geospatial"] - name: "neo4j" class: "Neo4jBackend" scalability: "enterprise" features: ["Cypher", "APOC", "GDS"] - name: "custom_graph" class: "MyCustomGraphDB" config: connection: "custom://localhost:7687" vector: - name: "faiss" class: "FaissBackend" scalability: "100M vectors" features: ["GPU acceleration", "HNSW"] - name: "qdrant" class: "QdrantBackend" scalability: "distributed" features: ["filtering", "payloads", "snapshots"] # Embedding model providers embedders: - name: "openai" class: "OpenAIEmbedder" models: ["text-embedding-3-large", "text-embedding-3-small"] - name: "local" class: "LocalEmbedder" models: ["all-MiniLM-L6-v2", "instructor-xl"] - name: "custom" class: "MyFineTunedEmbedder" model_path: "path/to/model" ``` ##### 5. Agent Skills for Code Intelligence Agents interact with the code intelligence system through specialized skills: ```yaml actors: code_explorer: type: tool config: tools: # Semantic code search - name: search_code_semantically code: | # Find code similar to a concept results = context.code_intelligence.vector_search( query="validate user input against schema", scope=input_data.get("scope", "all"), limit=input_data.get("limit", 10) ) # Enrich with graph context for result in results: dependencies = context.code_intelligence.get_dependencies( entity=result.entity_id, depth=2 ) result.context = dependencies return results # Structural analysis - name: analyze_dependencies code: | # Use graph store for dependency analysis module = input_data["module_path"] # SPARQL query for full dependency closure query = f''' PREFIX code: SELECT ?dep ?type WHERE {{ <{module}> code:imports* ?dep . ?dep a ?type . }} ''' deps = context.code_intelligence.graph_query(query) # Compute metrics return { "direct_deps": len([d for d in deps if d.distance == 1]), "transitive_deps": len(deps), "circular_deps": context.code_intelligence.find_circular_deps(module), "dependency_graph": deps } # Intelligent refactoring assistant - name: suggest_refactoring_targets code: | # Combine all three indices for comprehensive analysis # 1. Text search for TODO/FIXME/HACK comments todos = context.code_intelligence.text_search( pattern="(TODO|FIXME|HACK):", scope=input_data["scope"] ) # 2. Graph analysis for high complexity complex_functions = context.code_intelligence.graph_query(''' SELECT ?func ?complexity ?coverage WHERE { ?func a code:Function ; code:hasComplexity ?complexity ; code:hasTestCoverage ?coverage . FILTER(?complexity > 15 || ?coverage < 0.3) } ORDER BY DESC(?complexity) ''') # 3. Vector search for code smells smells = [] for pattern in ["duplicate code", "long method", "large class"]: similar = context.code_intelligence.find_similar_code( pattern=pattern, threshold=0.8 ) smells.extend(similar) # Synthesize recommendations return { "high_priority": complex_functions[:5], "technical_debt": todos, "code_smells": smells, "suggested_order": context.code_intelligence.rank_by_impact( complex_functions + todos + smells ) } ``` ##### 6. Real-time Index Synchronization The system maintains index freshness through immediate, proactive updates: ```python class IndexSynchronizer: def __init__(self): self.file_watcher = FileSystemWatcher() self.git_monitor = GitChangeMonitor() self.incremental_indexer = IncrementalIndexer() def on_resource_added(self, resource: Resource, project: Project): """When a resource is added to a project, index it immediately""" # Full initial indexing - happens once when resource is added with self.progress_reporter(f"Indexing {resource.name}") as progress: files = self.scan_resource(resource) total = len(files) # Parallel indexing for performance with ThreadPoolExecutor(max_workers=cpu_count()) as executor: futures = [] for i, file in enumerate(files): future = executor.submit(self.index_file_complete, file) futures.append(future) progress.update(i / total) # Wait for all indexing to complete for future in futures: future.result() # Now set up watchers for incremental updates self.setup_watchers(resource) # Mark resource as indexed and ready resource.indexing_status = "ready" self.notify_agents_index_ready(resource) # CRITICAL: Agents can now immediately search this resource # No "warming up" period - indices are complete and ready def index_file_complete(self, file_path: str): """Comprehensive initial indexing of a file""" # Parse file once ast = self.parse_file(file_path) # Update all indices immediately self.update_text_index(file_path, ast) self.update_vector_embeddings(file_path, ast) self.update_graph_triples(file_path, ast) # Extract and index all metadata self.index_symbols(file_path, ast) self.index_dependencies(file_path, ast) self.index_complexity_metrics(file_path, ast) def setup_watchers(self, project: Project): # File system watching for immediate updates self.file_watcher.watch( path=project.root_path, events=["create", "modify", "delete"], callback=self.on_file_change ) # Git monitoring for batch updates self.git_monitor.watch( repo=project.git_repo, events=["commit", "merge", "rebase"], callback=self.on_git_change ) def on_file_change(self, event: FileEvent): # Quick incremental update if event.type in ["create", "modify"]: # Parse changed file ast = self.parse_file(event.path) # Update indices self.update_text_index(event.path, ast) self.update_vector_embeddings(event.path, ast) self.update_graph_triples(event.path, ast) elif event.type == "delete": self.remove_from_indices(event.path) def on_git_change(self, event: GitEvent): # Batch update for git operations changed_files = event.get_changed_files() # Optimize batch processing with self.batch_updater() as updater: for file in changed_files: updater.queue_update(file) # Process in parallel updater.execute(parallel=True) def update_graph_triples(self, file_path: str, ast: AST): # Generate RDF triples from AST triples = [] # Module-level triples module_uri = self.uri_for_file(file_path) for import_stmt in ast.imports: imported_uri = self.resolve_import(import_stmt) triples.append((module_uri, "code:imports", imported_uri)) # Function-level triples for func in ast.functions: func_uri = self.uri_for_function(func) triples.append((module_uri, "code:contains", func_uri)) triples.append((func_uri, "a", "code:Function")) triples.append((func_uri, "code:hasComplexity", func.complexity)) # Call relationships for call in func.calls: called_uri = self.resolve_call(call) triples.append((func_uri, "code:calls", called_uri)) # Update graph store self.graph_store.update_triples(triples) ``` ##### 7. Fallback to Traditional Search When advanced features are unavailable, the system gracefully degrades: ```python class FallbackSearchProvider: def search(self, query: str, resources: List[Resource]) -> SearchResults: # Try advanced search first try: if self.vector_index.is_available(): return self.vector_search(query, resources) except ServiceUnavailable: pass # Fallback to graph search try: if self.graph_index.is_available(): return self.graph_search(query, resources) except ServiceUnavailable: pass # Ultimate fallback: grep-like text search return self.basic_text_search(query, resources) def basic_text_search(self, query: str, resources: List[Resource]): # Use ripgrep or similar for fast text search results = [] for resource in resources: matches = ripgrep.search( pattern=query, path=resource.path, context_lines=3 ) for match in matches: results.append(SearchResult( file=match.file, line=match.line, content=match.content, score=1.0 # Basic scoring )) return results ``` #### Index Lifecycle The system follows a clear lifecycle for index management: ```yaml Index Lifecycle: 1_resource_added: trigger: "agents project add-resource" action: "Immediate full indexing" duration: "Depends on size (10K files ~1 minute)" result: "All indices ready for instant search" 2_code_changed: trigger: "File modification detected" action: "Immediate incremental update" duration: "Milliseconds per file" result: "Indices stay synchronized" 3_resource_removed: trigger: "agents project remove-resource" action: "Immediate index cleanup" duration: "Seconds" result: "No stale data in indices" 4_maintenance: trigger: "Scheduled or manual" action: "Reindex for consistency" duration: "Background process" result: "Indices optimized and verified" Key Guarantees: - "No search happens on stale data" - "No 'index building' delays during agent execution" - "Changes visible in search immediately" - "Initial indexing is a one-time cost per resource" ``` #### Integration with Context Tiers The Code Intelligence system directly feeds into the three-tier context architecture: ```yaml Context Tier Integration: hot_tier: source: "Real-time results from code intelligence" content: - Currently edited files - Direct dependencies - Immediately relevant functions warm_tier: source: "Indexed embeddings and graph queries" content: - Recent search results - Cached graph traversals - Vector similarity matches - Active decision contexts cold_tier: source: "Historical indices and compressed data" content: - Previous plan analyses - Archived dependency graphs - Historical refactoring patterns - Learned codebase conventions ``` #### Performance Characteristics The system maintains pre-computed indices for instant search performance: ```yaml Performance Metrics: initial_indexing_speed: text_index: "10,000 files/minute" vector_index: "1,000 files/minute (with GPU)" graph_index: "5,000 files/minute" query_performance: text_search: "< 100ms for 1M files" vector_search: "< 200ms for 10M embeddings" graph_traversal: "< 500ms for 3-hop queries" storage_requirements: text_index: "~10% of source size" vector_index: "~1GB per 100K functions" graph_store: "~100MB per 10K files" scalability: max_files: "No hard limit (tested to 10M files)" max_graph_size: "1B+ triples" max_vectors: "100M+ embeddings" ``` #### Progressive Enhancement Path Organizations can adopt Code Intelligence features progressively. At each stage, existing resources are reindexed to take advantage of new capabilities: ```yaml adoption_stages: stage_1_basic: features: ["text search", "file watching"] requirements: ["ripgrep", "sqlite"] initial_setup: "Index all text content on resource add" benefit: "Instant exact-match search" stage_2_semantic: features: ["vector embeddings", "similarity search"] requirements: ["embedding model", "vector DB"] initial_setup: "Generate embeddings for all code (one-time cost)" benefit: "Instant semantic similarity search" stage_3_structural: features: ["RDF graph", "relationship queries"] requirements: ["graph database", "language analyzers"] initial_setup: "Parse and build complete knowledge graph" benefit: "Instant relationship queries" stage_4_intelligent: features: ["ML-driven ranking", "automated analysis"] requirements: ["GPU", "training data"] initial_setup: "Pre-compute ML features and rankings" benefit: "Instant intelligent suggestions" stage_5_custom: features: ["Domain-specific ontologies", "Custom analyzers"] requirements: ["Domain expertise", "Custom development"] initial_setup: "Build domain-specific indices" benefit: "Instant domain-aware intelligence" ``` This Code Intelligence & Context Discovery system ensures that CleverAgents can efficiently work with codebases of any size, providing agents with the contextual understanding they need to make intelligent decisions about code changes, refactoring, and feature development. **Summary of Timing**: - **Indexing**: Eager (happens immediately when resources are added/changed) - **Searching**: Instant (because indices are pre-computed and ready) - **Sandboxing**: Lazy (only when execution needs to modify a resource) - **Context Assembly**: Real-time (but fast because it queries ready indices) This design ensures agents never wait for index building during execution, providing a responsive and predictable experience even on massive codebases. ### Context **Note: This section describes the high-level context management system. For details on how context is discovered and indexed, see the Code Intelligence & Context Discovery section above.** Context in CleverAgents is not "dump all files into an LLM." It is a system that: * finds relevant information from resources, * injects appropriate subsets into each actor/node, * and scales to large repositories. #### Current Reality and Planned Improvements key concepts: * There is a global context concept, * But nodes only see what is injected into their prompts, * It's functional but not yet elegant, * A more advanced automated system is planned. #### Tiered Context Architecture (Hot/Warm/Cold) The system uses a sophisticated three-tier memory architecture that enables working with massive codebases without holding everything in memory: * **Hot context (hot cache)** The small set of immediately relevant chunks injected into the current actor prompt. When working on a 50,000 file codebase, hot context focuses on the immediate task (e.g., 10-20 files for a specific refactoring). * **Warm context** Recent decisions and their contexts from this plan tree - quickly accessible. Includes indexed embeddings, vector search results, and graph store representations. Maintains the decision chain that led to the current work. * **Cold storage** Historical decisions from past plans on this codebase - queryable but not in active memory. Long-term storage in SQLite or caching systems containing prior summaries, older plan artifacts, and historical patterns (e.g., "last time we refactored auth, we also had to update these services"). Promotion/demotion behavior: * System analyzes current query * Promotes relevant data upward (cold → warm → hot) * Demotes stale data out of hot to keep prompts tight * Preserves complete context snapshots for every decision This architecture leverages the key insight that software development is inherently local - even in huge codebases, individual changes typically touch a bounded set of files. The Decision Tree captures these localities. #### Actor-Specific Context Views A key missing feature identified in the notes is per-actor context views, filtering, relevance, and actor-aware context limits. The intended direction: * global context exists at plan level, * each actor gets a "view" of that context tuned to their role, * memory may be shared or per-plan depending on design choices. ##### Actor Context View Service (Proposed) A dedicated module/service that: * maintains actor-specific "context views," * tracks actor memory and relevance, * enforces actor-specific limits (tokens, file types, etc.). #### Initial Context vs Deep Context A practical approach mentioned: * Initial context is high-level (repo tree, language, overview), * Then the system searches for relevant details iteratively via RAG. This strongly suggests CleverAgents should define: * an "initial context recipe" per project type (codebase vs documents vs infra), * iterative context refinement loops during strategize/execute. ## Behavior ### Automation Levels Automation levels determine which phase transitions happen automatically. #### Automation Level Modes | Mode | Behavior | |------|----------| | **Manual** | User explicitly triggers: create → use → execute → apply. Every phase transition requires a command. Every decision point pauses for human input. User sees context, alternatives, and recommendation. User provides explicit choice or custom guidance. | | **Review-before-apply** | Strategize + Execute happen automatically. AI makes all decisions autonomously. System pauses before Apply to show diff and ask for approval. User can approve, reject, or correct specific decisions. | | **Full automation** | System runs all phases automatically end-to-end. AI makes all decisions. Execution proceeds through apply. Human notified of completion. Rollback available if issues detected. | #### Progressive Trust Building New users typically follow this progression: 1. Start with manual mode to understand system behavior 2. Move to review-before-apply as confidence builds 3. Enable full automation for specific task types 4. Gradually expand full automation scope #### Semantic Escalation Even in full automation mode, the system understands when it needs help: ```python class AutonomyController: def assess_decision_confidence(self, decision, context): factors = { 'past_success_rate': self.get_historical_success(decision.type), 'codebase_familiarity': self.get_familiarity_score(context.project), 'risk_assessment': self.evaluate_risk(decision), 'invariant_complexity': self.analyze_invariants(decision) } confidence = self.compute_confidence(factors) if confidence < self.threshold: if self.automation_level == 'full': # Even in full automation, critical decisions escalate return RequestHumanGuidance(decision, factors) return ProceedAutonomously(decision) ``` #### Automation Level Hierarchy Automation levels are determined using this precedence (highest to lowest): 1. **Plan-level**: Explicitly set when using an action on projects 2. **Session-level**: Set for the current session 3. **Global-level**: Persisted application configuration ```bash # Set global automation level (persists across sessions) agents config set automation-level review-before-apply # Set session automation level (overrides global for this session) agents session set automation-level full-automation # Use action with explicit automation level (overrides session and global) agents plan use local/my-action --project my-proj --automation-level manual ``` #### Automation Level Persistence Rules 1. **Global level**: Persisted in application configuration. Default is `manual` on fresh install. 2. **Session level**: Lives for the duration of the session. Not persisted. 3. **Plan level**: Once a plan's automation level is determined (at the moment of `use`), it is **locked to that plan**. Even if the session or global level changes later, the plan retains its original automation level. 4. **Explicit change**: A plan's automation level can be changed **explicitly** after creation: ```bash agents plan set-automation-level full-automation ``` #### Subplan Automation Levels **Subplans inherit the parent plan's automation level.** However, if the parent plan's automation level is changed explicitly mid-execution, new subplans will use the new level while already-completed subplans retain their original level. #### Granular Automation Flags For fine-grained control, additional flags can modify behavior: * `auto_strategize` - Automatically proceed from Action to Strategize * `auto_execute` - Automatically proceed from Strategize to Execute * `auto_apply` - Automatically proceed from Execute to Apply * `auto_retry_on_failure` - Automatically retry failed phases This allows combinations like: * Auto strategize, manual execute, manual apply (for risky infra tasks) * Auto strategize + execute, manual apply (the review-before-apply pattern) ### Validation and Guardrails #### Plan generation validation The validation logic is stubbed and must be implemented. The spec should require: * validate action schema * validate actor availability * validate required skills exist * validate permission policy * validate rollback feasibility (if enabled) * validate project resource accessibility This prevents "plan runs with fake providers" and other surprises. #### Cost / rate limits Future concerns: * API call limits * cost caps So CleverAgents should define: * per-plan budgets * per-session budgets * per-org budgets * per-actor max tool calls / max retries ### Correcting Plans (Core Feature) Correcting plans is where CleverAgents becomes more than "a fancy prompt runner." #### The Goal When a plan makes a wrong decision early, we want to: * correct the decision, * recompute only the affected subtree, * preserve unaffected work. This is explicitly described: "redo everything below that decision, not the entire code base." #### Decision Tree Representation Every plan records (see Decision Data Model section): * decisions (choice points) - created during Strategize * dependencies (which later work depended on that decision) * child plans spawned because of that decision - populated during Execute * artifacts generated under that branch This makes plan runs auditable and correctable. #### Two Correction Modes 1. **Revert-from-history correction (`--mode=revert`)** * Find the decision point in the tree * Roll back all changes (code and non-code) to that point * Re-run from that decision point forward * Keep old execution artifacts for comparison * Potentially expensive if high up in the tree 2. **Add-at-end correction (`--mode=append`)** * Leave history intact * Append a new plan at the end that fixes the outcome * Cheaper and safer sometimes * Does not rewrite history #### Correction Flow (Revert Mode) When user requests correction at Decision B: 1. **Mark for Correction** ``` Decision B.superseded_by = new_decision_id ``` 2. **Identify Downstream Impact** - Recursively collect all decisions that depend on Decision B - Collect all subplans spawned from those decisions - These form the "affected subtree" 3. **Rollback Resources** - For each affected decision's `artifacts_produced`: - Rollback to the checkpoint before that artifact was created - For affected subplans: - Rollback their sandboxes entirely 4. **Preserve for Comparison** - Archive the original subtree's artifacts - Create a `CorrectionAttempt` record linking old and new 5. **Re-execute from Decision Point** - Restore context to `Decision B.context_snapshot` - User provides new guidance/correction - Re-run execution from that point - New decisions get `is_correction: true, corrects_decision_id: B` 6. **Apply as Normal** - The corrected plan goes through normal Apply gating - Diff shows changes from the correction #### History Cleanup **History can only be flagged for cleanup after a plan is Applied.** Once a plan is applied: - It can no longer be rolled back - Old correction artifacts can be archived or deleted based on retention policy - The decision tree is preserved for audit purposes #### CLI Commands for Correction ```bash # View decision tree agents plan tree agents plan tree --format=json # For visualization tools # Inspect a specific decision agents plan explain # Shows: question, chosen option, alternatives, rationale, downstream impact # Correct via revert-and-replay agents plan correct --mode=revert --guidance "" # Re-executes from that point with the new guidance # Correct via append (add fix at end) agents plan correct --mode=append --guidance "" # Creates a new subplan to fix the outcome without rewriting history # For long guidance text, use a file agents plan correct --mode=revert --guidance-file ./correction.txt # Compare old vs new after correction agents plan diff ``` #### Correction Safety Corrections always: * Create a new attempt revision (increment `plan.attempt`) * Preserve old artifacts for diff/compare * Run execute in sandbox again * Require apply gating again * Never modify already-applied changes This keeps history reproducible and prevents accidental destructive edits. ### Human-in-the-Loop Collaboration Even though the direction is "more autonomous," the transcript explicitly recognizes that real workflows require engineers to collaborate with the system, editing code while it works, and using better UX integration (TUI/web/IDE). So CleverAgents should aim for: * visibility: what is it doing now? * interruptibility: pause/cancel/retry * editability: allow user to modify strategy before execute * reconciliation: detect if user changed sandbox files mid-run and handle it ### UI / Interaction Model #### CLI-first + TUI + Web App + IDE The system intends to be CLI-first, with: * a TUI built using Textual, * which can generate a web app "for free," * and later an IDE plugin that embeds the TUI in the IDE. This implies a "single UI codebase" model: * same underlying view logic, * multiple frontends. #### Plan Tree Visualization The TUI should show: * plan list * plan details * plan tree (ASCII) * diff view * approvals And should later allow exporting the tree as image (PNG) or JSON for other visualization tools. ## Additional Recommended Sections (Missing from Draft but Important) ### Storage and Persistence CleverAgents should define where each concept lives: * Actions: stored in a registry (local files or server DB) * Actors: stored similarly (config files + DB indexing) * Projects: stored locally or on server * Plans: stored in plan DB with full logs * Context indexes: vector store / graph store / SQLite * Artifacts: filesystem or object storage ### Observability To debug large plans: * every phase should emit events * every actor call should log prompt/context references * every skill call should log resource access * every checkpoint should be recorded ### Security Model * sandbox isolation * resource-level ACLs * prompt injection mitigations (server mode) * secret management (API keys, DB credentials) * audit logs for apply ### Extensibility * plugin system for skills * custom node types * action templates * actor templates (noted as missing currently) ## Summary of Key Intended Behaviors (If You Only Read One Section) * Plans follow **Action → Strategize → Execute → Apply** with automation levels controlling transitions. * **Strategize is read-only** and produces a strategy + blueprint. * **Execute happens in a sandbox**, can spawn subplans, and should support checkpoints/rollback when enabled. * **Apply commits changes** from sandbox to real project after review/validation. * Actors are **hierarchical**: an actor can be a single agent or an entire graph. * Graph nodes can be **actors, MCP skills, or custom tool nodes**. * Context should evolve toward **hot/warm/cold tiers** and **actor-specific context views**. * The system is designed for large tasks where the user can **correct a decision** and only recompute downstream work, visualizable as a plan decision tree. The system can handle Firefox-scale projects not through magic, but through: - **Hierarchical decomposition** breaking massive tasks into bounded work - **Persistent decision graphs** maintaining context across any scale - **Isolated execution** preventing cascading failures - **Semantic validation** catching errors before propagation - **Progressive automation** building trust through incremental success Future details to add to this document: * a canonical **JSON/YAML schema** for Actions, Actors, Projects, Plans, Skills, and Context Views, * a **CLI command reference** (every command, flags, examples), * and a set of **end-to-end example workflows** (single project, multi-project, infra task, paper-writing task) consistent with this spec. ## Work Remaining to Make CleverAgents Fully Functional This section describes what remains to be done to bring the current CleverAgents codebase up to the specification. **Last Updated:** February 6, 2026 ### Current State Assessment #### What's Implemented Based on analysis of the current codebase: 1. **Plan Lifecycle Foundation** - The 4-phase lifecycle (Action → Strategize → Execute → Apply) is partially implemented in `plan_lifecycle_service.py` 2. **Database Models** - Models exist for projects, plans, contexts, changes, and actors 3. **LangGraph Integration** - Graph-based workflow support exists but is not fully connected to the plan workflow 4. **Reactive System** - A reactive system with stream routing is present 5. **Actor System** - Actor models and services exist but lack full behavioral definitions 6. **Basic Context Analysis** - Simple context loading and analysis capabilities 7. **Change Tracking** - Basic `Change` and `ChangeSet` models exist #### What's Missing Critical gaps between the specification and current implementation: 1. **No Resource Abstraction** - The unified resource layer for files, databases, APIs is entirely missing 2. **No Sandboxing** - Execute phase writes directly to files without isolation 3. **No Decision Tree** - No decision tracking, storage, or correction mechanism 4. **No Skills/MCP Integration** - No skill abstraction or MCP protocol support 5. **Single-File Limitation** - Hard-coded to generate exactly one file per plan 6. **Text-Based Code Generation** - Still parsing LLM output instead of tool-based approach 7. **No Checkpointing** - No rollback or checkpoint capabilities 8. **No Code Intelligence** - Missing indexing, vector search, and RDF graph store 9. **No Context Tiers** - No hot/warm/cold memory architecture 10. **Limited Validation** - Basic stub validation instead of semantic checks ### Work Items by Priority ### 1) Implement Tool-Based Resource Modification (Critical Foundation) #### Problem The current system generates code as a single text blob that gets written to one file. The specification requires a modern tool-based approach where LLMs invoke discrete operations on resources. #### Implementation Steps 1. **Create Resource Abstraction Layer** ```python # New modules needed: src/cleveragents/domain/models/resources.py src/cleveragents/domain/resources/handlers.py src/cleveragents/domain/resources/sandbox.py ``` 2. **Implement Built-in Skills** ```python # Core file operations read_file(path: str) -> str write_file(path: str, content: str) -> None edit_file(path: str, changes: list[Edit]) -> None delete_file(path: str) -> None move_file(src: str, dst: str) -> None create_directory(path: str) -> None list_files(pattern: str) -> list[str] search_files(pattern: str, content_pattern: str) -> list[Match] ``` 3. **Connect Skills to ChangeSet** - Each skill invocation that modifies resources creates a `Change` record - ChangeSet accumulates these changes during execution - No more parsing LLM text output for code #### Estimated Effort: 2-3 weeks ### 2) Implement Sandbox Infrastructure (Critical for Safety) #### Problem Execute phase currently writes directly to the project. The specification requires all changes happen in an isolated sandbox that can be reviewed before applying. #### Implementation Steps 1. **Define Sandbox Interface** ```python class Sandbox(Protocol): def create() -> SandboxRef def read(path: str) -> Content def write(path: str, content: Content) -> Change def diff() -> DiffView def commit() -> None def rollback() -> None ``` 2. **Implement Sandbox Strategies** - `GitWorktreeSandbox` - For git repositories (preferred) - `FilesystemCopySandbox` - For non-git projects - `TransactionSandbox` - For databases - `NoOpSandbox` - For non-sandboxable resources 3. **Lazy Sandbox Creation** - Only create sandboxes when resources are accessed - Each plan gets its own sandbox namespace #### Estimated Effort: 2 weeks ### 3) Build Decision Tree System (Core Innovation) #### Problem No decision tracking exists. The specification's key innovation is recording every decision with full context, enabling correction without full re-execution. #### Implementation Steps 1. **Create Decision Models** ```sql -- New tables needed CREATE TABLE decisions ( decision_id TEXT PRIMARY KEY, -- ULID plan_id TEXT NOT NULL, parent_decision_id TEXT, decision_type TEXT NOT NULL, question TEXT NOT NULL, chosen_option TEXT NOT NULL, alternatives_considered TEXT, -- JSON array context_snapshot TEXT NOT NULL, -- JSON created_at TEXT NOT NULL ); ``` 2. **Implement Decision Recording** - Every choice during Strategize creates a Decision record - Capture complete context snapshot with each decision - Track downstream dependencies 3. **Build Correction Mechanism** ```bash agents plan correct --mode=revert --guidance "" ``` - Mark decision as superseded - Recompute only affected subtree - Preserve unaffected work #### Estimated Effort: 3 weeks ### 4) Implement Code Intelligence System (Scalability Enabler) #### Problem Current context is limited to a few hundred characters from a few files. Large codebases require intelligent context discovery. #### Implementation Steps 1. **Eager Indexing on Resource Add** ```python def on_resource_added(resource: Resource): # Index immediately when resource added to project index_text_content(resource) # Full-text search generate_embeddings(resource) # Vector embeddings build_knowledge_graph(resource) # RDF triples ``` 2. **Three-Index Architecture** - **Text Index**: Tantivy/SQLite FTS for exact matches - **Vector Index**: FAISS/Qdrant for semantic search - **Graph Store**: RDF store for relationships 3. **Tiered Context System** - **Hot**: Current working set (in LLM context) - **Warm**: Recent decisions and search results - **Cold**: Historical data and patterns #### Estimated Effort: 4 weeks ### 5) Complete Actor System with Behavioral Definitions #### Problem Actors exist but don't define behavior beyond model selection. The specification requires actors to be composable graphs with tools, memory, and context policies. #### Implementation Steps 1. **Extend Actor Configuration** ```yaml actors: my_strategist: type: graph config: provider: anthropic model: claude-3-opus tools: [read_file, search_code, analyze_dependencies] memory_policy: per_plan context_view: architect # High-level view routes: strategize: entry_point: analyze nodes: - name: analyze type: llm - name: plan type: llm ``` 2. **Implement Tool Access Policies** - Strategy actors: read-only tools - Execute actors: read/write within sandbox - Apply actors: commit tools 3. **Actor-Specific Context Views** - Strategist: Architecture, dependencies, patterns - Executor: Implementation details, specific files - Reviewer: Diffs, tests, risk analysis #### Estimated Effort: 2 weeks ### 6) MCP Integration and Skill System #### Problem No skill abstraction exists. The specification extends MCP for both read and write operations. #### Implementation Steps 1. **MCP Tool Adapter** ```python class MCPSkillAdapter: def wrap_mcp_tool(self, tool) -> Skill: # Add sandbox interception # Add change tracking # Add capability metadata ``` 2. **Skill Capability Metadata** ```python class SkillMetadata: read_only: bool write_scope: list[str] checkpointable: bool idempotent: bool side_effects: list[str] ``` 3. **External MCP Server Support** ```yaml mcp_servers: - name: github command: "npx @anthropic/mcp-github" env: {GITHUB_TOKEN: "${GITHUB_TOKEN}"} ``` #### Estimated Effort: 2 weeks ### 7) Implement Validation and Semantic Error Prevention #### Problem Current validation is a stub. The specification requires multi-layer semantic validation. #### Implementation Steps 1. **Decision-Time Validation** - Validate choices during Strategize - Check alternatives for feasibility - Record validation in decision metadata 2. **Execution-Time Guards** ```yaml nodes: - name: semantic_validator type: tool tools: - validate_api_compatibility - check_invariants - verify_test_coverage ``` 3. **Project-Specific Validation** ```bash agents project set-validation \ --project my-api \ --test-command "pytest" \ --lint-command "ruff check ." ``` #### Estimated Effort: 2 weeks ### 8) Connect LangGraph to Plan Lifecycle #### Problem The reactive/LangGraph infrastructure exists but isn't connected to the main plan workflow. #### Implementation Steps 1. **Create Unified Plan Graph** ```python class PlanLifecycleGraph: def strategize_subgraph(self) -> StateGraph def execute_subgraph(self) -> StateGraph def apply_subgraph(self) -> StateGraph ``` 2. **Wire Phase Transitions** - `use` command triggers strategize graph - `execute` command triggers execute graph - `apply` command triggers apply graph 3. **Remove Linear Pipeline** - Replace tell/build/apply with graph execution - Maintain backward compatibility at CLI level #### Estimated Effort: 1 week ### Implementation Roadmap #### Phase 1: Foundation (6-8 weeks) 1. Resource abstraction and built-in skills 2. Sandbox infrastructure 3. Connect LangGraph to lifecycle #### Phase 2: Core Features (8-10 weeks) 4. Decision tree system 5. Code intelligence (basic text/vector search) 6. MCP integration #### Phase 3: Advanced Features (4-6 weeks) 7. Complete actor behavioral system 8. Semantic validation layers 9. Advanced code intelligence (RDF graph) #### Phase 4: Production Hardening (4 weeks) 10. Checkpointing and rollback 11. Cost controls and rate limiting 12. Security fixes (remove eval, fix async) ### Total Estimated Timeline: 5-6 months ### Key Success Metrics 1. **Multi-file generation**: Can generate a REST API with routes/, models/, tests/ 2. **Safe execution**: All changes happen in sandbox, reviewed before apply 3. **Decision correction**: Can correct a decision and recompute only affected work 4. **Scale to large codebases**: Can work with 10K+ file projects efficiently 5. **Semantic safety**: Catches breaking changes before they're applied ### Migration Strategy The implementation can proceed incrementally: 1. Start with resource abstraction (enables everything else) 2. Add sandboxing to existing execute phase 3. Gradually replace text parsing with tool invocations 4. Build decision tree alongside existing flow 5. Enhance context as indexing comes online This allows the system to remain functional during development while progressively adding the architectural improvements described in the specification. ## CleverAgents Architecture FAQ ### Q: How does CleverAgents handle persistent repository knowledge beyond ephemeral context windows? **What exists today architecturally**: The specification defines a sophisticated multi-tier memory system that goes far beyond ephemeral context windows. At its core is the Decision Tree structure which provides a durable, queryable record of every choice made during planning, along with the complete context that informed those choices. **How the persistent model works in practice**: When a strategy actor analyzes a codebase during the Strategize phase, it doesn't just make decisions in isolation. Each decision creates a comprehensive Decision record that includes: ```yaml context_snapshot: hot_context_hash: str # Cryptographic hash of the exact context hot_context_ref: str # Pointer to the full stored snapshot relevant_resources: list[ResourceRef] # Every file/symbol that influenced this decision actor_state_ref: str # Complete LangGraph checkpoint ``` This means when the system decides "refactor the authentication module to use async patterns," it permanently records: - Which files were examined to make that decision - What symbols and dependencies were traced - The exact code state that was analyzed - The reasoning chain that led to this choice - Alternative approaches that were considered but rejected **The three-tier memory architecture enables scale**: 1. **Hot tier**: Immediate working context (what's in the current LLM context window) 2. **Warm tier**: Recent decisions and their contexts from this plan tree - quickly accessible 3. **Cold tier**: Historical decisions from past plans on this codebase - queryable but not in active memory When working on a 50,000 file codebase, the system doesn't need to hold all files in memory. Instead: - Hot context focuses on the immediate task (e.g., 10-20 files for a specific refactoring) - Warm context maintains the decision chain that got us here - Cold context provides historical patterns ("last time we refactored auth, we also had to update these services") **Why this scales to massive codebases**: The key insight is that software development is inherently local - even in huge codebases, individual changes typically touch a bounded set of files. The Decision Tree captures these localities. When converting Firefox to Rust (your example), the system would: 1. Make high-level architectural decisions (captured as root decision nodes) 2. Decompose into major subsystem conversions (each a decision spawning subplans) 3. Each subsystem plan makes decisions about its modules 4. Module plans make decisions about individual files At each level, only the relevant context is loaded. The persistent decision graph means we can always reconstruct why we're converting a particular module and what constraints apply from higher-level decisions. **Concrete example of persistence in action**: ``` Plan: Convert Firefox Renderer to Rust ├── [Decision] Architecture approach: Start with leaf modules, work inward │ Context: Analyzed module dependency graph, 2,847 modules total │ Resources: module_graph.json, architecture_docs.md │ ├── [Decision] Phase 1: Convert utility libraries (no external deps) │ └── [Subplan] Convert string_utils module │ ├── [Decision] Use Rust's String type, not custom implementation │ │ Context: Analyzed 47 string_utils.cpp functions │ │ Resources: string_utils.cpp, string_utils.h, 12 dependent files │ │ Rationale: Rust's String provides same guarantees with better ergonomics ``` Even months later, we can query: "Why did we use Rust's String type?" and get the exact context and reasoning, without reprocessing the entire codebase. ### Q: How does the system compute task-specific dependency closures for large-scale operations? **What exists today architecturally**: The specification defines multiple mechanisms for computing and maintaining minimal dependency closures. The execution blueprint produced during the Strategize phase doesn't just list steps - it includes a complete dependency graph with explicit scoping for each operation. **How dependency closure computation works**: During the Strategize phase, the strategy actor employs several mechanisms to compute precise dependency closures: 1. **Resource-aware analysis**: The actor uses specialized skills to trace dependencies: ```python # Pseudocode of what happens inside a strategy actor def compute_closure_for_refactoring(target_module): closure = ResourceClosure() # Direct file dependencies closure.add_files(find_imports(target_module)) closure.add_files(find_includes(target_module)) # Symbol dependencies for symbol in extract_exported_symbols(target_module): closure.add_files(find_symbol_usage(symbol, scope='project')) # Test dependencies closure.add_files(find_tests_for_module(target_module)) # Build system dependencies closure.add_files(find_build_references(target_module)) return closure ``` 2. **Hierarchical scoping**: When spawning subplans, each subplan receives: - An explicit `relevant_resources` list - A `sandbox_strategy` appropriate for those resources - Clear boundaries of what it can and cannot modify 3. **Decision-based tracking**: Each `subplan_spawn` decision records: ```yaml decision_type: subplan_spawn chosen_option: "Refactor authentication module" downstream_plan_ids: ["plan-auth-refactor-123"] artifacts_produced: - auth_module_files: ["auth.rs", "auth_test.rs", "auth_types.rs"] - api_updates: ["api/v2/login.rs", "api/v2/logout.rs"] ``` **Concrete example - Converting a subsystem to Rust**: Let's trace how the system handles "Convert Firefox's Network Stack to Rust": ``` STRATEGIZE PHASE: 1. Analyze network stack structure - Identifies 847 C++ files in netwerk/ directory - Traces public API surface (237 exported functions) - Maps internal dependencies (1,432 internal calls) 2. Compute minimal closure for Phase 1 (DNS resolver): - Core files: dns_resolver.cpp, dns_cache.cpp, dns_config.cpp (3 files) - Direct dependencies: 12 files in netwerk/base/ - Test files: 8 test files specific to DNS - Build files: 2 moz.build files - Total closure: 25 files (not 847!) 3. Generate execution blueprint with subplans: - convert-dns-types: Closure of 5 files (type definitions) - convert-dns-cache: Closure of 8 files (cache + tests) - convert-dns-resolver: Closure of 12 files (resolver + integration) ``` **Why this is tractable even for massive codebases**: The system leverages several key insights about real software: 1. **Modular boundaries exist**: Even in legacy codebases, there are natural boundaries 2. **Changes are incremental**: We don't convert 50,000 files atomically 3. **Dependencies are sparse**: Most modules depend on a small fraction of the codebase 4. **Interfaces are narrow**: Public APIs are much smaller than implementations The Firefox example would decompose into ~1,000 bounded subplans, each touching 10-100 files. The parent plan tracks the overall architecture, while each subplan maintains its focused closure. **How we prevent closure explosion**: - **Lazy expansion**: Dependencies are traced only as deep as needed for correctness - **Interface-based boundaries**: When possible, work against stable interfaces - **Incremental validation**: Each subplan validates its changes don't break dependents - **Hierarchical merge strategies**: Parent plans resolve conflicts between subplan changes ### Q: What mechanisms enforce global consistency during parallel execution across many files? **What exists today architecturally**: The sandbox model combined with hierarchical plan execution provides strong guarantees about consistency during parallel execution. This isn't just process isolation - it's semantic isolation with intelligent merge strategies. **How the coordination mechanism prevents compound errors**: 1. **Complete isolation during execution**: Each plan executes in its own sandbox, which means: ``` Plan A (refactoring auth module): - Sandbox A1: Contains only auth/*.cpp, auth_tests/*.cpp - Cannot see Plan B's intermediate states - Cannot accidentally depend on Plan B's half-done work Plan B (updating API endpoints): - Sandbox B1: Contains only api/*.cpp, api_tests/*.cpp - Makes changes assuming current auth interface - Protected from Plan A's intermediate refactoring ``` 2. **Resource-specific sandbox strategies provide natural coordination**: ```yaml Git repositories: - Strategy: git worktrees - Coordination: Git's three-way merge algorithm - Conflict detection: Built into Git - Rollback: git reset/checkout Databases: - Strategy: Transaction isolation - Coordination: MVCC (multi-version concurrency control) - Conflict detection: Serialization failures - Rollback: Transaction abort Cloud Infrastructure: - Strategy: Terraform workspaces - Coordination: State locking - Conflict detection: Resource conflicts in plan - Rollback: Previous state restoration ``` 3. **Hierarchical merge resolution**: When subplans complete, the parent plan performs intelligent merging: ```python def merge_subplan_results(subplan_results): # Group by resource type by_resource = group_by_resource_type(subplan_results) # Apply resource-specific merge strategies for resource_type, changes in by_resource: if resource_type == 'git_repo': merge_git_changes(changes) # Three-way merge elif resource_type == 'database': merge_db_changes(changes) # Sequential application elif resource_type == 'config_files': merge_config_changes(changes) # Smart JSON/YAML merge # Validate merged state run_integration_tests() ``` **Concrete example - Preventing cascading failures**: Consider refactoring a shared authentication library used by 15 services: ``` PARALLEL EXECUTION WITHOUT COORDINATION (what we prevent): - Service A refactors to async auth → breaks Service B - Service B compensates with workaround → breaks Service C - Service C changes error handling → breaks Services D, E, F - Cascade of failures! CLEVERAGENTS COORDINATED EXECUTION: Parent Plan: Refactor auth library ├── Subplan 1: Update auth library interface │ Sandbox: Only auth library files │ Output: New interface definition │ ├── Barrier: Wait for Subplan 1 completion │ ├── Parallel Subplans 2-16: Update each service │ Each sandbox: Only that service's files │ Each uses: New interface from Subplan 1 │ No inter-service dependencies during execution │ └── Merge Phase: - Collect all service updates - Apply to main branch in order - Run integration tests - If conflicts: Parent plan resolves using semantic understanding ``` **Advanced coordination patterns**: 1. **Optimistic concurrency with semantic conflict resolution**: ```yaml Two subplans both modify api/user.rs: - Plan A: Adds async fn get_user_profile() - Plan B: Adds fn validate_user_permissions() Merge strategy: - Git merge succeeds (different functions) - Semantic validation ensures both functions work together - Parent plan adds integration glue if needed ``` 2. **Checkpoint-based coordination**: ``` Execution timeline: T1: Subplan A creates checkpoint before major refactor T2: Subplan B creates checkpoint before API changes T3: Subplan A encounters error, rolls back to T1 T4: Subplan B completes successfully T5: Subplan A retries with knowledge of B's success ``` 3. **Resource locking for critical sections**: ```yaml When modifying shared schema files: - Acquire exclusive lock on schema resources - Make changes atomically - Release lock with new version - Other plans rebase on new schema ``` ### Q: How does the system proactively prevent semantic errors before they propagate? **What exists today architecturally**: The specification defines multiple layers of proactive error prevention that go far beyond traditional testing. This is a comprehensive defense-in-depth approach that catches semantic errors before they can propagate. **Layer 1: Decision-time validation during Strategize**: Every decision includes semantic validation: ```yaml Decision: Refactor payment module to async alternatives_considered: - "Convert to async/await patterns" (chosen) - "Use thread pool with channels" (rejected: doesn't integrate with async ecosystem) - "Keep synchronous with timeout" (rejected: doesn't solve core latency issue) confidence_score: 0.85 validation_performed: - Checked all payment API consumers can handle async - Verified database driver supports async operations - Confirmed no regulatory requirement for sync processing ``` **Layer 2: Execution-time semantic guards**: The execution actor configuration includes validation nodes that understand semantics: ```yaml actors: code_executor: type: graph nodes: - name: semantic_validator type: tool config: tools: - name: validate_api_compatibility code: | # Not just syntax checking - semantic validation old_api = extract_api_signature(previous_version) new_api = extract_api_signature(current_version) breaking_changes = find_breaking_changes(old_api, new_api) if breaking_changes: # Don't just fail - understand the impact affected_consumers = find_api_consumers(breaking_changes) migration_plan = generate_migration(breaking_changes) if can_auto_migrate(affected_consumers, migration_plan): apply_migration(migration_plan) else: raise SemanticError( "Breaking API changes require manual review", changes=breaking_changes, affected=affected_consumers ) ``` **Layer 3: Invariant enforcement through the type system**: ```python # The system maintains semantic invariants class RefactoringInvariants: # User-defined invariants for the codebase invariants = [ "All public APIs must maintain backward compatibility", "Database transactions must complete within 5 seconds", "Authentication must always use OAuth2", "Payment processing must be idempotent" ] def check_invariant_preservation(self, changes): for invariant in self.invariants: if not self.verify_invariant(invariant, changes): return InvariantViolation(invariant, changes) return Success() ``` **Layer 4: Predictive error prevention through pattern matching**: The system learns from past failures: ```yaml Error Pattern Database: - pattern: "Async conversion in payment module" historical_failures: - "Race condition in payment confirmation" - "Timeout handling breaks idempotency" preventive_checks: - "Add explicit transaction boundaries" - "Verify idempotency keys are preserved" - "Check distributed lock acquisition" ``` **Concrete example - Preventing a subtle distributed systems bug**: Scenario: Refactoring a service to use event sourcing: ``` PROACTIVE CONTAINMENT IN ACTION: 1. Strategy Phase Semantic Analysis: - Decision: "Convert order service to event sourcing" - Semantic check: "Event sourcing requires eventual consistency" - Identifies: 3 services assume immediate consistency - Adds decision: "Update dependent services for eventual consistency" 2. Execution Phase Invariant Checking: - Detects: PaymentService.chargeCard() called after OrderCreated event - Semantic issue: Payment before order confirmation violates business rules - Automatic fix: Insert OrderConfirmed event requirement 3. Validation Node Catches Edge Case: - Discovers: Audit service expects synchronous order numbers - Impact: Async events break compliance reporting - Resolution: Add audit event buffer with guaranteed ordering 4. Pre-Apply Semantic Verification: - Simulates production event flow - Detects: Under high load, events can arrive out of order - Adds: Event ordering guarantees via vector clocks ``` **Why this prevents issues that traditional testing misses**: Traditional tests check "does the code work?" Our semantic containment asks: - Does it preserve business invariants? - Does it maintain architectural patterns? - Does it respect distributed systems principles? - Does it handle the edge cases we've seen before? **Integration with Definition of Done (DoD)**: Each plan's DoD includes semantic requirements: ```yaml definition_of_done: must: - "All API changes maintain backward compatibility" - "No increase in p99 latency" - "Audit trail remains complete" should: - "Improve code coverage by 10%" - "Reduce cyclomatic complexity" may: - "Optimize for memory usage" ``` The validation nodes enforce these semantics, not just test passage. ### Q: How does the system balance human supervision with autonomous operation? **What exists today architecturally**: The specification defines a sophisticated gradation of automation levels that precisely controls when human intervention is needed. This isn't a binary human/AI split - it's a spectrum that can be adjusted per task, per project, or per organization. **How the automation levels work in practice**: ```yaml Manual Mode: - Every decision point pauses for human input - User sees: Context, alternatives, recommendation - User provides: Explicit choice or custom guidance - Use case: Critical production changes, learning new codebases Review-before-apply Mode: - AI makes all decisions autonomously - Execution completes in sandbox - Human reviews complete diff before apply - User can: Approve, reject, or correct specific decisions - Use case: Normal feature development, refactoring Full Automation Mode: - AI makes all decisions - Execution proceeds through apply - Human notified of completion - Rollback available if issues detected - Use case: Routine updates, test generation, documentation ``` **The decision correction mechanism enables progressive automation**: The `agents plan correct` command is crucial for building trust: ```bash # User observes AI made suboptimal choice agents plan tree # Sees: [Decision] "Use REST API for service communication" # User knows gRPC would be better for this use case agents plan correct --mode=revert \ --guidance "Use gRPC instead of REST. This service requires streaming updates and binary protocol efficiency. Set up protocol buffer definitions and generate client/server stubs." # System: # 1. Marks original decision as superseded # 2. Creates new decision with user guidance # 3. Recomputes ONLY affected downstream decisions # 4. Preserves all unrelated work ``` **Progressive trust building through automation levels**: New users typically follow this progression: 1. Start with manual mode to understand system behavior 2. Move to review-before-apply as confidence builds 3. Enable full automation for specific task types 4. Gradually expand full automation scope **Real autonomy through semantic understanding**: True autonomy isn't about removing humans - it's about the system understanding when it needs help: ```python class AutonomyController: def assess_decision_confidence(self, decision, context): factors = { 'past_success_rate': self.get_historical_success(decision.type), 'codebase_familiarity': self.get_familiarity_score(context.project), 'risk_assessment': self.evaluate_risk(decision), 'invariant_complexity': self.analyze_invariants(decision) } confidence = self.compute_confidence(factors) if confidence < self.threshold: if self.automation_level == 'full': # Even in full automation, critical decisions escalate return RequestHumanGuidance(decision, factors) return ProceedAutonomously(decision) ``` **Concrete example - Autonomous handling of a complex refactoring**: ``` Scenario: "Modernize legacy e-commerce system" INITIAL PLAN (Full Automation Mode): 1. System analyzes 50,000 line codebase 2. Identifies modernization opportunities 3. Creates plan with 47 subplans AUTONOMOUS EXECUTION WITH SMART ESCALATION: Subplan 1-15: Update utility functions (executes autonomously) - Confidence: 0.95 (straightforward transformations) - Result: Success Subplan 16: Refactor payment processing - Confidence: 0.4 (critical business logic) - Action: ESCALATES to human - Human provides: "Preserve exact penny rounding behavior" - Continues autonomously with constraint Subplan 17-30: UI component updates (executes autonomously) - Confidence: 0.9 (isolated changes) - Result: Success Subplan 31: Database schema migration - Detects: Would require 6-hour downtime - Action: ESCALATES to human - Human provides: "Use online migration with feature flags" - Re-plans with zero-downtime approach Subplan 32-47: Complete autonomously ``` **The path to greater autonomy**: The system becomes more autonomous through: 1. **Learning from corrections**: - Every correction teaches the system about user preferences - Patterns emerge: "This team always prefers gRPC for microservices" - Future decisions incorporate these learnings 2. **Building project-specific context**: - Each successful plan adds to project knowledge - System learns codebase patterns, team conventions, business rules - Confidence increases with familiarity 3. **Hierarchical delegation**: - Proven subplan patterns become fully autonomous - Human focuses on high-level decisions - System handles implementation details 4. **Semantic safety nets**: - Comprehensive invariant checking reduces risk - Rollback capabilities provide recovery path - Humans can trust system won't cause catastrophic failures ### Q: What's actually implemented today versus planned for the future? **Concrete implementations in the architecture**: 1. **Decision Tree with Complete Context Capture** - Full schema defined - Storage model specified - Correction mechanism detailed - Query patterns established 2. **Hierarchical Plan/Subplan System** - Spawning mechanism defined - Execution semantics specified - Merge strategies documented - Failure handling described 3. **Resource-Aware Sandbox Isolation** - Multiple strategies defined (git worktree, filesystem overlay, transactions) - Lazy sandboxing for efficiency - Resource-specific merge algorithms - Cleanup behavior specified 4. **Multi-Layer Error Prevention** - Decision validation during planning - Semantic validation nodes - Invariant enforcement - Definition of Done checking 5. **Graduated Automation Controls** - Three levels clearly defined - Decision correction without full re-execution - Confidence-based escalation - Progressive trust building **Near-term implementations** (architecture complete, engineering straightforward): 1. **Memory Tier Management** - Hot/warm/cold distinction clear - Context loading patterns defined - Just needs LRU cache and storage backend 2. **Cross-Plan Learning** - Decision history provides training data - Pattern extraction is standard ML - Confidence scoring is well-understood 3. **Cost/Risk Estimation** - Dedicated estimation actor role defined - Historical data provides baselines - Standard prediction problem 4. **Extended Validation Patterns** - Pluggable validation architecture - Project-specific rules as configuration - Industry patterns can be packaged **Research territory** (requires innovation but architecture supports): 1. **Optimal Context Selection for 100K+ file codebases** - Current: Heuristic-based selection - Research: ML-driven relevance ranking - Architecture supports: Any selection algorithm can plug in 2. **Automated Invariant Discovery** - Current: User-defined invariants - Research: Mining invariants from code patterns - Architecture supports: Invariants are just validation rules 3. **Cross-Project Knowledge Transfer** - Current: Project-specific learning - Research: Generalized pattern recognition - Architecture supports: Cold tier can span projects 4. **Fully Autonomous Recovery Strategies** - Current: Rollback and retry with guidance - Research: Automatic error understanding and fixing - Architecture supports: Recovery is just another plan type **Why we can confidently handle Firefox-scale projects**: The architecture doesn't require magical AI breakthroughs. It requires: - **Hierarchical decomposition**: ✓ Fully specified - **Bounded context operations**: ✓ Dependency closure computation defined - **Parallel execution with isolation**: ✓ Sandbox model complete - **Semantic validation**: ✓ Multi-layer approach specified - **Progressive automation**: ✓ Automation levels and correction defined The difference between handling a 1,000 file project and a 100,000 file project is: - More subplans (hierarchical decomposition handles this) - Larger cold storage (standard database scaling) - Better context selection (improves with use but works with heuristics) - More validation patterns (accumulate over time) This isn't speculative architecture astronautics - it's applying proven distributed systems principles to AI agent coordination. The innovation is in the integration, not in requiring fundamental breakthroughs.