# Automation Profiles Automation Profiles control how much autonomy the CleverAgents system has at each phase of plan execution. Each profile defines a set of **threshold values** and **safety requirements** that determine when the system may proceed automatically versus when it must wait for human approval. ## Threshold Semantics Each threshold field is a float in the range `[0.0, 1.0]`: | Value | Meaning | |-------|---------| | `0.0` | Fully automatic — no human gate required | | `1.0` | Always requires human approval | | `0.0 < v < 1.0` | Probabilistic — the system may proceed if its confidence exceeds the threshold | ### Threshold Fields | Field | Category | Description (spec § Automatable Tasks) | |-------|----------|-------------| | `decompose_task` | Phase transition | Automatically enter Strategize after `plan use` | | `create_tool` | Phase transition | Automatically proceed from Strategize to Execute | | `select_tool` | Phase transition | Automatically proceed from Execute to Apply | | `edit_code` | Decision autonomy | Automatically make decisions during Strategize | | `execute_command` | Decision autonomy | Automatically make decisions during Execute | | `create_file` | Self-repair | Automatically attempt to fix validation failures | | `delete_content` | Self-repair | Automatically revise strategy when Execute hits constraints | | `access_network` | Self-repair | Automatically revert from Apply (`constrained`) to Strategize | | `install_dependency` | Execution controls | Automatically spawn and execute child plans | | `modify_config` | Retry | Automatically retry on transient failures | | `approve_plan` | Checkpoint | Automatically restore from checkpoint on failure | ### Safety Profile (Composed Sub-Model) Each automation profile composes a `SafetyProfile` via its `safety` field. The safety profile is the single source of truth for all hard safety constraints. See [ADR-041](../adr/ADR-041-safety-profile-extraction.md) for the architectural rationale. A `SafetyProfile` may also be attached directly to an `Action` (via `safety_profile`) when only safety constraints are needed without full autonomy thresholds. | Field | Type | Default | Description | |-------|------|---------|-------------| | `safety.require_sandbox` | bool | `true` | Execution must happen in a sandbox | | `safety.require_checkpoints` | bool | `true` | Checkpoints must be created before writes | | `safety.allow_unsafe_tools` | bool | `false` | Tools flagged as `unsafe` may be invoked | | `safety.require_human_approval` | bool | `false` | Human approval required before each action step | | `safety.allowed_skill_categories` | list[str] | `[]` | Skill categories permitted (empty = all allowed) | | `safety.max_cost_per_plan` | float \| null | `null` | Maximum cost in USD per plan (null = no limit) | | `safety.max_total_cost` | float \| null | `null` | Maximum total cost in USD across all plans (null = no limit) | | `safety.max_retries_per_step` | int | `3` | Maximum retry attempts per action step (0–100) | **Relationship to Automation Guards**: `safety.max_total_cost` is a plan-level budget cap, while `guards.max_total_cost` is a per-invocation budget cap. Both may be active; the tighter constraint wins. ## Built-in Profiles Eight profiles ship with every CleverAgents installation: | Flag | manual | review | supervised | cautious | trusted | auto | ci | full-auto | |------|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:| | decompose_task | 1.0 | 0.0 | 0.0 | 0.7 | 0.0 | 0.0 | 0.0 | 0.0 | | create_tool | 1.0 | 0.0 | 1.0 | 0.7 | 0.0 | 0.0 | 0.0 | 0.0 | | select_tool | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 0.0 | 0.0 | | edit_code | 1.0 | 1.0 | 0.0 | 0.6 | 0.0 | 0.0 | 0.0 | 0.0 | | execute_command | 1.0 | 1.0 | 1.0 | 0.8 | 0.0 | 0.0 | 0.0 | 0.0 | | create_file | 1.0 | 1.0 | 1.0 | 0.7 | 0.0 | 0.0 | 0.0 | 0.0 | | delete_content | 1.0 | 1.0 | 1.0 | 0.8 | 1.0 | 0.0 | 0.0 | 0.0 | | access_network | 1.0 | 1.0 | 1.0 | 0.9 | 1.0 | 1.0 | 0.0 | 0.0 | | install_dependency | 1.0 | 0.0 | 1.0 | 0.7 | 0.0 | 0.0 | 0.0 | 0.0 | | modify_config | 1.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | | approve_plan | 1.0 | 1.0 | 1.0 | 0.6 | 1.0 | 0.0 | 0.0 | 0.0 | | **Safety Profile** | | | | | | | | | | safety.require_sandbox | true | true | true | true | true | true | true | false | | safety.require_checkpoints | true | true | true | true | true | true | true | false | | safety.allow_unsafe_tools | false | false | false | false | false | false | false | true | | safety.require_human_approval | false | false | false | false | false | false | false | false | | safety.allowed_skill_categories | [] | [] | [] | [] | [] | [] | [] | [] | | safety.max_cost_per_plan | null | null | null | null | null | null | null | null | | safety.max_total_cost | null | null | null | null | null | null | null | null | | safety.max_retries_per_step | 3 | 3 | 3 | 3 | 3 | 3 | 3 | 3 | ### Profile Descriptions - **manual** — Human approves every action. Maximum safety, minimum autonomy. - **review** — Strategy and execution proceed automatically; human reviews before apply. Good default for development. - **supervised** — Human reviews strategy and execution decisions. Strategy creation itself is automatic. - **cautious** — Probabilistic gates on most actions. The system proceeds when confident, asks when unsure. - **trusted** — Automatic for most phases, but human approval required for apply and reversion. Suitable for experienced teams. - **auto** — Fully automatic except reversion from apply. Suitable for well-tested pipelines. - **ci** — Designed for CI/CD pipelines. All thresholds at 0.0 but sandbox and checkpoints remain required. - **full-auto** — No gates, no sandbox, no checkpoints, unsafe tools allowed. Use with extreme caution. ## Resolution Precedence When determining which profile applies to a given plan execution, the system resolves profiles in the following order (highest priority first): 1. **Plan-level override** — A profile specified directly on the plan. 2. **Project-level setting** — The default profile configured for the project. 3. **Organization-level default** — The organization's default profile. 4. **System default** — Falls back to the `review` built-in profile. At each level, the profile may be specified by: - A built-in name (e.g. `cautious`) - A namespaced custom profile (e.g. `acme/strict`) ## Custom Profiles Custom profiles use a `namespace/name` naming convention: ```yaml name: acme/strict description: Strict profile for production deployments schema_version: "1.0" decompose_task: 0.9 create_tool: 0.9 select_tool: 1.0 edit_code: 0.8 execute_command: 0.9 create_file: 0.5 delete_content: 0.9 access_network: 1.0 install_dependency: 0.8 modify_config: 0.3 approve_plan: 0.7 # Safety profile (composed sub-model) safety: require_sandbox: true require_checkpoints: true allow_unsafe_tools: false allowed_skill_categories: [] max_cost_per_plan: 50.0 max_total_cost: 200.0 max_retries_per_step: 3 ``` See `docs/schema/automation_profile.schema.yaml` for the full YAML schema and `examples/profiles/` for example configurations. ## Automation Guards Guards are optional enforcement hooks attached to a profile that gate tool invocations at runtime. They operate independently of the phase-transition thresholds and provide fine-grained control over tool usage. ### Guard Fields | Field | Type | Default | Description | |-------|------|---------|-------------| | `max_tool_calls_per_step` | `int \| None` | `None` | Maximum tool invocations per step before requiring approval | | `max_total_cost` | `float \| None` | `None` | Budget cap before requiring approval | | `tool_allowlist` | `list[str] \| None` | `None` | Only these tools can be called automatically | | `tool_denylist` | `list[str] \| None` | `None` | These tools always require approval | | `require_approval_for_writes` | `bool` | `false` | Require human approval for write operations | | `require_approval_for_apply` | `bool` | `false` | Require human approval before apply phase | ### Guard Evaluation Order When `check_guard()` is called, guards are evaluated in this order: 1. **Denylist** — If the tool is on the denylist, it is blocked immediately. 2. **Allowlist** — If an allowlist is set and the tool is not on it, it is blocked. 3. **Max tool calls** — If the call count has reached the limit, approval is required. 4. **Cost budget** — If the cumulative cost has reached the cap, approval is required. 5. **Write approval** — If the operation is a write and write approval is required. 6. **Apply approval** — If the tool name is `__apply__` and apply approval is required. If no guard blocks the invocation, it is allowed. ### GuardResult The result of a guard evaluation is a `GuardResult` with three fields: - `allowed` (`bool`) — Whether the tool invocation may proceed. - `reason` (`str | None`) — Human-readable explanation when blocked. - `requires_approval` (`bool`) — Whether human approval can unblock the invocation. ### Example YAML with Guards ```yaml name: acme/guarded description: Profile with guard constraints schema_version: "1.0" decompose_task: 0.7 create_tool: 0.7 select_tool: 1.0 # Safety profile (plan-level constraints) safety: require_sandbox: true require_checkpoints: true allow_unsafe_tools: false # Guards (per-invocation constraints) guards: max_tool_calls_per_step: 10 max_total_cost: 5.0 tool_denylist: - shell_exec - file_delete require_approval_for_writes: true require_approval_for_apply: true ``` ## CLI Commands The `agents automation-profile` command group manages automation profiles from the terminal. ### `automation-profile add` Register a custom profile from a YAML configuration file: ```bash agents automation-profile add --config ./profiles/strict.yaml agents automation-profile add --config ./profiles/strict.yaml --update agents automation-profile add --config ./profiles/strict.yaml --format json ``` Options: - `--config / -c` — Path to the YAML configuration file (required). - `--update` — Overwrite an existing custom profile instead of raising a conflict error. - `--format / -f` — Output format: `json`, `yaml`, `plain`, `table`, or `rich` (default: `rich`). ### `automation-profile remove` Remove a custom profile by name (built-in profiles cannot be removed): ```bash agents automation-profile remove acme/strict --yes agents automation-profile remove acme/strict --yes --format json ``` Options: - `--yes / -y` — Skip the interactive confirmation prompt. - `--format / -f` — Output format. ### `automation-profile list` List all profiles with optional filtering: ```bash agents automation-profile list agents automation-profile list --namespace acme agents automation-profile list "caut.*" agents automation-profile list --format json ``` Options: - *regex* (positional) — Optional regex pattern to filter profile names. - `--namespace / -n` — Filter by the namespace portion of namespaced names. - `--format / -f` — Output format. ### `automation-profile show` Display full details for a single profile: ```bash agents automation-profile show manual agents automation-profile show acme/strict --format yaml ``` Options: - `--format / -f` — Output format. ## Deprecation: `--automation-level` The `--automation-level` option on `plan use` and the `plan set-automation-level` command are **deprecated** and will be removed in a future release. Legacy automation levels are mapped to built-in profiles: | Legacy level | Built-in profile | |-------------|-----------------| | `manual` | `manual` | | `supervised` | `supervised` | | `auto` | `auto` | | `full_auto` | `full-auto` | Use `--automation-profile ` instead of `--automation-level`. ## Autonomy Guardrails Autonomy guardrails extend the automation guard framework with runtime constraints that enforce step limits, tool budgets, and required human confirmations during plan execution. Unlike the phase-transition thresholds which gate entire phases, autonomy guardrails operate at the individual step and tool invocation level. ### Guardrail Fields | Field | Type | Default | Description | |-------|------|---------|-------------| | `max_steps` | `int \| None` | `None` | Maximum execution steps allowed. `None` means unlimited. | | `tool_budget` | `float \| None` | `None` | Maximum cumulative cost for tool invocations. `None` means unlimited. A value of `0.0` blocks any positive cost. | | `max_wall_clock_seconds` | `float \| None` | `None` | Maximum wall-clock time in seconds for plan execution. `None` means unlimited. | | `actor_limits` | `ActorLimits` | `ActorLimits()` | Per-actor runtime limits (see below). | | `required_confirmations` | `list[str]` | `[]` | Operation names requiring human confirmation before execution (case-insensitive). | | `step_count` | `int` | `0` | Current execution step counter (managed by the service). | | `budget_spent` | `float` | `0.0` | Current cumulative budget spent (managed by the service). | | `start_time` | `str \| None` | `None` | ISO-8601 timestamp of when plan execution started (for wall-clock tracking). | ### Actor Limits Per-actor constraints that limit tool calls within a single actor invocation: | Field | Type | Default | Description | |-------|------|---------|-------------| | `max_tool_calls_per_invocation` | `int \| None` | `None` | Maximum tool calls allowed per actor invocation. `None` means unlimited. | | `max_retries_per_failure` | `int \| None` | `None` | Maximum retries per tool failure within an actor step. `None` means unlimited. | ### Enforcement Behavior Guardrails are enforced at multiple points during plan execution: - **Pre-flight** — Before execution begins, `PlanExecutor._guard_execute()` calls `_enforce_guardrails()` which checks wall-clock time and the current step limit. If either check fails, execution is blocked with a `PlanError` before any decisions are processed. - **Per-decision** — Before each decision in the execution loop, `_enforce_guardrails_per_step()` increments the step counter, checks the step limit, and rechecks wall-clock time. If any check fails, execution halts with a `PlanError`. - **Start-time tracking** — On the first execution, `_enforce_guardrails()` automatically calls `mark_started()` on the guardrails model to record the wall-clock start time (if not already set). The individual guardrail checks are: 1. **Step limit** — The `step_count` is compared against `max_steps`. If equal or exceeded, the step is **blocked** and an audit entry is recorded. 2. **Tool budget** — The projected cost (`budget_spent + tool_cost`) is compared against `tool_budget`. If exceeded, the invocation is **blocked** and an audit entry is recorded. When allowed, the cost is recorded **before** the audit entry so the trail reflects the post-deduction budget state. 3. **Wall-clock time** — If `max_wall_clock_seconds` is set and `start_time` is recorded, the elapsed wall-clock time is checked. If the limit is exceeded, the step is **blocked**. Malformed `start_time` values are treated as a block (defense-in-depth) rather than raising an unhandled exception. 4. **Actor tool-call limit** — The actor's current tool call count is compared against `actor_limits.max_tool_calls_per_invocation`. If exceeded, the call is **blocked**. 5. **Retry-per-failure limit** — The current retry count is compared against `actor_limits.max_retries_per_failure`. If equal or exceeded, the retry is **blocked** and an audit entry is recorded. When `max_retries_per_failure` is `None`, retries are unlimited. 6. **Required confirmations** — The operation name is checked (case-insensitively) against `required_confirmations`. If matched, the operation is **blocked** pending human confirmation. If no guardrail blocks the invocation, it proceeds and the step counter and budget are updated. ### Input Validation - `start_time` is validated on assignment to ensure it is a parseable ISO-8601 timestamp. Malformed strings are rejected with a `ValidationError`. - `max_retries_per_failure` must be non-negative (`>= 0`) if set; `None` means unlimited. - `max_steps` must be positive (`>= 1`) if set. - `tool_budget` must be non-negative (`>= 0.0`) if set. ### Audit Trail Every guardrail enforcement event is recorded in a `GuardrailAuditTrail` which is persisted to plan metadata. The audit trail provides a bounded, ordered record of autonomy-related decisions. When the number of entries exceeds `max_entries` (default: 10,000), the oldest entries are evicted to prevent unbounded memory growth. #### Audit Entry Fields | Field | Type | Description | |-------|------|-------------| | `timestamp` | `str` | ISO-8601 timestamp of the enforcement event. | | `event_type` | `GuardrailEventType` | Type of guardrail event (see below). | | `guard_name` | `str` | Name of the guard that was evaluated. | | `result` | `GuardrailResult` | Outcome: `allowed` or `denied`. | | `reason` | `str \| None` | Human-readable reason for the decision. | | `context` | `dict` | Additional metadata about the enforcement event. | #### Event Types | Event Type | Description | |------------|-------------| | `step_blocked` | A step was blocked due to the step limit. | | `budget_blocked` | A tool invocation was blocked due to budget. | | `confirmation_required` | An operation was blocked pending confirmation. | | `step_allowed` | A step was allowed within the limit. | | `budget_allowed` | A tool invocation was allowed within budget. | | `confirmation_granted` | An operation was allowed (no confirmation needed). | | `time_blocked` | A step was blocked due to the wall-clock time limit. | | `time_allowed` | A step was allowed within the wall-clock time limit. | | `actor_limit_blocked` | An actor tool call was blocked due to per-invocation limits. | | `actor_limit_allowed` | An actor tool call was allowed within per-invocation limits. | | `retry_blocked` | A retry was blocked due to the per-failure retry limit. | | `retry_allowed` | A retry was allowed within the per-failure retry limit. | ### Plan Metadata Persistence The guardrail state and audit trail are serialized to plan metadata using two keys: - `autonomy_guardrails` — The current guardrail configuration and counters. - `guardrail_audit_trail` — The ordered list of audit entries. This allows guardrail state to survive across plan restarts and provides a complete audit record for post-mortem analysis. ### Example Configuration ```yaml name: acme/guarded-auto description: Auto profile with autonomy guardrails schema_version: "1.0" decompose_task: 0.0 create_tool: 0.0 select_tool: 1.0 guards: max_tool_calls_per_step: 10 max_total_cost: 100.0 # Autonomy guardrails (configured via service or CLI) # max_steps: 50 # tool_budget: 200.0 # required_confirmations: # - deploy # - delete ```