Compare commits
14 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| c077d3b492 | |||
| 8005385ace | |||
| 245d9eec44 | |||
| 53d15cbec3 | |||
| 86ea4e2ee1 | |||
| 1c6d1786d6 | |||
| 9a5e63ff20 | |||
| 6a4e8e96a0 | |||
| afaefe5fe3 | |||
| cc49e8f8f6 | |||
| 97f969fafa | |||
| fb37aa7c15 | |||
| 77b48a76df | |||
| a650d307e1 |
+14
-8
@@ -5,6 +5,13 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
|
||||
- **Multi-Session Tabs** (`feat(tui)`): Enhanced TUI with independent session management,
|
||||
session creation/switching/closing/renaming via keyboard bindings (Ctrl+N for new, Ctrl+W for close),
|
||||
and independent A2A binding support per session. Each session maintains its own state, persona selection,
|
||||
transcript history, and argument presets. (#8445)
|
||||
|
||||
### Fixed
|
||||
|
||||
- **Actor v3 YAML Schema Validation in CLI** (#5869): The `agents actor add --config`
|
||||
@@ -28,14 +35,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||||
correctly in all deployment modes: Docker containers, local pip installs
|
||||
(wheel or editable), and development environments.
|
||||
- **TDD Non-AssertionError Guard Visibility** (#8294): `apply_tdd_inversion` in
|
||||
- **bug-hunt-pool-supervisor Non-Blocking Tracking** (#8835): The automation-tracking-manager
|
||||
call in step 5 was blocking the main loop indefinitely, causing 3+ consecutive initialization
|
||||
failures. Step 5 now explicitly marks tracking as best-effort -- if the call does not complete
|
||||
within a reasonable time or fails, it is skipped and the supervisor continues to the next
|
||||
cycle. A new Rule 9 reinforces that tracking must never block the main loop; core
|
||||
functionality (module mapping, worker dispatch, monitoring) takes priority over status
|
||||
reporting.
|
||||
|
||||
`features/environment.py` now emits its non-assertion exception guard warning to
|
||||
both the structured logger and `stderr` via a new `_warning_with_stderr` helper.
|
||||
This makes the guard firing visible in standard Behave console output and CI log
|
||||
@@ -356,6 +355,13 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||||
references; persisted in `~/.config/cleveragents/personas/`.
|
||||
- **TUI session export/import** — full JSON round-trip and Markdown transcript export
|
||||
(`--format md`).
|
||||
- **PersonaRegistry** — YAML-based persona management system with full CRUD operations,
|
||||
atomic file operations with fcntl locking, and thread-safe implementation.
|
||||
- **TUI Web Mode** (`--web`, `--web-port`) — HTTP server for browser-based TUI access
|
||||
with HTML interface and WebSocket support placeholder for future enhancements.
|
||||
- **Multi-Session Tabs** — Enhanced TUI with independent session management, session
|
||||
creation/switching/closing/renaming via keyboard bindings (Ctrl+N for new, Ctrl+W for close),
|
||||
and independent A2A binding support per session.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -23,3 +23,4 @@ Below are some of the specific details of various contributions.
|
||||
* HAL 9000 has contributed automated bug fixes, including fix #7488 (store sandbox_path in checkpoint metadata to enable rollback).
|
||||
* This project was made possible thanks to considerable donation of time, money, and resources by CleverThis, Inc.
|
||||
* HAL 9000 has contributed automated bug fixes, CLI output formatting improvements, and ongoing maintenance as part of the CleverAgents automation system.
|
||||
* HAL 9000 has contributed the multi-session tabs feature (issue #8445): implemented independent session management with Ctrl+N/Ctrl+W keyboard bindings, session switching/closing/renaming, and per-session A2A binding isolation in the TUI application layer.
|
||||
|
||||
@@ -1,453 +0,0 @@
|
||||
# Configuration Reference Guide
|
||||
|
||||
This guide provides a comprehensive reference for configuring CleverAgents. Configuration is primarily managed through environment variables, with sensible defaults for most settings.
|
||||
|
||||
## Overview
|
||||
|
||||
CleverAgents uses a **Pydantic-based settings system** that reads configuration from environment variables with the `CLEVERAGENTS_` prefix. The configuration system is designed to be:
|
||||
|
||||
- **Environment-first**: All configuration comes from environment variables, making it ideal for containerized and cloud-native deployments
|
||||
- **Hierarchical**: Settings are organized by functional area (runtime, providers, storage, observability, etc.)
|
||||
- **Validated**: All configuration values are validated at startup with clear error messages
|
||||
- **Sensible defaults**: Most settings have reasonable defaults suitable for development and production
|
||||
|
||||
### Configuration Loading Order
|
||||
|
||||
1. **Environment variables** (highest priority) — `CLEVERAGENTS_*` prefixed variables
|
||||
2. **Provider-specific environment variables** — Native provider variables (e.g., `OPENAI_API_KEY`)
|
||||
3. **Default values** — Hardcoded defaults in the settings schema
|
||||
|
||||
When a setting is not explicitly configured, the system falls back to the next level in the order above.
|
||||
|
||||
### Configuration File Locations
|
||||
|
||||
CleverAgents stores configuration and data in the following locations:
|
||||
|
||||
| Purpose | Default Location | Environment Variable |
|
||||
|---------|------------------|----------------------|
|
||||
| User data (sessions, personas, etc.) | `~/.cleveragents/` | `CLEVERAGENTS_DATA_DIR` |
|
||||
| Database | `~/.cleveragents/cleveragents.db` | `CLEVERAGENTS_DATABASE_URL` |
|
||||
| Logs | `./logs/` | `CLEVERAGENTS_LOG_DIR` |
|
||||
| Storage (plans, resources, etc.) | `./data/` | `CLEVERAGENTS_STORAGE_BASE_PATH` |
|
||||
| Vector store | `./.cleveragents/vector_store/` | `CLEVERAGENTS_VECTOR_STORE_PATH` |
|
||||
|
||||
## Environment Variables Reference
|
||||
|
||||
All environment variables are organized by functional area. Each entry includes:
|
||||
|
||||
- **Variable name** — The full `CLEVERAGENTS_*` environment variable name
|
||||
- **Type** — The expected data type (string, integer, boolean, etc.)
|
||||
- **Default** — The default value if not set
|
||||
- **Description** — What the setting controls
|
||||
- **Example** — A practical example value
|
||||
|
||||
### Runtime & Server Configuration
|
||||
|
||||
#### `CLEVERAGENTS_ENV`
|
||||
- **Type:** String
|
||||
- **Default:** `development`
|
||||
- **Valid values:** `development`, `production`, `staging`, `test`
|
||||
- **Description:** The runtime environment. Controls logging verbosity, debug features, and validation strictness.
|
||||
- **Example:** `export CLEVERAGENTS_ENV=production`
|
||||
|
||||
#### `CLEVERAGENTS_DEBUG_ENABLED`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** Enable debug mode with verbose logging and additional runtime checks. Should never be enabled in production.
|
||||
- **Example:** `export CLEVERAGENTS_DEBUG_ENABLED=true`
|
||||
|
||||
#### `CLEVERAGENTS_SERVER_HOST`
|
||||
- **Type:** String (IP address)
|
||||
- **Default:** `0.0.0.0`
|
||||
- **Description:** The host address for the CleverAgents server to bind to.
|
||||
- **Example:** `export CLEVERAGENTS_SERVER_HOST=127.0.0.1`
|
||||
|
||||
#### `CLEVERAGENTS_SERVER_PORT`
|
||||
- **Type:** Integer
|
||||
- **Default:** `8080`
|
||||
- **Valid range:** 1-65535
|
||||
- **Description:** The port for the CleverAgents server to listen on.
|
||||
- **Example:** `export CLEVERAGENTS_SERVER_PORT=9000`
|
||||
|
||||
#### `CLEVERAGENTS_SERVER_RELOAD`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** Enable auto-reload on code changes (development only). Never use in production.
|
||||
- **Example:** `export CLEVERAGENTS_SERVER_RELOAD=true`
|
||||
|
||||
#### `CLEVERAGENTS_SERVER_URL`
|
||||
- **Type:** String (URL)
|
||||
- **Default:** `None` (local mode)
|
||||
- **Description:** URL of a remote CleverAgents server to connect to. When set, the CLI operates in server mode.
|
||||
- **Example:** `export CLEVERAGENTS_SERVER_URL=https://cleveragents.example.com`
|
||||
|
||||
#### `CLEVERAGENTS_SERVER_TOKEN`
|
||||
- **Type:** String
|
||||
- **Default:** `None`
|
||||
- **Description:** Authentication token for connecting to a remote CleverAgents server. Required when `CLEVERAGENTS_SERVER_URL` is set.
|
||||
- **Example:** `export CLEVERAGENTS_SERVER_TOKEN=sk-1234567890abcdef`
|
||||
|
||||
### AI Provider Configuration
|
||||
|
||||
#### Provider API Keys
|
||||
|
||||
Each AI provider requires its API key to be configured. CleverAgents supports multiple providers simultaneously and automatically selects the best available one.
|
||||
|
||||
| Provider | Primary Variable | Fallback Variables | Example |
|
||||
|----------|------------------|-------------------|---------|
|
||||
| OpenAI | `OPENAI_API_KEY` | `CLEVERAGENTS_OPENAI_API_KEY` | `sk-proj-...` |
|
||||
| Anthropic | `ANTHROPIC_API_KEY` | `CLEVERAGENTS_ANTHROPIC_API_KEY` | `sk-ant-...` |
|
||||
| Google | `GOOGLE_API_KEY` | `GOOGLE_GENAI_API_KEY`, `CLEVERAGENTS_GOOGLE_API_KEY` | `AIza...` |
|
||||
| Azure OpenAI | `AZURE_OPENAI_API_KEY` | `AZURE_API_KEY`, `CLEVERAGENTS_AZURE_API_KEY` | `...` |
|
||||
| OpenRouter | `OPENROUTER_API_KEY` | `CLEVERAGENTS_OPENROUTER_API_KEY` | `sk-or-...` |
|
||||
| Gemini | `GEMINI_API_KEY` | `GOOGLE_GEMINI_API_KEY`, `CLEVERAGENTS_GEMINI_API_KEY` | `AIza...` |
|
||||
| Cohere | `COHERE_API_KEY` | `CLEVERAGENTS_COHERE_API_KEY` | `...` |
|
||||
| Groq | `GROQ_API_KEY` | `CLEVERAGENTS_GROQ_API_KEY` | `gsk_...` |
|
||||
| Together | `TOGETHER_API_KEY` | `CLEVERAGENTS_TOGETHER_API_KEY` | `...` |
|
||||
| HuggingFace | `HF_TOKEN` | `HUGGINGFACEHUB_API_TOKEN`, `HUGGING_FACE_HUB_TOKEN` | `hf_...` |
|
||||
| Perplexity | `PERPLEXITY_API_KEY` | `CLEVERAGENTS_PERPLEXITY_API_KEY` | `pplx-...` |
|
||||
|
||||
#### `CLEVERAGENTS_DEFAULT_PROVIDER`
|
||||
- **Type:** String
|
||||
- **Default:** `None` (auto-detect from configured providers)
|
||||
- **Valid values:** `openai`, `anthropic`, `google`, `azure`, `openrouter`, `gemini`, `cohere`, `groq`, `together`, `huggingface`, `perplexity`
|
||||
- **Description:** Pin the default AI provider. When not set, the system automatically selects the first available configured provider in the fallback order.
|
||||
- **Example:** `export CLEVERAGENTS_DEFAULT_PROVIDER=anthropic`
|
||||
|
||||
#### `CLEVERAGENTS_DEFAULT_MODEL`
|
||||
- **Type:** String
|
||||
- **Default:** `None` (provider-specific default)
|
||||
- **Description:** Pin the default model for the selected provider. When not set, each provider's published default is used (e.g., `gpt-4o` for OpenAI, `claude-sonnet-4-20250514` for Anthropic).
|
||||
- **Example:** `export CLEVERAGENTS_DEFAULT_MODEL=gpt-4-turbo`
|
||||
|
||||
#### `CLEVERAGENTS_MOCK_PROVIDERS`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** Enable mock AI providers for testing. Must not be used in production. Useful for development and CI/CD pipelines.
|
||||
- **Example:** `export CLEVERAGENTS_MOCK_PROVIDERS=true`
|
||||
|
||||
#### `CLEVERAGENTS_TESTING_USE_MOCK_AI`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** Force the in-repo mock provider for Behave/Robot test suites to prevent external API calls.
|
||||
- **Example:** `export CLEVERAGENTS_TESTING_USE_MOCK_AI=true`
|
||||
|
||||
### Logging & Paths
|
||||
|
||||
#### `CLEVERAGENTS_LOG_LEVEL`
|
||||
- **Type:** String
|
||||
- **Default:** `INFO`
|
||||
- **Valid values:** `DEBUG`, `INFO`, `WARNING`, `ERROR`, `CRITICAL`
|
||||
- **Description:** The logging level for standard output and file logs.
|
||||
- **Example:** `export CLEVERAGENTS_LOG_LEVEL=DEBUG`
|
||||
|
||||
#### `CLEVERAGENTS_LOG_DIR`
|
||||
- **Type:** Path
|
||||
- **Default:** `./logs/`
|
||||
- **Description:** Directory where log files are written.
|
||||
- **Example:** `export CLEVERAGENTS_LOG_DIR=/var/log/cleveragents`
|
||||
|
||||
#### `CLEVERAGENTS_DATA_DIR`
|
||||
- **Type:** Path
|
||||
- **Default:** `~/.cleveragents/`
|
||||
- **Description:** Directory for user data, sessions, personas, and other persistent state.
|
||||
- **Example:** `export CLEVERAGENTS_DATA_DIR=/data/cleveragents`
|
||||
|
||||
#### `CLEVERAGENTS_STORAGE_BASE_PATH`
|
||||
- **Type:** Path
|
||||
- **Default:** `./data/`
|
||||
- **Description:** Base directory for plan storage, checkpoints, and other runtime data.
|
||||
- **Example:** `export CLEVERAGENTS_STORAGE_BASE_PATH=/storage/cleveragents`
|
||||
|
||||
### Database & Persistence
|
||||
|
||||
#### `CLEVERAGENTS_DATABASE_URL`
|
||||
- **Type:** String (database URL)
|
||||
- **Default:** `sqlite:///$HOME/.cleveragents/cleveragents.db`
|
||||
- **Supported formats:**
|
||||
- SQLite: `sqlite:///path/to/database.db`
|
||||
- PostgreSQL: `postgresql://user:password@host:port/database`
|
||||
- MySQL: `mysql+pymysql://user:password@host:port/database`
|
||||
- DuckDB: `duckdb:///path/to/database.duckdb`
|
||||
- **Description:** The database URL for storing sessions, plans, and other persistent data.
|
||||
- **Example:** `export CLEVERAGENTS_DATABASE_URL=postgresql://user:pass@localhost/cleveragents`
|
||||
|
||||
### Cost Controls & Budgets
|
||||
|
||||
#### `CLEVERAGENTS_SESSION_MAX_COST_USD`
|
||||
- **Type:** Float
|
||||
- **Default:** `None` (unlimited)
|
||||
- **Valid range:** >= 0.0
|
||||
- **Description:** Maximum USD spend per session across all plans. `None` means unlimited.
|
||||
- **Example:** `export CLEVERAGENTS_SESSION_MAX_COST_USD=100.00`
|
||||
|
||||
#### `CLEVERAGENTS_ORG_MAX_COST_USD`
|
||||
- **Type:** Float
|
||||
- **Default:** `None` (unlimited)
|
||||
- **Valid range:** >= 0.0
|
||||
- **Description:** Maximum USD spend per organization across all sessions. `None` means unlimited.
|
||||
- **Example:** `export CLEVERAGENTS_ORG_MAX_COST_USD=1000.00`
|
||||
|
||||
#### `CLEVERAGENTS_BUDGET_PER_PLAN`
|
||||
- **Type:** Float
|
||||
- **Default:** `None` (unlimited)
|
||||
- **Valid range:** >= 0.0
|
||||
- **Description:** Maximum USD spend per plan execution. `None` means unlimited.
|
||||
- **Example:** `export CLEVERAGENTS_BUDGET_PER_PLAN=50.00`
|
||||
|
||||
#### `CLEVERAGENTS_BUDGET_PER_DAY`
|
||||
- **Type:** Float
|
||||
- **Default:** `None` (unlimited)
|
||||
- **Valid range:** >= 0.0
|
||||
- **Description:** Maximum USD spend per calendar day across all plans. `None` means unlimited.
|
||||
- **Example:** `export CLEVERAGENTS_BUDGET_PER_DAY=500.00`
|
||||
|
||||
### Observability & Logging
|
||||
|
||||
#### LangSmith Tracing
|
||||
|
||||
#### `CLEVERAGENTS_LANGSMITH_ENABLED`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Aliases:** `LANGCHAIN_TRACING_V2`, `LANGSMITH_TRACING_V2`
|
||||
- **Description:** Enable LangSmith tracing for LangChain/LangGraph agents.
|
||||
- **Example:** `export CLEVERAGENTS_LANGSMITH_ENABLED=true`
|
||||
|
||||
#### `CLEVERAGENTS_LANGSMITH_API_KEY`
|
||||
- **Type:** String
|
||||
- **Default:** `None`
|
||||
- **Aliases:** `LANGSMITH_API_KEY`, `LANGCHAIN_API_KEY`
|
||||
- **Description:** API key for LangSmith. Required when tracing is enabled.
|
||||
- **Example:** `export CLEVERAGENTS_LANGSMITH_API_KEY=ls_...`
|
||||
|
||||
#### `CLEVERAGENTS_LANGSMITH_PROJECT`
|
||||
- **Type:** String
|
||||
- **Default:** `None`
|
||||
- **Aliases:** `LANGSMITH_PROJECT`, `LANGCHAIN_PROJECT`
|
||||
- **Description:** LangSmith project name. Required when tracing is enabled.
|
||||
- **Example:** `export CLEVERAGENTS_LANGSMITH_PROJECT=my-project`
|
||||
|
||||
### Cleanup & Retention Policies
|
||||
|
||||
#### `CLEVERAGENTS_CLEANUP_SANDBOX_MAX_AGE_HOURS`
|
||||
- **Type:** Integer
|
||||
- **Default:** `48`
|
||||
- **Valid range:** >= 1
|
||||
- **Description:** Maximum age (in hours) before stale sandboxes are eligible for cleanup.
|
||||
- **Example:** `export CLEVERAGENTS_CLEANUP_SANDBOX_MAX_AGE_HOURS=72`
|
||||
|
||||
#### `CLEVERAGENTS_CHECKPOINT_MAX`
|
||||
- **Type:** Integer
|
||||
- **Default:** `50`
|
||||
- **Valid range:** >= 2
|
||||
- **Description:** Maximum number of checkpoints to keep per plan. Oldest checkpoints are pruned first, but the first and most recent are always preserved.
|
||||
- **Example:** `export CLEVERAGENTS_CHECKPOINT_MAX=100`
|
||||
|
||||
### Audit Logging
|
||||
|
||||
#### `CLEVERAGENTS_AUDIT_RETENTION_DAYS`
|
||||
- **Type:** Integer
|
||||
- **Default:** `0` (keep indefinitely)
|
||||
- **Valid range:** >= 0
|
||||
- **Description:** Days to retain audit log entries before pruning. `0` means keep indefinitely (spec default for compliance).
|
||||
- **Example:** `export CLEVERAGENTS_AUDIT_RETENTION_DAYS=365`
|
||||
|
||||
#### `CLEVERAGENTS_AUDIT_ASYNC`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `true`
|
||||
- **Description:** When true, audit entries are written asynchronously via a write-behind queue on a background thread. Set to false for synchronous behavior (useful for debugging).
|
||||
- **Example:** `export CLEVERAGENTS_AUDIT_ASYNC=false`
|
||||
|
||||
### Metrics & Observability
|
||||
|
||||
#### `CLEVERAGENTS_METRICS_ENABLED`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `true`
|
||||
- **Description:** Enable structured metric collection and emission.
|
||||
- **Example:** `export CLEVERAGENTS_METRICS_ENABLED=false`
|
||||
|
||||
#### `CLEVERAGENTS_METRICS_EXPORT_PROMETHEUS`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** Enable Prometheus metrics export endpoint.
|
||||
- **Example:** `export CLEVERAGENTS_METRICS_EXPORT_PROMETHEUS=true`
|
||||
|
||||
### Retry & Circuit Breaker Configuration
|
||||
|
||||
#### `CLEVERAGENTS_RETRY_MAX_ATTEMPTS`
|
||||
- **Type:** Integer
|
||||
- **Default:** `3`
|
||||
- **Valid range:** 1-50
|
||||
- **Description:** Default maximum retry attempts for service operations.
|
||||
- **Example:** `export CLEVERAGENTS_RETRY_MAX_ATTEMPTS=5`
|
||||
|
||||
#### `CLEVERAGENTS_RETRY_BASE_DELAY`
|
||||
- **Type:** Float
|
||||
- **Default:** `1.0`
|
||||
- **Valid range:** 0.0-300.0
|
||||
- **Description:** Default base delay in seconds between retries.
|
||||
- **Example:** `export CLEVERAGENTS_RETRY_BASE_DELAY=0.5`
|
||||
|
||||
#### `CLEVERAGENTS_RETRY_MAX_DELAY`
|
||||
- **Type:** Float
|
||||
- **Default:** `60.0`
|
||||
- **Valid range:** 0.0-3600.0
|
||||
- **Description:** Default maximum delay in seconds between retries.
|
||||
- **Example:** `export CLEVERAGENTS_RETRY_MAX_DELAY=120.0`
|
||||
|
||||
#### `CLEVERAGENTS_RETRY_BACKOFF_STRATEGY`
|
||||
- **Type:** String
|
||||
- **Default:** `exponential`
|
||||
- **Valid values:** `exponential`, `linear`, `fixed`, `jitter`, `none`
|
||||
- **Description:** Default backoff strategy for retries.
|
||||
- **Example:** `export CLEVERAGENTS_RETRY_BACKOFF_STRATEGY=linear`
|
||||
|
||||
### Security & Secrets
|
||||
|
||||
#### `CLEVERAGENTS_SHOW_SECRETS`
|
||||
- **Type:** Boolean
|
||||
- **Default:** `false`
|
||||
- **Description:** When true, secrets are shown in CLI output and logs. Should never be enabled in production.
|
||||
- **Example:** `export CLEVERAGENTS_SHOW_SECRETS=true`
|
||||
|
||||
### Automation & Output
|
||||
|
||||
#### `CLEVERAGENTS_AUTOMATION_PROFILE`
|
||||
- **Type:** String
|
||||
- **Default:** `` (empty, manual mode)
|
||||
- **Valid values:** `manual`, `auto`, `full-auto` (or custom profile names)
|
||||
- **Description:** Default automation profile name controlling plan execution behavior.
|
||||
- **Example:** `export CLEVERAGENTS_AUTOMATION_PROFILE=auto`
|
||||
|
||||
#### `CLEVERAGENTS_FORMAT`
|
||||
- **Type:** String
|
||||
- **Default:** `None` (uses default format)
|
||||
- **Valid values:** `json`, `markdown`, `text`, `yaml`
|
||||
- **Description:** Default output format for CLI commands.
|
||||
- **Example:** `export CLEVERAGENTS_FORMAT=json`
|
||||
|
||||
## Configuration Best Practices
|
||||
|
||||
### Development Environment
|
||||
|
||||
For local development, create a `.env` file in your project root:
|
||||
|
||||
```bash
|
||||
# .env
|
||||
CLEVERAGENTS_ENV=development
|
||||
CLEVERAGENTS_DEBUG_ENABLED=true
|
||||
CLEVERAGENTS_LOG_LEVEL=DEBUG
|
||||
OPENAI_API_KEY=sk-proj-...
|
||||
CLEVERAGENTS_DATABASE_URL=sqlite:///./dev.db
|
||||
```
|
||||
|
||||
### Production Environment
|
||||
|
||||
For production deployments:
|
||||
|
||||
```bash
|
||||
# Use strong, randomly generated secrets
|
||||
export CLEVERAGENTS_ENV=production
|
||||
export CLEVERAGENTS_DEBUG_ENABLED=false
|
||||
export CLEVERAGENTS_LOG_LEVEL=WARNING
|
||||
export CLEVERAGENTS_SHOW_SECRETS=false
|
||||
|
||||
# Use a production database
|
||||
export CLEVERAGENTS_DATABASE_URL=postgresql://user:password@db.example.com/cleveragents
|
||||
|
||||
# Configure cost controls
|
||||
export CLEVERAGENTS_SESSION_MAX_COST_USD=100
|
||||
export CLEVERAGENTS_ORG_MAX_COST_USD=1000
|
||||
|
||||
# Enable observability
|
||||
export CLEVERAGENTS_METRICS_ENABLED=true
|
||||
export CLEVERAGENTS_LANGSMITH_ENABLED=true
|
||||
export CLEVERAGENTS_LANGSMITH_API_KEY=ls_...
|
||||
export CLEVERAGENTS_LANGSMITH_PROJECT=production
|
||||
```
|
||||
|
||||
### Multi-Provider Setup
|
||||
|
||||
To support multiple AI providers with automatic fallback:
|
||||
|
||||
```bash
|
||||
# Configure multiple providers
|
||||
export OPENAI_API_KEY=sk-proj-...
|
||||
export ANTHROPIC_API_KEY=sk-ant-...
|
||||
export GOOGLE_API_KEY=AIza...
|
||||
|
||||
# Set primary provider (optional)
|
||||
export CLEVERAGENTS_DEFAULT_PROVIDER=anthropic
|
||||
```
|
||||
|
||||
## Security Considerations
|
||||
|
||||
### Secrets Management
|
||||
|
||||
1. **Never commit secrets** — Use environment variables or a secrets manager (Vault, AWS Secrets Manager, etc.)
|
||||
2. **Use strong API keys** — Ensure all provider API keys are strong and rotated regularly
|
||||
3. **Restrict access** — Limit who can view or modify configuration in your deployment
|
||||
4. **Audit logging** — Enable audit logging to track configuration changes and API usage
|
||||
5. **Redaction** — CleverAgents automatically redacts secrets in logs and output
|
||||
|
||||
### Database Security
|
||||
|
||||
- **Use encrypted connections** — Always use TLS/SSL for database connections
|
||||
- **Strong credentials** — Use strong, randomly generated database passwords
|
||||
- **Network isolation** — Restrict database access to authorized services only
|
||||
|
||||
### API Key Security
|
||||
|
||||
- **Rotate keys regularly** — Implement automated key rotation (e.g., monthly)
|
||||
- **Use scoped keys** — If your provider supports it, use keys with minimal required permissions
|
||||
- **Monitor usage** — Track API usage and set up alerts for unusual activity
|
||||
|
||||
## Troubleshooting Configuration Issues
|
||||
|
||||
### No AI Provider Configured
|
||||
|
||||
**Error:** `ValueError: No AI providers configured and mock_providers is disabled.`
|
||||
|
||||
**Solution:**
|
||||
1. Set at least one provider API key: `export OPENAI_API_KEY=sk-proj-...`
|
||||
2. Or enable mock providers for testing: `export CLEVERAGENTS_MOCK_PROVIDERS=true`
|
||||
3. Verify the key is set: `echo $OPENAI_API_KEY`
|
||||
|
||||
### LangSmith Configuration Incomplete
|
||||
|
||||
**Error:** `LangSmith API key is required` or `LangSmith project name is required`
|
||||
|
||||
**Solution:**
|
||||
1. Set both required variables:
|
||||
```bash
|
||||
export CLEVERAGENTS_LANGSMITH_API_KEY=ls_...
|
||||
export CLEVERAGENTS_LANGSMITH_PROJECT=my-project
|
||||
```
|
||||
2. Or disable LangSmith: `export CLEVERAGENTS_LANGSMITH_ENABLED=false`
|
||||
|
||||
### Database Connection Failed
|
||||
|
||||
**Error:** `sqlalchemy.exc.OperationalError: (psycopg2.OperationalError) could not connect to server`
|
||||
|
||||
**Solution:**
|
||||
1. Verify the database URL: `echo $CLEVERAGENTS_DATABASE_URL`
|
||||
2. Check database credentials and host
|
||||
3. Ensure the database is running and accessible
|
||||
|
||||
### Configuration Not Taking Effect
|
||||
|
||||
**Problem:** Environment variable changes not reflected
|
||||
|
||||
**Solution:**
|
||||
1. Verify the variable is set: `env | grep CLEVERAGENTS`
|
||||
2. Restart the application (singleton settings are cached)
|
||||
3. Check for typos in variable names (case-sensitive)
|
||||
4. Ensure the variable is exported: `export CLEVERAGENTS_VAR=value`
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [Architecture Guide](../architecture.md) — System design and component overview
|
||||
- [API Reference - Configuration](../api/config.md) — Detailed API documentation
|
||||
- [Development Guide](../development/quality-automation.md) — Development setup and workflows
|
||||
- [Specification](../specification.md) — Authoritative design documentation
|
||||
- [FAQ](../faq.md) — Frequently asked questions
|
||||
+5494
-18
File diff suppressed because one or more lines are too long
@@ -0,0 +1,164 @@
|
||||
Feature: Budget enforcement in PlanExecutor
|
||||
As a platform operator
|
||||
I want PlanExecutor to halt execution when budget limits are exceeded
|
||||
So that autonomous agents cannot overspend configured limits
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# BudgetExceededError and PlanBudgetExceededError exceptions
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: BudgetExceededError has correct attributes
|
||||
When I create a BudgetExceededError with plan_id "plan-1" budget_type "daily" used 5.0 limit 3.0
|
||||
Then the BudgetExceededError plan_id should be "plan-1"
|
||||
And the BudgetExceededError budget_type should be "daily"
|
||||
And the BudgetExceededError used should be 5.0
|
||||
And the BudgetExceededError limit should be 3.0
|
||||
And the BudgetExceededError message should contain "budget"
|
||||
|
||||
Scenario: PlanBudgetExceededError has correct attributes
|
||||
When I create a PlanBudgetExceededError with plan_id "plan-2" used 10.0 limit 5.0
|
||||
Then the PlanBudgetExceededError plan_id should be "plan-2"
|
||||
And the PlanBudgetExceededError used should be 10.0
|
||||
And the PlanBudgetExceededError limit should be 5.0
|
||||
And the PlanBudgetExceededError message should contain "budget"
|
||||
|
||||
Scenario: BudgetExceededError is a subclass of PlanError
|
||||
When I create a BudgetExceededError with plan_id "plan-3" budget_type "session" used 1.0 limit 0.5
|
||||
Then the BudgetExceededError should be an instance of PlanError
|
||||
|
||||
Scenario: PlanBudgetExceededError is a subclass of PlanError
|
||||
When I create a PlanBudgetExceededError with plan_id "plan-4" used 2.0 limit 1.0
|
||||
Then the PlanBudgetExceededError should be an instance of PlanError
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# PlanExecutor budget enforcement - no cost tracker (no-op)
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: PlanExecutor without cost_tracker does not check budget
|
||||
Given a budget enforcement PlanExecutor without cost_tracker
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
When I run budget enforcement execute
|
||||
Then the budget enforcement execute should succeed without budget error
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# PlanExecutor budget enforcement - plan budget exceeded
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: PlanExecutor halts with PlanBudgetExceededError when plan budget exceeded
|
||||
Given a budget enforcement PlanExecutor with plan budget 0.0001
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the cost_metadata has total_cost 0.001
|
||||
When I run budget enforcement execute expecting budget error
|
||||
Then a PlanBudgetExceededError should be raised
|
||||
And the PlanBudgetExceededError plan_id should match the plan
|
||||
|
||||
Scenario: PlanExecutor saves plan state before halting on plan budget exceeded
|
||||
Given a budget enforcement PlanExecutor with plan budget 0.0001
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the cost_metadata has total_cost 0.001
|
||||
When I run budget enforcement execute expecting budget error
|
||||
Then the lifecycle _commit_plan should have been called with budget_halt details
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# PlanExecutor budget enforcement - daily budget exceeded
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: PlanExecutor halts with BudgetExceededError when daily budget exceeded
|
||||
Given a budget enforcement PlanExecutor with daily budget 0.0001
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the daily spend is 0.001
|
||||
When I run budget enforcement execute expecting budget error
|
||||
Then a BudgetExceededError should be raised
|
||||
And the BudgetExceededError budget_type should be "daily"
|
||||
|
||||
Scenario: PlanExecutor saves plan state before halting on daily budget exceeded
|
||||
Given a budget enforcement PlanExecutor with daily budget 0.0001
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the daily spend is 0.001
|
||||
When I run budget enforcement execute expecting budget error
|
||||
Then the lifecycle _commit_plan should have been called with budget_halt details
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# PlanExecutor budget enforcement - within budget (no halt)
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: PlanExecutor continues when plan budget is not exceeded
|
||||
Given a budget enforcement PlanExecutor with plan budget 100.0
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the cost_metadata has total_cost 0.001
|
||||
When I run budget enforcement execute
|
||||
Then the budget enforcement execute should succeed without budget error
|
||||
|
||||
Scenario: PlanExecutor continues when daily budget is not exceeded
|
||||
Given a budget enforcement PlanExecutor with daily budget 100.0
|
||||
And a budget enforcement plan in Execute-Queued state
|
||||
And the daily spend is 0.001
|
||||
When I run budget enforcement execute
|
||||
Then the budget enforcement execute should succeed without budget error
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# AutomationProfile budget fields
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: AutomationProfile has budget_per_plan field
|
||||
When I create an AutomationProfile with budget_per_plan 10.0
|
||||
Then the AutomationProfile budget_per_plan should be 10.0
|
||||
|
||||
Scenario: AutomationProfile has budget_per_session field
|
||||
When I create an AutomationProfile with budget_per_session 50.0
|
||||
Then the AutomationProfile budget_per_session should be 50.0
|
||||
|
||||
Scenario: AutomationProfile budget_per_plan defaults to None
|
||||
When I create an AutomationProfile with default budget fields
|
||||
Then the AutomationProfile budget_per_plan should be None
|
||||
And the AutomationProfile budget_per_session should be None
|
||||
|
||||
Scenario: AutomationProfile rejects negative budget_per_plan
|
||||
When I try to create an AutomationProfile with budget_per_plan -1.0
|
||||
Then a budget enforcement validation error should be raised
|
||||
|
||||
Scenario: AutomationProfile rejects negative budget_per_session
|
||||
When I try to create an AutomationProfile with budget_per_session -1.0
|
||||
Then a budget enforcement validation error should be raised
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# _check_budget method - direct unit tests
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: _check_budget is a no-op when cost_tracker is None
|
||||
Given a budget enforcement PlanExecutor without cost_tracker
|
||||
When I call _check_budget directly with plan_id "test-plan"
|
||||
Then no budget exception should be raised
|
||||
|
||||
Scenario: _check_budget raises PlanBudgetExceededError when plan budget exceeded
|
||||
Given a budget enforcement PlanExecutor with plan budget 0.0001
|
||||
And the cost_metadata has total_cost 0.001
|
||||
When I call _check_budget directly with plan_id "test-plan"
|
||||
Then a PlanBudgetExceededError should be raised
|
||||
|
||||
Scenario: _check_budget raises BudgetExceededError when daily budget exceeded
|
||||
Given a budget enforcement PlanExecutor with daily budget 0.0001
|
||||
And the daily spend is 0.001
|
||||
When I call _check_budget directly with plan_id "test-plan"
|
||||
Then a BudgetExceededError should be raised
|
||||
|
||||
Scenario: _check_budget creates CostMetadata when none provided
|
||||
Given a budget enforcement PlanExecutor with plan budget 100.0 and no cost_metadata
|
||||
When I call _check_budget directly with plan_id "test-plan"
|
||||
Then no budget exception should be raised
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# _save_plan_state_on_budget_halt - graceful halt
|
||||
# ------------------------------------------------------------------
|
||||
|
||||
Scenario: _save_plan_state_on_budget_halt persists budget details to plan
|
||||
Given a budget enforcement PlanExecutor without cost_tracker
|
||||
When I call _save_plan_state_on_budget_halt with plan_id "halt-plan" budget_type "plan" used 5.0 limit 3.0
|
||||
Then the lifecycle _commit_plan should have been called
|
||||
And the plan error_details should contain budget_halt true
|
||||
And the plan error_details should contain budget_type "plan"
|
||||
|
||||
Scenario: _save_plan_state_on_budget_halt is non-fatal on lifecycle error
|
||||
Given a budget enforcement PlanExecutor with failing lifecycle
|
||||
When I call _save_plan_state_on_budget_halt with plan_id "halt-plan" budget_type "daily" used 1.0 limit 0.5
|
||||
Then no exception should be raised from _save_plan_state_on_budget_halt
|
||||
@@ -0,0 +1,604 @@
|
||||
"""Step definitions for budget_enforcement_plan_executor.feature.
|
||||
|
||||
Tests for budget enforcement in PlanExecutor including:
|
||||
- BudgetExceededError and PlanBudgetExceededError exceptions
|
||||
- PlanExecutor halting on plan budget exceeded
|
||||
- PlanExecutor halting on daily budget exceeded
|
||||
- Graceful halt with plan state save
|
||||
- AutomationProfile budget fields
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import date
|
||||
from typing import Any
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
from behave import given, then, when
|
||||
from behave.runner import Context
|
||||
|
||||
from cleveragents.application.services.plan_executor import PlanExecutor
|
||||
from cleveragents.core.exceptions import (
|
||||
BudgetExceededError,
|
||||
PlanBudgetExceededError,
|
||||
PlanError,
|
||||
)
|
||||
from cleveragents.domain.models.core.automation_profile import AutomationProfile
|
||||
from cleveragents.domain.models.core.cost_metadata import CostMetadata
|
||||
from cleveragents.domain.models.core.plan import (
|
||||
PlanPhase,
|
||||
PlanTimestamps,
|
||||
ProcessingState,
|
||||
)
|
||||
from cleveragents.providers.cost_tracker import CostTracker
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Constants
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
_BUDGET_PLAN_ID = "01KBUDGET0PLAN000000000001"
|
||||
_BUDGET_ROOT_ID = "01KBUDGET0ROOT000000000001"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _make_budget_plan(
|
||||
*,
|
||||
phase: PlanPhase = PlanPhase.EXECUTE,
|
||||
state: ProcessingState = ProcessingState.QUEUED,
|
||||
definition_of_done: str = "Implement feature",
|
||||
decision_root_id: str = _BUDGET_ROOT_ID,
|
||||
) -> MagicMock:
|
||||
"""Build a mock plan for budget enforcement tests."""
|
||||
plan = MagicMock()
|
||||
plan.phase = phase
|
||||
plan.state = state
|
||||
plan.definition_of_done = definition_of_done
|
||||
plan.decision_root_id = decision_root_id
|
||||
plan.invariants = []
|
||||
plan.timestamps = PlanTimestamps()
|
||||
plan.changeset_id = None
|
||||
plan.sandbox_refs = []
|
||||
plan.error_details = None
|
||||
plan.read_only = False
|
||||
plan.project_links = []
|
||||
plan.subplan_statuses = []
|
||||
plan.identity = MagicMock()
|
||||
plan.identity.plan_id = _BUDGET_PLAN_ID
|
||||
return plan
|
||||
|
||||
|
||||
def _make_budget_lifecycle(plan: Any | None = None) -> MagicMock:
|
||||
"""Build a mock lifecycle service for budget tests."""
|
||||
lcs = MagicMock()
|
||||
if plan is not None:
|
||||
lcs.get_plan.return_value = plan
|
||||
lcs.start_execute = MagicMock()
|
||||
lcs.complete_execute = MagicMock()
|
||||
lcs.fail_execute = MagicMock()
|
||||
lcs._commit_plan = MagicMock()
|
||||
return lcs
|
||||
|
||||
|
||||
def _make_cost_tracker_with_plan_budget(budget: float) -> CostTracker:
|
||||
"""Create a CostTracker with a specific plan budget."""
|
||||
return CostTracker(budget_per_plan=budget)
|
||||
|
||||
|
||||
def _make_cost_tracker_with_daily_budget(budget: float) -> CostTracker:
|
||||
"""Create a CostTracker with a specific daily budget."""
|
||||
return CostTracker(budget_per_day=budget)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Exception creation steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@when(
|
||||
'I create a BudgetExceededError with plan_id "{pid}" budget_type "{btype}" used {used:f} limit {limit:f}'
|
||||
)
|
||||
def step_create_budget_exceeded_error(
|
||||
context: Context, pid: str, btype: str, used: float, limit: float
|
||||
) -> None:
|
||||
"""Create a BudgetExceededError with given attributes."""
|
||||
context.budget_exc = BudgetExceededError(
|
||||
f"Budget exceeded: {used} >= {limit}",
|
||||
plan_id=pid,
|
||||
budget_type=btype,
|
||||
used=used,
|
||||
limit=limit,
|
||||
)
|
||||
|
||||
|
||||
@then('the BudgetExceededError plan_id should be "{expected}"')
|
||||
def step_check_budget_exc_plan_id(context: Context, expected: str) -> None:
|
||||
"""Verify BudgetExceededError plan_id."""
|
||||
assert context.budget_exc.plan_id == expected, (
|
||||
f"Expected plan_id={expected!r}, got {context.budget_exc.plan_id!r}"
|
||||
)
|
||||
|
||||
|
||||
@then('the BudgetExceededError budget_type should be "{expected}"')
|
||||
def step_check_budget_exc_budget_type(context: Context, expected: str) -> None:
|
||||
"""Verify BudgetExceededError budget_type."""
|
||||
assert context.budget_exc.budget_type == expected, (
|
||||
f"Expected budget_type={expected!r}, got {context.budget_exc.budget_type!r}"
|
||||
)
|
||||
|
||||
|
||||
@then("the BudgetExceededError used should be {expected:f}")
|
||||
def step_check_budget_exc_used(context: Context, expected: float) -> None:
|
||||
"""Verify BudgetExceededError used."""
|
||||
assert context.budget_exc.used == expected, (
|
||||
f"Expected used={expected}, got {context.budget_exc.used}"
|
||||
)
|
||||
|
||||
|
||||
@then("the BudgetExceededError limit should be {expected:f}")
|
||||
def step_check_budget_exc_limit(context: Context, expected: float) -> None:
|
||||
"""Verify BudgetExceededError limit."""
|
||||
assert context.budget_exc.limit == expected, (
|
||||
f"Expected limit={expected}, got {context.budget_exc.limit}"
|
||||
)
|
||||
|
||||
|
||||
@then('the BudgetExceededError message should contain "{text}"')
|
||||
def step_check_budget_exc_message(context: Context, text: str) -> None:
|
||||
"""Verify BudgetExceededError message contains text."""
|
||||
assert text in str(context.budget_exc), (
|
||||
f"Expected '{text}' in '{context.budget_exc}'"
|
||||
)
|
||||
|
||||
|
||||
@then("the BudgetExceededError should be an instance of PlanError")
|
||||
def step_check_budget_exc_is_plan_error(context: Context) -> None:
|
||||
"""Verify BudgetExceededError is a PlanError."""
|
||||
assert isinstance(context.budget_exc, PlanError), (
|
||||
f"Expected PlanError, got {type(context.budget_exc).__name__}"
|
||||
)
|
||||
|
||||
|
||||
@when(
|
||||
'I create a PlanBudgetExceededError with plan_id "{pid}" used {used:f} limit {limit:f}'
|
||||
)
|
||||
def step_create_plan_budget_exceeded_error(
|
||||
context: Context, pid: str, used: float, limit: float
|
||||
) -> None:
|
||||
"""Create a PlanBudgetExceededError with given attributes."""
|
||||
context.plan_budget_exc = PlanBudgetExceededError(
|
||||
f"Plan budget exceeded: {used} >= {limit}",
|
||||
plan_id=pid,
|
||||
used=used,
|
||||
limit=limit,
|
||||
)
|
||||
|
||||
|
||||
@then('the PlanBudgetExceededError plan_id should be "{expected}"')
|
||||
def step_check_plan_budget_exc_plan_id(context: Context, expected: str) -> None:
|
||||
"""Verify PlanBudgetExceededError plan_id."""
|
||||
assert context.plan_budget_exc.plan_id == expected, (
|
||||
f"Expected plan_id={expected!r}, got {context.plan_budget_exc.plan_id!r}"
|
||||
)
|
||||
|
||||
|
||||
@then("the PlanBudgetExceededError used should be {expected:f}")
|
||||
def step_check_plan_budget_exc_used(context: Context, expected: float) -> None:
|
||||
"""Verify PlanBudgetExceededError used."""
|
||||
assert context.plan_budget_exc.used == expected, (
|
||||
f"Expected used={expected}, got {context.plan_budget_exc.used}"
|
||||
)
|
||||
|
||||
|
||||
@then("the PlanBudgetExceededError limit should be {expected:f}")
|
||||
def step_check_plan_budget_exc_limit(context: Context, expected: float) -> None:
|
||||
"""Verify PlanBudgetExceededError limit."""
|
||||
assert context.plan_budget_exc.limit == expected, (
|
||||
f"Expected limit={expected}, got {context.plan_budget_exc.limit}"
|
||||
)
|
||||
|
||||
|
||||
@then('the PlanBudgetExceededError message should contain "{text}"')
|
||||
def step_check_plan_budget_exc_message(context: Context, text: str) -> None:
|
||||
"""Verify PlanBudgetExceededError message contains text."""
|
||||
assert text in str(context.plan_budget_exc), (
|
||||
f"Expected '{text}' in '{context.plan_budget_exc}'"
|
||||
)
|
||||
|
||||
|
||||
@then("the PlanBudgetExceededError should be an instance of PlanError")
|
||||
def step_check_plan_budget_exc_is_plan_error(context: Context) -> None:
|
||||
"""Verify PlanBudgetExceededError is a PlanError."""
|
||||
assert isinstance(context.plan_budget_exc, PlanError), (
|
||||
f"Expected PlanError, got {type(context.plan_budget_exc).__name__}"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# PlanExecutor setup steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@given("a budget enforcement PlanExecutor without cost_tracker")
|
||||
def step_budget_executor_no_tracker(context: Context) -> None:
|
||||
"""Create a PlanExecutor without a cost tracker."""
|
||||
plan = _make_budget_plan()
|
||||
context.budget_lifecycle = _make_budget_lifecycle(plan)
|
||||
context.budget_plan = plan
|
||||
context.budget_plan_id = _BUDGET_PLAN_ID
|
||||
context.budget_executor = PlanExecutor(
|
||||
lifecycle_service=context.budget_lifecycle,
|
||||
cost_tracker=None,
|
||||
)
|
||||
|
||||
|
||||
@given("a budget enforcement plan in Execute-Queued state")
|
||||
def step_budget_plan_execute_queued(context: Context) -> None:
|
||||
"""Set up a plan in Execute-Queued state (already done in executor setup)."""
|
||||
|
||||
|
||||
@given("a budget enforcement PlanExecutor with plan budget {budget:f}")
|
||||
def step_budget_executor_with_plan_budget(context: Context, budget: float) -> None:
|
||||
"""Create a PlanExecutor with a plan budget limit."""
|
||||
plan = _make_budget_plan()
|
||||
context.budget_lifecycle = _make_budget_lifecycle(plan)
|
||||
context.budget_plan = plan
|
||||
context.budget_plan_id = _BUDGET_PLAN_ID
|
||||
context.budget_cost_tracker = _make_cost_tracker_with_plan_budget(budget)
|
||||
context.budget_cost_metadata = CostMetadata()
|
||||
context.budget_executor = PlanExecutor(
|
||||
lifecycle_service=context.budget_lifecycle,
|
||||
cost_tracker=context.budget_cost_tracker,
|
||||
cost_metadata=context.budget_cost_metadata,
|
||||
)
|
||||
|
||||
|
||||
@given("a budget enforcement PlanExecutor with daily budget {budget:f}")
|
||||
def step_budget_executor_with_daily_budget(context: Context, budget: float) -> None:
|
||||
"""Create a PlanExecutor with a daily budget limit."""
|
||||
plan = _make_budget_plan()
|
||||
context.budget_lifecycle = _make_budget_lifecycle(plan)
|
||||
context.budget_plan = plan
|
||||
context.budget_plan_id = _BUDGET_PLAN_ID
|
||||
context.budget_cost_tracker = _make_cost_tracker_with_daily_budget(budget)
|
||||
context.budget_cost_metadata = CostMetadata()
|
||||
context.budget_executor = PlanExecutor(
|
||||
lifecycle_service=context.budget_lifecycle,
|
||||
cost_tracker=context.budget_cost_tracker,
|
||||
cost_metadata=context.budget_cost_metadata,
|
||||
)
|
||||
|
||||
|
||||
@given(
|
||||
"a budget enforcement PlanExecutor with plan budget {budget:f} and no cost_metadata"
|
||||
)
|
||||
def step_budget_executor_plan_budget_no_metadata(
|
||||
context: Context, budget: float
|
||||
) -> None:
|
||||
"""Create a PlanExecutor with plan budget but no cost_metadata."""
|
||||
plan = _make_budget_plan()
|
||||
context.budget_lifecycle = _make_budget_lifecycle(plan)
|
||||
context.budget_plan = plan
|
||||
context.budget_plan_id = _BUDGET_PLAN_ID
|
||||
context.budget_cost_tracker = _make_cost_tracker_with_plan_budget(budget)
|
||||
context.budget_executor = PlanExecutor(
|
||||
lifecycle_service=context.budget_lifecycle,
|
||||
cost_tracker=context.budget_cost_tracker,
|
||||
cost_metadata=None,
|
||||
)
|
||||
|
||||
|
||||
@given("a budget enforcement PlanExecutor with failing lifecycle")
|
||||
def step_budget_executor_failing_lifecycle(context: Context) -> None:
|
||||
"""Create a PlanExecutor with a lifecycle that raises on get_plan."""
|
||||
lcs = MagicMock()
|
||||
lcs.get_plan.side_effect = RuntimeError("lifecycle failure")
|
||||
lcs._commit_plan = MagicMock()
|
||||
context.budget_lifecycle = lcs
|
||||
context.budget_plan_id = _BUDGET_PLAN_ID
|
||||
context.budget_executor = PlanExecutor(
|
||||
lifecycle_service=lcs,
|
||||
cost_tracker=None,
|
||||
)
|
||||
|
||||
|
||||
@given("the cost_metadata has total_cost {cost:f}")
|
||||
def step_set_cost_metadata_total_cost(context: Context, cost: float) -> None:
|
||||
"""Set the cost_metadata total_cost."""
|
||||
context.budget_cost_metadata.total_cost = cost
|
||||
|
||||
|
||||
@given("the daily spend is {spend:f}")
|
||||
def step_set_daily_spend(context: Context, spend: float) -> None:
|
||||
"""Set the daily spend by recording usage."""
|
||||
today_key = date.today().isoformat()
|
||||
with context.budget_cost_tracker._daily_costs_lock:
|
||||
context.budget_cost_tracker._daily_costs[today_key] = spend
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Execute steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@when("I run budget enforcement execute")
|
||||
def step_run_budget_execute(context: Context) -> None:
|
||||
"""Run the execute phase."""
|
||||
try:
|
||||
context.budget_exec_result = context.budget_executor.run_execute(
|
||||
context.budget_plan_id
|
||||
)
|
||||
context.budget_raised = None
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@when("I run budget enforcement execute expecting budget error")
|
||||
def step_run_budget_execute_expect_error(context: Context) -> None:
|
||||
"""Run the execute phase expecting a budget error."""
|
||||
try:
|
||||
context.budget_exec_result = context.budget_executor.run_execute(
|
||||
context.budget_plan_id
|
||||
)
|
||||
context.budget_raised = None
|
||||
except (BudgetExceededError, PlanBudgetExceededError) as exc:
|
||||
context.budget_raised = exc
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@then("the budget enforcement execute should succeed without budget error")
|
||||
def step_budget_execute_success(context: Context) -> None:
|
||||
"""Verify execute succeeded without budget error."""
|
||||
if context.budget_raised is not None and isinstance(
|
||||
context.budget_raised, (BudgetExceededError, PlanBudgetExceededError)
|
||||
):
|
||||
raise AssertionError(
|
||||
f"Expected no budget error, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
|
||||
|
||||
@then("a PlanBudgetExceededError should be raised")
|
||||
def step_check_plan_budget_error_raised(context: Context) -> None:
|
||||
"""Verify PlanBudgetExceededError was raised."""
|
||||
assert context.budget_raised is not None, (
|
||||
"Expected PlanBudgetExceededError but none was raised"
|
||||
)
|
||||
assert isinstance(context.budget_raised, PlanBudgetExceededError), (
|
||||
f"Expected PlanBudgetExceededError, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
|
||||
|
||||
@then("the PlanBudgetExceededError plan_id should match the plan")
|
||||
def step_check_plan_budget_error_plan_id(context: Context) -> None:
|
||||
"""Verify PlanBudgetExceededError has the correct plan_id."""
|
||||
assert isinstance(context.budget_raised, PlanBudgetExceededError)
|
||||
assert context.budget_raised.plan_id == context.budget_plan_id, (
|
||||
f"Expected plan_id={context.budget_plan_id!r}, "
|
||||
f"got {context.budget_raised.plan_id!r}"
|
||||
)
|
||||
|
||||
|
||||
@then("a BudgetExceededError should be raised")
|
||||
def step_check_budget_error_raised(context: Context) -> None:
|
||||
"""Verify BudgetExceededError was raised."""
|
||||
assert context.budget_raised is not None, (
|
||||
"Expected BudgetExceededError but none was raised"
|
||||
)
|
||||
assert isinstance(context.budget_raised, BudgetExceededError), (
|
||||
f"Expected BudgetExceededError, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
|
||||
|
||||
@then("the lifecycle _commit_plan should have been called with budget_halt details")
|
||||
def step_check_commit_plan_budget_halt(context: Context) -> None:
|
||||
"""Verify _commit_plan was called with budget_halt in error_details."""
|
||||
assert context.budget_lifecycle._commit_plan.called, (
|
||||
"Expected _commit_plan to be called"
|
||||
)
|
||||
plan = context.budget_lifecycle.get_plan.return_value
|
||||
if isinstance(plan.error_details, dict):
|
||||
assert "budget_halt" in plan.error_details, (
|
||||
f"Expected 'budget_halt' in error_details, got {plan.error_details}"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# _check_budget direct call steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@when('I call _check_budget directly with plan_id "{plan_id}"')
|
||||
def step_call_check_budget_directly(context: Context, plan_id: str) -> None:
|
||||
"""Call _check_budget directly."""
|
||||
try:
|
||||
context.budget_executor._check_budget(plan_id)
|
||||
context.budget_raised = None
|
||||
except (BudgetExceededError, PlanBudgetExceededError) as exc:
|
||||
context.budget_raised = exc
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@then("no budget exception should be raised")
|
||||
def step_no_budget_exception(context: Context) -> None:
|
||||
"""Verify no budget exception was raised."""
|
||||
if isinstance(
|
||||
context.budget_raised, (BudgetExceededError, PlanBudgetExceededError)
|
||||
):
|
||||
raise AssertionError(
|
||||
f"Expected no budget exception, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# _save_plan_state_on_budget_halt steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@when(
|
||||
'I call _save_plan_state_on_budget_halt with plan_id "{plan_id}" budget_type "{btype}" used {used:f} limit {limit:f}'
|
||||
)
|
||||
def step_call_save_plan_state(
|
||||
context: Context, plan_id: str, btype: str, used: float, limit: float
|
||||
) -> None:
|
||||
"""Call _save_plan_state_on_budget_halt directly."""
|
||||
plan = _make_budget_plan()
|
||||
context.budget_lifecycle.get_plan.return_value = plan
|
||||
context.budget_plan = plan
|
||||
try:
|
||||
context.budget_executor._save_plan_state_on_budget_halt(
|
||||
plan_id=plan_id,
|
||||
budget_type=btype,
|
||||
used=used,
|
||||
limit=limit,
|
||||
)
|
||||
context.budget_raised = None
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@then("the lifecycle _commit_plan should have been called")
|
||||
def step_check_commit_plan_called(context: Context) -> None:
|
||||
"""Verify _commit_plan was called."""
|
||||
assert context.budget_lifecycle._commit_plan.called, (
|
||||
"Expected _commit_plan to be called"
|
||||
)
|
||||
|
||||
|
||||
@then("the plan error_details should contain budget_halt true")
|
||||
def step_check_error_details_budget_halt(context: Context) -> None:
|
||||
"""Verify plan error_details has budget_halt."""
|
||||
plan = context.budget_plan
|
||||
assert isinstance(plan.error_details, dict), (
|
||||
f"Expected dict error_details, got {type(plan.error_details)}"
|
||||
)
|
||||
assert plan.error_details.get("budget_halt") == "true", (
|
||||
f"Expected budget_halt='true', got {plan.error_details.get('budget_halt')!r}"
|
||||
)
|
||||
|
||||
|
||||
@then('the plan error_details should contain budget_type "{expected}"')
|
||||
def step_check_error_details_budget_type(context: Context, expected: str) -> None:
|
||||
"""Verify plan error_details has correct budget_type."""
|
||||
plan = context.budget_plan
|
||||
assert isinstance(plan.error_details, dict)
|
||||
assert plan.error_details.get("budget_type") == expected, (
|
||||
f"Expected budget_type={expected!r}, got {plan.error_details.get('budget_type')!r}"
|
||||
)
|
||||
|
||||
|
||||
@then("no exception should be raised from _save_plan_state_on_budget_halt")
|
||||
def step_no_exception_from_save(context: Context) -> None:
|
||||
"""Verify no exception was raised from _save_plan_state_on_budget_halt."""
|
||||
assert context.budget_raised is None, (
|
||||
f"Expected no exception, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# AutomationProfile budget fields steps
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@when("I create an AutomationProfile with budget_per_plan {budget:f}")
|
||||
def step_create_profile_with_plan_budget(context: Context, budget: float) -> None:
|
||||
"""Create an AutomationProfile with budget_per_plan."""
|
||||
context.budget_profile = AutomationProfile(
|
||||
name="test-budget-profile",
|
||||
budget_per_plan=budget,
|
||||
)
|
||||
|
||||
|
||||
@when("I create an AutomationProfile with budget_per_session {budget:f}")
|
||||
def step_create_profile_with_session_budget(context: Context, budget: float) -> None:
|
||||
"""Create an AutomationProfile with budget_per_session."""
|
||||
context.budget_profile = AutomationProfile(
|
||||
name="test-budget-profile",
|
||||
budget_per_session=budget,
|
||||
)
|
||||
|
||||
|
||||
@when("I create an AutomationProfile with default budget fields")
|
||||
def step_create_profile_with_default_budget(context: Context) -> None:
|
||||
"""Create an AutomationProfile with default budget fields."""
|
||||
context.budget_profile = AutomationProfile(name="test-default-profile")
|
||||
|
||||
|
||||
@then("the AutomationProfile budget_per_plan should be {expected}")
|
||||
def step_check_profile_plan_budget(context: Context, expected: str) -> None:
|
||||
"""Verify AutomationProfile budget_per_plan."""
|
||||
if expected == "None":
|
||||
assert context.budget_profile.budget_per_plan is None, (
|
||||
f"Expected None, got {context.budget_profile.budget_per_plan}"
|
||||
)
|
||||
else:
|
||||
assert context.budget_profile.budget_per_plan == float(expected), (
|
||||
f"Expected {expected}, got {context.budget_profile.budget_per_plan}"
|
||||
)
|
||||
|
||||
|
||||
@then("the AutomationProfile budget_per_session should be {expected}")
|
||||
def step_check_profile_session_budget(context: Context, expected: str) -> None:
|
||||
"""Verify AutomationProfile budget_per_session."""
|
||||
if expected == "None":
|
||||
assert context.budget_profile.budget_per_session is None, (
|
||||
f"Expected None, got {context.budget_profile.budget_per_session}"
|
||||
)
|
||||
else:
|
||||
assert context.budget_profile.budget_per_session == float(expected), (
|
||||
f"Expected {expected}, got {context.budget_profile.budget_per_session}"
|
||||
)
|
||||
|
||||
|
||||
@when("I try to create an AutomationProfile with budget_per_plan {budget:f}")
|
||||
def step_try_create_profile_negative_plan_budget(
|
||||
context: Context, budget: float
|
||||
) -> None:
|
||||
"""Try to create an AutomationProfile with invalid budget_per_plan."""
|
||||
try:
|
||||
context.budget_profile = AutomationProfile(
|
||||
name="test-profile",
|
||||
budget_per_plan=budget,
|
||||
)
|
||||
context.budget_raised = None
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@when("I try to create an AutomationProfile with budget_per_session {budget:f}")
|
||||
def step_try_create_profile_negative_session_budget(
|
||||
context: Context, budget: float
|
||||
) -> None:
|
||||
"""Try to create an AutomationProfile with invalid budget_per_session."""
|
||||
try:
|
||||
context.budget_profile = AutomationProfile(
|
||||
name="test-profile",
|
||||
budget_per_session=budget,
|
||||
)
|
||||
context.budget_raised = None
|
||||
except Exception as exc:
|
||||
context.budget_raised = exc
|
||||
|
||||
|
||||
@then("a budget enforcement validation error should be raised")
|
||||
def step_check_budget_validation_error(context: Context) -> None:
|
||||
"""Verify a validation error was raised."""
|
||||
assert context.budget_raised is not None, (
|
||||
"Expected a validation error but none was raised"
|
||||
)
|
||||
assert (
|
||||
"validation" in type(context.budget_raised).__name__.lower()
|
||||
or "value" in str(context.budget_raised).lower()
|
||||
), (
|
||||
f"Expected validation error, got {type(context.budget_raised).__name__}: "
|
||||
f"{context.budget_raised}"
|
||||
)
|
||||
@@ -755,11 +755,6 @@ def step_try_create_plan_empty_desc(context: Context) -> None:
|
||||
context.pydantic_error = exc
|
||||
|
||||
|
||||
@then("a Pydantic validation error should be raised")
|
||||
def step_check_pydantic_error(context: Context) -> None:
|
||||
assert context.pydantic_error is not None, "Expected a Pydantic validation error"
|
||||
|
||||
|
||||
@when("I try to create an edge case plan with invalid phase value")
|
||||
def step_try_create_plan_invalid_phase(context: Context) -> None:
|
||||
"""Attempt to create a plan with a non-existent phase value."""
|
||||
@@ -777,6 +772,11 @@ def step_try_create_plan_invalid_phase(context: Context) -> None:
|
||||
context.pydantic_error = exc
|
||||
|
||||
|
||||
@then("a Pydantic validation error should be raised")
|
||||
def step_check_pydantic_error(context: Context) -> None:
|
||||
assert context.pydantic_error is not None, "Expected a Pydantic validation error"
|
||||
|
||||
|
||||
@when('I try to parse a namespaced name with special characters "{full_name}"')
|
||||
def step_try_parse_ns_special_chars(context: Context, full_name: str) -> None:
|
||||
context.pydantic_error = None
|
||||
|
||||
@@ -0,0 +1,287 @@
|
||||
"""Step definitions for TUI multi-session tabs feature."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import datetime
|
||||
|
||||
from behave import given, then, when
|
||||
|
||||
from cleveragents.tui.app import SessionView
|
||||
from cleveragents.tui.persona.registry import PersonaRegistry
|
||||
from cleveragents.tui.persona.state import PersonaState
|
||||
|
||||
|
||||
class MockCommandRouter:
|
||||
"""Mock command router for testing."""
|
||||
|
||||
def handle(self, raw: str, *, session_id: str) -> str:
|
||||
"""Mock command handler."""
|
||||
return f"Mock response for {raw} in session {session_id}"
|
||||
|
||||
|
||||
@given("a TUI app is initialized with multi-session support")
|
||||
def step_init_tui_app(context: object) -> None:
|
||||
"""Initialize a TUI app with multi-session support."""
|
||||
context.registry = PersonaRegistry() # type: ignore
|
||||
context.persona_state = PersonaState(registry=context.registry) # type: ignore
|
||||
context.router = MockCommandRouter() # type: ignore
|
||||
# Note: We can't instantiate _TextualCleverAgentsTuiApp directly without Textual
|
||||
# So we'll test the session management logic separately
|
||||
|
||||
|
||||
@when("the TUI app is created")
|
||||
def step_create_tui_app(context: object) -> None:
|
||||
"""Create a TUI app instance."""
|
||||
# Create a mock app with session management
|
||||
context.app = type("MockApp", (), {})() # type: ignore
|
||||
context.app._sessions = [ # type: ignore
|
||||
SessionView(
|
||||
session_id="default",
|
||||
transcript=[],
|
||||
name="Default",
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
]
|
||||
context.app._active_session_index = 0 # type: ignore
|
||||
|
||||
|
||||
@then("the app should have exactly {count:d} session")
|
||||
def step_check_session_count(context: object, count: int) -> None:
|
||||
"""Check the number of sessions."""
|
||||
assert len(context.app._sessions) == count # type: ignore
|
||||
|
||||
|
||||
@then("the active session should have session_id {session_id}")
|
||||
def step_check_active_session_id(context: object, session_id: str) -> None:
|
||||
"""Check the active session ID."""
|
||||
active = context.app._sessions[context.app._active_session_index] # type: ignore
|
||||
assert active.session_id == session_id
|
||||
|
||||
|
||||
@then("the active session should have name {name}")
|
||||
def step_check_active_session_name(context: object, name: str) -> None:
|
||||
"""Check the active session name."""
|
||||
active = context.app._sessions[context.app._active_session_index] # type: ignore
|
||||
assert active.name == name
|
||||
|
||||
|
||||
@when("I create a new session with name {name}")
|
||||
def step_create_session(context: object, name: str) -> None:
|
||||
"""Create a new session."""
|
||||
import uuid
|
||||
|
||||
session_id = str(uuid.uuid4())[:8]
|
||||
new_session = SessionView(
|
||||
session_id=session_id,
|
||||
transcript=[],
|
||||
name=name,
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
context.app._sessions.append(new_session) # type: ignore
|
||||
context.app._active_session_index = len(context.app._sessions) - 1 # type: ignore
|
||||
|
||||
|
||||
@then("the new session should have an independent session_id")
|
||||
def step_check_new_session_id(context: object) -> None:
|
||||
"""Check that the new session has a unique ID."""
|
||||
sessions = context.app._sessions # type: ignore
|
||||
session_ids = [s.session_id for s in sessions]
|
||||
assert len(session_ids) == len(set(session_ids)) # All unique
|
||||
|
||||
|
||||
@given("the TUI app has {count:d} session")
|
||||
def step_setup_sessions(context: object, count: int) -> None:
|
||||
"""Set up the TUI app with a specific number of sessions."""
|
||||
context.app = type("MockApp", (), {})() # type: ignore
|
||||
context.app._sessions = [] # type: ignore
|
||||
for i in range(count):
|
||||
if i == 0:
|
||||
session_id = "default"
|
||||
name = "Default"
|
||||
else:
|
||||
import uuid
|
||||
|
||||
session_id = str(uuid.uuid4())[:8]
|
||||
name = f"Session {i + 1}"
|
||||
session = SessionView(
|
||||
session_id=session_id,
|
||||
transcript=[],
|
||||
name=name,
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
context.app._sessions.append(session) # type: ignore
|
||||
context.app._active_session_index = 0 # type: ignore
|
||||
|
||||
|
||||
@given("the first session has session_id {session_id}")
|
||||
def step_check_first_session_id(context: object, session_id: str) -> None:
|
||||
"""Verify the first session has the expected ID."""
|
||||
assert context.app._sessions[0].session_id == session_id # type: ignore
|
||||
|
||||
|
||||
@given("the second session has session_id {session_id}")
|
||||
def step_check_second_session_id(context: object, session_id: str) -> None:
|
||||
"""Verify the second session has the expected ID."""
|
||||
assert context.app._sessions[1].session_id == session_id # type: ignore
|
||||
|
||||
|
||||
@when("I switch to session {session_id}")
|
||||
def step_switch_session(context: object, session_id: str) -> None:
|
||||
"""Switch to a specific session."""
|
||||
for idx, session in enumerate(context.app._sessions): # type: ignore
|
||||
if session.session_id == session_id:
|
||||
context.app._active_session_index = idx # type: ignore
|
||||
return
|
||||
raise ValueError(f"Session {session_id} not found")
|
||||
|
||||
|
||||
@when("I close the session with session_id {session_id}")
|
||||
def step_close_session(context: object, session_id: str) -> None:
|
||||
"""Close a session."""
|
||||
if len(context.app._sessions) <= 1: # type: ignore
|
||||
context.close_failed = True # type: ignore
|
||||
return
|
||||
for idx, session in enumerate(context.app._sessions): # type: ignore
|
||||
if session.session_id == session_id:
|
||||
context.app._sessions.pop(idx) # type: ignore
|
||||
if context.app._active_session_index >= len(context.app._sessions): # type: ignore
|
||||
context.app._active_session_index = len(context.app._sessions) - 1 # type: ignore
|
||||
context.close_failed = False # type: ignore
|
||||
return
|
||||
raise ValueError(f"Session {session_id} not found")
|
||||
|
||||
|
||||
@when("I try to close the session with session_id {session_id}")
|
||||
def step_try_close_session(context: object, session_id: str) -> None:
|
||||
"""Try to close a session (may fail)."""
|
||||
context.close_failed = False # type: ignore
|
||||
if len(context.app._sessions) <= 1: # type: ignore
|
||||
context.close_failed = True # type: ignore
|
||||
return
|
||||
for idx, session in enumerate(context.app._sessions): # type: ignore
|
||||
if session.session_id == session_id:
|
||||
context.app._sessions.pop(idx) # type: ignore
|
||||
if context.app._active_session_index >= len(context.app._sessions): # type: ignore
|
||||
context.app._active_session_index = len(context.app._sessions) - 1 # type: ignore
|
||||
return
|
||||
|
||||
|
||||
@then("the close operation should fail")
|
||||
def step_check_close_failed(context: object) -> None:
|
||||
"""Check that the close operation failed."""
|
||||
assert context.close_failed # type: ignore
|
||||
|
||||
|
||||
@when("I rename the session to {new_name}")
|
||||
def step_rename_session(context: object, new_name: str) -> None:
|
||||
"""Rename the active session."""
|
||||
active = context.app._sessions[context.app._active_session_index] # type: ignore
|
||||
active.name = new_name
|
||||
|
||||
|
||||
@given("the active session has name {name}")
|
||||
def step_check_active_session_has_name(context: object, name: str) -> None:
|
||||
"""Verify the active session has a specific name."""
|
||||
active = context.app._sessions[context.app._active_session_index] # type: ignore
|
||||
assert active.name == name
|
||||
|
||||
|
||||
@given("the first session is active")
|
||||
def step_first_session_active(context: object) -> None:
|
||||
"""Make the first session active."""
|
||||
context.app._active_session_index = 0 # type: ignore
|
||||
|
||||
|
||||
@when("I set persona {persona_name} for the first session")
|
||||
def step_set_persona_first(context: object, persona_name: str) -> None:
|
||||
"""Set persona for the first session."""
|
||||
session_id = context.app._sessions[0].session_id # type: ignore
|
||||
context.persona_state.active_by_session[session_id] = persona_name # type: ignore
|
||||
|
||||
|
||||
@when("I set persona {persona_name} for the second session")
|
||||
def step_set_persona_second(context: object, persona_name: str) -> None:
|
||||
"""Set persona for the second session."""
|
||||
session_id = context.app._sessions[1].session_id # type: ignore
|
||||
context.persona_state.active_by_session[session_id] = persona_name # type: ignore
|
||||
|
||||
|
||||
@when("I switch back to the first session")
|
||||
def step_switch_back_to_first(context: object) -> None:
|
||||
"""Switch back to the first session."""
|
||||
context.app._active_session_index = 0 # type: ignore
|
||||
|
||||
|
||||
@then("the first session should have active persona {persona_name}")
|
||||
def step_check_first_session_persona(context: object, persona_name: str) -> None:
|
||||
"""Check the first session's active persona."""
|
||||
session_id = context.app._sessions[0].session_id # type: ignore
|
||||
assert context.persona_state.active_by_session.get(session_id) == persona_name # type: ignore
|
||||
|
||||
|
||||
@then("the second session should have active persona {persona_name}")
|
||||
def step_check_second_session_persona(context: object, persona_name: str) -> None:
|
||||
"""Check the second session's active persona."""
|
||||
session_id = context.app._sessions[1].session_id # type: ignore
|
||||
assert context.persona_state.active_by_session.get(session_id) == persona_name # type: ignore
|
||||
|
||||
|
||||
@when("I add message {message} to the first session")
|
||||
def step_add_message_first(context: object, message: str) -> None:
|
||||
"""Add a message to the first session."""
|
||||
context.app._sessions[0].transcript.append(message) # type: ignore
|
||||
|
||||
|
||||
@when("I add message {message} to the second session")
|
||||
def step_add_message_second(context: object, message: str) -> None:
|
||||
"""Add a message to the second session."""
|
||||
context.app._sessions[1].transcript.append(message) # type: ignore
|
||||
|
||||
|
||||
@then("the first session transcript should contain {message}")
|
||||
def step_check_first_transcript_contains(context: object, message: str) -> None:
|
||||
"""Check that the first session transcript contains a message."""
|
||||
assert message in context.app._sessions[0].transcript # type: ignore
|
||||
|
||||
|
||||
@then("the first session transcript should not contain {message}")
|
||||
def step_check_first_transcript_not_contains(context: object, message: str) -> None:
|
||||
"""Check that the first session transcript does not contain a message."""
|
||||
assert message not in context.app._sessions[0].transcript # type: ignore
|
||||
|
||||
|
||||
@then("the second session transcript should contain {message}")
|
||||
def step_check_second_transcript_contains(context: object, message: str) -> None:
|
||||
"""Check that the second session transcript contains a message."""
|
||||
assert message in context.app._sessions[1].transcript # type: ignore
|
||||
|
||||
|
||||
@then("the second session transcript should not contain {message}")
|
||||
def step_check_second_transcript_not_contains(context: object, message: str) -> None:
|
||||
"""Check that the second session transcript does not contain a message."""
|
||||
assert message not in context.app._sessions[1].transcript # type: ignore
|
||||
|
||||
|
||||
@when("I create a new session")
|
||||
def step_create_new_session(context: object) -> None:
|
||||
"""Create a new session."""
|
||||
import uuid
|
||||
|
||||
session_id = str(uuid.uuid4())[:8]
|
||||
new_session = SessionView(
|
||||
session_id=session_id,
|
||||
transcript=[],
|
||||
name=f"Session {len(context.app._sessions) + 1}", # type: ignore
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
context.app._sessions.append(new_session) # type: ignore
|
||||
context.app._active_session_index = len(context.app._sessions) - 1 # type: ignore
|
||||
context.new_session = new_session # type: ignore
|
||||
|
||||
|
||||
@then("the new session should have a created_at timestamp in ISO format")
|
||||
def step_check_created_at_timestamp(context: object) -> None:
|
||||
"""Check that the new session has a valid ISO format timestamp."""
|
||||
timestamp = context.new_session.created_at # type: ignore
|
||||
# Try to parse it as ISO format
|
||||
datetime.fromisoformat(timestamp)
|
||||
@@ -0,0 +1,37 @@
|
||||
"""Behave steps for TUI persona cycling."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
from behave import given, then, when
|
||||
from behave.runner import Context
|
||||
|
||||
from cleveragents.tui.persona.registry import PersonaRegistry
|
||||
from cleveragents.tui.persona.schema import Persona
|
||||
from cleveragents.tui.persona.state import PersonaState
|
||||
|
||||
|
||||
def _registry_for_temp_dir(path: Path) -> PersonaRegistry:
|
||||
return PersonaRegistry(config_dir=path)
|
||||
|
||||
|
||||
@given('I save TUI persona "{name}" with actor "{actor}" and cycle order {cycle:d}')
|
||||
def step_save_persona_cycle(
|
||||
context: Context, name: str, actor: str, cycle: int
|
||||
) -> None:
|
||||
persona = Persona(name=name, actor=actor, cycle_order=cycle)
|
||||
context.tui_registry.save(persona)
|
||||
|
||||
|
||||
@when('I cycle persona for session "{session_id}"')
|
||||
def step_cycle_persona(context: Context, session_id: str) -> None:
|
||||
if not hasattr(context, "tui_state"):
|
||||
context.tui_state = PersonaState(registry=context.tui_registry)
|
||||
context.tui_state.cycle_persona(session_id)
|
||||
|
||||
|
||||
@then("the registry last persona should be set to {persona_name}")
|
||||
def step_registry_last_persona(context: Context, persona_name: str) -> None:
|
||||
last = context.tui_registry.get_last_persona()
|
||||
assert last == persona_name
|
||||
@@ -236,7 +236,7 @@ def step_verify_session_active_persona(context, session_id, expected):
|
||||
assert context.state.active_by_session[session_id] == expected
|
||||
|
||||
|
||||
@then('the registry last persona should be set to "{expected}"')
|
||||
@then('the mock registry last persona should be set to "{expected}"')
|
||||
def step_verify_last_persona_set(context, expected):
|
||||
context.mock_registry.set_last_persona.assert_called_with(expected)
|
||||
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
Feature: TUI Multi-Session Tabs with Independent A2A Bindings
|
||||
The TUI supports multiple session tabs, each with independent A2A bindings,
|
||||
persona selection, and conversation history.
|
||||
|
||||
Background:
|
||||
Given a TUI app is initialized with multi-session support
|
||||
|
||||
Scenario: TUI starts with a default session
|
||||
When the TUI app is created
|
||||
Then the app should have exactly 1 session
|
||||
And the active session should have session_id "default"
|
||||
And the active session should have name "Default"
|
||||
|
||||
Scenario: Create a new session
|
||||
Given the TUI app has 1 session
|
||||
When I create a new session with name "Session 2"
|
||||
Then the app should have exactly 2 sessions
|
||||
And the active session should have name "Session 2"
|
||||
And the new session should have an independent session_id
|
||||
|
||||
Scenario: Switch between sessions
|
||||
Given the TUI app has 2 sessions
|
||||
And the first session has session_id "default"
|
||||
And the second session has session_id "sess-2"
|
||||
When I switch to session "default"
|
||||
Then the active session should have session_id "default"
|
||||
When I switch to session "sess-2"
|
||||
Then the active session should have session_id "sess-2"
|
||||
|
||||
Scenario: Close a session
|
||||
Given the TUI app has 2 sessions
|
||||
When I close the session with session_id "sess-2"
|
||||
Then the app should have exactly 1 session
|
||||
And the active session should have session_id "default"
|
||||
|
||||
Scenario: Cannot close the last session
|
||||
Given the TUI app has 1 session
|
||||
When I try to close the session with session_id "default"
|
||||
Then the close operation should fail
|
||||
And the app should still have exactly 1 session
|
||||
|
||||
Scenario: Rename a session
|
||||
Given the TUI app has 1 session
|
||||
And the active session has name "Default"
|
||||
When I rename the session to "My Session"
|
||||
Then the active session should have name "My Session"
|
||||
|
||||
Scenario: Each session has independent persona tracking
|
||||
Given the TUI app has 2 sessions
|
||||
And the first session is active
|
||||
When I set persona "analyst" for the first session
|
||||
And I switch to the second session
|
||||
And I set persona "coder" for the second session
|
||||
And I switch back to the first session
|
||||
Then the first session should have active persona "analyst"
|
||||
When I switch to the second session
|
||||
Then the second session should have active persona "coder"
|
||||
|
||||
Scenario: Each session has independent transcript
|
||||
Given the TUI app has 2 sessions
|
||||
And the first session is active
|
||||
When I add message "Hello from session 1" to the first session
|
||||
And I switch to the second session
|
||||
And I add message "Hello from session 2" to the second session
|
||||
Then the first session transcript should contain "Hello from session 1"
|
||||
And the first session transcript should not contain "Hello from session 2"
|
||||
And the second session transcript should contain "Hello from session 2"
|
||||
And the second session transcript should not contain "Hello from session 1"
|
||||
|
||||
Scenario: Session creation includes timestamp
|
||||
When I create a new session
|
||||
Then the new session should have a created_at timestamp in ISO format
|
||||
@@ -0,0 +1,51 @@
|
||||
Feature: TUI Persona Cycling
|
||||
Personas can be cycled through in order using cycle_order field.
|
||||
|
||||
Scenario: cycle_persona cycles through personas with cycle_order > 0
|
||||
Given a temporary TUI persona registry
|
||||
And I save TUI persona "first" with actor "local/mock-default" and cycle order 1
|
||||
And I save TUI persona "second" with actor "local/mock-default" and cycle order 2
|
||||
And I save TUI persona "third" with actor "local/mock-default" and cycle order 3
|
||||
When I set active persona to "first" for session "s1"
|
||||
And I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "second"
|
||||
When I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "third"
|
||||
When I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "first"
|
||||
|
||||
Scenario: cycle_persona returns current persona when no cyclic personas exist
|
||||
Given a temporary TUI persona registry
|
||||
And I save TUI persona "noncyclic" with actor "local/mock-default" and cycle order 0
|
||||
When I set active persona to "noncyclic" for session "s1"
|
||||
And I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "noncyclic"
|
||||
|
||||
Scenario: cycle_persona starts from first when current is not in cycle
|
||||
Given a temporary TUI persona registry
|
||||
And I save TUI persona "cyclic1" with actor "local/mock-default" and cycle order 1
|
||||
And I save TUI persona "noncyclic" with actor "local/mock-default" and cycle order 0
|
||||
When I set active persona to "noncyclic" for session "s1"
|
||||
And I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "cyclic1"
|
||||
|
||||
Scenario: cycle_persona respects cycle_order field ordering
|
||||
Given a temporary TUI persona registry
|
||||
And I save TUI persona "alpha" with actor "local/mock-default" and cycle order 3
|
||||
And I save TUI persona "beta" with actor "local/mock-default" and cycle order 1
|
||||
And I save TUI persona "gamma" with actor "local/mock-default" and cycle order 2
|
||||
When I set active persona to "beta" for session "s1"
|
||||
And I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "gamma"
|
||||
When I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "alpha"
|
||||
When I cycle persona for session "s1"
|
||||
Then active persona for session "s1" should be "beta"
|
||||
|
||||
Scenario: cycle_persona updates last persona in registry
|
||||
Given a temporary TUI persona registry
|
||||
And I save TUI persona "p1" with actor "local/mock-default" and cycle order 1
|
||||
And I save TUI persona "p2" with actor "local/mock-default" and cycle order 2
|
||||
When I set active persona to "p1" for session "s1"
|
||||
And I cycle persona for session "s1"
|
||||
Then the registry last persona should be set to "p2"
|
||||
@@ -33,7 +33,7 @@ Feature: TUI Persona State Coverage
|
||||
When I set persona "coder" for session "sess-6"
|
||||
Then the returned persona name should be "coder"
|
||||
And session "sess-6" should have active persona "coder"
|
||||
And the registry last persona should be set to "coder"
|
||||
And the mock registry last persona should be set to "coder"
|
||||
|
||||
Scenario: set_active_persona skips preset init when session already has one
|
||||
Given the preset for session "sess-6b" is already set to "turbo"
|
||||
|
||||
@@ -24,8 +24,6 @@ nav:
|
||||
- AI Providers: api/providers.md
|
||||
- TUI: api/tui.md
|
||||
- ACMS / UKO: api/acms.md
|
||||
- Guides:
|
||||
- Configuration Reference: guides/configuration-reference.md
|
||||
- Modules:
|
||||
- Shell Safety: modules/shell-safety.md
|
||||
- UKO Provenance Tracking: modules/uko-provenance.md
|
||||
|
||||
@@ -37,8 +37,14 @@ from cleveragents.application.services.plan_execution_context import (
|
||||
RuntimeExecuteActor,
|
||||
RuntimeExecuteResult,
|
||||
)
|
||||
from cleveragents.core.exceptions import PlanError, ValidationError
|
||||
from cleveragents.core.exceptions import (
|
||||
BudgetExceededError,
|
||||
PlanBudgetExceededError,
|
||||
PlanError,
|
||||
ValidationError,
|
||||
)
|
||||
from cleveragents.domain.models.core.change import ChangeSetStore
|
||||
from cleveragents.domain.models.core.cost_metadata import CostMetadata
|
||||
from cleveragents.domain.models.core.estimation import EstimationResult
|
||||
from cleveragents.domain.models.core.plan import (
|
||||
PlanInvariant,
|
||||
@@ -54,6 +60,7 @@ from cleveragents.infrastructure.sandbox.checkpoint import (
|
||||
CheckpointManager,
|
||||
SandboxCheckpoint,
|
||||
)
|
||||
from cleveragents.providers.cost_tracker import BudgetStatus, CostTracker
|
||||
from cleveragents.tool.builtins.changeset import ChangeSet, ChangeSetCapture
|
||||
from cleveragents.tool.runner import ToolRunner
|
||||
|
||||
@@ -321,6 +328,8 @@ class PlanExecutor:
|
||||
fix_revalidate_orchestrator: FixThenRevalidateOrchestrator | None = None,
|
||||
subplan_service: SubplanService | None = None,
|
||||
subplan_execution_service: SubplanExecutionService | None = None,
|
||||
cost_tracker: CostTracker | None = None,
|
||||
cost_metadata: CostMetadata | None = None,
|
||||
) -> None:
|
||||
"""Initialize the plan executor.
|
||||
|
||||
@@ -368,6 +377,8 @@ class PlanExecutor:
|
||||
self._fix_revalidate_orchestrator = fix_revalidate_orchestrator
|
||||
self._subplan_service = subplan_service
|
||||
self._subplan_execution_service = subplan_execution_service
|
||||
self._cost_tracker = cost_tracker
|
||||
self._cost_metadata = cost_metadata
|
||||
self._strategize_actor = strategize_actor or StrategizeStubActor()
|
||||
self._execute_actor = execute_actor or ExecuteStubActor()
|
||||
self._logger = logger.bind(service="plan_executor")
|
||||
@@ -915,6 +926,97 @@ class PlanExecutor:
|
||||
)
|
||||
if not self._guardrail_service.check_wall_clock(plan_id):
|
||||
raise PlanError(f"Guardrail wall-clock limit exceeded for plan {plan_id}")
|
||||
self._check_budget(plan_id)
|
||||
|
||||
def _check_budget(self, plan_id: str) -> None:
|
||||
"""Check budget limits before each execution step.
|
||||
|
||||
Checks both per-plan and session/daily budget limits using the
|
||||
configured ``CostTracker``. If a budget is exceeded, saves the
|
||||
plan state gracefully before raising the appropriate exception.
|
||||
|
||||
- Per-plan budget exceeded: raises :class:`PlanBudgetExceededError`
|
||||
- Session/daily budget exceeded: raises :class:`BudgetExceededError`
|
||||
|
||||
Args:
|
||||
plan_id: The plan identifier.
|
||||
|
||||
Raises:
|
||||
PlanBudgetExceededError: When the per-plan budget is exceeded.
|
||||
BudgetExceededError: When the session or daily budget is exceeded.
|
||||
"""
|
||||
if self._cost_tracker is None:
|
||||
return
|
||||
|
||||
cost_metadata = self._cost_metadata
|
||||
if cost_metadata is None:
|
||||
cost_metadata = CostMetadata()
|
||||
|
||||
# Check per-plan budget
|
||||
plan_result = self._cost_tracker.check_plan_budget(cost_metadata)
|
||||
if plan_result.status == BudgetStatus.EXCEEDED:
|
||||
self._save_plan_state_on_budget_halt(
|
||||
plan_id,
|
||||
budget_type="plan",
|
||||
used=plan_result.used,
|
||||
limit=plan_result.limit or 0.0,
|
||||
)
|
||||
raise PlanBudgetExceededError(
|
||||
f"Plan budget exceeded for plan {plan_id}: "
|
||||
f"${plan_result.used:.4f} >= ${plan_result.limit or 0.0:.4f}",
|
||||
plan_id=plan_id,
|
||||
used=plan_result.used,
|
||||
limit=plan_result.limit or 0.0,
|
||||
)
|
||||
|
||||
# Check session/daily budget
|
||||
daily_result = self._cost_tracker.check_daily_budget()
|
||||
if daily_result.status == BudgetStatus.EXCEEDED:
|
||||
self._save_plan_state_on_budget_halt(
|
||||
plan_id,
|
||||
budget_type="daily",
|
||||
used=daily_result.used,
|
||||
limit=daily_result.limit or 0.0,
|
||||
)
|
||||
raise BudgetExceededError(
|
||||
f"Daily budget exceeded for plan {plan_id}: "
|
||||
f"${daily_result.used:.4f} >= ${daily_result.limit or 0.0:.4f}",
|
||||
plan_id=plan_id,
|
||||
budget_type="daily",
|
||||
used=daily_result.used,
|
||||
limit=daily_result.limit or 0.0,
|
||||
)
|
||||
|
||||
def _save_plan_state_on_budget_halt(
|
||||
self,
|
||||
plan_id: str,
|
||||
budget_type: str,
|
||||
used: float,
|
||||
limit: float,
|
||||
) -> None:
|
||||
"""Save plan state gracefully before halting due to budget exceeded."""
|
||||
try:
|
||||
plan = self._lifecycle.get_plan(plan_id)
|
||||
plan.error_details = {
|
||||
"budget_halt": "true",
|
||||
"budget_type": budget_type,
|
||||
"budget_used": str(used),
|
||||
"budget_limit": str(limit),
|
||||
}
|
||||
self._lifecycle._commit_plan(plan)
|
||||
self._logger.warning(
|
||||
"Plan halted due to budget exceeded",
|
||||
plan_id=plan_id,
|
||||
budget_type=budget_type,
|
||||
used=used,
|
||||
limit=limit,
|
||||
)
|
||||
except Exception:
|
||||
self._logger.debug(
|
||||
"Failed to save plan state on budget halt (non-fatal)",
|
||||
plan_id=plan_id,
|
||||
exc_info=True,
|
||||
)
|
||||
|
||||
def _run_execute_with_runtime(
|
||||
self,
|
||||
|
||||
@@ -18,9 +18,23 @@ def tui_callback(
|
||||
help="Run a one-shot headless startup check instead of full UI loop.",
|
||||
),
|
||||
] = False,
|
||||
web: Annotated[
|
||||
bool,
|
||||
typer.Option(
|
||||
"--web",
|
||||
help="Launch TUI in web mode accessible via browser.",
|
||||
),
|
||||
] = False,
|
||||
web_port: Annotated[
|
||||
int,
|
||||
typer.Option(
|
||||
"--web-port",
|
||||
help="Port for web server (default: 8000).",
|
||||
),
|
||||
] = 8000,
|
||||
) -> None:
|
||||
"""Launch the CleverAgents TUI."""
|
||||
# Import lazily so non-TUI commands avoid Textual startup cost.
|
||||
from cleveragents.tui.commands import run_tui
|
||||
|
||||
raise typer.Exit(run_tui(headless=headless))
|
||||
raise typer.Exit(run_tui(headless=headless, web=web, web_port=web_port))
|
||||
|
||||
@@ -293,6 +293,78 @@ class PlanError(DomainError):
|
||||
pass
|
||||
|
||||
|
||||
class BudgetExceededError(PlanError):
|
||||
"""Raised when a session or daily budget limit is exceeded during plan execution.
|
||||
|
||||
Halts plan execution gracefully after saving plan state.
|
||||
|
||||
Attributes:
|
||||
plan_id: The plan that was halted.
|
||||
budget_type: The type of budget that was exceeded ('daily' or 'session').
|
||||
used: Amount spent so far (USD).
|
||||
limit: The budget limit that was exceeded (USD).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
message: str,
|
||||
plan_id: str = "",
|
||||
budget_type: str = "session",
|
||||
used: float = 0.0,
|
||||
limit: float = 0.0,
|
||||
details: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
"""Initialize with budget context.
|
||||
|
||||
Args:
|
||||
message: Human-readable error message.
|
||||
plan_id: The plan identifier.
|
||||
budget_type: Type of budget exceeded ('daily' or 'session').
|
||||
used: Amount spent so far in USD.
|
||||
limit: The budget limit in USD.
|
||||
details: Optional additional context.
|
||||
"""
|
||||
super().__init__(message, details)
|
||||
self.plan_id = plan_id
|
||||
self.budget_type = budget_type
|
||||
self.used = used
|
||||
self.limit = limit
|
||||
|
||||
|
||||
class PlanBudgetExceededError(PlanError):
|
||||
"""Raised when a per-plan budget limit is exceeded during plan execution.
|
||||
|
||||
Halts plan execution gracefully after saving plan state.
|
||||
|
||||
Attributes:
|
||||
plan_id: The plan that was halted.
|
||||
used: Amount spent so far (USD).
|
||||
limit: The per-plan budget limit (USD).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
message: str,
|
||||
plan_id: str = "",
|
||||
used: float = 0.0,
|
||||
limit: float = 0.0,
|
||||
details: dict[str, Any] | None = None,
|
||||
) -> None:
|
||||
"""Initialize with plan budget context.
|
||||
|
||||
Args:
|
||||
message: Human-readable error message.
|
||||
plan_id: The plan identifier.
|
||||
used: Amount spent so far in USD.
|
||||
limit: The per-plan budget limit in USD.
|
||||
details: Optional additional context.
|
||||
"""
|
||||
super().__init__(message, details)
|
||||
self.plan_id = plan_id
|
||||
self.used = used
|
||||
self.limit = limit
|
||||
|
||||
|
||||
class DecisionPhaseViolationError(BusinessRuleViolation):
|
||||
"""Raised when a decision type is invalid for the plan's current phase.
|
||||
|
||||
@@ -328,6 +400,7 @@ class ExecutionError(CleverAgentsError):
|
||||
__all__ = [
|
||||
"AuthenticationError",
|
||||
"AuthorizationError",
|
||||
"BudgetExceededError",
|
||||
"BusinessRuleViolation",
|
||||
"CleverAgentsError",
|
||||
"ConfigurationError",
|
||||
@@ -345,6 +418,7 @@ __all__ = [
|
||||
"ModelNotAvailableError",
|
||||
"NetworkError",
|
||||
"NotFoundError",
|
||||
"PlanBudgetExceededError",
|
||||
"PlanError",
|
||||
"ProviderError",
|
||||
"RateLimitError",
|
||||
|
||||
@@ -221,6 +221,24 @@ class AutomationProfile(BaseModel):
|
||||
description="Optional enforcement hooks for runtime constraints",
|
||||
)
|
||||
|
||||
# -- Budget limits (YAML-configurable) ---------------------------------
|
||||
|
||||
budget_per_plan: float | None = Field(
|
||||
default=None,
|
||||
ge=0.0,
|
||||
description=(
|
||||
"Maximum USD spend per plan execution. None means unlimited. "
|
||||
"When set, PlanExecutor halts with PlanBudgetExceededError if exceeded."
|
||||
),
|
||||
)
|
||||
budget_per_session: float | None = Field(
|
||||
default=None,
|
||||
ge=0.0,
|
||||
description=(
|
||||
"Maximum USD spend per session. None means unlimited. "
|
||||
"When set, PlanExecutor halts with BudgetExceededError if exceeded."
|
||||
),
|
||||
)
|
||||
# -- Name validation ---------------------------------------------------
|
||||
|
||||
@field_validator("name")
|
||||
|
||||
+105
-9
@@ -1,10 +1,12 @@
|
||||
"""Textual TUI application shell."""
|
||||
"""Textual TUI application shell - Multi-session support."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
import uuid
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import datetime
|
||||
from typing import TYPE_CHECKING, Any, ClassVar, Protocol
|
||||
|
||||
from cleveragents.tui.first_run import create_default_persona_for_actor, is_first_run
|
||||
@@ -54,10 +56,12 @@ def textual_available() -> bool:
|
||||
|
||||
@dataclass(slots=True)
|
||||
class SessionView:
|
||||
"""Minimal per-session TUI view model."""
|
||||
"""Per-session TUI view model with independent A2A binding."""
|
||||
|
||||
session_id: str
|
||||
transcript: list[str]
|
||||
transcript: list[str] = field(default_factory=list)
|
||||
name: str = "" # User-friendly session name
|
||||
created_at: str = "" # ISO format timestamp
|
||||
|
||||
|
||||
class _CommandRouter(Protocol):
|
||||
@@ -93,6 +97,8 @@ if _TEXTUAL_AVAILABLE:
|
||||
("ctrl+q", "quit", "Quit"),
|
||||
("f1", "help", "Help"),
|
||||
("ctrl+t", "cycle_preset", "Cycle Preset"),
|
||||
("ctrl+n", "new_session", "New Session"),
|
||||
("ctrl+w", "close_session", "Close Session"),
|
||||
]
|
||||
|
||||
def __init__(
|
||||
@@ -104,7 +110,80 @@ if _TEXTUAL_AVAILABLE:
|
||||
super().__init__()
|
||||
self._command_router = command_router
|
||||
self._persona_state = persona_state
|
||||
self._session = SessionView(session_id="default", transcript=[])
|
||||
# Initialize with default session
|
||||
default_session = SessionView(
|
||||
session_id="default",
|
||||
transcript=[],
|
||||
name="Default",
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
self._sessions: list[SessionView] = [default_session]
|
||||
self._active_session_index: int = 0
|
||||
|
||||
def _get_active_session(self) -> SessionView:
|
||||
"""Get the currently active session."""
|
||||
if 0 <= self._active_session_index < len(self._sessions):
|
||||
return self._sessions[self._active_session_index]
|
||||
# Fallback to first session if index is invalid
|
||||
if self._sessions:
|
||||
self._active_session_index = 0
|
||||
return self._sessions[0]
|
||||
# Create default session if none exist
|
||||
default_session = SessionView(
|
||||
session_id="default",
|
||||
transcript=[],
|
||||
name="Default",
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
self._sessions = [default_session]
|
||||
self._active_session_index = 0
|
||||
return default_session
|
||||
|
||||
def _create_session(self, name: str = "") -> SessionView:
|
||||
"""Create a new session with independent A2A binding."""
|
||||
session_id = str(uuid.uuid4())[:8]
|
||||
session_name = name or f"Session {len(self._sessions) + 1}"
|
||||
new_session = SessionView(
|
||||
session_id=session_id,
|
||||
transcript=[],
|
||||
name=session_name,
|
||||
created_at=datetime.utcnow().isoformat(),
|
||||
)
|
||||
self._sessions.append(new_session)
|
||||
return new_session
|
||||
|
||||
def _switch_session(self, session_id: str) -> SessionView | None:
|
||||
"""Switch to a session by ID."""
|
||||
for idx, session in enumerate(self._sessions):
|
||||
if session.session_id == session_id:
|
||||
self._active_session_index = idx
|
||||
return session
|
||||
return None
|
||||
|
||||
def _close_session(self, session_id: str) -> bool:
|
||||
"""Close a session by ID. Returns False if it's the last session."""
|
||||
if len(self._sessions) <= 1:
|
||||
return False
|
||||
for idx, session in enumerate(self._sessions):
|
||||
if session.session_id == session_id:
|
||||
self._sessions.pop(idx)
|
||||
# Adjust active index if needed
|
||||
if self._active_session_index >= len(self._sessions):
|
||||
self._active_session_index = len(self._sessions) - 1
|
||||
return True
|
||||
return False
|
||||
|
||||
def _rename_session(self, session_id: str, new_name: str) -> bool:
|
||||
"""Rename a session by ID."""
|
||||
for session in self._sessions:
|
||||
if session.session_id == session_id:
|
||||
session.name = new_name
|
||||
return True
|
||||
return False
|
||||
|
||||
def _list_sessions(self) -> list[SessionView]:
|
||||
"""Get all sessions."""
|
||||
return self._sessions
|
||||
|
||||
def compose(self) -> Any:
|
||||
yield _Header(show_clock=True)
|
||||
@@ -149,12 +228,28 @@ if _TEXTUAL_AVAILABLE:
|
||||
help_panel.toggle(context_name)
|
||||
|
||||
def action_cycle_preset(self) -> None:
|
||||
self._persona_state.cycle_preset(self._session.session_id)
|
||||
session = self._get_active_session()
|
||||
self._persona_state.cycle_preset(session.session_id)
|
||||
self._refresh_persona_bar()
|
||||
|
||||
def action_new_session(self) -> None:
|
||||
"""Create a new session (Ctrl+N)."""
|
||||
self._create_session()
|
||||
self._active_session_index = len(self._sessions) - 1
|
||||
self._refresh_persona_bar()
|
||||
|
||||
def action_close_session(self) -> None:
|
||||
"""Close the current session (Ctrl+W)."""
|
||||
session = self._get_active_session()
|
||||
if not self._close_session(session.session_id):
|
||||
# Cannot close the last session
|
||||
return
|
||||
self._refresh_persona_bar()
|
||||
|
||||
def _refresh_persona_bar(self) -> None:
|
||||
persona = self._persona_state.active_persona(self._session.session_id)
|
||||
preset = self._persona_state.current_preset(self._session.session_id)
|
||||
session = self._get_active_session()
|
||||
persona = self._persona_state.active_persona(session.session_id)
|
||||
preset = self._persona_state.current_preset(session.session_id)
|
||||
scope_count = len(persona.scoped_projects) + len(persona.scoped_plans)
|
||||
scope_text = f"{scope_count} scope refs"
|
||||
bar = self.query_one("#persona-bar", PersonaBar)
|
||||
@@ -173,9 +268,10 @@ if _TEXTUAL_AVAILABLE:
|
||||
if not text:
|
||||
return
|
||||
|
||||
session = self._get_active_session()
|
||||
mode_router = InputModeRouter(
|
||||
command_handler=lambda raw: self._command_router.handle(
|
||||
raw, session_id=self._session.session_id
|
||||
raw, session_id=session.session_id
|
||||
),
|
||||
shell_confirm=lambda _cmd: (
|
||||
os.environ.get("CLEVERAGENTS_ALLOW_DANGEROUS_SHELL", "").strip()
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import json
|
||||
from collections import defaultdict
|
||||
from collections.abc import Callable
|
||||
@@ -223,8 +224,144 @@ class TuiCommandRouter:
|
||||
return f"Import failed: {exc}"
|
||||
|
||||
|
||||
def run_tui(*, headless: bool = False) -> int:
|
||||
"""Run the Textual TUI app or a headless startup check."""
|
||||
def _get_tui_web_html(port: int) -> str:
|
||||
"""Generate HTML for TUI web mode.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
port:
|
||||
Port number for the web server.
|
||||
|
||||
Returns
|
||||
-------
|
||||
HTML content as string.
|
||||
"""
|
||||
return """<!DOCTYPE html>
|
||||
<html>
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>CleverAgents TUI</title>
|
||||
<style>
|
||||
body {
|
||||
margin: 0;
|
||||
padding: 0;
|
||||
font-family: monospace;
|
||||
background-color: #1e1e1e;
|
||||
color: #f8f8f2;
|
||||
}
|
||||
#tui-container {
|
||||
width: 100%;
|
||||
height: 100vh;
|
||||
overflow: hidden;
|
||||
}
|
||||
.loading {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
height: 100vh;
|
||||
font-size: 18px;
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div id="tui-container">
|
||||
<div class="loading">Loading CleverAgents TUI...</div>
|
||||
</div>
|
||||
<script>
|
||||
// WebSocket connection to TUI app
|
||||
// Note: This is a placeholder. Full implementation would require
|
||||
// a WebSocket server in the TUI app to handle real-time rendering.
|
||||
console.log("TUI Web mode loaded. WebSocket support coming soon.");
|
||||
</script>
|
||||
</body>
|
||||
</html>"""
|
||||
|
||||
|
||||
def _run_tui_web(app: CleverAgentsTuiApp, *, port: int = 8000) -> int: # type: ignore[valid-type]
|
||||
"""Run the TUI app in web mode via HTTP server.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
app:
|
||||
The Textual TUI app instance.
|
||||
port:
|
||||
Port for the web server.
|
||||
|
||||
Returns
|
||||
-------
|
||||
Exit code (0 for success, non-zero for failure).
|
||||
"""
|
||||
try:
|
||||
# Import web server dependencies
|
||||
import threading
|
||||
import webbrowser
|
||||
from http.server import BaseHTTPRequestHandler, HTTPServer
|
||||
|
||||
class TuiWebHandler(BaseHTTPRequestHandler):
|
||||
"""HTTP request handler for TUI web mode."""
|
||||
|
||||
def do_GET(self) -> None:
|
||||
"""Handle GET requests."""
|
||||
if self.path == "/" or self.path == "/index.html":
|
||||
self.send_response(200)
|
||||
self.send_header("Content-type", "text/html")
|
||||
self.end_headers()
|
||||
html = _get_tui_web_html(port)
|
||||
self.wfile.write(html.encode("utf-8"))
|
||||
else:
|
||||
self.send_response(404)
|
||||
self.end_headers()
|
||||
|
||||
def log_message(self, format: str, *args: Any) -> None:
|
||||
"""Suppress default logging."""
|
||||
pass
|
||||
|
||||
# Create and start HTTP server
|
||||
server = HTTPServer(("127.0.0.1", port), TuiWebHandler)
|
||||
server_thread = threading.Thread(target=server.serve_forever, daemon=True)
|
||||
server_thread.start()
|
||||
|
||||
# Print startup message
|
||||
url = f"http://127.0.0.1:{port}"
|
||||
print(f"TUI Web mode started at {url}")
|
||||
print("Press Ctrl+C to stop")
|
||||
|
||||
# Try to open browser
|
||||
with contextlib.suppress(Exception):
|
||||
webbrowser.open(url)
|
||||
|
||||
# Run the app in headless mode (web driver will handle rendering)
|
||||
try:
|
||||
app.run(headless=True)
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
finally:
|
||||
server.shutdown()
|
||||
|
||||
return 0
|
||||
|
||||
except Exception as exc:
|
||||
print(f"Error starting web mode: {exc}")
|
||||
return 1
|
||||
|
||||
|
||||
def run_tui(*, headless: bool = False, web: bool = False, web_port: int = 8000) -> int:
|
||||
"""Run the Textual TUI app, headless check, or web mode.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
headless:
|
||||
Run a one-shot headless startup check instead of full UI loop.
|
||||
web:
|
||||
Launch TUI in web mode accessible via browser.
|
||||
web_port:
|
||||
Port for web server (default: 8000).
|
||||
|
||||
Returns
|
||||
-------
|
||||
Exit code (0 for success, non-zero for failure).
|
||||
"""
|
||||
container = get_container()
|
||||
registry = container.persona_registry()
|
||||
state = container.persona_state(registry=registry)
|
||||
@@ -243,5 +380,9 @@ def run_tui(*, headless: bool = False) -> int:
|
||||
return 0
|
||||
|
||||
app = CleverAgentsTuiApp(command_router=router, persona_state=state)
|
||||
|
||||
if web:
|
||||
return _run_tui_web(app, port=web_port)
|
||||
|
||||
app.run()
|
||||
return 0
|
||||
|
||||
@@ -79,23 +79,25 @@ class PersonaRegistry:
|
||||
return result
|
||||
|
||||
def resolve_export_path(self, output_path: Path) -> Path:
|
||||
"""Resolve export path, accepting both absolute and relative paths."""
|
||||
resolved = output_path.resolve()
|
||||
# Allow absolute paths directly
|
||||
if output_path.is_absolute():
|
||||
raise ValueError(
|
||||
"Export path must be relative to current working directory"
|
||||
)
|
||||
return resolved
|
||||
# For relative paths, ensure they stay within working directory
|
||||
base = Path.cwd().resolve()
|
||||
resolved = (base / output_path).resolve()
|
||||
if not resolved.is_relative_to(base):
|
||||
raise ValueError("Export path must stay within working directory")
|
||||
return resolved
|
||||
|
||||
def resolve_import_path(self, input_path: Path) -> Path:
|
||||
"""Resolve import path, accepting both absolute and relative paths."""
|
||||
resolved = input_path.resolve()
|
||||
# Allow absolute paths directly
|
||||
if input_path.is_absolute():
|
||||
raise ValueError(
|
||||
"Import path must be relative to current working directory"
|
||||
)
|
||||
return resolved
|
||||
# For relative paths, ensure they stay within working directory
|
||||
base = Path.cwd().resolve()
|
||||
resolved = (base / input_path).resolve()
|
||||
if not resolved.is_relative_to(base):
|
||||
raise ValueError("Import path must stay within working directory")
|
||||
return resolved
|
||||
|
||||
@@ -63,6 +63,32 @@ class PersonaState:
|
||||
self.preset_by_session[session_id] = next_name
|
||||
return next_name
|
||||
|
||||
def cycle_persona(self, session_id: str) -> Persona:
|
||||
"""Cycle to the next persona in cycle_order sequence.
|
||||
|
||||
Only personas with cycle_order > 0 are included in the cycle.
|
||||
If no cyclic personas exist, returns the current active persona.
|
||||
"""
|
||||
personas = self.registry.list_personas()
|
||||
cyclic = sorted(
|
||||
[p for p in personas if p.cycle_order > 0], key=lambda p: p.cycle_order
|
||||
)
|
||||
|
||||
if not cyclic:
|
||||
return self.active_persona(session_id)
|
||||
|
||||
current = self.active_name(session_id)
|
||||
current_names = [p.name for p in cyclic]
|
||||
|
||||
if current not in current_names:
|
||||
# Current persona is not in cycle, start from first
|
||||
next_persona = cyclic[0]
|
||||
else:
|
||||
idx = current_names.index(current)
|
||||
next_persona = cyclic[(idx + 1) % len(cyclic)]
|
||||
|
||||
return self.set_active_persona(session_id, next_persona.name)
|
||||
|
||||
def effective_arguments(self, session_id: str) -> dict[str, object]:
|
||||
persona = self.active_persona(session_id)
|
||||
preset = self.current_preset(session_id)
|
||||
|
||||
Submodule
+1
Submodule work/repo added at 435e409df9
Reference in New Issue
Block a user