Luis Mendes 1ffb0fddbe
CI / benchmark-publish (pull_request) Has been skipped
CI / build (pull_request) Successful in 21s
CI / lint (pull_request) Successful in 3m18s
CI / typecheck (pull_request) Successful in 4m11s
CI / security (pull_request) Successful in 4m13s
CI / integration_tests (pull_request) Successful in 5m56s
CI / coverage (pull_request) Successful in 10m2s
CI / quality (pull_request) Successful in 3m40s
CI / unit_tests (pull_request) Successful in 6m6s
CI / docker (pull_request) Successful in 1m2s
CI / status-check (pull_request) Successful in 1s
CI / e2e_tests (pull_request) Successful in 8m19s
CI / build (push) Successful in 14s
CI / lint (push) Successful in 3m35s
CI / typecheck (push) Successful in 3m53s
CI / benchmark-regression (push) Has been skipped
CI / integration_tests (push) Successful in 5m54s
CI / unit_tests (push) Successful in 7m17s
CI / e2e_tests (push) Successful in 9m39s
CI / coverage (push) Successful in 11m22s
CI / benchmark-regression (pull_request) Successful in 58m8s
CI / benchmark-publish (push) Successful in 32m8s
CI / quality (push) Successful in 3m29s
CI / security (push) Successful in 3m40s
CI / docker (push) Successful in 51s
CI / status-check (push) Successful in 1s
feat(plan): decision correction recomputes only affected subtree
Enhanced the CorrectionService to properly isolate correction scope:

1. analyze_impact() now populates excluded_decisions by computing
   the set difference between all plan decisions and the affected
   subtree, ensuring root and sibling decisions are tracked.

2. Added rollback_tier computation that counts parent hops from
   the target decision to the root, enabling depth-aware rollback
   strategies (tier 0 = root targeted, tier N = N levels deep).

3. Enhanced dry-run report with excluded decisions list, rollback
   tier, and tier-0 warning when entire tree is affected.

4. Added subtree isolation validation confirming root exclusion
   and sibling non-contamination for non-root corrections.

5. Fixed status state-machine regression in execute_revert() where
   analyze_impact() overwrote status back to ANALYZING from EXECUTING.
   Guard now only transitions to ANALYZING when status is PENDING.
   execute_revert() now transitions through ANALYZING before EXECUTING
   for correct lifecycle ordering.

6. Fixed validate_subtree_isolation() to use structural-only BFS for
   the sibling invariant check, so that influence-DAG-caused sibling
   reachability is not misreported as an isolation violation (per spec
   § Affected Subtree Computation).

7. Fixed false-positive cycle-detection warnings from convergent
   (diamond) topologies by replacing per-node seen_this_round with a
   global enqueued set in BFS, preventing duplicate queue insertions
   from different parents.

8. Added dry_run enforcement guard in _assert_executable() to prevent
   execution of dry-run-only corrections per spec (§ plan correct
   --dry-run: "Show impact without executing").

9. Fixed generate_dry_run_report() to preserve request status so that
   generating a preview does not advance the correction lifecycle
   (dry-run is non-mutating per spec).  Status restoration now uses
   try/finally to guarantee recovery even when analyze_impact() raises
   after transitioning the status.

10. Improved cycle-detection log message accuracy to cover both
    structural tree and influence DAG sources.

11. Extracted cost/time estimation constants (_COST_PER_DECISION,
    _RECOMPUTE_SECONDS_PER_DECISION) from magic numbers.

12. Added terminal-state guard in analyze_impact() to reject
    re-analysis after execution (APPLIED/FAILED/CANCELLED/REJECTED),
    preventing audit-data corruption.  Promoted terminal-status set
    to a module-level _TERMINAL_STATUSES frozenset constant.

13. Added mode validation in execute_revert()/execute_append() to
    prevent mode-mismatched execution (e.g. calling execute_revert
    on an APPEND correction).

14. Fixed _collect_all_decisions() to always include the target
    decision in the universe, preventing broken partition invariant
    for isolated single-node plans.

15. Fixed tier-0 dry-run warning to only trigger when the target
    is genuinely in the structural tree (avoiding false warnings for
    nodes not present in the tree).

16. Removed duplicate as_cli_dict regression scenarios in
    resource_type_deferred_physical.feature.

ISSUES CLOSED: #845
2026-03-23 18:41:53 +00:00
2026-03-22 23:45:54 -04:00
2024-01-25 23:10:04 -05:00
2024-01-25 23:10:04 -05:00
2024-01-25 23:10:04 -05:00
2024-01-25 23:10:04 -05:00
2026-02-23 22:20:03 +00:00
2024-01-25 23:10:04 -05:00
2024-01-25 23:10:04 -05:00
2024-01-25 23:14:00 -05:00
2024-01-25 23:14:00 -05:00

CleverAgents Core

CleverAgents is a Python-first automation platform. It provides a unified agents CLI, embedded runtime, and service orchestration tools while embracing modern Python tooling.

Highlights

  • Unified CLI entry points: cleveragents and agents
  • Fast Typer/Click-based interface with parity for help/version behavior
  • Behavior-driven coverage via Behave and Robot Framework
  • Nox automation for linting, typing, testing, docs, builds, and benchmarks
  • MkDocs-powered documentation with CleverAgents branding

Quick Start

# clone the CleverAgents core repository
git clone https://git.cleverthis.com/cleveragents/core.git
cd core

# install dependencies
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev,tests,docs]"

# set up pre-commit hooks and verify tooling
bash scripts/setup-dev.sh

# verify the CLI
agents --help
agents --version

Developing

Pre-commit hooks run automatically on every git commit (formatting, linting, type checking, security scanning). To run checks manually:

# core validation
nox -s format            # ruff auto-formatting
nox -s lint              # ruff linting
nox -s typecheck         # pyright type checking
nox -s unit_tests        # behave unit tests
nox -s integration_tests # robot integration tests

# quality & security
nox -s security_scan     # bandit security scanning
nox -s dead_code         # vulture dead code detection
nox -s complexity        # radon complexity analysis
nox -s pre_commit        # run all pre-commit hooks
nox -s adr_compliance    # verify ADR compliance

For the full quality automation guide, see docs/development/quality-automation.md.

Documentation

nox -s docs
nox -s serve_docs

Tests

Behave feature scenarios live under features/ and Robot suites under robot/. Use the Nox sessions above to execute them in parity with the implementation plan.

Observability

LangSmith tracing is optional and off by default. Enable it by exporting CLEVERAGENTS_LANGSMITH_ENABLED=true along with a project name and API key (CLEVERAGENTS_LANGSMITH_PROJECT, CLEVERAGENTS_LANGSMITH_API_KEY). The settings module automatically mirrors these values to LANGCHAIN_TRACING_V2, LANGCHAIN_PROJECT, and LANGCHAIN_API_KEY, so LangChain/LangGraph agents emit traces without extra wiring. Additional knobs such as CLEVERAGENTS_LANGSMITH_ENDPOINT, CLEVERAGENTS_LANGSMITH_USER_ID, and CLEVERAGENTS_LANGSMITH_TAGS are documented in docs/observability.md.

LLM provider configuration

CleverAgents ships with a LangChain/LangGraph powered provider registry that discovers whichever API keys you export and automatically selects the best available provider. The CLI now uses actors: select an actor with --actor <name> (or set a default via agents actor set-default). Actors embed provider/model choices; built-in actors are seeded from CLEVERAGENTS_DEFAULT_PROVIDER / CLEVERAGENTS_DEFAULT_MODEL, then fall back to the built-in order (openai → anthropic → google → azure → openrouter → groq → together → cohere → gemini).

Required environment variables

Provider Primary variables
OpenAI OPENAI_API_KEY
Anthropic ANTHROPIC_API_KEY
Google GOOGLE_API_KEY or GOOGLE_GENAI_API_KEY
Azure OpenAI AZURE_OPENAI_API_KEY plus AZURE_OPENAI_ENDPOINT/AZURE_OPENAI_DEPLOYMENT
OpenRouter OPENROUTER_API_KEY (+ optional CLEVERAGENTS_OPENROUTER_ORGANIZATION for sanitized headers)
Gemini GEMINI_API_KEY or GOOGLE_GEMINI_API_KEY
Cohere COHERE_API_KEY
Groq GROQ_API_KEY
Together TOGETHER_API_KEY

Set CLEVERAGENTS_DEFAULT_PROVIDER to pin the global provider (for example export CLEVERAGENTS_DEFAULT_PROVIDER=openai) and CLEVERAGENTS_DEFAULT_MODEL to lock in a model ID. When unset, the registry picks the first configured provider and uses its published default model such as gpt-4o for OpenAI or claude-sonnet-4-20250514 for Anthropic.

Diagnostics and testing shortcuts

  • agents diagnostics prints whether the registry can see your credentials and which actor/provider is selected.
  • agents tell and agents build require --actor <name> unless a default actor is set; use agents actor set-default <name> to configure one.
  • Built-in actors (<provider>/<model>) are immutable, custom actors must be named local/<id>, and the default actor cannot be removed. Use --unsafe when adding/updating configs marked unsafe; runtime only warns when invoking unsafe actors.
  • CLEVERAGENTS_TESTING_USE_MOCK_AI=true forces the in-repo mock provider so Behave/Robot suites never hit external APIs.
  • The full capability matrix (streaming, tool calls, JSON mode, etc.) is documented in docs/reference/providers.md.
S
Description
test
Readme 84 MiB
Languages
Python 75.8%
Gherkin 18.3%
RobotFramework 4.8%
TypeScript 0.9%
Shell 0.2%