docs: document LLM provider fallback behavior (Anthropic Sonnet fallback) [AUTO-DOCS-7]
CI / push-validation (pull_request) Successful in 22s
CI / helm (pull_request) Successful in 32s
CI / build (pull_request) Successful in 3m47s
CI / lint (pull_request) Successful in 3m54s
CI / quality (pull_request) Successful in 4m20s
CI / typecheck (pull_request) Successful in 4m34s
CI / security (pull_request) Successful in 4m36s
CI / integration_tests (pull_request) Successful in 6m53s
CI / e2e_tests (pull_request) Successful in 6m58s
CI / unit_tests (pull_request) Successful in 7m20s
CI / docker (pull_request) Successful in 1m29s
CI / coverage (pull_request) Successful in 13m37s
CI / benchmark-regression (push) Waiting to run
CI / benchmark-publish (push) Waiting to run
CI / status-check (pull_request) Successful in 2s
CI / push-validation (push) Successful in 25s
CI / helm (push) Successful in 41s
CI / build (push) Successful in 3m50s
CI / lint (push) Successful in 3m57s
CI / quality (push) Successful in 4m23s
CI / typecheck (push) Successful in 4m30s
CI / security (push) Successful in 4m53s
CI / e2e_tests (push) Successful in 6m47s
CI / integration_tests (push) Successful in 7m49s
CI / unit_tests (push) Successful in 9m3s
CI / docker (push) Successful in 1m42s
CI / coverage (push) Successful in 15m19s
CI / status-check (push) Successful in 3s
CI / benchmark-publish (pull_request) Has been skipped
CI / benchmark-regression (pull_request) Failing after 13m19s

This commit was merged in pull request #10214.
This commit is contained in:
2026-04-17 06:09:57 +00:00
parent f8233000cc
commit 4e86ffe596
3 changed files with 226 additions and 132 deletions
+193
View File
@@ -0,0 +1,193 @@
# LLM Provider Fallback Behavior
## Overview
CleverAgents implements automatic LLM provider fallback to ensure resilience when the
primary provider (typically OpenAI) exhausts its API quota. When a quota error is
detected, the system transparently switches to **Anthropic Claude Sonnet
(`claude-sonnet-4-20250514`)** as the fallback provider, allowing plan strategize
operations to continue without interruption.
This behavior was introduced in commit `f5712787` (feat: add fallback to Anthropic
Sonnet when OpenAI quota is exhausted) and is implemented in
`StrategyActor._execute_with_llm()` located in
`src/cleveragents/application/services/strategy_actor.py`.
---
## Trigger Conditions
The fallback activates when the primary LLM call raises an exception whose string
representation matches any of the following patterns (case-insensitive):
| Pattern | Typical Source |
|---------|---------------|
| `insufficient_quota` | OpenAI quota exhaustion |
| `429` | HTTP 429 Too Many Requests |
| `quota` | Generic quota error from any provider |
| `rate_limit` | Rate-limit errors from any provider |
Detection is performed by the module-level helper function `_is_quota_error(exc)` in
`src/cleveragents/application/services/strategy_actor.py`.
---
## Fallback Provider
| Constant | Value | Description |
|----------|-------|-------------|
| `_FALLBACK_PROVIDER` | `"anthropic"` | Provider identifier |
| `_FALLBACK_MODEL` | `"claude-sonnet-4-20250514"` | Model identifier |
| `_QUOTA_RECOVERY_INTERVAL` | `300` (seconds) | Time before re-attempting primary provider |
The fallback LLM instance is created lazily on first use and then **cached** as
`StrategyActor._fallback_llm` to avoid per-call re-initialization overhead.
---
## Fallback State Machine
`StrategyActor` tracks fallback state using three instance variables:
| Variable | Type | Description |
|----------|------|-------------|
| `_fallback_llm` | `Any \| None` | Cached fallback LLM instance |
| `_last_quota_error_time` | `float \| None` | Unix timestamp of the last quota error |
| `_using_fallback` | `bool` | Whether the actor is currently in fallback mode |
### State Transitions
```
Normal mode (_using_fallback=False)
|
| Primary provider raises quota error
v
Fallback mode (_using_fallback=True, _last_quota_error_time=<now>)
|
| _QUOTA_RECOVERY_INTERVAL (300 s) elapses
v
Recovery attempt (_using_fallback=False, primary re-tried)
|
+-- Primary succeeds --> Normal mode
+-- Primary raises quota error again --> Fallback mode
```
### Recovery Logic
Once in fallback mode, `_execute_with_llm()` skips the primary provider call for
subsequent invocations until `_QUOTA_RECOVERY_INTERVAL` seconds have elapsed since
the last quota error. This prevents hammering the primary provider with repeated
quota errors. After the interval, the primary provider is re-attempted automatically.
---
## Behavior from the User Perspective
From the user's perspective, the fallback is **transparent**:
- Plan strategize operations continue to produce valid strategy trees.
- Response quality and format are identical — both OpenAI and Anthropic Claude Sonnet
support the structured JSON output required by the strategy prompt.
- No user action is required to trigger or disable the fallback.
- The fallback is **temporary**: the system automatically attempts to recover to the
primary provider after 5 minutes.
### When Both Providers Are Unavailable
If the fallback provider (Anthropic Claude Sonnet) also fails, the original quota
exception from the primary provider is re-raised (chained from the fallback
exception). The `StrategyActor.execute()` method catches this and falls back to the
deterministic stub strategy (`_execute_stub()`), ensuring the plan lifecycle is never
blocked by LLM unavailability.
In E2E test contexts, when both providers are unavailable, tests fail with a clear
message indicating that the test outcome cannot be verified — rather than silently
skipping and creating false confidence in test coverage.
---
## Logging
The fallback mechanism emits structured log events at the following levels:
| Level | Event | Key Fields |
|-------|-------|-----------|
| `WARNING` | Quota error detected on primary provider | `plan_id`, `original_error` |
| `WARNING` | Creating fallback LLM instance | `plan_id` |
| `WARNING` | Reusing cached fallback LLM instance | `plan_id` |
| `WARNING` | Fallback recovery successful | `plan_id`, `fallback_provider`, `fallback_model` |
| `ERROR` | Fallback provider also failed | `plan_id`, `fallback_provider`, `fallback_model`, `fallback_error`, `fallback_error_type` |
| `DEBUG` | Still in quota fallback mode (with time remaining) | `plan_id` |
| `INFO` | Quota recovery interval elapsed, attempting primary | `plan_id`, `time_since_error` |
All log events include `plan_id` for correlation. Enable `DEBUG` level logging to
observe fallback mode transitions in detail.
---
## Configuration
The fallback behavior is controlled by module-level constants in
`src/cleveragents/application/services/strategy_actor.py`. These are not currently
exposed as environment variables or configuration keys. To change the fallback
provider or recovery interval, modify the constants directly:
```python
# src/cleveragents/application/services/strategy_actor.py
_FALLBACK_PROVIDER = "anthropic" # Provider identifier
_FALLBACK_MODEL = "claude-sonnet-4-20250514" # Model identifier
_QUOTA_RECOVERY_INTERVAL = 300 # Seconds before re-trying primary
```
> **Note:** The fallback provider requires a valid `ANTHROPIC_API_KEY` environment
> variable to be set. If the key is absent, the fallback LLM creation will fail and
> the original quota exception will be re-raised.
---
## Prerequisites
To enable the Anthropic fallback, ensure the following environment variable is set:
```bash
export ANTHROPIC_API_KEY=<your-anthropic-api-key>
```
The primary provider (e.g. OpenAI) is configured as usual:
```bash
export OPENAI_API_KEY=<your-openai-api-key>
```
See the [LLM provider configuration](../../README.md#llm-provider-configuration)
section of the README for the full list of supported providers and environment
variables.
---
## Implementation Reference
| Symbol | Location | Description |
|--------|----------|-------------|
| `_is_quota_error(exc)` | `src/cleveragents/application/services/strategy_actor.py` | Detects quota errors by string matching |
| `StrategyActor._execute_with_llm()` | `src/cleveragents/application/services/strategy_actor.py` | Orchestrates primary call, quota detection, and fallback |
| `StrategyActor._fallback_llm` | `src/cleveragents/application/services/strategy_actor.py` | Cached fallback LLM instance |
| `StrategyActor._last_quota_error_time` | `src/cleveragents/application/services/strategy_actor.py` | Timestamp of last quota error |
| `StrategyActor._using_fallback` | `src/cleveragents/application/services/strategy_actor.py` | Fallback mode flag |
| `_FALLBACK_PROVIDER` | `src/cleveragents/application/services/strategy_actor.py` | `"anthropic"` |
| `_FALLBACK_MODEL` | `src/cleveragents/application/services/strategy_actor.py` | `"claude-sonnet-4-20250514"` |
| `_QUOTA_RECOVERY_INTERVAL` | `src/cleveragents/application/services/strategy_actor.py` | `300` seconds |
**Introduced in:** commit `f5712787` — *feat: add fallback to Anthropic Sonnet when
OpenAI quota is exhausted*
**Fixes:** issue #10042
---
## Related Documentation
- [Retry Policy](retry_policy.md) — service-level retry and circuit breaker configuration
- [Error Recovery](error_recovery.md) — plan-level error classification and recovery actions
- [Error Handling](error_handling.md) — error codes, classification, and secret redaction
- [Observability](observability.md) — structured logging and tracing configuration