Files
cleveragents-core/features/retry_patterns.feature
CoreRasurae 4d3499dcfb
CI / benchmark-publish (pull_request) Has been skipped
CI / lint (pull_request) Successful in 16s
CI / build (pull_request) Successful in 17s
CI / quality (pull_request) Successful in 18s
CI / security (pull_request) Successful in 35s
CI / typecheck (pull_request) Successful in 38s
CI / unit_tests (pull_request) Successful in 2m43s
CI / integration_tests (pull_request) Successful in 3m20s
CI / docker (pull_request) Successful in 51s
CI / coverage (pull_request) Successful in 6m18s
CI / build (push) Successful in 14s
CI / lint (push) Successful in 21s
CI / quality (push) Successful in 29s
CI / typecheck (push) Successful in 35s
CI / security (push) Successful in 36s
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Successful in 3m17s
CI / docker (push) Successful in 9s
CI / integration_tests (push) Successful in 3m29s
CI / coverage (push) Successful in 7m3s
CI / benchmark-publish (push) Successful in 19m50s
CI / benchmark-regression (pull_request) Successful in 40m6s
feat(async): wire retry policies into services
Wire per-service retry policies and circuit breakers into the service
layer via ServiceRetryWiring, backed by ServiceRetryPolicyRegistry and
configurable through Settings environment variables.

Production hardening from code review:
- Fix TOCTOU race in CircuitBreaker._on_success (half-open state)
- Add half-open probe limit to prevent unbounded concurrent requests
- Track all exception types for circuit breaker failure counting
- Detect async callables wrapped in functools.partial and callable objects
- Enforce spec-compliant 2s minimum for linear backoff strategy
- Enforce 0.1s floor for fixed backoff strategy
- Add retry amplification guard via contextvars nesting depth tracking
- Cap total retry wall-clock time at 300s (MAX_RETRY_TOTAL_TIMEOUT)
- Sanitize exception messages in retry logs to prevent secret leakage
- Fix wrap_service_method TOCTOU by holding cache lock for full operation
- Deep-copy default policies to prevent cross-policy mutation
- Warn on unknown override keys in apply_overrides
- Guard apply_overrides against non-dict and deeply nested JSON values
- Read circuit breaker state under lock in is_circuit_open
- Catch RecursionError in JSON config parsing
- Add total_timeout + nesting guard to retry_service_operation decorator
- Extend secret sanitization to Authorization headers, private_key,
  connection_string, and access_key patterns
- Enforce 0.1s floor on jitter backoff strategy
- Cache wait strategies per service in ServiceRetryWiring (M3)
- Reset failure_count to 0 when entering half-open from open (M6)
- Use cached _get_wait_strategy() in execute()/async_execute()
- Move circuit-open logging out of _on_failure lock scope to prevent
  holding the lock during potentially slow I/O (F1)
- Pass total_timeout=MAX_RETRY_TOTAL_TIMEOUT to wrap_service_method
  retry_service_operation call for consistency with execute() (F4)
- Capture failure_count into local variable inside lock scope before
  logging outside the lock, preventing stale reads from concurrent
  threads in CircuitBreaker.call() and async_call() (F1)
- Deep-copy module-level DEFAULT_DATABASE_RETRY and DEFAULT_CIRCUIT_BREAKER
  in ServiceRetryPolicyRegistry.get() for auto-generated unknown service
  policies, preventing shared mutable state corruption (F1)
- Unify CircuitBreaker to a single threading.Lock for sync and async (P1-1)
- Restore BaseException permit in half-open path to prevent permit leak (P1-6)
- Prevent CircuitBreakerOpen cascading into failure_count (S2)
- Protect all logger calls with contextlib.suppress (S3, S4)
- Replace time.time() with time.monotonic() for monotonic timing (S5)
- Add distinct log events for half-open and closed transitions (S11, S12)
- Track pre-existing services so second apply_settings_defaults only
  targets newly registered services (P1-2)
- Lazy circuit breaker creation via _get_or_create_cb() (P1-3)
- Reject async callables in sync execute() with TypeError (P1-5)
- Strengthen retry predicate to retry_if_exception_type(Exception) &
  retry_if_not_exception_type(CircuitBreakerOpen) (S1)
- Add lock on _get_wait_strategy cache access (P2-16)
- Truncate raw JSON to 80 chars in override warning (P2-17)
- Warn on non-dict JSON overrides (P2-29)
- Debug log for nesting guard bypass (S13)
- Deep-copy from get() and all_policies() in registry (P1-4)
- Thread-safe registry with threading.Lock (P2-15)
- Robust exception handling in apply_overrides get() (P2-18)
- Log ValidationError details on override failure (P2-19)
- Sanitize service_name via _safe_service_name() (P2-28)
- Warn on non-dict sub-key values in overrides (P2-30)
- Allowlist for is_read_only_plan_operation phases (P2-10)
- Cap retry_auto_debug sleep at 60s (P2-11)
- Use is-not-None instead of falsy checks for error values (P2-12)
- Extend secret regex with bearer, session_id, auth_token,
  refresh_token, client_secret patterns (P2-25)
- Pre-truncate error messages to 2000 chars before regex (P2-26)
- Add upper bounds on retry Settings fields (P2-7)
- Add cross-field validator max_delay >= base_delay (P2-8)
- Case-insensitive backoff strategy validation (P2-21)
- Add half_open_max_successes setting (S10)
- Remove phantom ContextFragment from services __all__ (ImportError fix)
- Export ServiceRetryWiring from application.services package
- Include sanitised error context in TypeError logging fallback
- Initialise RetryContext.attempt_count to 1 for bare context-manager usage
- Introduce CircuitBreakerState StrEnum replacing raw string literals
- Fix vacuous CircuitBreakerOpen propagation assertions in BDD steps
- Replace tautological logging test with structlog capture verification
- Assert circuit breaker existence instead of silently skipping on None
- Add Unicode control-char rejection validator to ServiceRetryPolicy.service_name
- Add name parameter with service= in all log calls
- Add extra="forbid" to all 3 Pydantic models
- Deep-copy _SERVICE_DEFAULTS construction
- Key normalisation (.strip()) in get() and apply_overrides()
- Add cooldown <= recovery_timeout validator
- Async guard on RetryContext.execute()
- Nesting guard on RetryContext.execute()/async_execute()
- stop_after_delay(300.0) on RetryContext
- retry_auto_debug async-only guard, dict result fix, sleep guard
- Retry-attempt logging in RetryContext
- Module-level docs for contextlib.suppress(TypeError) rationale
- Exhaustion log on retry failure
- Startup log in __init__; name=service_name to CircuitBreaker
- log_after_retry guarded to not fire on first-attempt success
- get_retry_decorator now includes logging callbacks
- Changed retry_backoff_strategy from str to RetryStrategy StrEnum

Closes #313
2026-03-11 17:42:13 +00:00

378 lines
16 KiB
Gherkin

@retry @patterns
Feature: Retry Patterns Implementation
As a developer
I want retry patterns implemented with tenacity
So that operations can be resilient to transient failures
Background:
Given I have the retry patterns module imported
# -------------------------------------------------------------------
# Basic Retry Strategies
# -------------------------------------------------------------------
@exponential
Scenario: Basic exponential backoff retry works
Given I have a function that fails 2 times then succeeds
When I apply exponential backoff retry with max 3 attempts
Then the function should eventually succeed
And the function should be called 3 times
@exponential @async
Scenario: Async exponential backoff retry works
Given I have an async function that fails 2 times then succeeds
When I apply async exponential backoff retry with max 3 attempts
Then the async function should eventually succeed
And the function should be called 3 times
@timeout
Scenario: Retry with timeout retries transient failures without waiting
Given I have a function that fails 1 times then succeeds
When I apply retry with timeout of 3 attempts and 0.01 seconds
Then the function should eventually succeed
And the function should be called 2 times
@jitter
Scenario: Retry with jitter prevents thundering herd
Given I have multiple concurrent operations
When I apply retry with jitter
Then the retries should have random delays
And operations should not retry simultaneously
@result-retry
Scenario: Retry on result retries when predicate flags response
Given I have sequential result payloads requiring a retry
And I have a retry predicate that checks for retry flag
When I apply retry on result with max 3 attempts
Then the decorated function should return the successful payload
And the function should be called 2 times
# -------------------------------------------------------------------
# Network / Provider / File Retry Patterns
# -------------------------------------------------------------------
@network
Scenario: Network operation retry handles connection errors
Given I have a function that raises NetworkError twice
When I apply network retry pattern
Then the function should retry on NetworkError
And the function should be called 3 times maximum
@provider
Scenario: Provider operation retry handles rate limits
Given I have a function that raises RateLimitError
When I apply provider retry pattern
Then the function should retry with exponential jitter
And the function should respect rate limit backoff
@file
Scenario: File operation retry recovers from transient OSError
Given I have a file operation that fails with OSError once
When I apply file retry pattern
Then the file operation should eventually succeed
And the file operation should be invoked 2 times
# -------------------------------------------------------------------
# Circuit Breaker
# -------------------------------------------------------------------
@circuit-breaker
Scenario: Circuit breaker opens after threshold
Given I have a circuit breaker with threshold 3
When a function fails 3 times
Then the circuit breaker should open
And subsequent calls should fail immediately
@circuit-breaker
Scenario: Circuit breaker resets after timeout
Given I have an open circuit breaker
When the recovery timeout expires
Then the circuit breaker should enter half-open state
And successful calls should close the circuit
@circuit-breaker
Scenario: Circuit breaker resets failure counters after closed success
Given I have a circuit breaker with threshold 2
And the circuit breaker has recorded failures
When I execute a successful call while the circuit is closed
Then the circuit breaker should reset failure tracking after success
@circuit-breaker
Scenario: Circuit breaker half-open closes after consecutive successes
Given I have a half-open circuit breaker with threshold 2
When I record a successful half-open call
And I record another successful half-open call
Then the circuit breaker should close after consecutive successes
@circuit-breaker
Scenario: Circuit breaker half-open reopens after failure
Given I have a half-open circuit breaker with threshold 2
When I record a successful half-open call
And I simulate a half-open failure
Then the circuit breaker should reopen after half-open failure
@circuit-breaker
Scenario: Circuit breaker rejects invalid half_open_max_successes
When I create a CircuitBreaker with half_open_max_successes 0
Then a ValueError should be raised for invalid half_open_max_successes
@circuit-breaker
Scenario: Circuit breaker _should_attempt_reset returns False with no prior failures
Given I have a fresh circuit breaker with no failures
When I check _should_attempt_reset on the fresh circuit breaker
Then _should_attempt_reset should return False
# -------------------------------------------------------------------
# Async Circuit Breaker
# -------------------------------------------------------------------
@circuit-breaker @async
Scenario: Async circuit breaker raises immediately while open
Given I have an async circuit breaker with threshold 1
When I trigger an async failure on the circuit breaker
Then the async circuit breaker should raise immediately while open
@circuit-breaker @async
Scenario: Async circuit breaker closes after recovery and successful calls
Given I have an async circuit breaker with threshold 1
When I trigger an async failure on the circuit breaker
And the recovery timeout expires
And I perform two async successful calls after recovery
Then the async circuit breaker should close after async successes
@circuit-breaker @async
Scenario: Async circuit breaker half-open probe limit blocks excess calls
Given I have an async circuit breaker in half-open with 0 permits
When I make an async call on the half-open circuit breaker
Then a CircuitBreakerOpen should be raised for async probe limit
@circuit-breaker @async
Scenario: Async circuit breaker restores permits on CancelledError
Given I have an async circuit breaker in half-open with 1 permit
When the async call is cancelled mid-execution
Then the half-open permit should be restored
# -------------------------------------------------------------------
# Retry Context
# -------------------------------------------------------------------
@retry-context
Scenario: Retry context tracks attempts
Given I have a retry context for "test operation"
When I execute a function that fails twice
Then the retry context should track 3 attempts
And the errors should be recorded
@retry-context
Scenario: Retry context manager records exceptions raised in block
Given I have a retry context for "context block"
When I execute a failing block under the retry context manager
Then the retry context manager should capture the exception
@retry-context @async
Scenario: Async retry context manager records exceptions raised in block
Given I have a retry context for "async context block"
When I execute a failing block under the async retry context manager
Then the async retry context manager should capture the exception
@retry-context
Scenario: Retry context manager leaves errors untouched on success
Given I have a retry context for "successful context block"
When I execute a successful block under the retry context manager
Then the retry context manager should not record errors
@retry-context @async
Scenario: Async retry context manager leaves errors untouched on success
Given I have a retry context for "async successful context block"
When I execute a successful block under the async retry context manager
Then the async retry context manager should not record errors
@retry-context @async
Scenario: Retry context async execute handles transient async failures
Given I have a retry context for "async execute operation"
And I have an async function that fails 2 times then succeeds
When I execute the async function with the retry context
Then the async retry context execute should succeed after 3 attempts
# -------------------------------------------------------------------
# should_retry_result
# -------------------------------------------------------------------
@result-retry
Scenario: should_retry_result handles error payloads
Given I have a response payload with error "timeout detected"
When I evaluate the retry decision
Then the retry decision should be True
@result-retry
Scenario: should_retry_result handles server error status codes
Given I have a response payload with status code 503
When I evaluate the retry decision
Then the retry decision should be True
@result-retry
Scenario: should_retry_result returns False for None payloads
Given I have a result payload set to None
When I evaluate the retry decision
Then the retry decision should be False
@result-retry
Scenario: should_retry_result returns False for non-error dictionaries
Given I have a response payload with status code 200
When I evaluate the retry decision
Then the retry decision should be False
@result-retry
Scenario: should_retry_result returns False for non-dictionary results
Given I have a result payload set to the string "ok"
When I evaluate the retry decision
Then the retry decision should be False
# -------------------------------------------------------------------
# Retry Categories & Logging
# -------------------------------------------------------------------
@categories
Scenario: Retry categories have correct configurations
Given I have the retry categories defined
Then network operations should retry 5 times
And provider operations should retry 3 times
And database operations should retry 3 times
And file operations should retry 3 times
@categories
Scenario: Unknown retry category falls back to default decorator
When I request a retry decorator for "unknown-category"
And I decorate a function returning "decorated result"
Then the decorated function should execute successfully
@categories
Scenario: Known retry category returns a configured decorator
When I request a retry decorator for "network"
And I decorate a function returning "network result"
Then the decorated function should execute successfully
@logging
Scenario: Retry logging handles loggers without keyword support
Given the retry logger rejects keyword arguments
When I invoke logging callbacks for failed and successful attempts
Then the fallback logging messages should be recorded
@logging
Scenario: configure_retry_logging sets tenacity logger levels
When I configure retry logging to warning level
Then the tenacity loggers should be set to warning
# -------------------------------------------------------------------
# Auto-Debug Retry
# -------------------------------------------------------------------
@auto-debug
Scenario: Auto-debug retry pattern works
Given I have a function that fails with errors
And I have a debug callback that fixes issues
When I apply auto-debug retry
Then the function should retry with debug fixes
And eventually succeed after debugging
@auto-debug
Scenario: Auto-debug retry raises the last error message when unresolved
Given I have a function that returns persistent error payloads
And I have a debug callback that never fixes issues
When I apply auto-debug retry expecting failure
Then the auto-debug retry should raise the last error message
@auto-debug
Scenario: Auto-debug retry logs debug callback failures
Given I have a function that returns persistent error payloads
And I have a debug callback that raises errors
When I apply auto-debug retry expecting failure
Then the auto-debug retry should raise the debug callback failure
And the debug callback failures should be counted
@auto-debug
Scenario: Auto-debug retry returns result for non-dict success
Given I have an async function that returns a non-dict value
When I apply auto-debug retry to the non-dict function
Then the auto-debug should return the non-dict value directly
@auto-debug
Scenario: Auto-debug retry returns None when no errors occur and no last_error
Given I have an async function that returns error-free dict then stops
And I have no debug callback
When I apply auto-debug retry to the error-free function
Then the auto-debug should return the dict result
@auto-debug
Scenario: Auto-debug with fix callback that resolves the error
Given I have an async function that returns error then succeeds after fix
And I have a debug callback that returns fixed
When I apply auto-debug retry with fix callback
Then the auto-debug should eventually succeed
# -------------------------------------------------------------------
# Retry Service Operation Decorators
# -------------------------------------------------------------------
@decorator @async
Scenario: Async retry_service_operation decorator retries correctly
Given I have an async function that fails once then returns ok
When I apply retry_service_operation as async and call
Then the async decorated function should succeed
And the async decorated function should have been called 2 times
@decorator @async
Scenario: Async retry_service_operation nesting guard skips retries
When I call nested async retry_service_operation decorators
Then the inner async decorator should execute without retries
@decorator
Scenario: Sync decorator with circuit breaker handles open
Given I have a sync function guarded by retry_service_operation with CB
When the circuit breaker opens during sync decorated retry
Then the sync decorator should propagate CircuitBreakerOpen
@decorator @async
Scenario: Async retry_service_operation with circuit breaker and CB open
Given I have an async function guarded by retry_service_operation with CB
When the circuit breaker opens during async decorated retry
Then the async decorator should propagate CircuitBreakerOpen
@decorator
Scenario: Sync decorator nesting guard delegates to circuit breaker
When I call a sync decorated function at max nesting depth with CB
Then the sync decorator nesting guard should delegate to CB
@decorator
Scenario: retry_service_operation non-idempotent returns function as-is
When I create a non-idempotent retry_service_operation decorator
Then the decorator should return the original function unchanged
@decorator @async
Scenario: Async decorator nesting guard delegates to circuit breaker
When I call an async decorated function at max nesting depth with CB
Then the async nesting CB guard should delegate successfully
# -------------------------------------------------------------------
# _is_async_callable
# -------------------------------------------------------------------
@async
Scenario: _is_async_callable detects partial wrapping async callable object
When I check _is_async_callable with partial wrapping an async callable object
Then _is_async_callable should return True for partial async callable
@async
Scenario: _is_async_callable detects partial wrapping plain async function
When I check _is_async_callable with partial wrapping a plain async function
Then _is_async_callable should return True for partial plain async
# -------------------------------------------------------------------
# Structured Log Fallbacks
# -------------------------------------------------------------------
@logging
Scenario: Structured log TypeError fallback for service retry
When I trigger the structlog TypeError fallback for service retry logging
Then the fallback logging should not raise any exception