4d3499dcfb
CI / benchmark-publish (pull_request) Has been skipped
CI / lint (pull_request) Successful in 16s
CI / build (pull_request) Successful in 17s
CI / quality (pull_request) Successful in 18s
CI / security (pull_request) Successful in 35s
CI / typecheck (pull_request) Successful in 38s
CI / unit_tests (pull_request) Successful in 2m43s
CI / integration_tests (pull_request) Successful in 3m20s
CI / docker (pull_request) Successful in 51s
CI / coverage (pull_request) Successful in 6m18s
CI / build (push) Successful in 14s
CI / lint (push) Successful in 21s
CI / quality (push) Successful in 29s
CI / typecheck (push) Successful in 35s
CI / security (push) Successful in 36s
CI / benchmark-regression (push) Has been skipped
CI / unit_tests (push) Successful in 3m17s
CI / docker (push) Successful in 9s
CI / integration_tests (push) Successful in 3m29s
CI / coverage (push) Successful in 7m3s
CI / benchmark-publish (push) Successful in 19m50s
CI / benchmark-regression (pull_request) Successful in 40m6s
Wire per-service retry policies and circuit breakers into the service layer via ServiceRetryWiring, backed by ServiceRetryPolicyRegistry and configurable through Settings environment variables. Production hardening from code review: - Fix TOCTOU race in CircuitBreaker._on_success (half-open state) - Add half-open probe limit to prevent unbounded concurrent requests - Track all exception types for circuit breaker failure counting - Detect async callables wrapped in functools.partial and callable objects - Enforce spec-compliant 2s minimum for linear backoff strategy - Enforce 0.1s floor for fixed backoff strategy - Add retry amplification guard via contextvars nesting depth tracking - Cap total retry wall-clock time at 300s (MAX_RETRY_TOTAL_TIMEOUT) - Sanitize exception messages in retry logs to prevent secret leakage - Fix wrap_service_method TOCTOU by holding cache lock for full operation - Deep-copy default policies to prevent cross-policy mutation - Warn on unknown override keys in apply_overrides - Guard apply_overrides against non-dict and deeply nested JSON values - Read circuit breaker state under lock in is_circuit_open - Catch RecursionError in JSON config parsing - Add total_timeout + nesting guard to retry_service_operation decorator - Extend secret sanitization to Authorization headers, private_key, connection_string, and access_key patterns - Enforce 0.1s floor on jitter backoff strategy - Cache wait strategies per service in ServiceRetryWiring (M3) - Reset failure_count to 0 when entering half-open from open (M6) - Use cached _get_wait_strategy() in execute()/async_execute() - Move circuit-open logging out of _on_failure lock scope to prevent holding the lock during potentially slow I/O (F1) - Pass total_timeout=MAX_RETRY_TOTAL_TIMEOUT to wrap_service_method retry_service_operation call for consistency with execute() (F4) - Capture failure_count into local variable inside lock scope before logging outside the lock, preventing stale reads from concurrent threads in CircuitBreaker.call() and async_call() (F1) - Deep-copy module-level DEFAULT_DATABASE_RETRY and DEFAULT_CIRCUIT_BREAKER in ServiceRetryPolicyRegistry.get() for auto-generated unknown service policies, preventing shared mutable state corruption (F1) - Unify CircuitBreaker to a single threading.Lock for sync and async (P1-1) - Restore BaseException permit in half-open path to prevent permit leak (P1-6) - Prevent CircuitBreakerOpen cascading into failure_count (S2) - Protect all logger calls with contextlib.suppress (S3, S4) - Replace time.time() with time.monotonic() for monotonic timing (S5) - Add distinct log events for half-open and closed transitions (S11, S12) - Track pre-existing services so second apply_settings_defaults only targets newly registered services (P1-2) - Lazy circuit breaker creation via _get_or_create_cb() (P1-3) - Reject async callables in sync execute() with TypeError (P1-5) - Strengthen retry predicate to retry_if_exception_type(Exception) & retry_if_not_exception_type(CircuitBreakerOpen) (S1) - Add lock on _get_wait_strategy cache access (P2-16) - Truncate raw JSON to 80 chars in override warning (P2-17) - Warn on non-dict JSON overrides (P2-29) - Debug log for nesting guard bypass (S13) - Deep-copy from get() and all_policies() in registry (P1-4) - Thread-safe registry with threading.Lock (P2-15) - Robust exception handling in apply_overrides get() (P2-18) - Log ValidationError details on override failure (P2-19) - Sanitize service_name via _safe_service_name() (P2-28) - Warn on non-dict sub-key values in overrides (P2-30) - Allowlist for is_read_only_plan_operation phases (P2-10) - Cap retry_auto_debug sleep at 60s (P2-11) - Use is-not-None instead of falsy checks for error values (P2-12) - Extend secret regex with bearer, session_id, auth_token, refresh_token, client_secret patterns (P2-25) - Pre-truncate error messages to 2000 chars before regex (P2-26) - Add upper bounds on retry Settings fields (P2-7) - Add cross-field validator max_delay >= base_delay (P2-8) - Case-insensitive backoff strategy validation (P2-21) - Add half_open_max_successes setting (S10) - Remove phantom ContextFragment from services __all__ (ImportError fix) - Export ServiceRetryWiring from application.services package - Include sanitised error context in TypeError logging fallback - Initialise RetryContext.attempt_count to 1 for bare context-manager usage - Introduce CircuitBreakerState StrEnum replacing raw string literals - Fix vacuous CircuitBreakerOpen propagation assertions in BDD steps - Replace tautological logging test with structlog capture verification - Assert circuit breaker existence instead of silently skipping on None - Add Unicode control-char rejection validator to ServiceRetryPolicy.service_name - Add name parameter with service= in all log calls - Add extra="forbid" to all 3 Pydantic models - Deep-copy _SERVICE_DEFAULTS construction - Key normalisation (.strip()) in get() and apply_overrides() - Add cooldown <= recovery_timeout validator - Async guard on RetryContext.execute() - Nesting guard on RetryContext.execute()/async_execute() - stop_after_delay(300.0) on RetryContext - retry_auto_debug async-only guard, dict result fix, sleep guard - Retry-attempt logging in RetryContext - Module-level docs for contextlib.suppress(TypeError) rationale - Exhaustion log on retry failure - Startup log in __init__; name=service_name to CircuitBreaker - log_after_retry guarded to not fire on first-attempt success - get_retry_decorator now includes logging callbacks - Changed retry_backoff_strategy from str to RetryStrategy StrEnum Closes #313
76 lines
3.8 KiB
Gherkin
76 lines
3.8 KiB
Gherkin
Feature: Retry Policy Wiring – Settings and Async Coverage
|
||
As a developer
|
||
I want settings overrides and async edge-cases tested
|
||
So that service retry wiring is fully covered
|
||
|
||
Background:
|
||
Given I have the retry policy wiring module imported
|
||
|
||
Scenario: Structured logging captures actual log output on retry
|
||
Given I have a ServiceRetryWiring instance
|
||
And I have a function that fails once then succeeds
|
||
When I execute the function through the wiring with structlog captured
|
||
Then the captured log should contain a retry warning message
|
||
|
||
Scenario: Settings apply individual retry max_attempts override
|
||
Given I have Settings with non-default retry_max_attempts 5
|
||
When I create a ServiceRetryWiring from those max-attempts Settings
|
||
Then the plan_service policy should have max_attempts 5
|
||
|
||
Scenario: Settings apply circuit breaker failure_threshold override
|
||
Given I have Settings with non-default circuit_breaker_failure_threshold 10
|
||
When I create a ServiceRetryWiring from those cb-threshold Settings
|
||
Then the plan_service circuit breaker should have failure_threshold 10
|
||
|
||
Scenario: is_circuit_open returns False for unknown service
|
||
Given I have a ServiceRetryWiring instance
|
||
When I check is_circuit_open for "nonexistent_service"
|
||
Then the circuit should report not open
|
||
|
||
Scenario: Async ServiceRetryWiring execute with circuit breaker retries
|
||
Given I have a ServiceRetryWiring instance
|
||
And I have an async function that fails once then succeeds
|
||
When I async_execute the function with circuit breaker enabled
|
||
Then the async function should succeed after retry
|
||
|
||
Scenario: Settings apply retry base_delay override
|
||
Given I have Settings with non-default retry_base_delay 2.5
|
||
When I create a ServiceRetryWiring from those base-delay Settings
|
||
Then the plan_service policy should have base_delay 2.5
|
||
|
||
Scenario: Settings apply circuit breaker recovery_timeout override
|
||
Given I have Settings with non-default circuit_breaker_recovery_timeout 120.0
|
||
When I create a ServiceRetryWiring from those recovery-timeout Settings
|
||
Then the plan_service circuit breaker should have recovery_timeout 120.0
|
||
|
||
Scenario: Settings apply circuit breaker cooldown override
|
||
Given I have Settings with non-default circuit_breaker_cooldown 60.0
|
||
When I create a ServiceRetryWiring from those cooldown Settings
|
||
Then the plan_service circuit breaker should have cooldown_seconds 60.0
|
||
|
||
Scenario: _build_cached_wait handles string backoff strategy
|
||
Given I have a ServiceRetryWiring instance with string backoff policy
|
||
When I request the wait strategy for the string-backoff service
|
||
Then the wait strategy should be built successfully
|
||
|
||
Scenario: Settings apply retry max_delay override
|
||
Given I have Settings with non-default retry_max_delay 120.0
|
||
When I create a ServiceRetryWiring from those max-delay Settings
|
||
Then the plan_service policy should have max_delay 120.0
|
||
|
||
Scenario: Settings apply retry jitter override
|
||
Given I have Settings with non-default retry_jitter False
|
||
When I create a ServiceRetryWiring from those jitter Settings
|
||
Then the plan_service policy should have jitter False
|
||
|
||
Scenario: Async ServiceRetryWiring execute without circuit breaker
|
||
Given I have a ServiceRetryWiring instance for a service without circuit breaker
|
||
And I have an async function that fails once then succeeds
|
||
When I async_execute the function for the unprotected service
|
||
Then the async function should succeed after retry
|
||
|
||
Scenario: Sync ServiceRetryWiring nesting guard without circuit breaker
|
||
Given I have a ServiceRetryWiring instance for a service without circuit breaker
|
||
When I execute a nested sync operation for the unprotected service
|
||
Then the nested inner function should execute directly without CB
|