Files

90 lines
4.6 KiB
Markdown

# Observability — LLM Trace & Operational Metrics
## Overview
The observability subsystem records telemetry for every LLM provider
call and exposes aggregated operational metrics per plan. Trace data
drives cost analysis, latency profiling, and optional forwarding to
LangSmith for external analysis.
## LLMTrace Model
`cleveragents.domain.models.observability.llm_trace.LLMTrace`
Each trace captures one LLM invocation:
| Field | Type | Description |
|:-------------------|:------------------|:-----------------------------------------|
| `trace_id` | `str` (ULID) | Unique identifier for the trace |
| `plan_id` | `str` (ULID) | Plan that owns this trace |
| `decision_id` | `str \| None` | Decision that triggered the call |
| `actor` | `str` | Actor name |
| `provider` | `str` | Provider identifier (e.g. `openai`) |
| `model` | `str` | Model name (e.g. `gpt-4o`) |
| `prompt_tokens` | `int` | Prompt token count |
| `completion_tokens`| `int` | Completion token count |
| `cost_usd` | `float` | Estimated USD cost |
| `latency_ms` | `float` | Wall-clock latency (ms) |
| `tool_calls` | `list[dict]` | Tool call descriptors from the model |
| `context_hash` | `str \| None` | SHA-256 of the context window |
| `streaming` | `bool` | Whether streaming was used |
| `retry_count` | `int` | Retries before success |
| `error` | `str \| None` | Error message on failure |
| `timestamp` | `datetime` | When the trace was recorded |
The model is **frozen** (immutable) once created.
## Operational Metrics
`cleveragents.domain.models.observability.metrics.OperationalMetricKey`
14 metric keys grouped by subsystem:
| Key | Subsystem | Description |
|:---------------------------|:----------|:-----------------------------------|
| `PLAN_DURATION_MS` | Plan | Total plan execution time |
| `PLAN_TOTAL_COST_USD` | Plan | Aggregate cost across all LLM calls|
| `PLAN_DECISION_COUNT` | Plan | Number of decisions made |
| `SUBPLAN_COUNT` | Plan | Number of subplans spawned |
| `ACTOR_INVOCATION_COUNT` | Actor | Actor call count |
| `ACTOR_LATENCY_MS` | Actor | Actor invocation latency |
| `TOOL_INVOCATION_COUNT` | Tool | Tool execution count |
| `TOOL_ERROR_RATE` | Tool | Tool error occurrences |
| `CONTEXT_BUILD_TIME_MS` | Context | Context assembly time |
| `CONTEXT_TOKEN_COUNT` | Context | Context token count |
| `LLM_CALL_COUNT` | LLM | Total LLM calls |
| `LLM_TOTAL_TOKENS` | LLM | Sum of all tokens |
| `LLM_TOTAL_COST_USD` | LLM | Sum of all costs |
| `LLM_AVG_LATENCY_MS` | LLM | Average call latency |
## TraceService
`cleveragents.application.services.trace_service.TraceService`
Registered in the DI container as `trace_service`.
### Methods
| Method | Description |
|:------------------------|:------------------------------------------------|
| `record_trace(trace)` | Persist trace and optionally forward to LangSmith|
| `get_traces(plan_id)` | List traces for a plan |
| `get_trace(trace_id)` | Retrieve a single trace |
| `compute_metrics(plan_id)` | Compute LLM-level metrics from traces |
| `on_plan_start(plan_id)` | Lifecycle hook: plan start |
| `on_actor_invocation(...)` | Lifecycle hook: actor invocation |
| `on_tool_execution(...)` | Lifecycle hook: tool execution |
### LangSmith Forwarding
When `LANGCHAIN_TRACING_V2=true` is set in the environment,
`record_trace` automatically forwards each trace to LangSmith
via the `langsmith` SDK. Forwarding is best-effort: failures
are logged but do not raise exceptions.
## Database
Table `llm_traces` with indexes on `plan_id`, `decision_id`,
`actor`, and `provider`. Repository:
`cleveragents.infrastructure.database.llm_trace_repository.LLMTraceRepository`.