90 lines
4.6 KiB
Markdown
90 lines
4.6 KiB
Markdown
# Observability — LLM Trace & Operational Metrics
|
|
|
|
## Overview
|
|
|
|
The observability subsystem records telemetry for every LLM provider
|
|
call and exposes aggregated operational metrics per plan. Trace data
|
|
drives cost analysis, latency profiling, and optional forwarding to
|
|
LangSmith for external analysis.
|
|
|
|
## LLMTrace Model
|
|
|
|
`cleveragents.domain.models.observability.llm_trace.LLMTrace`
|
|
|
|
Each trace captures one LLM invocation:
|
|
|
|
| Field | Type | Description |
|
|
|:-------------------|:------------------|:-----------------------------------------|
|
|
| `trace_id` | `str` (ULID) | Unique identifier for the trace |
|
|
| `plan_id` | `str` (ULID) | Plan that owns this trace |
|
|
| `decision_id` | `str \| None` | Decision that triggered the call |
|
|
| `actor` | `str` | Actor name |
|
|
| `provider` | `str` | Provider identifier (e.g. `openai`) |
|
|
| `model` | `str` | Model name (e.g. `gpt-4o`) |
|
|
| `prompt_tokens` | `int` | Prompt token count |
|
|
| `completion_tokens`| `int` | Completion token count |
|
|
| `cost_usd` | `float` | Estimated USD cost |
|
|
| `latency_ms` | `float` | Wall-clock latency (ms) |
|
|
| `tool_calls` | `list[dict]` | Tool call descriptors from the model |
|
|
| `context_hash` | `str \| None` | SHA-256 of the context window |
|
|
| `streaming` | `bool` | Whether streaming was used |
|
|
| `retry_count` | `int` | Retries before success |
|
|
| `error` | `str \| None` | Error message on failure |
|
|
| `timestamp` | `datetime` | When the trace was recorded |
|
|
|
|
The model is **frozen** (immutable) once created.
|
|
|
|
## Operational Metrics
|
|
|
|
`cleveragents.domain.models.observability.metrics.OperationalMetricKey`
|
|
|
|
14 metric keys grouped by subsystem:
|
|
|
|
| Key | Subsystem | Description |
|
|
|:---------------------------|:----------|:-----------------------------------|
|
|
| `PLAN_DURATION_MS` | Plan | Total plan execution time |
|
|
| `PLAN_TOTAL_COST_USD` | Plan | Aggregate cost across all LLM calls|
|
|
| `PLAN_DECISION_COUNT` | Plan | Number of decisions made |
|
|
| `SUBPLAN_COUNT` | Plan | Number of subplans spawned |
|
|
| `ACTOR_INVOCATION_COUNT` | Actor | Actor call count |
|
|
| `ACTOR_LATENCY_MS` | Actor | Actor invocation latency |
|
|
| `TOOL_INVOCATION_COUNT` | Tool | Tool execution count |
|
|
| `TOOL_ERROR_RATE` | Tool | Tool error occurrences |
|
|
| `CONTEXT_BUILD_TIME_MS` | Context | Context assembly time |
|
|
| `CONTEXT_TOKEN_COUNT` | Context | Context token count |
|
|
| `LLM_CALL_COUNT` | LLM | Total LLM calls |
|
|
| `LLM_TOTAL_TOKENS` | LLM | Sum of all tokens |
|
|
| `LLM_TOTAL_COST_USD` | LLM | Sum of all costs |
|
|
| `LLM_AVG_LATENCY_MS` | LLM | Average call latency |
|
|
|
|
## TraceService
|
|
|
|
`cleveragents.application.services.trace_service.TraceService`
|
|
|
|
Registered in the DI container as `trace_service`.
|
|
|
|
### Methods
|
|
|
|
| Method | Description |
|
|
|:------------------------|:------------------------------------------------|
|
|
| `record_trace(trace)` | Persist trace and optionally forward to LangSmith|
|
|
| `get_traces(plan_id)` | List traces for a plan |
|
|
| `get_trace(trace_id)` | Retrieve a single trace |
|
|
| `compute_metrics(plan_id)` | Compute LLM-level metrics from traces |
|
|
| `on_plan_start(plan_id)` | Lifecycle hook: plan start |
|
|
| `on_actor_invocation(...)` | Lifecycle hook: actor invocation |
|
|
| `on_tool_execution(...)` | Lifecycle hook: tool execution |
|
|
|
|
### LangSmith Forwarding
|
|
|
|
When `LANGCHAIN_TRACING_V2=true` is set in the environment,
|
|
`record_trace` automatically forwards each trace to LangSmith
|
|
via the `langsmith` SDK. Forwarding is best-effort: failures
|
|
are logged but do not raise exceptions.
|
|
|
|
## Database
|
|
|
|
Table `llm_traces` with indexes on `plan_id`, `decision_id`,
|
|
`actor`, and `provider`. Repository:
|
|
`cleveragents.infrastructure.database.llm_trace_repository.LLMTraceRepository`.
|