Files
cleveragents-core/features/session_tell_llm.feature
hurui200320 87a7ce35d7
CI / benchmark-regression (push) Has been skipped
CI / helm (push) Successful in 38s
CI / lint (push) Successful in 1m14s
CI / build (push) Successful in 1m10s
CI / push-validation (push) Successful in 49s
CI / quality (push) Successful in 1m35s
CI / typecheck (push) Successful in 1m39s
CI / security (push) Successful in 1m45s
CI / integration_tests (push) Successful in 3m44s
CI / e2e_tests (push) Successful in 4m29s
CI / unit_tests (push) Successful in 5m10s
CI / docker (push) Successful in 1m55s
CI / coverage (push) Successful in 11m8s
CI / status-check (push) Successful in 3s
CI / benchmark-publish (push) Successful in 1h20m21s
CI / benchmark-publish (pull_request) Has been skipped
CI / benchmark-regression (pull_request) Failing after 1m35s
CI / docker (pull_request) Successful in 1m26s
CI / unit_tests (pull_request) Successful in 6m40s
CI / push-validation (pull_request) Successful in 1m24s
CI / quality (pull_request) Successful in 4m4s
CI / integration_tests (pull_request) Successful in 5m44s
CI / e2e_tests (pull_request) Failing after 6m5s
CI / helm (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m39s
CI / lint (pull_request) Successful in 2m58s
CI / typecheck (pull_request) Successful in 3m56s
CI / security (pull_request) Successful in 3m56s
CI / coverage (pull_request) Successful in 10m52s
CI / status-check (pull_request) Failing after 3s
feat(session): implement real LLM actor invocation in session tell
Replace the M3 stub in `agents session tell` with real orchestrator actor
invocation via `SessionWorkflow.tell()`. The stub echoed a canned
`Acknowledged: ...` response without calling any LLM or actor; this
commit wires up the full pipeline:

Architecture:
- New `SessionWorkflow` (Application layer) orchestrates LLM invocation
  for session tell. Accepts a `ProviderRegistry` and `ToolRegistry`;
  falls back to `FakeListLLM` when no provider is configured.
- `LangChainSessionCaller` implements the `LLMCaller` protocol, building
  history-aware LangChain message lists from the session conversation and
  invoking the LLM via `ToolCallingRuntime.run_tool_loop()`.
- `TellResult` (Pydantic BaseModel) carries the assistant response plus
  token usage (input_tokens, output_tokens, cost_usd, duration_ms).

Domain / service layer:
- `SessionActorNotConfiguredError` added to the session domain model;
  raised when tell is invoked with no actor on the session and no
  `--actor` override. Clear message; CLI exits with code 1.
- `SessionService.get_messages()` abstract method added (+ implementation
  in `PersistentSessionService`) to load ordered message history.

A2A facade:
- `A2aLocalFacade` gains `message/send` and `message/stream` standard
  A2A operation handlers that route to `SessionWorkflow.tell()`.
  Total supported operations count: 42 → 44.

CLI:
- `session tell` command delegates to `_build_session_workflow()`
  (patchable factory) instead of directly calling `SessionService`.
- Non-streaming: `SessionWorkflow.tell()` returns `TellResult`; output
  includes a Usage panel (Rich/Plain) or a `usage` object (JSON/YAML).
- Streaming: `SessionWorkflow.tell_stream()` yields tokens; CLI prints
  them via `console.print` (not raw `sys.stdout.write`).
- Token usage recorded via `SessionService.update_token_usage()` in
  both paths.

Tests:
- New Behave feature `session_tell_llm.feature` (4 scenarios): real LLM
  response persisted; streaming yields tokens; no-actor exits code 1;
  `--actor` override resolves correctly.
- New Robot suite `session_tell_llm.robot` (4 tests): end-to-end with
  stub LLM injected via monkey-patching `_resolve_llm`.
- Updated all existing `session tell` test steps to patch
  `_build_session_workflow` returning a mock `TellResult` so tests
  remain isolated from the LLM layer.
- Updated operation-count assertions (42 → 44) in
  `a2a_cli_facade_integration`, `consolidated_misc`,
  `m6_autonomy_acceptance` feature files and steps.

Cycle 7 (Review ID 8088) fixes:
- Blocker 4: Split `session_workflow.py` (was 801 lines) into two files:
  `session_workflow.py` (467 lines) and `session_caller.py` (318 lines).
  Extracted: `LangChainSessionCaller`, `extract_content`,
  `extract_token_usage`, `estimate_cost`, `estimate_tokens`,
  `history_to_langchain_messages`, and stub classes.
- Blocker 5: Route CLI streaming path through
  `_facade_dispatch("message/stream", ...)` instead of directly
  calling `workflow.tell_stream()`. The facade's
  `_handle_message_stream` falls back to non-streaming with
  `streamed: false` (acceptable for this milestone per the spec).
- Updated streaming test assertion to validate full response presence
  (no longer checks token-by-token word positions, since the facade
  fallback returns a complete message).

ISSUES CLOSED: #5784
2026-05-11 04:39:29 +00:00

59 lines
2.5 KiB
Gherkin

Feature: Real LLM actor invocation in session tell
As a user of CleverAgents
I want `agents session tell` to invoke the real orchestrator actor
So that I get meaningful AI responses instead of stub acknowledgements
Background:
Given a session tell LLM mock environment
@session_tell_llm
Scenario: Real LLM response is returned and persisted
Given a session with actor "openai/gpt-4" exists
When I invoke session tell with prompt "What can you do?"
Then the tell command returns the LLM response
And the assistant message is persisted to the session
And token usage is recorded
@session_tell_llm
Scenario: --stream flag yields incremental tokens
Given a session with actor "openai/gpt-4" exists
When I invoke session tell with --stream and prompt "Hello"
Then the streamed output contains the LLM response tokens
And the assistant message is persisted after streaming
@session_tell_llm
Scenario: No actor configured raises clear error with exit code 1
Given a session with no actor exists
When I invoke session tell with prompt "Hello"
Then the tell command exits with code 1
And the error output mentions actor configuration
@session_tell_llm
Scenario: --actor flag overrides the session's bound actor
Given a session with actor "anthropic/claude-3-haiku" exists
When I invoke session tell with --actor "openai/gpt-4" and prompt "Hello"
Then the tell command uses the override actor "openai/gpt-4"
And the tell command returns the LLM response
@session_tell_llm
Scenario: --format json output includes usage object
Given a session with actor "openai/gpt-4" exists
When I invoke session tell with --format json and prompt "What can you do?"
Then the tell output is valid JSON with a data envelope
And the data section contains the session_id
And the data section contains a usage object with expected keys
@session_tell_llm
Scenario: --stream --format json produces valid JSON with usage
Given a session with actor "openai/gpt-4" exists
When I invoke session tell with --stream --format json and prompt "Hello"
Then the tell output is valid JSON with a data envelope
And the data section contains a usage object with expected keys
@session_tell_llm
Scenario: Usage panel appears in Rich output
Given a session with actor "openai/gpt-4" exists
When I invoke session tell with prompt "What can you do?"
Then the tell command returns the LLM response
And the output contains a Usage panel