87a7ce35d7
CI / benchmark-regression (push) Has been skipped
CI / helm (push) Successful in 38s
CI / lint (push) Successful in 1m14s
CI / build (push) Successful in 1m10s
CI / push-validation (push) Successful in 49s
CI / quality (push) Successful in 1m35s
CI / typecheck (push) Successful in 1m39s
CI / security (push) Successful in 1m45s
CI / integration_tests (push) Successful in 3m44s
CI / e2e_tests (push) Successful in 4m29s
CI / unit_tests (push) Successful in 5m10s
CI / docker (push) Successful in 1m55s
CI / coverage (push) Successful in 11m8s
CI / status-check (push) Successful in 3s
CI / benchmark-publish (push) Successful in 1h20m21s
CI / benchmark-publish (pull_request) Has been skipped
CI / benchmark-regression (pull_request) Failing after 1m35s
CI / docker (pull_request) Successful in 1m26s
CI / unit_tests (pull_request) Successful in 6m40s
CI / push-validation (pull_request) Successful in 1m24s
CI / quality (pull_request) Successful in 4m4s
CI / integration_tests (pull_request) Successful in 5m44s
CI / e2e_tests (pull_request) Failing after 6m5s
CI / helm (pull_request) Successful in 1m10s
CI / build (pull_request) Successful in 2m39s
CI / lint (pull_request) Successful in 2m58s
CI / typecheck (pull_request) Successful in 3m56s
CI / security (pull_request) Successful in 3m56s
CI / coverage (pull_request) Successful in 10m52s
CI / status-check (pull_request) Failing after 3s
Replace the M3 stub in `agents session tell` with real orchestrator actor
invocation via `SessionWorkflow.tell()`. The stub echoed a canned
`Acknowledged: ...` response without calling any LLM or actor; this
commit wires up the full pipeline:
Architecture:
- New `SessionWorkflow` (Application layer) orchestrates LLM invocation
for session tell. Accepts a `ProviderRegistry` and `ToolRegistry`;
falls back to `FakeListLLM` when no provider is configured.
- `LangChainSessionCaller` implements the `LLMCaller` protocol, building
history-aware LangChain message lists from the session conversation and
invoking the LLM via `ToolCallingRuntime.run_tool_loop()`.
- `TellResult` (Pydantic BaseModel) carries the assistant response plus
token usage (input_tokens, output_tokens, cost_usd, duration_ms).
Domain / service layer:
- `SessionActorNotConfiguredError` added to the session domain model;
raised when tell is invoked with no actor on the session and no
`--actor` override. Clear message; CLI exits with code 1.
- `SessionService.get_messages()` abstract method added (+ implementation
in `PersistentSessionService`) to load ordered message history.
A2A facade:
- `A2aLocalFacade` gains `message/send` and `message/stream` standard
A2A operation handlers that route to `SessionWorkflow.tell()`.
Total supported operations count: 42 → 44.
CLI:
- `session tell` command delegates to `_build_session_workflow()`
(patchable factory) instead of directly calling `SessionService`.
- Non-streaming: `SessionWorkflow.tell()` returns `TellResult`; output
includes a Usage panel (Rich/Plain) or a `usage` object (JSON/YAML).
- Streaming: `SessionWorkflow.tell_stream()` yields tokens; CLI prints
them via `console.print` (not raw `sys.stdout.write`).
- Token usage recorded via `SessionService.update_token_usage()` in
both paths.
Tests:
- New Behave feature `session_tell_llm.feature` (4 scenarios): real LLM
response persisted; streaming yields tokens; no-actor exits code 1;
`--actor` override resolves correctly.
- New Robot suite `session_tell_llm.robot` (4 tests): end-to-end with
stub LLM injected via monkey-patching `_resolve_llm`.
- Updated all existing `session tell` test steps to patch
`_build_session_workflow` returning a mock `TellResult` so tests
remain isolated from the LLM layer.
- Updated operation-count assertions (42 → 44) in
`a2a_cli_facade_integration`, `consolidated_misc`,
`m6_autonomy_acceptance` feature files and steps.
Cycle 7 (Review ID 8088) fixes:
- Blocker 4: Split `session_workflow.py` (was 801 lines) into two files:
`session_workflow.py` (467 lines) and `session_caller.py` (318 lines).
Extracted: `LangChainSessionCaller`, `extract_content`,
`extract_token_usage`, `estimate_cost`, `estimate_tokens`,
`history_to_langchain_messages`, and stub classes.
- Blocker 5: Route CLI streaming path through
`_facade_dispatch("message/stream", ...)` instead of directly
calling `workflow.tell_stream()`. The facade's
`_handle_message_stream` falls back to non-streaming with
`streamed: false` (acceptable for this milestone per the spec).
- Updated streaming test assertion to validate full response presence
(no longer checks token-by-token word positions, since the facade
fallback returns a complete message).
ISSUES CLOSED: #5784
106 lines
5.5 KiB
Gherkin
106 lines
5.5 KiB
Gherkin
Feature: Session workflow coverage boost
|
|
Coverage boost for session_workflow.py helper functions (M1, n4).
|
|
|
|
Background:
|
|
Given the session workflow coverage environment is set up
|
|
|
|
# _extract_content
|
|
Scenario: _extract_content extracts text from content attribute
|
|
Given coverage boost a mock response with content "hello world"
|
|
When coverage boost _extract_content is called
|
|
Then coverage boost the extracted result should be "hello world"
|
|
|
|
Scenario: _extract_content falls back to text attribute
|
|
Given coverage boost a mock response with text "from text attr" and no content
|
|
When coverage boost _extract_content is called
|
|
Then coverage boost the extracted result should be "from text attr"
|
|
|
|
Scenario: _extract_content handles list content
|
|
Given coverage boost a mock response with list content containing text dicts and plain strings
|
|
When coverage boost _extract_content is called
|
|
Then coverage boost the result should contain the concatenated texts
|
|
|
|
Scenario: _extract_content falls back to str for unknown types
|
|
Given coverage boost a mock response with no content or text attribute
|
|
When coverage boost _extract_content is called
|
|
Then coverage boost the result should be a string
|
|
|
|
# _extract_token_usage
|
|
Scenario: _extract_token_usage reads from response_metadata.usage
|
|
Given coverage boost a mock response with response_metadata usage input_tokens=100 output_tokens=50
|
|
When coverage boost _extract_token_usage is called
|
|
Then coverage boost input tokens should be 100 and output tokens should be 50
|
|
|
|
Scenario: _extract_token_usage reads from response_metadata.token_usage
|
|
Given coverage boost a mock response with response_metadata token_usage input_tokens=200 output_tokens=100
|
|
When coverage boost _extract_token_usage is called
|
|
Then coverage boost input tokens should be 200 and output tokens should be 100
|
|
|
|
Scenario: _extract_token_usage reads prompt_tokens and completion_tokens
|
|
Given coverage boost a mock response with response_metadata usage prompt_tokens=300 completion_tokens=150
|
|
When coverage boost _extract_token_usage is called
|
|
Then coverage boost input tokens should be 300 and output tokens should be 150
|
|
|
|
Scenario: _extract_token_usage reads from usage_metadata
|
|
Given coverage boost a mock response with usage_metadata input_tokens=400 output_tokens=200
|
|
When coverage boost _extract_token_usage is called
|
|
Then coverage boost input tokens should be 400 and output tokens should be 200
|
|
|
|
Scenario: _extract_token_usage returns zeros for no metadata
|
|
Given coverage boost a mock response with no usage metadata
|
|
When coverage boost _extract_token_usage is called
|
|
Then coverage boost input tokens should be 0 and output tokens should be 0
|
|
|
|
# _estimate_cost
|
|
Scenario: _estimate_cost computes cost from token counts
|
|
Given coverage boost input tokens 1000 and output tokens 500
|
|
When coverage boost _estimate_cost is called
|
|
Then coverage boost the estimated cost should be positive
|
|
|
|
# _history_to_langchain_messages
|
|
Scenario: _history_to_langchain_messages converts all roles
|
|
Given coverage boost session messages with roles SYSTEM, USER, ASSISTANT, and TOOL
|
|
When coverage boost _history_to_langchain_messages is called
|
|
Then coverage boost the result should contain SystemMessage, HumanMessage, AIMessage, and ToolMessage
|
|
|
|
Scenario: _history_to_langchain_messages treats unknown role as human
|
|
Given coverage boost a session message with an unknown role
|
|
When coverage boost _history_to_langchain_messages is called
|
|
Then coverage boost the result should contain a HumanMessage
|
|
|
|
Scenario: _history_to_langchain_messages handles empty list
|
|
Given coverage boost an empty list of session messages
|
|
When coverage boost _history_to_langchain_messages is called
|
|
Then coverage boost the result should be an empty list
|
|
|
|
# LangChainSessionCaller.invoke() tool_results branch
|
|
Scenario: LangChainSessionCaller.invoke appends tool results
|
|
Given coverage boost a LangChainSessionCaller with a stub LLM and empty history
|
|
When coverage boost invoke is called with tool_results containing one success and one failure
|
|
Then coverage boost the accumulated messages should include tool result messages
|
|
|
|
# LangChainSessionCaller.invoke() with tool_calls in response
|
|
Scenario: LangChainSessionCaller.invoke extracts tool calls from response
|
|
Given coverage boost a LangChainSessionCaller with a stub LLM that returns tool calls
|
|
When coverage boost invoke is called for the first time
|
|
Then coverage boost the LLMResponse should contain the extracted tool calls
|
|
|
|
# _build_lc_messages_from_history
|
|
Scenario: _build_lc_messages_from_history adds system prompt when absent
|
|
Given coverage boost a SessionWorkflow with a stub service and no registry
|
|
And coverage boost session history without a system message
|
|
When coverage boost _build_lc_messages_from_history is called with a prompt
|
|
Then coverage boost the first message should be a SystemMessage with the session system prompt
|
|
|
|
# _MinimalStubLLM
|
|
Scenario: _MinimalStubLLM.invoke returns stub response
|
|
Given coverage boost a _MinimalStubLLM instance
|
|
When coverage boost invoke on the stub is called
|
|
Then coverage boost the stub response content should be "(no LLM configured)"
|
|
And coverage boost the stub response should have empty tool_calls
|
|
|
|
Scenario: _MinimalStubLLM.stream yields stub chunk
|
|
Given coverage boost a _MinimalStubLLM instance
|
|
When coverage boost stream on the stub is called
|
|
Then coverage boost it should yield a chunk with content "(no LLM configured)"
|