Files
cleveractors-core/features/execute_stream.feature
T
CoreRasurae a8f68b8262
CI / lint (pull_request) Successful in 50s
CI / typecheck (pull_request) Successful in 55s
CI / security (pull_request) Successful in 53s
CI / quality (pull_request) Successful in 38s
CI / unit_tests (pull_request) Successful in 3m15s
CI / integration_tests (pull_request) Successful in 1m14s
CI / build (pull_request) Successful in 39s
CI / coverage (pull_request) Successful in 3m16s
CI / status-check (pull_request) Successful in 3s
CI / lint (push) Successful in 37s
CI / typecheck (push) Successful in 56s
CI / security (push) Successful in 54s
CI / quality (push) Successful in 37s
CI / unit_tests (push) Successful in 3m15s
CI / integration_tests (push) Successful in 1m14s
CI / build (push) Successful in 39s
CI / coverage (push) Successful in 3m16s
CI / status-check (push) Successful in 3s
fix(graph): enforce USD budget at each agent invocation with pre-flight gate
Enforces the USD budget (max_cost_usd) as a hard pre-flight gate before
every node execution and at each main-agent ainvoke round inside the
multi-turn tool loop, with pruning passes silently skipped when the
budget is exhausted.

Pre-flight gate (pure_graph.py):
  _check_budget_pre_flight() raises ExecutionError(kind='cost',
  reason='budget_exhausted') immediately before node execution in both
  _execute_from_node and all three branches of _stream_from_node.

In-node budget gating (llm.py):
  _check_budget_before_invoke() gates every _retry_ainvoke() call
  inside _execute_tool_loop, process_message, and stream_message.
  _should_skip_pruning() skips pruning passes when budget exhausted.
  _accumulate_cost_after_invoke() computes USD cost from token counts
  using the pricing table and advances current_accumulated_cost so
  subsequent budget checks see the updated cost.

ContextVar plumbing (retry.py, pure_graph.py):
  current_accumulated_cost, current_max_cost_usd, and current_pricing
  ContextVars carry per-node budget state from PureLangGraph into
  LLMAgent. Set before node.execute() via _set_budget_context_for_node,
  reset in finally blocks via _reset_budget_context.

Reviewer fixes (rui.hu #80):
  - Move cost accumulation inside try block in terminal streaming branch
    to prevent ContextVar leak (was set after finally reset).
  - Standardize log precision to $%.6f across all budget messages.
  - Remove orphaned 'already exhausted' step (criterion unreachable).
  - Add dedicated streaming budget enforcement scenario.
  - Remove redundant = None default on _reset_budget_context pricing_token.

ISSUES CLOSED: #76
2026-07-12 19:51:39 +01:00

1112 lines
69 KiB
Gherkin

Feature: Executor.execute_stream() - token-by-token streaming delivery
As the CleverThis router
I want to call executor.execute_stream(message) and receive tokens one-by-one
So that I can stream partial responses to end-users for a better UX on long LLM calls
Background:
Given the execute_stream test context is initialised
# ── Executor.last_result attribute ───────────────────────────────────────
Scenario: Executor.last_result is None before any execute_stream call
Given an Executor with a basic llm config (stream)
Then executor.last_result should be None (stream)
# ── LLM actor streaming ──────────────────────────────────────────────────
Scenario: execute_stream on LLM actor yields string tokens and populates last_result
Given an Executor with a basic llm config (stream)
And a mock astream yielding tokens "Hello", " World", "!" with no usage (stream)
When I call execute_stream with message "Hi" (stream)
Then the collected tokens should be ["Hello", " World", "!"] (stream)
And executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "Hello World!" (stream)
Scenario: execute_stream populates last_result token counts from final chunk usage_metadata
Given an Executor with a basic llm config (stream)
And a mock astream yielding single token "Hi" with usage prompt=5 completion=10 (stream)
When I call execute_stream with message "test" (stream)
Then executor.last_result.prompt_tokens should be 5 (stream)
And executor.last_result.completion_tokens should be 10 (stream)
Scenario: execute_stream last_result token counts fall back to 0 when no usage_metadata
Given an Executor with a basic llm config (stream)
And a mock astream yielding single token "response" with no usage (stream)
When I call execute_stream with message "test" (stream)
Then executor.last_result.prompt_tokens should be 0 (stream)
And executor.last_result.completion_tokens should be 0 (stream)
Scenario: execute_stream last_result is cleared to None when a new stream starts
Given an Executor with a basic llm config (stream)
And a mock astream yielding tokens "Hello", " World" with no usage (stream)
And a prior successful execute_stream has populated last_result (stream)
When I start iterating execute_stream and stop after first token (stream)
Then executor.last_result should be None (stream)
# ── M1 fix: timeout_ms enforced on LLM streaming path (AC5) ──────────────
Scenario: execute_stream on LLM actor raises ExecutionError when timeout_ms is exceeded
Given an LLM Executor with timeout_ms 50 and a slow mock astream (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for bool timeout_ms
Given an LLM Executor with bool_true timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for zero timeout_ms
Given an LLM Executor with zero timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for non-numeric timeout_ms
Given an LLM Executor with string timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream on LLM actor completes normally when within timeout_ms
Given an LLM Executor with timeout_ms 2000 and a fast mock astream (stream)
When I call execute_stream with message "test" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── M1 fix: max_cost_usd enforced on LLM streaming path (AC5) ────────────
Scenario: execute_stream on LLM actor raises ExecutionError when max_cost_usd is exceeded
Given an LLM Executor with max_cost_usd 0.0 and pricing and a token-bearing mock astream (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for bool max_cost_usd
Given an LLM Executor with bool max_cost_usd and pricing (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for non-numeric max_cost_usd
Given an LLM Executor with non-numeric max_cost_usd and pricing (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for missing provider pricing
Given an LLM Executor with max_cost_usd 1.0 and missing provider pricing (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for missing model pricing
Given an LLM Executor with max_cost_usd 1.0 and missing model pricing (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for incomplete pricing entry
Given an LLM Executor with max_cost_usd 1.0 and incomplete pricing entry (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor raises ExecutionError for invalid pricing rate
Given an LLM Executor with max_cost_usd 1.0 and invalid pricing rate (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on LLM actor completes normally when cost is within max_cost_usd
Given an LLM Executor with max_cost_usd 1.0 and pricing and a token-bearing mock astream (stream)
When I call execute_stream with message "test" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── m1 fix: executor.last_result populated when create_agent raises on LLM path ─
Scenario: execute_stream on LLM actor populates last_result with no_llm placeholder when create_agent raises ConfigurationError
Given an LLM Executor where create_agent raises ConfigurationError (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
And executor.last_result.nodes should contain a no_llm placeholder (stream)
# ── m2 fix: billing integrity when build_chat_model raises ConfigurationError ─
Scenario: execute_stream on LLM actor populates last_result with zero tokens when build_chat_model raises ConfigurationError during lazy init
Given an LLM Executor where build_chat_model raises ConfigurationError during lazy init (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
And executor.last_result token counts should be zero (stream)
# ── Graph actor streaming ─────────────────────────────────────────────────
Scenario: execute_stream on graph actor yields tokens from terminal node
Given an Executor with a graph config that has a single agent node (stream)
And a mock astream yielding tokens "Graph", " response" with no usage (stream)
When I call execute_stream with message "hello graph" (stream)
Then the collected tokens should be ["Graph", " response"] (stream)
And executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "Graph response" (stream)
Scenario: execute_stream on graph actor populates last_result with node usages
Given an Executor with a graph config that has a single agent node (stream)
And a mock astream yielding single token "Answer" with no usage (stream)
When I call execute_stream with message "test graph" (stream)
Then executor.last_result should have at least one NodeUsage entry (stream)
# NOTE: This scenario verifies the end-state only (final token list and
# concatenated response), not the internal buffering behaviour. Both the
# fast path (immediate yield) and the slow path (buffer-then-yield) produce
# the same final token list, so a regression that removes the buffering code
# and takes the fast path would not be caught here. The buffering path is
# exercised by the _all_edges_unconditional=False branch in pure_graph.py;
# the correctness of that branch is verified by the C1 fix assertion
# (executor.last_result.response == "Buffered") which would fail if
# full_response were not set from the LLM response.
Scenario: execute_stream on terminal AGENT node with conditional edge buffers tokens before yielding
Given an Executor with a graph config that has a terminal agent node with conditional edge (stream)
And a mock astream yielding tokens "Buf", "fered" with no usage (stream)
When I call execute_stream with message "test conditional" (stream)
Then the collected tokens should be ["Buf", "fered"] (stream)
And executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "Buffered" (stream)
# ── LLMAgent.stream_message() ──────────────────────────────────────────────
Scenario: LLMAgent.stream_message yields chunks from astream
Given an LLMAgent with a mock astream that yields "tok1", "tok2", "tok3"
When I call stream_message with a string message
Then the yielded tokens should be ["tok1", "tok2", "tok3"]
Scenario: LLMAgent.stream_message captures token counts from final chunk usage_metadata
Given an LLMAgent with a mock astream whose last chunk has usage_metadata prompt=7 completion=13
When I call stream_message with a string message
Then _last_token_usage should be (7, 13) after stream_message
And last_token_usage_var was (7, 13) inside the async step
Scenario: LLMAgent.stream_message resets token counts to (0,0) at the start
Given an LLMAgent with stale _last_token_usage (100, 200)
And a mock astream yielding empty sequence (stream-agent)
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
Scenario: LLMAgent.stream_message falls back to (0,0) with warning when no usage_metadata
Given an LLMAgent with a mock astream that yields a chunk with no usage_metadata
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
And a no-usage warning was emitted during stream_message
# N1: usage_metadata present but empty dict {}
Scenario: LLMAgent.stream_message falls back to (0,0) when usage_metadata is empty dict
Given an LLMAgent with a mock astream that yields a chunk with empty usage_metadata dict
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
And a no-usage warning was emitted during stream_message
# N2: response_metadata present but not a dict
Scenario: LLMAgent.stream_message falls back to (0,0) when response_metadata is not a dict
Given an LLMAgent with a mock astream that yields a chunk with non-dict response_metadata
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
And a no-usage warning was emitted during stream_message
# N2: response_metadata is a dict but has no token_usage key
Scenario: LLMAgent.stream_message falls back to (0,0) when response_metadata has no token_usage key
Given an LLMAgent with a mock astream that yields a chunk with response_metadata missing token_usage
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
And a no-usage warning was emitted during stream_message
# n5: _CAUSE_RESPONSE_METADATA_MISSING branch — chunk has neither usage_metadata
# nor response_metadata attributes (SimpleNamespace with content only)
Scenario: LLMAgent.stream_message falls back to (0,0) when chunk has no response_metadata attribute
Given an LLMAgent with a mock astream that yields a chunk with no response_metadata attribute
When I call stream_message with a string message
Then _last_token_usage should be zero after stream_message
And a no-usage warning was emitted during stream_message
# ── Limit enforcement for streaming ───────────────────────────────────────
Scenario: execute_stream on graph respects timeout_ms limit
Given an Executor with a graph config and timeout_ms 50 (stream)
And a mock astream that takes too long (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream on graph respects max_model_calls limit
Given an Executor with a graph config and max_model_calls 0 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "model_calls" should be raised (stream)
# ── Tool/multi-actor not supported ─────────────────────────────────────────
Scenario: execute_stream raises ConfigurationError for tool actor type
Given an Executor with a tool actor config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised about unsupported streaming type (stream)
Scenario: execute_stream raises ConfigurationError for multi_actor type
Given an Executor with a multi_actor config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised about unsupported streaming type (stream)
# ── Non-AGENT node path in _stream_from_node() ───────────────────────────
Scenario: execute_stream on graph with non-LLM terminal function node yields output
Given an Executor with a pure function graph config (stream)
When I call execute_stream with message "test function" (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result should have a no_llm placeholder node (stream)
Scenario: execute_stream on graph with function then agent node exercises non-agent path
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "Final answer" with no usage (stream)
When I call execute_stream with message "test mixed" (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "Final answer" (stream)
# ── execute_stream with conversation history (messages param) ─────────────
Scenario: execute_stream on LLM actor with conversation history messages
Given an Executor with a basic llm config (stream)
And a mock astream yielding single token "Reply" with no usage (stream)
When I run execute_stream with history on executor (stream)
Then the collected tokens should be ["Reply"] (stream)
And executor.last_result should be an ActorResult (stream)
# ── Node.stream_agent() error and non-LLM paths ──────────────────────────
Scenario: Node._stream_agent() propagates exception when streaming fails
Given a Node stream_agent that raises during streaming (stream)
Then the exception should be propagated from stream_agent (stream)
# ── _execute_llm_stream() with conversation history ───────────────────────
Scenario: _execute_llm_stream with messages conversation history sets last_result
Given an Executor with a basic llm config (stream)
And a mock astream yielding single token "ConvReply" with no usage (stream)
When I run execute_stream with messages param on executor (stream)
Then executor.last_result should be an ActorResult (stream)
# ── execute_stream with initial_state ────────────────────────────────────
Scenario: execute_stream on graph with initial_state restores graph state
Given an Executor with a graph config that has a single agent node (stream)
And a mock astream yielding single token "Resumed" with no usage (stream)
When I run execute_stream with initial_state on graph executor (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "Resumed" (stream)
# ── Invalid timeout_ms in execute_stream ─────────────────────────────────
Scenario: execute_stream with bool timeout_ms raises ExecutionError kind timeout
Given an Executor with a graph config and bool_true timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream with zero timeout_ms raises ExecutionError kind timeout
Given an Executor with a graph config and zero timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
# ── TOOL node limit enforcement in streaming ──────────────────────────────
Scenario: execute_stream on graph with TOOL node respects max_tool_calls 0 limit
Given an Executor with a tool-node graph and max_tool_calls 0 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "tool_calls" should be raised (stream)
# ── _execute_llm_stream() error paths ────────────────────────────────────
Scenario: _execute_llm_stream raises ExecutionError when stream_message raises
Given an Executor with a basic llm config (stream)
And a mock astream that raises RuntimeError (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised from _execute_llm_stream (stream)
# ── Node.stream_agent() without current_message ──────────────────────────
Scenario: Node.stream_agent() uses last user message when current_message absent
Given a Node stream_agent without current_message in state (stream)
Then stream_agent yields tokens from last user message context (stream)
# ── execute_stream with graph and conversation_history ───────────────────
Scenario: execute_stream on graph with messages parameter passes history
Given an Executor with a graph config that has a single agent node (stream)
And a mock astream yielding single token "WithHistory" with no usage (stream)
When I run execute_stream on graph with messages param (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "WithHistory" (stream)
# ── Two-AGENT sequential (intermediate AGENT path) ───────────────────────
Scenario: execute_stream on two-AGENT sequential graph yields terminal tokens only
Given an Executor with a two-agent sequential graph config (stream)
And a mock astream yielding single token "TerminalToken" with no usage (stream)
When I call execute_stream with message "test sequential" (stream)
Then the collected tokens should be ["TerminalToken"] (stream)
And executor.last_result should be an ActorResult (stream)
# ── Parallel non-AGENT execution path ────────────────────────────────────
Scenario: execute_stream on parallel function graph exercises parallel non-agent path
Given an Executor with a parallel function graph config (stream)
When I call execute_stream with message "test parallel fn" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── execute_stream() timing paths ────────────────────────────────────────
Scenario: execute_stream completes before timeout yields tokens normally
Given an Executor with a graph config and timeout_ms 2000 (stream)
And a mock astream yielding single token "FastToken" with no usage (stream)
When I call execute_stream with message "fast" (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "FastToken" (stream)
# ── _execute_llm_stream() config_block format ────────────────────────────
Scenario: _execute_llm_stream with config_block format covers fallback paths
Given an Executor with a config_block format llm config (stream)
And a mock astream yielding single token "BlockReply" with no usage (stream)
When I call execute_stream with message "test block" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── _execute_llm_stream() invalid config errors ───────────────────────────
Scenario: _execute_llm_stream with invalid temperature raises ConfigurationError
Given an Executor with invalid temperature llm config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
Scenario: _execute_llm_stream with invalid max_tokens raises ConfigurationError
Given an Executor with invalid max_tokens llm config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
# ── _execute_llm_stream() cleanup exception path ──────────────────────────
Scenario: _execute_llm_stream logs warning when cleanup raises exception
Given an Executor with a basic llm config (stream)
And a mock astream yielding single token "CleanupTest" with no usage (stream)
And the mock agent cleanup raises RuntimeError (stream)
When I call execute_stream with message "test cleanup" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── M2 fix: LLM path — executor.last_result populated on exception ────────
Scenario: _execute_llm_stream populates executor.last_result when ExecutionError is raised mid-stream
Given an Executor with a basic llm config (stream)
And a mock astream that raises ExecutionError mid-stream with usage prompt=10 completion=20 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised (stream)
And executor.last_result should be an ActorResult with llm token counts after the error (stream)
# ── execute() with initial_state ─────────────────────────────────────────
Scenario: PureLangGraph.execute() with initial_state restores metadata
Given a PureLangGraph with a single agent node and initial_state (stream)
And a mock astream that yields nothing (just for setup) (stream)
When I execute the graph with initial_state and a message (stream)
Then the execution result should contain the restored state key (stream)
# ── LLMAgent.stream_message() memory_enabled path ────────────────────────
Scenario: stream_message() with memory_enabled config uses memory for history
Given an LLMAgent with memory_enabled config and mock astream (stream)
When I call stream_message with a string message
Then the yielded tokens should be ["mem_token"]
# ── More streaming depth/limit error paths ───────────────────────────────
Scenario: execute_stream with non-numeric timeout_ms raises ExecutionError kind timeout
Given an Executor with a graph config and string timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
Scenario: execute_stream with valid max_depth limit executes normally
Given an Executor with a graph config and max_depth 100 (stream)
And a mock astream yielding single token "DepthOk" with no usage (stream)
When I call execute_stream with message "test depth" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream with bool max_depth raises ExecutionError kind depth
Given an Executor with a graph config and bool max_depth (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "depth" should be raised (stream)
Scenario: execute_stream with non-numeric max_depth raises ExecutionError kind depth
Given an Executor with a graph config and string max_depth (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "depth" should be raised (stream)
# ── AGENT parallel streaming path ────────────────────────────────────────
Scenario: execute_stream on parallel AGENT graph exercises AGENT parallel path
Given an Executor with a parallel AGENT graph config (stream)
And a mock astream yielding single token "ParallelAgentToken" with no usage (stream)
When I call execute_stream with message "test parallel agent" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── Node.stream_agent() non-LLM agent path ───────────────────────────────
Scenario: Node.stream_agent() with ToolAgent falls back to process_message
Given a Node stream_agent with a ToolAgent (stream)
Then stream_agent yields the ToolAgent response (stream)
# ── _execute_llm_stream() create_agent exception paths ───────────────────
Scenario: _execute_llm_stream raises ConfigurationError when create_agent fails
Given an Executor with a basic llm config (stream)
And the mock factory raises ConfigurationError on create_agent (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
# ── LLMAgent.stream_message() template path ──────────────────────────────
Scenario: stream_message() with template config renders template for message
Given an LLMAgent with template config and mock renderer (stream)
When I call stream_message with a string message
Then the yielded tokens should be ["template_token"]
# ── _execute_graph_stream() actors key and alternate config ──────────────
Scenario: _execute_graph_stream with context in config sets global_context
Given an Executor with a graph config with actor context (stream)
And a mock astream yielding single token "ContextToken" with no usage (stream)
When I call execute_stream with message "test context" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── execute_stream() already-running guard ──────────────────────────────
Scenario: execute_stream raises RuntimeError when graph is already running
Given a PureLangGraph already in running state (stream)
When I call execute_stream on the running graph (stream)
Then a RuntimeError should be raised from execute_stream (stream)
# ── _stream_from_node() node not found path ─────────────────────────────
Scenario: execute_stream on graph with missing node logs error and continues
Given a PureLangGraph with a dangling edge to nonexistent node (stream)
When I call execute_stream on the dangling graph (stream)
Then the stream completes without error (stream)
# ── _stream_from_node() AGENT streaming exception path ──────────────────
Scenario: _stream_from_node re-raises exception from stream_agent as ExecutionError
Given a PureLangGraph with AGENT node that has no agent configured (stream)
When I call execute_stream on the bad-agent graph (stream)
Then an ExecutionError should be raised (stream)
# ── Node.stream_agent() edge cases ──────────────────────────────────────
Scenario: stream_agent raises ValueError when no agent configured in node
Given a Node stream_agent with no agent configured (stream)
Then ValueError is raised from stream_agent (stream)
Scenario: stream_agent raises ValueError when agent not found in agents dict
Given a Node stream_agent with agent key missing from agents dict (stream)
Then ValueError is raised from stream_agent (stream)
Scenario: stream_agent handles empty state messages
Given a Node stream_agent with empty state messages (stream)
Then stream_agent yields from empty input (stream)
Scenario: stream_agent handles long history with truncation
Given a Node stream_agent with long conversation history (stream)
Then stream_agent processes truncated history (stream)
Scenario: stream_agent handles nested context dict in state metadata
Given a Node stream_agent with nested context in metadata (stream)
Then stream_agent processes nested context tokens (stream)
# ── _execute_llm_stream() other exception in create_agent ───────────────
Scenario: _execute_llm_stream wraps RuntimeError from create_agent in ExecutionError
Given an Executor with a basic llm config (stream)
And the mock factory raises RuntimeError on create_agent (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised from _execute_llm_stream (stream)
# ── execute_stream() with last_context fallback ─────────────────────────
Scenario: execute_stream uses last_context when no global_context provided
Given a PureLangGraph with last_context populated (stream)
And a mock astream yielding single token "CtxToken" with no usage (stream)
When I call execute_stream on graph with last context (stream)
Then the stream completes with result (stream)
# ── execute_stream() with context_manager ───────────────────────────────
Scenario: execute_stream with context_manager uses context_manager for setup
Given a PureLangGraph with context_manager set (stream)
And a mock astream yielding single token "CmToken" with no usage (stream)
When I call execute_stream on graph with context manager (stream)
Then the stream completes with result (stream)
# ── execute_stream() conversation_history empty content skip ────────────
Scenario: execute_stream skips empty content entries in conversation_history
Given a PureLangGraph with context_manager set (stream)
And a mock astream yielding single token "HistToken" with no usage (stream)
When I call execute_stream with empty content history entry (stream)
Then the stream completes with result (stream)
# ── _execute_llm_stream() with system_prompt at top level ────────────────
Scenario: _execute_llm_stream with top-level system_prompt covers sp branch
Given an Executor with system_prompt at top level llm config (stream)
And a mock astream yielding single token "SpToken" with no usage (stream)
When I call execute_stream with message "sp test" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── LLMAgent.stream_message() system prompt render exception ─────────────
Scenario: stream_message falls back to original system_message on render exception
Given an LLMAgent with failing render_string mock (stream)
When I call stream_message with a string message
Then the yielded tokens should be ["render_ex_tok"]
# ── _execute_llm_stream() ExecutionError re-raise path ────────────────────
Scenario: _execute_llm_stream re-raises ExecutionError from stream_message
Given an Executor with a basic llm config (stream)
And a mock astream that raises ExecutionError (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised from _execute_llm_stream (stream)
# ── Intermediate AGENT node coverage (M5 ainvoke path edge cases) ──────────
Scenario: execute_stream on two-AGENT sequential graph with non-dict last_msg covers branch
Given an Executor with a two-agent sequential graph config (stream)
And a mock astream yielding single token "TerminalToken2" with no usage (stream)
And the intermediate agent returns a non-dict last_msg (stream)
When I call execute_stream with message "test non-dict" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on two-AGENT sequential graph with no messages in result covers branch
Given an Executor with a two-agent sequential graph config (stream)
And a mock astream yielding single token "TerminalToken3" with no usage (stream)
And the intermediate agent returns a result with no messages key (stream)
When I call execute_stream with message "test no-messages" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on two-AGENT sequential graph with non-dict ainvoke result covers branch
Given an Executor with a two-agent sequential graph config (stream)
And a mock astream yielding single token "TerminalToken4" with no usage (stream)
And the intermediate agent returns a non-dict ainvoke result (stream)
When I call execute_stream with message "test non-dict-result" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on two-AGENT sequential graph with ainvoke exception covers branch
Given an Executor with a two-agent sequential graph config (stream)
And a mock astream yielding single token "TerminalToken5" with no usage (stream)
And the intermediate agent ainvoke raises RuntimeError (stream)
When I call execute_stream with message "test ainvoke-error" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on function-then-agent graph with non-dict last_msg covers non-agent branch
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "FnAgentToken" with no usage (stream)
And the function node returns a non-dict last_msg in messages (stream)
When I call execute_stream with message "test fn-non-dict" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on function-then-agent graph with no messages in function result covers branch
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "FnAgentToken2" with no usage (stream)
And the function node returns a result with no messages key (stream)
When I call execute_stream with message "test fn-no-messages" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on function-then-agent graph with non-dict function result covers branch
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "FnAgentToken3" with no usage (stream)
And the function node returns a non-dict result (stream)
When I call execute_stream with message "test fn-non-dict-result" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on function-then-agent graph with function exception covers branch
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "FnAgentToken4" with no usage (stream)
And the function node raises RuntimeError (stream)
When I call execute_stream with message "test fn-exception" (stream)
Then executor.last_result should be an ActorResult (stream)
Scenario: execute_stream on function-then-agent graph with function node token usage covers branch
Given an Executor with a function-then-agent graph config (stream)
And a mock astream yielding single token "FnAgentToken5" with no usage (stream)
And the function node returns a result with _node_token_usage (stream)
When I call execute_stream with message "test fn-token-usage" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── stream_message() memory truncation and billing integrity (M2/m1 coverage) ─
Scenario: stream_message() with memory_enabled truncates history when over max_history
Given an LLMAgent with memory_enabled and long history for truncation test (stream)
When I call stream_message with a string message
Then the yielded tokens should be ["mem_trunc_tok"]
Scenario: stream_message() billing integrity preserves tokens when post-stream step raises
Given an LLMAgent with memory_enabled that raises on update_memory (stream)
When I call stream_message with a string message
Then _last_token_usage should be (5, 10) after stream_message
# ── _temperature_override support in stream_message() (n5 / M1 coverage) ──
Scenario: stream_message() applies _temperature_override from context
Given an LLMAgent with temperature 0.7 and mock astream for override test (stream)
When I call stream_message with _temperature_override 0.1 in context (stream)
Then the temperature was 0.1 during streaming and restored to 0.7 afterwards (stream)
Scenario: stream_message() with invalid _temperature_override raises ConfigurationError
Given an LLMAgent with temperature 0.7 and mock astream for override test (stream)
When I call stream_message with _temperature_override "bad" in context (stream)
Then a ConfigurationError is raised from stream_message (stream)
# ── response_metadata fallback in stream_message() (n6 / M3 coverage) ─────
Scenario: stream_message() falls back to response_metadata token_usage when usage_metadata absent
Given an LLMAgent with a mock astream whose final chunk has response_metadata token_usage (stream)
When I call stream_message with a string message
Then _last_token_usage should be (3, 7) after stream_message
# ── PureLangGraph.execute() with non-numeric timeout_ms ──────────────────
Scenario: PureLangGraph.execute() with non-numeric timeout_ms raises ExecutionError
Given a PureLangGraph with string timeout_ms limit (stream)
When I call execute on the pure graph with a message (stream)
Then an ExecutionError should be raised from execute (stream)
# ── C1: auto_finish_active bypass in streaming loop detection ────────────
Scenario: execute_stream with auto_finish_active bypasses loop detection for repeated node visits
Given a PureLangGraph with auto_finish_active set in state and a repeated-visit agent node (stream)
And a mock astream yielding single token "SectionToken" with no usage (stream)
When I call execute_stream on the auto_finish graph (stream)
Then the stream completes with result (stream)
# ── C2: Router-agent ping-pong detection in streaming path ───────────────
Scenario: execute_stream detects router-agent ping-pong and stops streaming
Given a PureLangGraph with router-agent ping-pong setup for streaming (stream)
When I call execute_stream on the ping-pong graph (stream)
Then the stream completes without error (stream)
# ── C3: No routing command → return to user guard in streaming path ───────
Scenario: execute_stream on agent-then-router graph with no routing command returns output to user
Given an Executor with an agent-then-router graph config and no routing command (stream)
And a mock astream yielding single token "DirectAnswer" with no usage (stream)
When I call execute_stream with message "hello" (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.response should equal "DirectAnswer" (stream)
# ── M2: executor.last_result populated on exception path ─────────────────
Scenario: execute_stream populates executor.last_result even when ExecutionError is raised mid-stream
Given an Executor with a two-agent sequential graph and max_model_calls 1 (stream)
And a mock astream yielding single token "PartialToken" with usage prompt=10 completion=20 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "model_calls" should be raised (stream)
And executor.last_result should be an ActorResult with token counts after the error (stream)
Scenario: execute_stream handles state-capture failure gracefully on exception path
Given an Executor with a graph that raises and has broken state capture (stream)
When I attempt execute_stream with broken state capture expecting an error (stream)
Then an ExecutionError should be raised (stream)
# ── M3: Cost-limit enforcement in streaming path ─────────────────────────
Scenario: execute_stream raises ExecutionError when cost limit is exceeded
Given an Executor with a graph config and max_cost_usd 0.0 with pricing (stream)
And a mock astream yielding single token "CostToken" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when pricing rate is invalid (non-numeric)
Given an Executor with a graph config and invalid pricing rate (stream)
And a mock astream yielding single token "RateToken" with usage prompt=100 completion=100 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when max_cost_usd is invalid (non-numeric string)
Given an Executor with a graph config and non-numeric max_cost_usd (stream)
And a mock astream yielding single token "BadCostToken" with usage prompt=100 completion=100 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
# ── M4: _collect_stream_tokens forwards depth ────────────────────────────
# The M4 fix forwards depth+1 to _collect_stream_tokens so that parallel
# children receive the correct depth and can trigger the depth limit.
# With max_depth=1 and a graph: start → agent_start → {agent_a, agent_b} → end,
# agent_start is at depth 1 (≤ max_depth=1, OK), but agent_a and agent_b are
# at depth 2 (> max_depth=1), triggering the depth error in the parallel branch.
Scenario: execute_stream on parallel AGENT graph with max_depth 1 raises depth error in parallel children
Given an Executor with a parallel AGENT graph config and max_depth 1 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "depth" should be raised (stream)
# ── C1 additional: loop detection stop path (auto_finish_active=False) ────
Scenario: _stream_from_node stops when node visited twice without auto_finish_active
Given a PureLangGraph with repeated-visit agent node and no auto_finish_active (stream)
When I call _stream_from_node on the repeated-visit graph (stream)
Then the stream completes without error (stream)
# ── C2 additional: ping-pong stop path ────────────────────────────────────
Scenario: _stream_from_node stops on router-agent ping-pong without auto_finish_active
Given a PureLangGraph with router-agent ping-pong setup for streaming (stream)
When I call execute_stream on the ping-pong graph (stream)
Then the stream completes without error (stream)
# ── C3 additional: no-routing-command guard in intermediate AGENT branch ──
Scenario: execute_stream on intermediate-agent-then-router graph with no routing command returns output
Given a PureLangGraph with intermediate agent then router and no routing command (stream)
And a mock astream yielding single token "IntermediateAnswer" with no usage (stream)
When I call execute_stream on the intermediate-router graph (stream)
Then the stream completes with result (stream)
# ── M3 additional: cost enforcement with missing pricing entry ────────────
Scenario: execute_stream raises ExecutionError when pricing entry is missing for provider
Given an Executor with a graph config and max_cost_usd 1.0 with missing pricing (stream)
And a mock astream yielding single token "CostToken2" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
# ── C3 intermediate branch: agent with non-END static successor routes to router ─
# Note: agent1 has edges to both "end" (static) and "router" (conditional:always).
# Because "router" is a non-END static successor, _statically_terminal=False for
# agent1, so this exercises the intermediate AGENT branch where the C3 guard fires.
Scenario: _stream_from_node C3 guard fires in intermediate AGENT branch with conditional router edge
Given a PureLangGraph with terminal agent and conditional router edge (stream)
And a mock astream yielding single token "ConditionalAnswer" with no usage (stream)
When I call execute_stream on the conditional-router graph (stream)
Then the stream completes with result (stream)
# ── C1 nested context: auto_finish_active in nested context dict ──────────
Scenario: _stream_from_node C1 bypass works when auto_finish_active is in nested context
Given a PureLangGraph with auto_finish_active in nested context and repeated-visit node (stream)
When I call _stream_from_node on the repeated-visit graph (stream)
Then the stream completes without error (stream)
# ── C2 nested context: auto_finish_active in nested context for ping-pong ─
Scenario: _stream_from_node C2 bypass works when auto_finish_active is in nested context
Given a PureLangGraph with auto_finish_active nested context and ping-pong setup (stream)
When I call execute_stream on the ping-pong graph (stream)
Then the stream completes without error (stream)
# ── Intermediate AGENT dynamically terminal path ─────────────────────────
Scenario: _stream_from_node intermediate AGENT yields when conditional edge resolves to END
Given a PureLangGraph with intermediate agent and conditional edge to END (stream)
When I call execute_stream on the conditional-end graph (stream)
Then the stream completes with result (stream)
And the collected tokens should be ["DynTerminalToken"] (stream)
# ── Streaming bool/invalid limit enforcement ─────────────────────────────
Scenario: execute_stream raises ExecutionError for bool max_model_calls in streaming
Given an Executor with a graph config and bool max_model_calls (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "model_calls" should be raised (stream)
Scenario: execute_stream raises ExecutionError for bool max_tool_calls in streaming
Given an Executor with a tool-node graph and bool max_tool_calls (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "tool_calls" should be raised (stream)
# ── M3 cost enforcement: missing model pricing entry ─────────────────────
Scenario: execute_stream raises ExecutionError when model pricing entry is missing
Given an Executor with a graph config and max_cost_usd 1.0 with missing model pricing (stream)
And a mock astream yielding single token "CostToken3" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
# ── Streaming invalid (non-numeric) limit enforcement ────────────────────
Scenario: execute_stream raises ExecutionError for string max_model_calls in streaming
Given an Executor with a graph config and string max_model_calls (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "model_calls" should be raised (stream)
Scenario: execute_stream raises ExecutionError for string max_tool_calls in streaming
Given an Executor with a tool-node graph and string max_tool_calls (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "tool_calls" should be raised (stream)
# ── M3 cost enforcement: incomplete and invalid pricing ──────────────────
Scenario: execute_stream raises ExecutionError when pricing entry is incomplete
Given an Executor with a graph config and max_cost_usd 1.0 with incomplete pricing (stream)
And a mock astream yielding single token "CostToken4" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError for bool max_cost_usd in streaming
Given an Executor with a graph config and bool max_cost_usd with pricing (stream)
And a mock astream yielding single token "CostToken5" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
# ── Intermediate AGENT parallel streaming path ───────────────────────────
Scenario: execute_stream on parallel intermediate AGENT graph yields tokens
Given an Executor with a parallel intermediate AGENT graph config (stream)
And a mock astream yielding single token "ParallelIntToken" with no usage (stream)
When I call execute_stream with message "test parallel int" (stream)
Then executor.last_result should be an ActorResult (stream)
# ── Major #1 fix: max_cost_usd enforced for intermediate AGENT nodes ──────
Scenario: Streaming graph exhausts budget via post-node check on first node
Given an Executor with a two-agent sequential graph and max_cost_usd 0.40 with pricing (stream)
And a mock astream yielding single token "BudgetNodeCost" with usage prompt=3000000 completion=0 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" and reason "budget_exhausted" should be raised (stream kind+reason)
Scenario: Second streaming node blocked by pre-flight when first node exactly hits limit
Given an Executor with a two-agent sequential graph and max_cost_usd 0.30 at 0.15/M prompt pricing (stream)
And a mock astream yielding single token "ExactPreFlight" with usage prompt=2000000 completion=0 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" and reason "budget_exhausted" should be raised (stream kind+reason)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node exceeds max_cost_usd
Given an Executor with a two-agent sequential graph and max_cost_usd 0.0 with pricing (stream)
And a mock astream yielding single token "IntCostToken" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has missing provider pricing
Given an Executor with a two-agent sequential graph and max_cost_usd 1.0 with missing provider pricing (stream)
And a mock astream yielding single token "IntCostToken2" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has missing model pricing
Given an Executor with a two-agent sequential graph and max_cost_usd 1.0 with missing model pricing (stream)
And a mock astream yielding single token "IntCostToken3" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has incomplete pricing
Given an Executor with a two-agent sequential graph and max_cost_usd 1.0 with incomplete pricing (stream)
And a mock astream yielding single token "IntCostToken4" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has invalid pricing rate
Given an Executor with a two-agent sequential graph and max_cost_usd 1.0 with invalid pricing rate (stream)
And a mock astream yielding single token "IntCostToken5" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has bool max_cost_usd
Given an Executor with a two-agent sequential graph and bool max_cost_usd with pricing (stream)
And a mock astream yielding single token "IntCostToken6" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
Scenario: execute_stream raises ExecutionError when intermediate AGENT node has non-numeric max_cost_usd
Given an Executor with a two-agent sequential graph and non-numeric max_cost_usd with pricing (stream)
And a mock astream yielding single token "IntCostToken7" with usage prompt=1000 completion=1000 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" should be raised (stream)
# ── Major #2 fix: state not polluted on agent failure ────────────────────
Scenario: execute_stream does not persist user input as assistant message when streaming fails
Given a PureLangGraph with terminal AGENT node that fails during streaming (stream)
When I call execute_stream on the failing-agent graph (stream)
Then the graph state messages should not contain the user input as an assistant message (stream)
# ── M1 fix: graph path — executor.last_result populated for unexpected exceptions ──
Scenario: execute_stream on graph actor populates executor.last_result when unexpected RuntimeError is raised from execute_stream
Given an Executor with a graph config that has a single agent node (stream)
And the graph execute_stream raises RuntimeError unexpectedly (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised (stream)
And executor.last_result should be an ActorResult with no_llm placeholder (stream)
Scenario: execute_stream on graph actor handles broken state capture gracefully on unexpected exception
Given an Executor with a graph config that has a single agent node (stream)
And the graph execute_stream raises RuntimeError with broken state capture (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised (stream)
And executor.last_result should be an ActorResult with no_llm placeholder (stream)
# ── M2 fix: graph path — executor.last_result populated when AgentCreationError ──
Scenario: execute_stream on graph actor populates executor.last_result with no_llm placeholder when AgentFactory.create_agent raises ConfigurationError
Given an Executor with a graph config that has a single agent node (stream)
And the mock graph factory raises ConfigurationError on create_agent (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult with no_llm placeholder (stream)
# ── m4 fix: empty astream with memory_enabled=True ───────────────────────
Scenario: stream_message with memory_enabled and empty astream stores empty assistant entry
Given an LLMAgent with memory_enabled and empty astream (stream)
When I call stream_message with a string message
Then the assistant memory entry should have empty content (stream)
# ── m5 fix: partial-stream abandonment leaves _last_token_usage at (0, 0) ─
Scenario: execute_stream abandonment leaves _last_token_usage at (0, 0)
Given an Executor with a basic llm config (stream)
And a mock astream yielding tokens "Hello", " World" with no usage (stream)
When I start iterating execute_stream and stop after first token (stream)
Then executor.last_result should be None (stream)
And the agent _last_token_usage should be (0, 0) after abandonment (stream)
# ── m6 fix: executor.last_result.state verified in graph streaming success ─
Scenario: execute_stream on graph with initial_state populates last_result.state
Given an Executor with a graph config that has a single agent node (stream)
And a mock astream yielding single token "Resumed" with no usage (stream)
When I run execute_stream with initial_state on graph executor (stream)
Then executor.last_result should be an ActorResult (stream)
And executor.last_result.state should be populated with initial_state keys (stream)
# ── Issue 6: _execute_llm_stream generic-exception path (billing integrity) ─
# The except Exception block in _execute_llm_stream handles unexpected
# non-ExecutionError exceptions and must populate executor.last_result before
# re-raising as ExecutionError. This mirrors the graph path's coverage.
Scenario: _execute_llm_stream populates executor.last_result when unexpected RuntimeError is raised during streaming
Given an Executor with a basic llm config (stream)
And a mock astream that raises RuntimeError (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
# ── Issue 1: billing-integrity for early config-validation errors in _execute_llm_stream ─
# The temperature/max_tokens/timeout_ms validations fire before the main
# try/except block. executor.last_result must be populated even for these
# early errors to satisfy the billing-integrity guarantee.
Scenario: _execute_llm_stream with invalid temperature populates executor.last_result
Given an Executor with invalid temperature llm config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: _execute_llm_stream with invalid max_tokens populates executor.last_result
Given an Executor with invalid max_tokens llm config (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
Scenario: _execute_llm_stream with bool timeout_ms populates executor.last_result
Given an LLM Executor with bool_true timeout_ms (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "timeout" should be raised (stream)
And executor.last_result should be an ActorResult (stream)
# ── Issue 2: billing-integrity for early config-validation errors in _execute_graph_stream ─
# The node/edge validation loop fires before the main try/except block.
# executor.last_result must be populated even for these early errors.
Scenario: _execute_graph_stream with invalid node definition populates executor.last_result
Given an Executor with a graph config that has an invalid node definition (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult with no_llm placeholder (stream)
Scenario: _execute_graph_stream with duplicate node ID populates executor.last_result
Given an Executor with a graph config that has a duplicate node ID (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then a ConfigurationError should be raised (stream)
And executor.last_result should be an ActorResult with no_llm placeholder (stream)
# ── Issue 3: GOTO_/ROUTE_ routing commands parsed in streaming path ───────
# The streaming path must parse GOTO_*/ROUTE_* prefixes from agent output
# and store the target in state.metadata["next_node"], mirroring
# _execute_from_node. Without this, graphs relying on LLM-emitted routing
# commands route correctly under execute() but incorrectly under
# execute_stream().
Scenario: _stream_from_node parses GOTO_ routing command from terminal AGENT output
Given a PureLangGraph with terminal AGENT node that emits a GOTO_ routing command (stream)
When I call execute_stream on the GOTO-routing graph (stream)
Then state.metadata next_node should be set from the GOTO_ command (stream)
Scenario: _stream_from_node parses GOTO_ routing command from intermediate AGENT output
Given a PureLangGraph with intermediate AGENT node that emits a GOTO_ routing command (stream)
When I call execute_stream on the intermediate GOTO-routing graph (stream)
Then state.metadata next_node should be set from the intermediate GOTO_ command (stream)
# ── Issue 1 (review round 7): GOTO_/ROUTE_ parsing in non-AGENT branch ────
# The non-AGENT branch of _stream_from_node() was missing GOTO_/ROUTE_ parsing
# that exists in _execute_from_node() for all node types. A function/tool node
# that emits a routing command (e.g. "GOTO_validation:...") must set
# state.metadata["next_node"] under execute_stream() just as it does under
# execute().
Scenario: _stream_from_node parses GOTO_ routing command from non-AGENT function node output
Given a PureLangGraph with non-AGENT function node that emits a GOTO_ routing command (stream)
When I call execute_stream on the non-AGENT GOTO-routing graph (stream)
Then state.metadata next_node should be set from the non-AGENT GOTO_ command (stream)
# ── Issue 2 (review round 7): str(None) guard in intermediate AGENT branch ─
# When an intermediate AGENT node returns None content (e.g. last_msg["content"]
# is None), the two yield sites in _stream_from_node() must yield "" instead of
# the literal string "None".
Scenario: _stream_from_node intermediate AGENT with None content yields empty string not "None"
Given a PureLangGraph with intermediate AGENT node that returns None content (stream)
When I call execute_stream on the None-content intermediate graph (stream)
Then the collected tokens should not contain the literal string "None" (stream)
# ── Issue 5 (review round 7): LangChainException arm in stream_message() ──
# The except LangChainException arm added in review round 5 has no test.
# A regression removing this arm would not be caught. This scenario injects
# a LangChainException into astream() and asserts:
# (a) ExecutionError is raised, (b) the captured log contains the streaming
# error message.
Scenario: stream_message raises ExecutionError when astream raises LangChainException
Given an Executor with a basic llm config (stream)
And a mock astream that raises LangChainException (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError should be raised (stream)
And executor.last_result should be an ActorResult (stream)
# ── Issue 4 (review round 7): resource-leak test verifies cleanup() called ─
# The existing abandonment test only checks last_result is None and
# _last_token_usage == (0, 0) — both true regardless of whether cleanup() was
# called. This scenario adds an explicit assertion that agent.cleanup was
# awaited after the caller abandoned the stream.
Scenario: execute_stream abandonment calls agent.cleanup() promptly
Given an Executor with a basic llm config (stream)
And a mock astream yielding tokens "Hello", " World" with no usage (stream)
When I start iterating execute_stream and stop after first token verifying cleanup (stream)
Then agent.cleanup should have been called (stream)
# ── Issue 8 (review round 7): ExecutionError.reason asserted in cost scenarios ─
# The existing cost-error step only asserts error.kind. A regression that
# always raises ExecutionError(kind="cost", reason="budget_exhausted") would
# pass all cost scenarios. This scenario explicitly checks the reason field.
Scenario: execute_stream raises ExecutionError with reason budget_exhausted when cost limit exceeded
Given an Executor with a basic llm config and zero cost limit (stream)
And a mock astream yielding single token "Hi" with usage prompt=10 completion=5 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" and reason "budget_exhausted" should be raised (stream kind+reason)
Scenario: execute_stream raises ExecutionError with reason missing_pricing_entry when provider missing
Given an Executor with a basic llm config and pricing but missing provider entry (stream)
And a mock astream yielding single token "Hi" with usage prompt=10 completion=5 (stream)
When I attempt execute_stream expecting an error with message "test" (stream)
Then an ExecutionError with kind "cost" and reason "missing_pricing_entry" should be raised (stream kind+reason)