The `shell` and `http_request` built-in tools are the only two whose
execution actually honors the tool agent's `timeout` config field
(ToolAgent._execute_shell_command, ToolAgent._http_request_tool), but
their LLM-facing schemas (_BUILTIN_TOOL_SCHEMAS in llm_tools.py) never
told the model a timeout applied, and neither tool's timeout error
named the config field that controls it.
- _BUILTIN_TOOL_SCHEMAS["shell"] and ["http_request"] descriptions now
state that execution is bounded by the agent's configured `timeout`
(seconds), matching the existing precedent of documenting `file_read`'s
`max_chars` behavior in its schema.
- _execute_shell_command's timeout ExecutionError now names `timeout`
as the adjustable config field, in addition to the elapsed seconds it
already reported.
- _http_request_tool now catches asyncio.TimeoutError distinctly from
the generic `except Exception`, so a request timeout produces a
message naming the config field instead of the previous generic
"HTTP request failed" text (which gave no indication a timeout was
the cause). Non-timeout HTTP failures are unaffected — same
exception handling, same message shape as before.
No change to enforcement mechanics: `self.timeout` remains the sole
config field (default 1s, Actor Configuration Standard §4.5). This is
schema/message content only, within the extension precedent already
established by ADR-2030 (D-1 schema registry, D-4/D-6 error-quality
improvements) — no new ADR was needed for this change.
Tests: extended features/tool_agent.feature (updated the two existing
shell/http_request timeout scenarios to also assert the new config-field
hint) and features/llm_tools_coverage.feature (new scenarios asserting
both schema descriptions mention the timeout budget). Full suite:
141 features / 2782 scenarios / 13172 steps passed; coverage 96.8%
(>=96.5% threshold). Robot integration_tests: 27/27 suites passed.
ASV benchmark_regression: no significant change, as expected for a
messaging-only change.
ISSUES CLOSED: #82
Implements the complete LLM agent tool-calling pipeline — the tool-calling
path was broken in multiple ways: tools from agent config were not passed
to the LLM, tool execution was single-pass (the model could not see tool
results or make follow-up calls), tool errors were discarded, and shell/
python_exec tools were never registered.
Key changes:
- Multi-turn tool-call loop: passes tools to the LLM via ainvoke(tools=...)
and re-invokes after each batch of ToolResults, with a configurable
round limit (tool_max_rounds config / TOOL_MAX_ROUNDS env var, default 20).
- Tool schema module (llm_tools.py): normalize_tool_entry() converts
string tool names and config dicts to OpenAI function-calling format
with full parameter schemas and LLM-facing guidance (max_chars for
file_read, sandbox builtins for python_exec, etc.).
- file_read enhancements: directory listing with file sizes and type
indicators, max_chars truncation to prevent context overflow, shell
command detection with redirect to shell tool, file-not-found
recovery hints with parent directory listing.
- python_exec sandbox: stdout capture so print() produces visible
output, NameError guidance directing LLM to file_read/file_write
instead of using open().
- shell/python_exec tool registration in builtin_tools when
allow_shell / exec_python is enabled.
- create_subprocess_shell for shell commands (was create_subprocess_exec,
which blocks pipes/redirections/chains).
- Tool error propagation: ExecutionError/ConfigurationError details
included in ToolMessages for LLM self-correction.
- Stuck-model recovery: injects synthesizing prompt when model
exhausts tool rounds without producing content.
- LLM provider error details preserved in ExecutionError messages
instead of generic "LLM processing failed".
- Empty message filtering in _prepare_conversation_history() prevents
whitespace-only messages from polluting downstream conversation
history in multi-agent pipelines.
Closes#59