fix(actors): llm token consumption undercounted #66
Merged
CoreRasurae
merged 1 commits from 2026-07-01 14:44:39 +00:00
bugfix/m2-lllm-token-consumption-undercounted into master
Dismiss Review
Are you sure you want to dismiss this review?
Labels
Clear labels
auto/blocked-by-deps
auto/ci-timeout
auto/claimed-implementer
auto/claimed-merge
auto/claimed-reviewer
auto/driver-down
auto/invariant-violation
auto/last-attempt-tier-0
auto/last-attempt-tier-1
auto/last-attempt-tier-2
auto/last-attempt-tier-min
Automation Tracking
auto/needs-conflict-resolution
auto/needs-implementer
auto/postmortem
auto/ready-to-merge
auto/restart-throttled
auto/revert
auto/sentinel
auto/stale-inactivity
auto/unstable
Blocked
Needs Feedback
Signed-off: Owner
Signed-off: Scrum Master
Signed-off: Tech Lead
Spike
PR blocked by an open issue dependency. Operator must close the dep (or remove the dependency link) before the merge driver can act. Auto-cleared by merge_drive when no open deps remain.
Most recent merge cycle hit CI timeout. Driver excludes this PR while last merge_cycle row is < 30 min old; label persists thereafter as visible history.
Currently being processed by an implementer worker.
Currently being processed by the merge driver.
Currently being processed by a reviewer worker.
Merge driver heartbeat stale; pipeline halted. Closed automatically on next clean tick.
Detected master commit violating the strict merge invariant. Tracked as an issue (not a PR label); kept here for label completeness.
In-cycle escalation: most recent attempt ran at the Tier 0 slot (`tier-0`). Slot's model defined in .opencode/models/tiers.yaml.
In-cycle escalation: most recent attempt ran at the Tier 1 slot (`tier-1`). Slot's model defined in .opencode/models/tiers.yaml.
In-cycle escalation: most recent attempt ran at the Tier 2 slot (`tier-2`). Slot's model defined in .opencode/models/tiers.yaml. Gated behind IMPLEMENTER_ESCALATION_TIER2_ENABLED.
In-cycle escalation: most recent attempt ran at the Tier -1 slot (`tier-min`). Slot's model defined in .opencode/models/tiers.yaml. Suffix is ``-min`` (not ``--1``) so the Forgejo UI reads naturally.
Tracking issues used by the AI Automation system for agents to communicate and report.
Rebase conflict needs LLM conflict-resolver.
Failing CI needs implementer attention.
Documenting a driver incident or rollback.
Reviewer has APPROVED this PR and no later REQUEST_CHANGES is outstanding. The merge driver requires this label to even consider a PR for merging. Set by the reviewer worker on APPROVE; cleared on REQUEST_CHANGES.
Train repeatedly lost master-tempo races. Driver excludes via merge_cycle until cooldown elapses; label persists as visible history.
Revert PR backing out an invariant violation. Fast-tracked through the merge driver.
Sentinel PR duplicated from upstream into a personal fork by tools/duplicate_prs_to_fork.py for pipeline testing. Lives only in the fork; the canonical pipeline never sees it.
No implementer activity for N days. Flagged for human review. Auto-cleared on next push to head branch.
Repeatedly fails on current master (>= 3 ci-fail-on-rebased-sha releases in 12 h). Excluded from driver until human triage.
A ticket in a blocked state and unable to complete until some other task is completed first.
Bounty
$100
A bounty of $100 for any open-source contributor who provides a MR that solves this issue
Bounty
$1000
A bounty of $1000 for any open-source contributor who provides a MR that solves this issue
Bounty
$10000
A bounty of $10000 for any open-source contributor who provides a MR that solves this issue
Bounty
$20
A bounty of $20 for any open-source contributor who provides a MR that solves this issue
Bounty
$2000
A bounty of $2000 for any open-source contributor who provides a MR that solves this issue
Bounty
$250
A bounty of $250 for any open-source contributor who provides a MR that solves this issue
Bounty
$50
A bounty of $50 for any open-source contributor who provides a MR that solves this issue
Bounty
$500
A bounty of $500 for any open-source contributor who provides a MR that solves this issue
Bounty
$5000
A bounty of $5000 for any open-source contributor who provides a MR that solves this issue
Bounty
$750
A bounty of $750 for any open-source contributor who provides a MR that solves this issue
MoSCoW
Could have
Could have feature in order to satisfy the epic/legendary.
MoSCoW
Must have
Must have feature in order to satisfy the epic/legendary.
MoSCoW
Should have
Should have feature in order to satisfy the epic/legendary.
There are questions in the ticket that can not be completed until the project owner provides clarity.
Points
1
1 man-hours worth of work for an expert with no learning curve.
Points
13
13 man-hours worth of work for an expert with no learning curve.
Points
2
2 man-hours worth of work for an expert with no learning curve.
Points
21
21 man-hours worth of work for an expert with no learning curve.
Points
3
3 man-hours worth of work for an expert with no learning curve.
Points
34
34 man-hours worth of work for an expert with no learning curve.
Points
5
5 man-hours worth of work for an expert with no learning curve.
Points
55
55 man-hours worth of work for an expert with no learning curve.
Points
8
8 man-hours worth of work for an expert with no learning curve.
Points
88
88 man-hours worth of work for an expert with no learning curve.
Priority
Backlog
This ticket has backlogged priority and is not to be worked on yet
Priority
CI Blocker
Critical priority issue that blocks CI/CD pipeline and prevents PR merges
Priority
Critical
The priority is critical
Priority
High
The priority is high
Priority
Low
The priority is low
Priority
Medium
The priority is medium
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
When an epic or legendary is in review it must be signed off by owner, tech lead, and scrum master before being marked as completed.
A ticket for learning a tool or technology that is needed to be able to do future planning and design.
State
Completed
The ticket has been fully implemented, completed, and merged with the source code. This label should only be applied once a ticket is closed.
State
Duplicate
A ticket that represents the same content as an existing ticket.
State
In Progress
A ticket that is actively being developed.
State
In Review
A ticket that has had some code completed to implement but is waiting to pass peer review and is not yet merged in.
State
Paused
This ticket's work started but wasn't finished. It's on hold (likely in a feature branch) and will be resumed later, either due to a blocker or a delay.
State
Unverified
All new tickets start in this state. A developer may set it to show the ticket is unverified. This means we haven't agreed to work on it. It will either move to a verified state or be closed as wontdo.
State
Verified
The issue has been verified by a developer as legitimate. It will be worked on and verified tickets are now considered part of the backlog.
State
Wont Do
This ticket has been decided it wont be done. This may mean the bug has been determined to not be real (cant verify) or the feature is one we have decided we dont want to adopt.
Type
Automation
Any edits or discussion about the AI automated coding system.
Type
Bug
Something that doesnt work as intended.
Type
Discussion
Anytime a ticket represents a discussion about a subject and doesnt fall into one of the other categories.
Type
Documentation
An error or improvement needed in the documentation.
Type
Epic
Any first tier epic. That is, an epic which contains only issues as children and will not have sub-epics.
Type
Feature
Some new functionality not present.
Type
Legendary
A type of Epic which will contain other Epics.
Type
Refactor
A code change that restructures existing code without changing its external behavior.
Type
Support
Someone needs help using the project.
Type
Task
A generic task that doesnt fit into the other type categories.
Type
Testing
Work exclusively focusing on fixing or expanding testing.
No Label
Milestone
No items
No Milestone
Projects
Clear projects
No project
Assignees
aditya (Aditya Chhabra)
aleenaumair (Aleena Umair)
brent.edwards (Brent Edwards)
CoreRasurae (Luis Mendes)
drew (Drew Morris)
eugen.thaci (Eugen Thaci)
freemo (Jeffrey Phillips Freeman)
HAL9000 (HAL 9000)
HAL9001 (HAL9001)
hamza.khyari (Hamza Khyari)
hurui200320 (Rui Hu)
justin.morris
khird (Kyle Hird)
org.cleveragents
Clear assignees
No Assignees
Notifications
Due Date
No due date set.
Blocks
#65 bug(agents): LLM token consumption undercounted — pruning model calls untracked, earlier loop rounds overwritten
cleveragents/cleveractors-core
Reference: cleveragents/cleveractors-core#66
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Delete Branch "bugfix/m2-lllm-token-consumption-undercounted"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Three defects in LLMAgent.process_message() caused token consumption to be undercounted in NodeUsage/ActorResult:
_run_pruning_pass()now returns usage metadata alongside parsed content_estimate_token_count()blends actuals with heuristic for tool-call JSON overheadTemporary Limitations (deviation from ADR-2031)
pruning_modelmust matchmodel—ConfigurationErrorraised at agent creation if they differ, because cross-model token tracking is not yet plumbed throughNodeUsage.usage_metadata, aRuntimeErroris raised immediately (no silent fallback to(0, 0)).Closes #65.
Dependencies
None.
432b161ae0tocf2a32ce92b6cbc57908to80a5d80e06305bf2503etoe88332ea3ae88332ea3ato59d6931e0759d6931e07to6059287fbeCode Review — PR #66
Solid fix for all three gaps from issue #65.
_extract_token_counts() helper
Clean extraction of the three-tier fallback logic that was duplicated. Consistent with the existing _safe_int / _log_no_usage_metadata patterns.
_run_pruning_pass return-type change (str -> tuple[str, int, int])
Split into two phases (LLM call vs. parse + validate) is the right structure. Hard RuntimeError on missing usage metadata is correct — pruning is a billing-visible call and silent (0, 0) would defeat the fix.
One concern: Phase 2 (after the try/except block) has no exception handler. _parse_pruning_response is probably pure string ops and won't raise, but the old code silently caught all Phase 2 errors and fell back to raw output. The new code lets a _parse_pruning_response exception propagate to the caller. Worth confirming that method is truly exception-safe, or wrapping Phase 2 with a narrower except that does NOT catch RuntimeError.
Minor nit:
if _pp == 0 and _pc == 0is used as the proxy for "no metadata present". Correct for every real LLM, but technically ambiguous — a provider could return genuine zero-token metadata. Low-risk in practice; could use a sentinel flag from _extract_token_counts instead of relying on zero as the signal.Tool-loop accumulation
All six ainvoke() sites are now covered: budget-exhaustion synthesis, main tool-call loop rounds, stuck-model synthesis, synthesis final tool-call round, and both pruning-pass sites (via returned counts). Comprehensive.
Design tradeoff: On round 1, _accumulated_prompt + _accumulated_completion == 0, so the budget guard never triggers on the first call (was a heuristic estimate before). This is actually better — the heuristic gave false precision. Noting it so reviewers don't mistake it for a regression.
ConfigurationError for pruning_model != model
Correct temporary guard. Error message references issue #65 and explains the limitation clearly.
_estimate_token_count removal
Right call — with actuals available from round 2 onward, the character-count heuristic has no remaining justification.
BDD tests
New scenarios cover: mismatch ConfigurationError, matching model success, token counts > 0 from pruning pass, hard failure on missing metadata, token_budget_percent type validation, missing end tag fallback, invalid tool_max_rounds, and direct-format tool calls. Good coverage.
Benchmark mock update
usage_metadata added to pruning mock — required by the new return signature.
Verdict: Logic correct, all gaps addressed, good test coverage. The Phase 2 exception safety question is worth a quick look before merge, but not a blocker if _parse_pruning_response is confirmed safe. Ready otherwise.
PR Review: cleveragents/cleveractors-core!66 (Ticket #65)
Verdict: Request Changes
The fix correctly addresses the three token-undercounting gaps called out in issue #65 — pruning-pass usage is now captured, tool-loop rounds are accumulated instead of overwritten, and the
_extract_token_countshelper extracts the existing three-tier fallback logic into a reusable method. However, the PR removes_estimate_token_countwithout updating its remaining callers inbenchmarks/token_budget_benchmark.py, which will breaknox -e benchmark/nox -e benchmark_regression. There is also a billing-accuracy regression in the new RuntimeError failure path: tokens already consumed by previous loop rounds are silently zeroed out.Critical Issues
_estimate_token_countremoved but callers not updated (benchmarks/token_budget_benchmark.py, lines 60–78)The
TokenBudgetBenchmarkclass still callsself.agent._estimate_token_count(msgs)in four benchmark methods (time_estimate_small_messages,time_estimate_medium_messages,time_estimate_large_messages,time_estimate_very_large_messages). The method was removed in this PR (CHANGELOG says "Estimate removed"), so all four will raiseAttributeError: 'LLMAgent' object has no attribute '_estimate_token_count'as soon asnox -e benchmarkornox -e benchmark_regressionruns. Per the project's "every stage passes" rule this is a build-integrity blocker.Recommendation: Either delete the four
time_estimate_*methods and theTokenBudgetBenchmarkclass entirely, or replace them with benchmarks against the new_extract_token_countshelper / the per-round accumulation path.Major Issues
Failure-case billing data is wiped to
(0, 0)(src/cleveractors/agents/llm.py,_run_pruning_passat ~lines 583–589, in conjunction with theexcept Exceptionblock at lines 1344–1357 ofprocess_message())When the pruning pass raises
RuntimeErrorbecause the response lacksusage_metadata, the local accumulators_accumulated_prompt/_accumulated_completion(lines 815–816) still hold the real token counts from prior loop rounds — those tokens were charged by the provider. But theexcept Exceptionhandler inprocess_message()checks_captured_prompt is not None(line 1352), and the sentinel is only assigned AFTER the tool loop completes (lines 1293–1296). At the point the RuntimeError propagates,_captured_promptis stillNone, so the handler resets_last_token_usage = (0, 0)andlast_token_usage_var.set((0, 0)), discarding the real costs. The PR's stated goal is billing accuracy; this regresses it precisely in the new failure mode the PR introduces.Recommendation: Set
_captured_prompt = _accumulated_promptand_captured_completion = _accumulated_completionimmediately before the Phase-2 RuntimeError raise (or right after the Phase-1 try/except), so the existing except handler preserves the counts.if _pp == 0 and _pc == 0is ambiguous as a "missing metadata" signal (src/cleveractors/agents/llm.py, line 583)_safe_intalready returns 0 for missing keys, malformed values, booleans, and negative integers (with a warning). A provider that legitimately returnsusage_metadata = {"input_tokens": 0, "output_tokens": 0}for an empty generation, or a buggy provider returning negative counts, will trigger the RuntimeError with a misleading "did not include token usage metadata" message. Graa's existing review notes this; the PR author has not responded. Relying on the zero-tuple as a sentinel conflates "no metadata" with "metadata that resolved to zero/zero".Recommendation: Change
_extract_token_countsto returntuple[int, int] | None(or(int, int, bool)with a "metadata present" flag) so callers can distinguish "metadata was absent" from "metadata said zero". The RuntimeError message can then be precise.Phase-1 description doesn't match implementation —
build_chat_modelerrors no longer fall back (src/cleveractors/agents/llm.py, lines 522–531 vs. 568–577)The PR description says "Phase 1: Make the pruning LLM call. Failures here fall back to raw output (non-functional model, network error, etc.)." The implementation only wraps
await prune_model.ainvoke(prune_messages)in the try/except —build_chat_model(...)and message construction are now OUTSIDE the try block. The original code wrapped the whole phase, so a missing API key, invalid model name, or template-rendering exception would have fallen back to(raw_output, 0, 0)with a warning. Now those errors propagate asExecutionError("LLM processing failed: …")and abort the agent call. This is a behavioral regression vs. the pre-PR semantics that contradicts the PR description.Recommendation: Either widen the try/except to include
build_chat_modeland message construction (matching the description), or update the PR description to clarify that model-construction errors now fail loud. The current state is internally inconsistent.Minor Issues
RuntimeErroris too generic for a domain-specific failure (src/cleveractors/agents/llm.py, line 584)Callers cannot distinguish "pruning model returned no token usage metadata" from any other RuntimeError. A dedicated
TokenTrackingError/MissingUsageMetadataErrorincleveractors/core/exceptions.pywould let callers (and tests) catch this specifically. The_extract_token_countshelper is already structured around typed exceptions (ConfigurationError,ExecutionError); this one outlier is inconsistent._safe_intswallows negative token counts into the same "missing metadata" bucket (src/cleveractors/agents/llm.py, lines 1768+ in_safe_int, used by_extract_token_counts)A malicious or buggy provider returning negative counts will be reported as a
RuntimeError("did not include token usage metadata"). The warning emitted by_safe_intis observable in logs but the exception message points the operator at the wrong root cause. Consider including a hint in the RuntimeError when the underlying cause was_safe_intwarnings.Nits
Redundant fallback
self._pruning_model or self.model(src/cleveractors/agents/llm.py, line 522)Given the new
ConfigurationErrorraised at construction (lines 228–233) when_pruning_model != self.model,_pruning_modelis now guaranteed to beNoneor equal toself.model. Theor self.modelfallback is dead code. Not worth a separate commit, but worth cleaning up next time this block is touched.Scope creep in tests (
features/llm_agent_tool_calling.feature)The PR also adds scenarios that are unrelated to issue #65:
LLMAgent validates token_budget_percent must be numeric,Parse pruning response with missing end tag falls back to full text,LLMAgent rejects invalid tool_max_rounds on process_message,LLMAgent handles tool_calls in direct format without function key. These are coverage wins for pre-existing code paths, but they broaden the PR's scope beyond the stated "fix token undercounting" goal. Fine to keep; flagged only because ticket-driven commits are expected to be atomic to the ticket.Existing reviewer note on Phase-2 exception safety still unanswered (Graa's review, 2026-06-30)
Graa explicitly asked whether
_parse_pruning_responseis truly exception-safe now that it's no longer wrapped in try/except, and noted that a narrowerexcept(catching everything exceptRuntimeError) would be safer if there's any doubt. The PR author has not responded on the PR. Looking at the code,_parse_pruning_responseis pure string ops (lines 470–502) andstr(prune_response.content)should always succeed, so Graa's concern is theoretical — but the lack of acknowledgment is worth noting.Summary
The fix correctly closes all three gaps from issue #65 and matches the spec/ADR's intent for actual-usage tracking. The
_extract_token_countsextraction is a clean refactor. But the PR is not ready to merge:_pp == 0 and _pc == 0ambiguity and the Phase-1/Phase-2 description mismatch are real but smaller; addressing them now keeps the fix internally consistent.Once the critical benchmark breakage and the failure-case billing preservation are addressed, this is a solid fix. Recommend re-review at that point.
6059287fbetoc8c7e9b8ddReply to Review — PR #66 (force-pushed to
c8c7e9b8)Thank you both for the thorough reviews. The branch was force-pushed to incorporate all substantive feedback. Below is the disposition of each item.
What was done
hurui200320 — Critical: Broken ASV benchmarks
[92mDone[0m:
TokenBudgetBenchmark(which called the removed_estimate_token_count) has been replaced withTokenExtractionBenchmarkthat benchmarks_extract_token_countsacross the three metadata paths. The_messages_with_sizehelper was removed along with the fourtime_estimate_*methods.hurui200320 — Major: Billing data wiped to (0,0) on MissingUsageMetadataError
[92mDone[0m: Both pruning call sites (budget-exhaustion path and main tool-loop path) now capture
_accumulated_prompt/_accumulated_completioninto_captured_prompt/_captured_completionimmediately before raisingMissingUsageMetadataError. The existingexcept Exceptionhandler preserves these counts.hurui200320 — Major:
_pp == 0 and _pc == 0ambiguity[92mDone[0m:
_extract_token_countsnow returns(int, int, bool)with a_metadata_presentflag. The Phase-2 guard checksnot _metadata_presentinstead of the zero-tuple.hurui200320 — Major: Phase-1/Phase-2 description mismatch
[92mDone[0m:
build_chat_model, message construction, andainvokeare all inside thetry:block. A missing API key, invalid model name, or template-rendering exception falls back to(raw_output, 0, 0)with a warning, matching the PR description.Graa — Minor nit: zero-tuple as missing-metadata proxy
[92mDone[0m: Same fix as above — the
_metadata_presentboolean resolves this.hurui200320 — Minor #1:
RuntimeErrortoo generic[92mDone[0m:
MissingUsageMetadataError(CleverAgentsException)added tocleveractors/core/exceptions.py. The phase-2 guard raises this instead ofRuntimeError.What was NOT done (with justification)
hurui200320 — Minor #2:
_safe_intswallows negative tokens[93mNot changed[0m: With the
_metadata_presentflag the paths are now cleanly separated: metadata present (including negatives) ⇒(0, 0)accumulated (undercount, but the provider returning negatives is a separate defect); metadata absent ⇒MissingUsageMetadataErrorraised. The_safe_intwarnings remain in logs to detect buggy providers.hurui200320 — Nit #1: Redundant
self._pruning_model or self.model[93mAcknowledged[0m: The construction guard ensures
_pruning_modelisNoneor equal toself.model, makingor self.modeldead code. Noted for cleanup in the next patch touching this area.hurui200320 — Nit #2: Scope creep in tests
[93mAcknowledged[0m: The additional coverage scenarios exercise pre-existing code paths. Keeping them improves overall coverage density.
Graa — Phase-2 exception safety of
_parse_pruning_response[92mConfirmed safe[0m: Pure string ops only (
str.find,str.strip, concatenation).str(prune_response.content)is safe for all standard LangChain message types. Phase 2 will not raise unexpectedly.Summary
All review feedback has been addressed. The PR is ready for re-review.
c8c7e9b8ddto7e427de2b1PR Review: cleveragents/cleveractors-core!66 (Ticket #65)
Verdict: Approve
The fix correctly addresses all three token-undercounting gaps from issue #65. The previous review's critical and major issues have been resolved: the ASV benchmarks were updated, the failure-case billing-data wipe is now prevented by the
_captured_prompt = _accumulated_promptpattern in the pruning call sites, the(0, 0)ambiguity is resolved via the_metadata_presentbool flag, thebuild_chat_modelis properly inside Phase 1's try block, and a dedicatedMissingUsageMetadataErrorreplaces the genericRuntimeError. The PR author has not posted any comments responding to the prior reviews, so nothing has been deferred or marked out-of-scope.Critical Issues
None
Major Issues
None
Minor Issues
or self.modelfallback forprune_model_name—src/cleveractors/agents/llm.py, line 534ConfigurationErrorraised at construction (lines 229–234) when_pruning_model != self.model,_pruning_modelis now guaranteed to beNoneor equal toself.model. Theor self.modelfallback is dead code. Cleaning this up toprune_model_name = self.modelwould be tidier but is not a correctness issue.Nits
_log_no_usage_metadatais emitted even whenMissingUsageMetadataErrorwill be raised —src/cleveractors/agents/llm.py, line 591. The helper logs a generic warning before_run_pruning_passraises. Operators will see both the warning and the exception. Acceptable for debuggability._log_no_usage_metadatafires for every main-modelainvoke()that lacks metadata —src/cleveractors/agents/llm.py, lines 877, 1017, 1025, 1182, 1278. The five main-model call sites discard_metadata_presentwith_, so any model lackingusage_metadata(some open-source providers, test mocks) will produce one warning per loop round. Noise concern, not a correctness one.Spec deviation note could be more explicit in the changelog —
CHANGELOG.md, line 18. The changelog does not explicitly note thatpruning_modelhas been effectively disabled (must equalmodel). The constructor's error message and PR description cover this at runtime, but a single changelog line would help downstream consumers.Summary
All three gaps from issue #65 are correctly addressed:
_run_pruning_pass()now returns(parsed, prompt_tokens, completion_tokens), and the caller accumulates these counts. The previous concern about Phase-2 exception safety is moot in practice —_parse_pruning_responseis pure string operations andstr(prune_response.content)is unlikely to raise.ainvoke()call sites inprocess_message()plus two pruning call sites all feed into_accumulated_prompt/_accumulated_completion, which are then written to_last_token_usage/last_token_usage_var. The failure-path billing wipe concern is fully addressed: the innertry/except MissingUsageMetadataErrorblocks at lines 976–981 and 1126–1129 set_captured_promptand_captured_completionbefore re-raising, so the outerexcept Exceptionhandler preserves the counts._estimate_token_countheuristic is removed. The budget check at line 846 now uses actual provider-reported consumption. The broken benchmarks concern is resolved:TokenExtractionBenchmarkreplacesTokenBudgetBenchmarkand exercises the new helper.The temporary limitation (
pruning_modelmust equalmodel) is well-documented and the path forward (per-model token tracking inNodeUsage) is identified in issue #65.MissingUsageMetadataErrorpropagation throughprocess_message()is correctly wrapped asExecutionError, while direct calls to_run_pruning_pass()still surface the typed exception. Type annotations, docstrings, and CHANGELOG are all updated.Recommendation: Approve and merge. The minor issues are stylistic and do not affect correctness or spec compliance.
7e427de2b1tod7e951402ad7e951402ato8a7e353767