The step "I initialise the overlay with real slash command specs"
was calling set_commands("", ...) which triggers the hide-on-empty-query
early return, leaving _text empty and failing all four assertions in
the "Overlay uses real slash_command_specs from catalog" scenario.
Use query "se" instead: all session:* entries (9) plus settings (1)
start with "se", giving 10 matches within the 12-entry display cap,
so both " /session:create" / "Create a new session tab" and
" /settings" / "Open settings" appear in the rendered overlay.
ISSUES CLOSED: #6450
The test at line 35-39 was checking that descriptions are NOT shown in the
slash overlay, which was the old behavior before implementing ADR-046. Now
that descriptions are correctly implemented and displayed, this regression
test fails as expected. Remove it since the feature is complete.
Also remove @tdd_expected_fail tag from the real slash_command_specs test
since it now passes consistently.
ISSUES CLOSED: #6450
The step that loads test commands into the overlay was directly assigning
_commands and _visible without ensuring selected_index was initialized.
This caused navigate_down tests to fail because selected_index might not
be properly reset for each scenario.
Fixes: features/tdd_slash_overlay_keyboard_nav.feature:30,55
Replace getattr(tui_app_module, '_strip_pending_reference_token')
with direct attribute access. Ruff B009 flags this as unnecessary
because the attribute is a compile-time constant — normal access
is equally safe and preferred.
The previous attempt left three CI-breaking issues:
- features/steps/config_cli_safety_net_coverage_steps.py imported
_env_var_for_key/_normalize_key from config.py, but those helpers
moved to _config_helpers.py and were never re-exported. Import
from _config_helpers directly (matches where the symbols live).
- features/steps/config_get_spec_output_steps.py defined two when-
steps whose default-parse {key} placeholders made the longer step
ambiguous against the shorter one (behave raised AmbiguousStep on
load). Switch those two patterns to the re matcher with [^"]+ so
the quoted argument can't swallow ` with format "..."`.
- The same step file cached parsed JSON on context._spec_json with
no invalidation. Behave layers context attributes — a value set at
one scenario's layer remained visible to later scenarios, so the
assertion read stale JSON from an earlier scenario's get. Parse
fresh each call.
- src/cleveragents/cli/commands/config.py built a spec-correct JSON
envelope and then passed it through format_output, which wraps its
input in another envelope — the outer envelope's data field was
the inner envelope, so data.type / data.source / data.overridden
were absent. Emit the manually-built envelope directly via
json.dumps / yaml.safe_dump so the started timing field and all
spec data fields appear at the correct depth.
ISSUES CLOSED: #3423
- Import _env_var_for_key from _config_helpers (where it was moved
during refactor) instead of config module, fixing ImportError that
blocked all unit_tests from running
- Restore --verbose / -v as a deprecated no-op parameter on config_get
so existing tests and user scripts that pass --verbose continue to work;
resolution chain is always shown per spec regardless of the flag
- Fix CliRunner(mix_stderr=False) in config_get_spec_output_steps.py to
CliRunner() — mix_stderr is not accepted by this version of Typer
ISSUES CLOSED: #3423
Implements all spec-required output for `agents config get`:
Rich output:
- Rename panel title from 'Configuration Value' to 'Config' per spec
- Add 'Overridden' field to the Config panel
- Add 'Origin' panel showing File, Line, and Default fields
- Add 'Resolution Chain' panel (always shown, not just with --verbose)
- Add 'Winner' indicator to the Resolution Chain panel
- Use spec-required type strings (string/boolean/integer) instead of
Python type names (str/bool/int)
- Add '✓ OK Config read' confirmation message
JSON output:
- Wrap result in standard envelope (command, status, exit_code, data,
timing, messages)
- Add 'overridden' field to data
- Add 'origin' nested object (file, line, default)
- Add 'winner' nested object (source, level)
- Use human-readable source names (CLI flag, Env var, Config file,
Default) instead of internal enum values
- Add 'started' timestamp to timing
BDD:
- Add features/config_get_spec_output.feature with 20 scenarios
covering all Rich panels and JSON envelope fields
- Update existing tests to match new spec-compliant output
ISSUES CLOSED: #3423
integration_tests was throttled to `min(2, _default_processes())` (commit
92e258535, a flaky-stabilization measure). Evidence from a cached CI run shows
the suite does ~47min of work serialized through 2 pabot lanes on a 32-core
runner -> ~24min wall (exactly 2x == 2 lanes), the chronic ~21-30min seen
across recent runs. It is linearly parallelism-bound.
Raise the cap to `min(6, _default_processes())`: ~4x more lanes -> projected
~6-8min wall, while staying well short of per-core fan-out (the runner reports
32 cores; behave already uses all of them). The cap is kept (not removed)
because the suite makes live LLM calls — ~149 HTTP 429 rate-limit responses
were observed even at 2-way, and unbounded fan-out would spike 429/OOM flakes
(the exact failure 92e258535 was masking). 6 is a deliberate middle ground;
TEST_PROCESSES / --processes still override.
Updates the now-stale "<=2 processes" docstrings on integration_tests and
slow_integration_tests. Behave/unit_tests parallelism (_default_processes(),
uncapped) is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
integration_tests (Robot/pabot) failed on robot/coverage_threshold.robot
"Noxfile Contains Coverage Threshold Constant": it asserted the literal
`COVERAGE_THRESHOLD = 96.5` in noxfile.py, but the coverage_report session now
delegates the constant to pyproject's single source
(`COVERAGE_THRESHOLD = _read_coverage_fail_under()`).
This was the last WS5 reconciliation gap (the plan enumerated behave feature
steps + ci.yml + guidelines prose, but not the Robot suite). Assert the
delegation plus the pyproject floor (`fail_under = 96.5`) instead of the literal,
mirroring the behave step fix in b499834c0. The other assertions in the suite
(--fail-under= substring, [tool.coverage.run], branch=true, source=["src"]) are
unaffected; verified no other robot/behave/pytest test depends on the literal.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The coverage_report noxfile session now delegates COVERAGE_THRESHOLD to the
pyproject single source (`COVERAGE_THRESHOLD = _read_coverage_fail_under()`),
so the AST value is a Call, not a Constant. The
coverage_threshold_enforcement.feature step "I parse the COVERAGE_THRESHOLD
constant from noxfile.py" raised ValueError (errored scenario) because it only
handled a literal constant — this was the missed third WS5 step file (the
config + consolidated steps were already reconciled).
Add the same pyproject `[tool.coverage.report].fail_under` fallback used by
coverage_threshold_config_steps.py, and fix the phantom-97 in the feature
description prose. Verified: the four coverage feature files now pass
(136 scenarios, 0 errored); previously unit_tests errored on
coverage_threshold_enforcement.feature:7.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Master passed the 96.5 floor only by rounding (96.465 -> 96.5). Add genuine
behave coverage + omit a structurally-uncoverable module so the floor has real
headroom. fail_under stays 96.5 (ratchet rule: raise only after coverage holds).
- omit src/cleveragents/application/services/__init__.py (noxfile + pyproject):
100% of its "missing" lines are inside an `if TYPE_CHECKING:` block (never
executes at runtime) and slipcover has no per-line pragma to exclude them.
- features/coverage_validation_error_paths.feature (+ steps): 16 scenarios
exercising the defensive error-path branches in core/validation.py that the
existing structural_validation.feature does not reach (non-dict node, dup
decision_id, non-list children, invalid child ULID, missing decision fields,
wrong-typed confidence/parent/sequence, malformed structured-output, unknown
dispatcher target). Routed through the public
validate_structured_component_output dispatcher; pure functions, no DB/CLI.
Full engine run: NOX_EXIT=0, no dead chunks, 96.623% (validation.py now fully
covered).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ports the WS1 fast/parallel coverage engine onto master and reconciles the
coverage gate to a single source. Local full-suite run on master's suite:
NOX_EXIT=0, all 12 chunks alive (no dead chunks), peak 658MB/chunk
(ceiling ~2.6GB vs the old ~1.4GB single-process that OOMs/reaper-kills),
96.465% -> 96.5 (rounded) >= 96.5 floor.
noxfile.py — coverage_report full-suite path now fans out K concurrent
slipcover processes (K=COVERAGE_PROCESSES, default 4) over N bin-packed
chunks (by scenario count), failing loud on any dead chunk (never merges
survivors) and merging per-chunk JSON. Bounds per-process peak RSS, killing
the single-process OOM/reaper collapse. Targeted (.feature posargs) runs keep
the single-process path. COVERAGE_THRESHOLD now reads
pyproject [tool.coverage.report].fail_under. Features are enumerated by direct
glob (NOT by importing run_behave_parallel, whose top level imports
behave/behave_parallel and is unavailable in the nox orchestrator process).
pyproject.toml — adds [tool.coverage.report].fail_under = 96.5 as the single
source of truth, with the evidence-gated ratchet rule (objective 97%).
.forgejo/workflows/ci.yml — coverage job: carries the skip_coverage operator
valve; propagates nox's exit EXPLICITLY (set -uo pipefail + PIPESTATUS +
exit $rc) instead of relying on the runner's implicit bash -eo pipefail;
adds timeout-minutes: 30 so a hang fails cleanly with diagnostics; deletes
the dead threshold=50 "Surface coverage summary" step; fixes the phantom-97
step label. Gating remains nox's --fail-under (sourced from pyproject).
coverage_threshold_config_steps.py — the fail-under feature step resolves the
floor from pyproject when COVERAGE_THRESHOLD delegates to the reader.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Resolve Forgejo label name to integer ID before PUT /issues/{n}/labels
(both create_uat_tracking_issue and create_uat_announcement_issue);
string names cause 422 / silent no-op, breaking the cleanup filter
- Add rm -rf /tmp/uat-tester-* and rm -rf /tmp/docs-* to bash allowlist
so worker clone cleanup can actually execute
- Replace multi-line python3 -c calls in Pool Supervisor with jq:
EXISTING_WORKERS lookup now uses jq ltrimstr/startswith filter;
SESSION_ID creation now uses jq -r '.id'
Worker Mode requires git clone, git config, git pull, git checkout,
git add, git commit, git push, and uv commands in bash. These were all
missing from the frontmatter permission block, causing every Worker Mode
instance to fail immediately with permission denied on the first command.
- Delete scripts/fix_uat_tester.py: one-off patch script with a
hardcoded /tmp path that no longer exists. Per CONTRIBUTING.md,
scripts/ is for ongoing-purpose utilities, not development artifacts.
- Remove update_examples_json() from create_documentation_pr(): each
per-PR call was modifying examples.json inside the git clone, causing
merge conflicts when parallel UAT workers created docs PRs
simultaneously. Option A from reviewer: move the call outside the PR
creation function entirely.
- Add batch update_examples_json(documented_examples) call after the
inner docs-PR creation loop completes, guarded by `if documented_examples`.
This runs once per cycle instead of once per PR, eliminating the
examples.json conflict surface for parallel workers.
- Expand CHANGELOG entry to reflect the examples.json fix.
ISSUES CLOSED: #4374
Fixed two critical bugs in the uat-tester agent's create_documentation_pr() function:
1. **Branch-existence guard bug**: Changed existing_branch.get('error') to existing_branch.get('errors') to correctly detect when a branch does not exist. The Forgejo API returns a JSON error body with an errors field (not error) when a branch is absent.
2. **Open-PR duplicate check bug**: Changed head={branch_name} to head={owner}:{branch_name} in the Forgejo API query. The API expects the owner prefix to properly filter PRs by head branch.
These fixes ensure parallel workers can safely coordinate without creating conflicting PRs.
Closes#4374
The session create command already emitted a spec-compliant JSON envelope
with messages[].text populated, but list, show, delete, export, and
import did not — they either passed no `messages` to `format_output()`
(producing an envelope with an empty messages array) or short-circuited
to a Rich-styled `console.print(...)` line that breaks JSON parsing
entirely.
Failing scenarios in features/session_cli.feature (unit_tests gate)
asserted `messages[0].text == "<command-specific success>"`. Wire each
command's machine-readable path to emit a structured envelope with the
expected message:
- list (empty) → "0 sessions listed"
- list (populated) → "<N> sessions listed"
- show → "Session details loaded"
- delete --format json|yaml|plain → "Session deleted"
- export --output-format json|yaml|plain → "Export completed"
- import --format json|yaml|plain → "Import completed"
The export command gains a new `--output-format` flag distinct from the
existing `--format` (which selects export content format: json or md).
When the new flag is non-rich, the raw export content is suppressed from
stdout so the envelope remains the only thing emitted, and Rich panels
are skipped.
The import command gains a `--format` flag. The delete command already
had a `--format` option but its non-rich branch emitted Rich-styled text
instead of an envelope; that branch now splits cleanly: `--format color`
keeps the human-readable line, and json/yaml/plain emit an envelope.
Also addresses the `Session create initializes MCP logger` and
`Session list initializes MCP logger` scenarios in
features/session_cli_mcp_logger_simple_execution.feature, which inspect
the create()/list_sessions() source via inspect.getsource() and assert
the literal string `logging.getLogger(_MCP_LOGGER_NAME)` appears. Both
functions held the inlined literal `"cleveragents.mcp"` instead of the
module-level `_MCP_LOGGER_NAME` constant; substitute the constant
reference in both call sites.
CHANGELOG entry extended to document the envelope coverage across all
session commands.
ISSUES CLOSED: #6441
Cover the two uncovered code paths introduced by the session create
JSON envelope fix:
1. `_resolve_actor_details(None)` → early `return None` (no actor bound)
2. Actor found in registry path (lines 271-297) — two scenarios:
- Full config actor (options.temperature + direct context_window)
- Graph descriptor actor (graph_descriptor.context_window)
Also remove dead redundant isinstance checks:
- `config_blob.get("options") if isinstance(config_blob, dict) else None`
simplified to `config_blob.get("options")` since config_blob is always
a dict (assigned on the previous line via ternary with {} fallback).
- Outer `if isinstance(config_blob, dict):` wrapper around context_window
logic removed for the same reason.
ISSUES CLOSED: #6441
- Fix `security_template_coverage_boost.feature` assertion for "Session
export to stdout outputs JSON": the export path outputs raw JSON with
a `session_id` key, not a `data` envelope, so revert the erroneous
`"data"` assertion back to `"session_id"`.
- Move all deferred `from cleveragents.application.container import
get_container` imports in session.py to the module-level import
block, consistent with every other CLI command file (action.py,
actor.py, config.py, plan.py, etc.). No circular import exists.
- Update `session_cli_uncovered_branches_steps.py` to patch
`cleveragents.cli.commands.session.get_container` directly (the
correct target after a top-level import) instead of replacing
`sys.modules["cleveragents.application.container"]`.
- Add CHANGELOG.md entry for the session create JSON envelope fix.
ISSUES CLOSED: #6441
Update feature file assertions to match the new nested JSON structure
for session create output. Changed assertions from 'session_id:' to 'id:'
to reflect the new data.session.id structure per spec #6441.
The Robot Framework `Catenate` keyword treats 2+ consecutive spaces as
field separators and strips them, which breaks Python's mandatory
indentation when a nested `assert` follows a `for` line. As a result
the "TUI Help Command Groups By Namespace" test produced a script with
a de-indented `assert` body, causing an IndentationError under
`python -c` and the test to fail with rc=1.
Rewrite the namespace-presence check as a single-line list
comprehension over the expected group headers, mirroring the pattern
already used by the sibling "Lists All Catalogued Commands" test in
the same file. The check is semantically equivalent — both fail with
the same diagnostic message when a group header is missing — but the
flattened form survives Robot's argument tokenisation intact.
ISSUES CLOSED: #3434
Replace the hardcoded help string in TuiCommandRouter.handle() with a
dynamic lookup against SLASH_COMMAND_SPECS from slash_catalog.py.
Changes:
- Add _help_command(), _help_list_all(), _help_for_command() methods to
TuiCommandRouter
- /help (no args): iterates SLASH_COMMAND_SPECS, groups commands by
namespace (sorted alphabetically), renders all 70 commands with
descriptions in colon-namespaced format (e.g. persona:list)
- /help <command>: looks up the given command in SLASH_COMMAND_SPECS and
renders its full help (group, description)
- /help <unknown>: returns 'Unknown command: /<cmd>' message
- /help /persona:list (with leading slash): strips the slash and resolves
correctly
- Import defaultdict and SLASH_COMMAND_SPECS at module level
Tests:
- Update tui_commands_coverage.feature: replace old exact-match scenario
for help text with new dynamic-listing assertions
- Add tui_commands_coverage_steps.py: new 'should contain' step definition
- Add tui_help_command_full_catalog.feature: 12 BDD scenarios covering
/help no-args, /help <command>, /help <unknown>, namespace grouping,
colon-namespaced format, and regression against old hardcoded string
- Add tui_help_command_full_catalog_steps.py: step definitions for the
new feature (all-commands check, not-equal assertion)
- Add robot/tui_help_command.robot: 5 Robot Framework integration tests
verifying the help command via direct Python invocation and headless
TUI startup
Closes#3434
---
**Automated by CleverAgents Bot**
Supervisor: Implementation | Agent: ca-issue-worker
Add ContextTierService mock to step_m5_invoke_project_context_show
so the CLI command can call _get_context_tier_service() without
failing. Also patch _load_policy_json to prevent JSON decode errors
from MagicMock session objects. Update CHANGELOG.md and CONTRIBUTORS.md.
ISSUES CLOSED: #6323
The step assertion checked for 'logging.getLogger("cleveragents.mcp")'
literally, but session.py uses the _MCP_LOGGER_NAME constant so
inspect.getsource() returns 'logging.getLogger(_MCP_LOGGER_NAME)'.
Update the assertion to match the actual source text.
ISSUES CLOSED: #6457
The delete command's output branch used `fmt == OutputFormat.RICH` which
caused --format color to emit a JSON envelope instead of Rich panels.
Align with all other session commands by checking
`fmt not in (OutputFormat.RICH, OutputFormat.COLOR)` for machine-readable
paths. Add BDD scenario and step for --format color to close the coverage gap.
ISSUES CLOSED: #6457
Extend the JSON/YAML envelope messages[].text fix to cover the remaining
three session subcommands that were still producing plain Rich output
instead of structured envelopes for non-rich format paths:
- session delete: route non-rich formats through format_output() with
messages=[{"level": "ok", "text": "Session deleted"}]
- session export: add --output-format/-f option; emit structured envelope
with session_export/contents/integrity data and "Export completed" message
- session import: add --format/-f option; emit structured envelope with
session_import/validation/merge data and "Import completed" message
Also add BDD scenarios to features/session_cli.feature for each new
envelope path, add corresponding step definitions, add CHANGELOG entry
under [Unreleased] Fixed, and assign milestone v3.2.0 to the PR.
ISSUES CLOSED: #6457
Extend the devcontainer cleanup Behave scenario for container-instance resources to assert that the stop command delegates to the lifecycle stop mock, providing regression coverage that both container-instance and devcontainer-instance remain stoppable.
ISSUES CLOSED: #6457