Files
cleveragents-core/features/security_template_coverage_boost.feature
hurui200320 2005b8ef82
CI / push-validation (push) Successful in 18s
CI / build (push) Successful in 19s
CI / helm (push) Successful in 24s
CI / lint (push) Successful in 29s
CI / security (push) Successful in 1m11s
CI / e2e_tests (push) Successful in 2m56s
CI / quality (push) Successful in 3m40s
CI / typecheck (push) Successful in 3m59s
CI / integration_tests (push) Successful in 4m3s
CI / unit_tests (push) Successful in 4m55s
CI / docker (push) Successful in 10s
CI / coverage (push) Successful in 10m44s
CI / status-check (push) Successful in 1s
CI / benchmark-regression (push) Has been skipped
CI / benchmark-publish (push) Successful in 1h13m28s
feat(tests): replace all @skip tags with proper @tdd_expected_fail tags or remove them across the entire codebase (#7221)
## Summary

Replaces all 234 bare `@skip` occurrences across 82 Behave feature files with the correct TDD issue-capture tagging system described in CONTRIBUTING.md § Bug Fix Workflow.

Previously, the noxfile ran Behave with `--tags=not @skip`, silently excluding all `@skip`-tagged scenarios from every CI run. These tests never ran, never inverted results via the `@tdd_expected_fail` mechanism, and never contributed to coverage — defeating the purpose of TDD issue-capture testing. Every `@skip` occurrence had a commented-out hint line immediately above it showing the intended proper tags (e.g., `# @tdd_issue @tdd_issue_4272 @tdd_expected_fail @skip`), confirming they were all intended for conversion.

## Changes

### Mechanical conversion (234 replacements across 82 files)
- Extracted the proper TDD tags from the comment hint above each `@skip` line, removed `@skip` from the tag set, and replaced the `@skip` line with those tags.
- Removed the now-redundant comment hint lines alongside each replacement.

### Bug-fixed scenarios — `@tdd_expected_fail` removed (84 scenarios)
- After conversion, ran `nox -s unit_tests` to identify which newly-enabled `@tdd_expected_fail` scenarios now **pass** (their referenced bugs have already been fixed). Removed `@tdd_expected_fail` from those 84 scenarios and their corresponding feature-level tags, leaving only the permanent `@tdd_issue @tdd_issue_<N>` regression-guard tags.
- Affected features include: `tdd_tool_runner_env_precedence`, `tdd_automation_profile_session_leak`, `tls_certificate_check`, `project_create_persist`, `resource_type_bootstrap_*`, and 18 others.

### Noxfile cleanup
- Removed all four `--tags=not @skip` arguments from `noxfile.py` (unit_tests and coverage sessions). With zero `@skip` tags remaining in the codebase, this filter was dead code and its presence would mislead future maintainers into thinking `@skip` is still a supported escape mechanism.

### Regression guard files
- Split the regression guards into two focused files:
  - `tdd_regression_guards_exec_env.feature` for bug #4281 (exec-env precedence)
  - `tdd_regression_guards_session_list.feature` for bug #4271 (session list summary)
- Each file carries only its own `@tdd_issue` tags at the feature level, avoiding cross-contamination via Behave tag inheritance. The `Background` step (`session-list-summary mock`) only appears in the session-list file where it is actually needed.

### Duplicate tag cleanup
- Removed duplicate `@tdd_issue @tdd_issue_4287` tag lines in `tdd_skill_add_regression.feature` (lines 20 and 29).

### Inline comment for retained `@tdd_expected_fail`
- Added inline comment in `ci_workflow_validation.feature:134` explaining why this specific #4227 scenario retains `@tdd_expected_fail` despite #4227 being closed (CI YAML does not encode threshold as a machine-readable value).

### Known edge cases — `@tdd_expected_fail` retained (closed issues, fix on master, scenarios still fail)
The following issues are **closed** and their fixes **are on master**, but the specific test assertions still fail because the fixes address other aspects of the bugs. The `@tdd_expected_fail` tags are functionally correct and must remain until the specific scenario assertions pass:
- `tdd_exec_env_resolution_precedence.feature` — bug #1080 (closed 2026-03-31). The precedence-level-2-vs-4 scenario still fails.
- `session_list_summary_dedup.feature` — bug #3046 (closed 2026-04-05). The dedup-consistency scenarios still fail.
- `actor_add_update_enforcement.feature` — bug #2609 (closed 2026-04-05). The enforcement scenarios still fail.
- `ci_workflow_validation.feature:134` — #4227 (closed 2026-04-08). The CI YAML threshold assertion still fails.

## Verification

- `grep -r "@skip" features/ --include="*.feature"` → **zero results** ✓
- `grep -n "tags=not @skip" noxfile.py` → **zero results** ✓
- `nox -s unit_tests` → **629 features passed, 0 failed** ✓ (up from ~545 before this PR)
- CI all green (coverage ≥ 97%) ✓
- Integration tests (Robot Framework) do not use `@skip` — confirmed no action needed ✓
- E2E tests (Robot Framework) do not use `@skip` — confirmed no action needed ✓
- `CHANGELOG.md` updated with entry for this change ✓
- `CONTRIBUTORS.md` — Rui Hu already listed ✓

## Issues Addressed

Closes #7025

Co-authored-by: CleverThis <hal9000@cleverthis.com>
Reviewed-on: #7221
Reviewed-by: HAL 9000 <HAL9000@cleverthis.com>
Reviewed-by: HAL9001 <hal9001@cleverthis.com>
Co-authored-by: Rui Hu <rui.hu@cleverthis.com>
Co-committed-by: Rui Hu <rui.hu@cleverthis.com>
2026-04-13 04:56:01 +00:00

227 lines
9.7 KiB
Gherkin

@coverage_boost
Feature: Coverage boost for security template branch
As a developer ensuring 97%+ test coverage
I want to exercise uncovered code paths across the codebase
So that the coverage threshold is comfortably met
# =========================================================================
# skills/registry.py validate_skill with tool registry, inline tools,
# and includes (lines 200-213)
# =========================================================================
@tdd_issue @tdd_issue_4267 @tdd_expected_fail
Scenario: Validate skill with tool registry reports missing tool refs
Given a skill registry with a mock tool registry
And a skill definition with tool ref "nonexistent-tool"
When I validate the skill definition via the registry
Then the skill validation errors should mention "not found in tool registry"
Scenario: Validate skill with inline tool missing description
Given a skill registry with no tool registry
And a skill definition with an inline tool missing its description
When I validate the skill definition via the registry
Then the skill validation errors should mention "missing a description"
Scenario: Validate skill referencing unknown included skill
Given a skill registry with no tool registry
And a skill definition that includes "unknown/skill"
When I validate the skill definition via the registry
Then the skill validation errors should mention "is not registered"
Scenario: Validate skill with all valid references
Given a skill registry with no tool registry
And a skill definition with no tool refs or includes
When I validate the skill definition via the registry
Then the skill validation errors should be empty
# =========================================================================
# cli/formatting.py — format_output_session, _format_table edge cases
# (lines 121, 128, 197, 231-236)
# =========================================================================
Scenario: Format output with color format type
Given sample formatting data with key "status" and value "ok"
When I format the data with format type "color"
Then the formatted output should contain "status"
Scenario: Format output session with list of dicts
Given sample formatting data as a list of two items
When I format the data via format_output_session with format "plain"
Then the formatted session output should not be empty
Scenario: Format output session with dict input
Given sample formatting data with key "name" and value "test"
When I format the data via format_output_session with format "json"
Then the formatted session output should contain "name"
Scenario: Format table with rows having extra keys
Given sample formatting data as a list with mismatched keys
When I format the data with format type "table"
Then the formatted output should contain "extra_col"
Scenario: Format table with empty list
Given sample formatting data as an empty list
When I format the data with format type "table"
Then the formatted output should be "(empty)"
# =========================================================================
# cli/commands/system.py — health check functions
# (lines 108-118, 146-147, 174-175, 193-194, 258-278, 325-334)
# =========================================================================
Scenario: Check config file when it exists and is readable
Given a system health check environment
And a config file that exists and is readable
When I run the config file health check
Then the health check status should be "ok"
And the health check details should be "readable"
Scenario: Check data dir when it does not exist
Given a system health check environment
And a data dir that does not exist
When I run the data dir health check
Then the health check status should be "warn"
And the health check details should contain "missing"
Scenario: Check disk space with plenty of space
Given a system health check environment
When I run the disk space health check
Then the health check status should be "ok"
Scenario: Check git availability
Given a system health check environment
When I run the git availability health check
Then the health check status should be "ok"
Scenario: Check file permissions on writable data dir
Given a system health check environment
And a writable data dir
When I run the file permissions health check
Then the health check status should be "ok"
And the health check details should be "data dir r/w"
Scenario: Check file permissions on nonexistent data dir
Given a system health check environment
And a data dir that does not exist
When I run the file permissions health check
Then the health check status should be "warn"
And the health check details should contain "does not exist"
Scenario: Build info data with existing database file
Given a system health check environment
And a database file that exists
When I build info data
Then the info data should contain a db_size value
Scenario: Build info data with log directory
Given a system health check environment
And a log directory with files
When I build info data
Then the info data should contain a logs value
# =========================================================================
# cli/commands/session.py — session CLI commands
# (lines 56-69, 75, 157-159, 267, 277-278, 320-323, 381-382, 482-485)
# =========================================================================
Scenario: Get session service lazily initialises from container
Given a session CLI test environment
When I call get_session_service
Then a session service should be returned
Scenario: Reset session service clears the cached instance
Given a session CLI test environment
When I call reset_session_service
Then the session service cache should be cleared
Scenario: Session create with non-rich format outputs formatted data
Given a session CLI test environment
And a mock session service that returns a created session
When I invoke session create with format "json"
Then the covboost session output should contain "session_id"
Scenario: Session create error shows session not found message
Given a session CLI test environment
And a mock session service that raises SessionNotFoundError on create
When I invoke session create expecting an error
Then the session CLI should have exited with error
Scenario: Session list with no sessions shows empty message
Given a session CLI test environment
And a mock session service that returns no sessions
When I invoke session list with format "rich"
Then the covboost session output should contain "No sessions found"
Scenario: Session delete without confirmation aborts
Given a session CLI test environment
And a mock session service that returns a session
When I invoke session delete without confirming
Then the session CLI should have been aborted
Scenario: Session delete with yes flag succeeds
Given a session CLI test environment
And a mock session service that returns a session
When I invoke session delete with yes flag
Then the covboost session output should contain "deleted"
Scenario: Session export to stdout outputs JSON
Given a session CLI test environment
And a mock session service that can export
When I invoke session export to stdout
Then the covboost session output should contain "session_id"
Scenario: Session export to file that exists without force fails
Given a session CLI test environment
And a mock session service that can export
And a temporary output file that already exists
When I invoke session export to that file without force
Then the session CLI should have exited with error
# =========================================================================
# cli/commands/cleanup.py — error handling paths
# (lines 49-58, 94, 129-131)
# =========================================================================
Scenario: Active plan detection with OSError returns empty set
Given a cleanup CLI test environment
And a cleanup service that raises OSError on sandbox scan
When I detect active plans
Then the active plans set should be empty
Scenario: Cleanup prune command runs successfully
Given a cleanup CLI test environment
And a mock cleanup service
When I invoke cleanup prune
Then the cleanup should have run without error
# =========================================================================
# cli/commands/audit.py — audit service creation
# (lines 29-32, 110-111)
# =========================================================================
Scenario: Audit service is created from settings
Given an audit CLI test environment
When I create the audit service
Then an audit service instance should be returned
Scenario: Audit list command with no entries
Given an audit CLI test environment
And a mock audit service with no entries
When I invoke audit list
Then the covboost audit output should contain "No audit"
# =========================================================================
# application/container.py — container init paths
# (lines 66-69, 125-130)
# =========================================================================
Scenario: Container initialises database from settings
Given a container test environment
When I get the application container
Then the container should provide a database session
Scenario: Container clears its singleton on reset
Given a container test environment
When I reset the container singleton
Then the container cache should be cleared