Commit Graph

761 Commits

Author SHA1 Message Date
aditya eff446f5e8 feat(acms): integrate FAISS into ACMS vector backend protocol
Add FAISS-backed ACMS read and write adapters on top of the shared VectorStoreService, wire them through the DI container, and cover indexing, scoped search, removal, and benchmark behavior with Behave and ASV.

Address review feedback by keeping the commit scoped to the FAISS backend work and replacing the loose vector-store cache typing with explicit FAISS store protocols instead of Any.

ISSUES CLOSED: #871
2026-03-30 13:19:08 +00:00
CoreRasurae 007af498b8 refactor(autonomy): rename automation profile task flags to spec names
Renamed all 11 task-type confidence threshold fields in AutomationProfile
from phase-transition semantics to spec-defined task-type semantics.
Updated all 8 built-in profiles, CLI formatting, YAML schema, services,
and all Behave/Robot tests referencing the old field names.

Post-review fixes:
- Fixed 24 stale old field names in M6 fixture files
  (automation_profiles.json, autonomy_guardrails.json)
- Added model_validator(mode='before') to detect legacy field names
  and raise actionable ValueError with rename mapping
- Added semantic bridge comments in PlanLifecycleService mapping
  task-type thresholds to phase-transition gates
- Added threshold_field to structured log messages for observability
- Restored categorised CLI automation-profile show output to match
  spec (Phase Transitions / Decision Automation / Self-Repair /
  Execution Controls) instead of flat list
- Added missing access_network field to spec show output examples
  (Rich, Plain, JSON, YAML variants)
- Aligned ADR-017 profile fields table to all 11 fields with
  descriptions matching spec Automatable Tasks table
- Aligned automation_profiles.md threshold descriptions with spec
- Added spec section references in phase_reversion.md, error_recovery.md,
  and plan_execute.md for field naming context
- Extended repository roundtrip test to assert all 11 threshold fields
- Fixed benchmark _make_profile() passing safety fields as top-level
  kwargs instead of via SafetyProfile sub-model (incompatible with
  extra="forbid")
- Aligned CLI JSON/YAML output structure for automation-profile show
  with the specification grouped format (phase_transitions,
  decision_automation, self_repair, execution_controls)
- Moved safety boolean fields into the Execution Controls section
  of Rich output per spec examples
- Reverted auto profile description to "Fully automatic except apply"
  per specification (line 16703, line 28406)
- Improved bridge comments in test steps with semantic context for
  threshold-to-gate mappings

ISSUES CLOSED: #902
2026-03-30 13:18:07 +01:00
hamza.khyari c97ec273a9 feat(lsp): add missing LspServerConfig model fields
Add specification-required fields to LspServerConfig:

- description: str (max_length=1000, default='')
- transport: LspTransport enum (stdio/tcp/pipe, default=stdio)
- initialization: dict[str, Any] (LSP initializationOptions, default={})
- workspace_settings: dict[str, Any] (workspace/didChangeConfiguration, default={})

All fields have defaults for backward compatibility with existing
configs. LspTransport enum exported from lsp package.

Tests: 17 Behave scenarios, 5 Robot integration tests.
Existing LSP tests: 250/250 pass unchanged.

ISSUES CLOSED: #835
2026-03-30 11:51:27 +00:00
freemo 9cace15d6e test(e2e): workflow example 18 — container with remote repo clone (trusted profile)
E2E test for Specification Workflow Example 18: Container with Remote
Repo Clone using the trusted automation profile.  Exercises the full
container-instance workflow per the spec:

- Registers a container-instance resource with --clone-into flag (new
  CLI flag added to resource add for REPO_URL:CONTAINER_PATH).
- Two-step project creation: project create + project link-resource.
- Dynamic LLM actor selection (Anthropic/OpenAI) matching available
  API keys.
- Plan use with --execution-environment container,
  --execution-env-priority fallback, and --automation-profile trusted.
- Full plan lifecycle: strategize, execute, diff, apply, status.
- Positive assertions: non-empty output, plan ID presence, terminal
  state verification, container/clone evidence logging.
- Negative assertions: Traceback and INTERNAL error checks.
- ULID regex uses Crockford base32 character set.
- Unique suffix for resource/project names (parallel CI safe).
- Skip If No LLM Keys guard for graceful degradation.
- CHANGELOG.md updated.

ISSUES CLOSED: #764
2026-03-30 19:17:41 +08:00
aditya 4e64544aae feat(acms): projects with 10,000+ files index without timeout
Add explicit indexing timeout bounds and propagate timeout_seconds through the repo indexing service, utility walker, and CLI entrypoint so large repositories fail predictably instead of hanging.

Add 10K-scale Behave/Robot verification and a dedicated benchmark path to validate that large-project indexing completes, persists status correctly, and remains measurable for regressions.

Address review feedback by splitting oversized Behave and Robot support files into focused modules, trimming repo_indexing_service.py back under the 500-line limit, and rebasing the branch onto current master.

ISSUES CLOSED: #851

# Conflicts:
#	CHANGELOG.md
#	robot/helper_m5_e2e_verification.py
2026-03-30 10:18:19 +00:00
aditya f0442e835d feat(acms): plan execution leverages ACMS context for LLM calls
Integrate ACMS execute-phase context assembly into LLMExecuteActor and inject assembled context into execute prompts with resilient fallback when assembly fails or returns empty output.

Wire the plan CLI executor to use an ACMS-backed execute context assembler, add Behave coverage for context injection/fallback/empty context, and extend Robot M5 verification helpers to assert execute-phase ACMS context usage.

ISSUES CLOSED: #850
2026-03-30 09:21:51 +00:00
aditya ef8f9640fa feat(acms): context analysis produces meaningful summaries
Add phase-wise context analysis summaries for strategize, execute, and apply, and integrate them into `agents project context inspect` and `agents project context simulate` with per-phase resource counts, size and token utilization, and narrowing diagnostics.\n\nAdd Behave and Robot coverage for empty, single-resource, multi-resource, and budget-constrained context analysis, and tighten the M5 smoke coverage after review feedback with exact summary assertions and a missing-path regression scenario for `agents context add`.\n\nISSUES CLOSED: #849
2026-03-30 08:42:48 +00:00
hurui200320 2651e15854 fix(cli): replace non-existent Container.resolve() with named provider calls
The three plan.py call sites (tree, explain, correct) that originally used
container.resolve(DecisionService) were already fixed in master. This commit
completes the fix by aligning the Robot Framework test mocks to match the
corrected production code:

- robot/helper_m3_decision_validation_smoke.py: mock_container.resolve →
  mock_container.decision_service (plan correct dry-run path)
- robot/helper_m4_correction_subplan_smoke.py: container.resolve →
  container.decision_service (correction subplan path)

Both helpers previously passed silently because MagicMock auto-creates any
attribute, masking the mismatch between mock setup and actual code. Added
spec=Container to both mock containers so that accessing non-existent
attributes raises AttributeError, matching real container behavior.

Updated comment wording in features/container_resolve_crash.feature to
clarify that the bug is fixed and the scenarios serve as permanent
regression guards. The @tdd_expected_fail tag was already removed by the
TDD branch before this fix branch was created; the permanent @tdd_issue
and @tdd_issue_647 tags remain per TDD tag lifecycle rules.

Added CHANGELOG.md entry describing the fix.

ISSUES CLOSED: #647
2026-03-30 04:45:56 +00:00
freemo 845cf61b47 test(e2e): workflow example 4 — multi-project dependency update (supervised profile)
Implement WF04 E2E test coverage for multi-project dependency update
using the supervised automation profile. A parent plan targets four
projects (common-lib + 3 services), spawns child subplans per project,
executes in dependency order, validates per project, and applies in
dependency order.

Key components:
- Robot test (robot/e2e/wf04_multi_project.robot): full supervised
  workflow exercising plan use, strategize, execute, validate, apply,
  and final status. JSON output parsing via raw_decode for resilience.
  Test-level guard skips the entire test (visible SKIPPED in CI) when
  the LLM produces zero subplans, preventing silent AC-4/5/6/7 bypass.
  Count Decision Nodes keyword delegates to the Python helper to avoid
  duplicating recursive tree-walking logic.
- Snapshot helper (robot/e2e/wf04_snapshot_helper.py): collects
  deterministic parent/subplan metadata from the lifecycle service.
  Timestamps normalised to UTC. Depth-guarded count_decision_nodes()
  function (max_depth=50) for safe reuse from Robot keywords.
  sys.path.append (not insert) to avoid shadowing the standard library.
- Behave unit tests (features/wf04_snapshot_helper.feature + steps):
  17 scenarios covering _iso(), _enum_value(), count_decision_nodes(),
  and _build_snapshot() with mocked lifecycle service. Includes tests
  for no-subplans (verifies plan_id, project_scopes, validation_summary),
  with-subplans (verifies concrete mapped project names and subplan_count),
  child-plan-is-None (verifies empty-string defaults), unmapped resources,
  and exact-match aware-datetime UTC normalisation.

ISSUES CLOSED: #750
2026-03-30 11:47:37 +08:00
Luis Mendes 7f078f75a5 fix(sandbox): make commit_all atomic per specification
Changed SandboxManager.commit_all() from partial-commit semantics to
all-or-nothing atomic operation per specification requirement. On
partial failure, already-committed sandboxes are rolled back. Error
reporting indicates which sandbox failed and what was rolled back.
Added Behave scenarios verifying atomicity guarantee.

Hardened atomicity guarantees after code review:

- commit_all and _rollback_committed catch Exception (not just
  SandboxError) so unexpected errors cannot bypass rollback.
  Non-SandboxError exceptions are wrapped in a new AtomicCommitError
  (chaining the original as __cause__) that carries rolled_back_ids
  and failed_rollback_ids attributes so callers can programmatically
  determine rollback outcomes.

- _rollback_committed returns both rolled_back_ids and
  failed_rollback_ids; the error result metadata now carries
  both "rolled_back" and "rollback_failed" keys.

- _rollback_committed iterates in reverse (LIFO) order following
  the standard transaction-log undo pattern.  Clarified in
  docstring that this is distinct from the specification DAG-based
  "top-down" rollback ordering (line 24632).

- TransactionSandbox.rollback() from COMMITTED now raises
  SandboxRollbackError (database commits are irreversible)
  instead of silently transitioning to ROLLED_BACK.

- TransactionSandbox is now classified as non-rollbackable in
  commit_all alongside NoSandbox and committed last in the batch,
  since database COMMIT is irreversible.  Docstring corrected to
  match the raising behavior.

- CopyOnWriteSandbox and OverlaySandbox rollback-from-COMMITTED
  uses rename-based safe_restore() to prevent data loss when
  the copytree step fails after the original was removed.

- CopyOnWriteSandbox and OverlaySandbox rollback() now catches
  Exception (not just OSError), matching the broader catch used
  in _rollback_committed, so non-OSError exceptions from
  safe_restore set the status to ERRORED correctly.

- Extracted shared _fs_utils module (backup_directory,
  safe_restore, compute_diff) with symlink, permission, and
  timestamp preservation (including directory timestamps),
  replacing duplicated per-class _backup_directory and
  _compute_diff methods.

- backup_directory defers directory permissions and timestamps
  to a bottom-up post-walk pass, fixing incorrect mtime
  preservation (POSIX file creation inside a directory overwrites
  its mtime) and preventing restrictive source permissions from
  blocking backup writes.

- backup_directory skips non-regular files (FIFOs, sockets,
  device files) with a warning to prevent hangs on special files.

- CopyOnWriteSandbox and OverlaySandbox commit() now attempts
  to restore the original from the pre-commit backup when the
  file-copy phase fails midway, preventing partial corruption.
  If the restore itself fails the backup is preserved for manual
  recovery (cleanup() still removes it).

- CopyOnWriteSandbox and OverlaySandbox commit() error handler
  now catches Exception (not just OSError) so that unexpected
  errors during the file-copy phase also trigger pre-commit
  backup restoration, preventing partial corruption of the
  original directory.

- Pre-commit backup exception handler catches Exception (not
  just OSError) preventing temp directory leaks on non-OSError
  failures from backup_directory.

- Fixed _pre_commit_backup assignment timing: the backup
  reference is now assigned only AFTER backup_directory()
  succeeds, preventing safe_restore() from corrupting an intact
  original with a partial backup when backup_directory() fails
  (e.g. disk full).

- Rollback from COMMITTED with no pre-commit backup (no changes
  were applied) is now a no-op instead of raising
  SandboxRollbackError, preventing false rollback-failure
  reports in commit_all error metadata.

- commit_all logs a warning when NoSandbox or TransactionSandbox
  instances are present in the batch since their changes cannot
  be rolled back, which breaks the atomicity guarantee.

- Pre-commit backup is skipped when compute_diff returns no
  changes, avoiding a full directory copy for no-op commits.

- GitWorktreeSandbox clears _pre_merge_commit on commit failure
  so the stale value cannot be used by future code.

- commit_all docstring documents Raises clause for AtomicCommitError
  exception wrapping behavior.

- GitWorktreeSandbox.rollback() docstring warns about
  multi-worktree safety when rolling back from COMMITTED.

- Updated SandboxStatus transition diagram in protocol.py to
  clearly show the COMMITTED -> ROLLED_BACK path.

- Added spec-contradiction note (line 45938 vs 19193) in
  commit_all docstring.

- OverlaySandbox rollback from COMMITTED now properly remounts
  OverlayFS for real overlay (unmount, clean upper/work dirs,
  remount) and uses dirs_exist_ok=True for userspace fallback
  to prevent FileExistsError if rmtree silently fails.  The
  merged directory is reset from the restored original,
  preventing stale pre-rollback data from being exposed on
  re-activation via get_path() (which allows ROLLED_BACK status).

- OverlaySandbox rollback from COMMITTED now raises
  SandboxRollbackError if the OverlayFS unmount fails,
  preventing a double-mount attempt that would leave the
  sandbox in an inconsistent state.

- OverlaySandbox rollback from ACTIVE now uses dirs_exist_ok=True
  for userspace fallback to prevent FileExistsError when rmtree
  with ignore_errors=True silently fails.

- CopyOnWriteSandbox rollback from ACTIVE now uses
  dirs_exist_ok=True in copytree to prevent FileExistsError
  when rmtree with ignore_errors=True silently fails, matching
  the fix already applied to OverlaySandbox.

- Non-rollbackable sandboxes (NoSandbox, TransactionSandbox) are
  committed last in the batch so that all rollbackable sandboxes
  commit first; if any rollbackable sandbox fails, none of the
  non-rollbackable sandboxes will have committed yet.

- Moved NoSandbox and TransactionSandbox imports to module level
  in manager.py (no circular dependency exists).

- Pre-commit backups are now created on the same filesystem as
  the original directory (using dir= argument to mkdtemp),
  avoiding cross-device copy overhead and ensuring os.rename
  compatibility.

- safe_restore now renames the target into the mkdtemp directory
  instead of removing the mkdtemp dir first, eliminating the
  residual TOCTOU window between rmdir and rename.

- safe_restore catches BaseException (not just OSError) to
  ensure the original directory is always renamed back on
  unexpected errors, preventing the original from being left
  in the renamed-aside state.

- Added AtomicCommitError exception class to protocol.py
  carrying rolled_back_ids and failed_rollback_ids attributes.

- Exported AtomicCommitError from sandbox package __init__.py
  so callers can import it from the public API.

- Added BDD scenarios: LIFO rollback order, AtomicCommitError
  wrapping with RuntimeError cause and rollback metadata,
  _fs_utils backup/restore coverage, no-change commit rollback
  success, directory timestamp preservation, OverlaySandbox
  merged dir reset after COMMITTED rollback,
  CopyOnWriteSandbox rollback from COMMITTED restores original,
  GitWorktreeSandbox rollback from COMMITTED undoes merge,
  TransactionSandbox rollback from COMMITTED raises
  SandboxRollbackError about irreversible commit.

ISSUES CLOSED: #925

Post-review hardening (PR #1146 review findings):

- OverlaySandbox rollback from COMMITTED with no backup (no-op)
  now skips the merged directory reset entirely, preventing
  unnecessary unmount/remount or re-copy that could fail and
  turn a harmless no-op rollback into a SandboxRollbackError
  during commit_all atomic recovery.

- OverlaySandbox rollback no longer double-wraps
  SandboxRollbackError: the outer except Exception handler
  now has a preceding except SandboxRollbackError clause that
  re-raises directly, avoiding a confusing double-wrapped
  error chain.

- CopyOnWriteSandbox.get_path() now accepts ROLLED_BACK status
  for consistency with OverlaySandbox and the protocol status
  transition table (ROLLED_BACK -> ACTIVE).

- CopyOnWriteSandbox rollback from COMMITTED now resets the
  sandbox copy from the restored original via rmtree+copytree,
  preventing stale pre-rollback modifications from being
  exposed on re-activation.

- rollback_all now catches Exception (not just SandboxError)
  so that unexpected rollback errors do not prevent remaining
  sandboxes from being rolled back, consistent with the pattern
  already used in _rollback_committed.

- commit_all docstring now documents a thread-safety warning:
  the method is not safe for concurrent calls on the same
  plan_id since sandbox commit/rollback runs outside the lock.

- Fixed CHANGELOG.md whitespace inconsistencies (double leading
  spaces on two lines).

Post-review hardening round 2 (PR #1146 automated review):

- safe_restore now uses os.rename (O(1) atomic rename) instead
  of shutil.copytree (O(n) recursive copy) for the main restore
  path, since backup and target are always on the same
  filesystem.  This eliminates the ENOTEMPTY bug where a partial
  copytree failure left target_path partially populated, causing
  the recovery os.rename to fail and strand the original in the
  stale temp directory.

- OverlaySandbox.get_path() now transitions ROLLED_BACK to
  ACTIVE, matching CopyOnWriteSandbox and the protocol
  transition table (ROLLED_BACK -> ACTIVE).

- GitWorktreeSandbox.get_path() now accepts ROLLED_BACK status
  for consistency with all other sandbox implementations and
  the protocol transition table (ROLLED_BACK -> ACTIVE).

- rollback_all now also handles sandboxes in COMMITTED status
  (not just ACTIVE), consistent with the state machine allowing
  COMMITTED -> ROLLED_BACK.

- cleanup_all now catches Exception (not just SandboxError) so
  a single unexpected error does not abort cleanup of remaining
  sandboxes, consistent with _rollback_committed and
  rollback_all.

- Restructured CHANGELOG entry from a single ~90-line paragraph
  into structured sub-bullets for readability.

- Added BDD scenarios: no-op rollback from COMMITTED for
  CopyOnWriteSandbox and OverlaySandbox (zero-change commit),
  commit ordering verification (rollbackable before
  non-rollbackable).

Post-review hardening round 3 (PR #1146 deep automated review):

- cleanup_abandoned now catches Exception (not just SandboxError)
  so that unexpected errors (e.g. raw OSError, PermissionError)
  do not crash the loop and prevent remaining abandoned sandboxes
  from being cleaned up, consistent with cleanup_all,
  rollback_all, and _rollback_committed.

- OverlaySandbox._mount_overlay() now catches
  subprocess.TimeoutExpired (in addition to CalledProcessError
  and OSError), preventing create() from leaving the sandbox in
  PENDING status when mount hangs beyond the timeout.

- OverlaySandbox._unmount_overlay() now catches
  subprocess.TimeoutExpired (in addition to CalledProcessError
  and OSError), preventing cleanup() from leaving the sandbox
  in a zombie state when umount hangs beyond the timeout.

- OverlaySandbox._mount_overlay() validates that overlay paths
  do not contain commas, which would corrupt the OverlayFS mount
  options string (comma is the mount option delimiter).

- GitWorktreeSandbox.commit() now checks git diff return code
  so that a failed diff command raises CalledProcessError instead
  of silently concluding there are no changes and skipping the
  merge.

- safe_restore cleanup of the temporary rollback container now
  runs in a finally block, preventing a temp directory leak
  when the rename fails and the exception is re-raised.

- Fixed misleading BDD step name: "backup path that will cause
  copytree to fail" renamed to "backup path that will cause
  rename to fail" since safe_restore now uses os.rename.
2026-03-29 18:57:15 +01:00
Luis Mendes f678d611bb feat(db): add correction_attempts table per specification DDL
Added CorrectionAttemptModel SQLAlchemy model with all spec-defined columns
(correction_attempt_id, plan_id, original_decision_id, new_decision_id,
mode, guidance, archived_artifacts_path, state, created_at, completed_at).
Added FK constraints to v3_plans and decisions tables. Created Alembic
migration and idx_corrections_plan index. Added repository layer for
CRUD operations.

Addressed code review feedback (rounds 1-12):
- Replaced Any type annotations with typed CorrectionAttemptRecord
  signatures using TYPE_CHECKING imports.
- Replaced fragile time.sleep with deterministic timestamps.
- Added cascade deletion test with PRAGMA foreign_keys=ON.
- Changed update_state() to accept typed enum and datetime params.
- Added guidance non-empty validator with max_length=10_000.
- Added spec-aligned server_default for created_at column.
- Fixed timezone handling in to_domain() and from_domain().
- Added spec-defined lifecycle state transition validation via
  validate_correction_state_transition() domain function.
- Improved FK-violation error messages in create() and update_state().
- Normalised timestamp to millisecond precision matching SQLite
  server_default strftime('%f') output.
- Added auto-set completed_at on terminal transitions.
- Added CorrectionAttemptRecord field validators (strip, non-empty).
- Changed original_decision_id FK from CASCADE to RESTRICT matching
  spec DDL default, preserving correction audit trail.
- Added input validation in update_state() for new_decision_id and
  archived_artifacts_path (empty/whitespace rejection).
- Fixed dirty-session bug by moving validation before ORM mutations.
- Extracted _SQLITE_TIMESTAMP_MS_LEN constant for timestamp truncation.

Addressed thirteenth code review feedback:
- Changed InvalidCorrectionStateTransitionError base class from
  DatabaseError to BusinessRuleViolation per CONTRIBUTING.md exception
  semantics (state transition is a business rule, not a database error;
  prevents incorrect retries by @database_retry decorator).
- Changed new_decision_id FK from SET NULL to RESTRICT matching the
  spec DDL default (no ON DELETE clause) and consistent with the
  RESTRICT approach used for original_decision_id.
- Changed update_state() input validation for new_decision_id and
  archived_artifacts_path from DatabaseError to ValueError per
  CONTRIBUTING.md argument validation guidelines.
- Defensive to_domain() coercion now defaults corrupted state to
  'failed' (terminal) instead of 'pending', preventing re-execution
  of completed/failed corrections with corrupted DB values.
- Extracted format_sqlite_timestamp() helper and SQLITE_TIMESTAMP_MS_LEN
  public constant, removing duplicated timestamp formatting logic
  between from_domain() and update_state().
- Added code comment explaining CASCADE on plan_id FK as a codebase
  convention deviation from spec DDL default.
- Added BDD scenario verifying RESTRICT FK on original_decision_id
  blocks decision deletion.
- Replaced weak cross-plan isolation test with stronger two-plan
  scenario verifying list_by_plan returns only each plan's attempts.
- Fixed hardcoded assertion in step_check_archived_path to use
  context variable.
- 45 BDD scenarios and 5 Robot integration tests.

Addressed fourteenth code review feedback:
- Added ORM-level relationship(cascade="all, delete-orphan") on
  LifecyclePlanModel for CorrectionAttemptModel, consistent with all
  other v3_plans child tables, ensuring ORM-level cascade deletes
  work even when SQLite FK enforcement is disabled.
- Added defensive to_domain() coercion for corrupted guidance column
  (defaults to "[corrupted]" with warning log), consistent with
  existing mode/state coercion pattern.
- Added ValueError guard in format_sqlite_timestamp() rejecting naive
  datetimes per CONTRIBUTING.md fail-fast argument validation.
- Fixed stale spec DDL line reference in CorrectionAttemptModel
  docstring.
- Fixed duplicated docstring on SQLITE_TIMESTAMP_MS_LEN constant.

Addressed fifteenth code review feedback:
- Fixed update_state() to defensively handle corrupted DB state values
  via try/except ValueError coercion to FAILED terminal state with
  warning log, consistent with to_domain() defensive coercion pattern.
- Strengthened RESTRICT FK BDD assertion to verify exception type
  (IntegrityError/DatabaseError) instead of only checking presence.
- Split multi-When/Then cross-plan isolation BDD scenario into
  idiomatic single-When/Then scenarios per Gherkin best practice.
- 53 BDD scenarios (was 45) and 5 Robot integration tests.

ISSUES CLOSED: #920
2026-03-29 16:23:04 +01:00
brent.edwards 3abf25f17f test: add TDD bug-capture test for #1038 — validation add --required flag (#1133)
## Summary

Adds TDD bug-capture tests for bug #1038 — the `agents validation add` command is missing `--required`/`--informational` flags that the specification describes at line 22334 and in numerous walkthrough examples.

### Behave BDD Tests

Four Behave BDD scenarios verify the correct expected behavior:

1. **`--required` flag is accepted** — invokes `validation add --config <file> --required` and asserts exit code 0 and mode `"required"` in output.
2. **`--informational` flag is accepted** — same pattern with `--informational`.
3. **`--required` overrides YAML config mode** — creates a YAML config with `mode: informational`, invokes with `--required`, and asserts the CLI output and service layer both reflect `mode: "required"`.
4. **`--informational` overrides YAML config mode** — creates a YAML config with `mode: required`, invokes with `--informational`, and asserts the CLI output and service layer both reflect `mode: "informational"`.

### Robot Framework Integration Tests

Four Robot Framework integration tests exercise the same CLI paths as the Behave scenarios via a subprocess helper script:

1. **TDD Validation Add Required Flag Accepted** — verifies `--required` is recognised by the CLI.
2. **TDD Validation Add Informational Flag Accepted** — verifies `--informational` is recognised.
3. **TDD Validation Add Required Flag Overrides YAML Config** — verifies `--required` overrides a YAML config with `mode: informational`.
4. **TDD Validation Add Informational Flag Overrides YAML Config** — verifies `--informational` overrides a YAML config with `mode: required`.

The helper script (`robot/helper_tdd_validation_required_flag.py`) uses `typer.testing.CliRunner` with mocked services, following the same pattern as other TDD Robot helpers (e.g., `helper_tdd_plan_apply_yes_flag.py`).

### Tags

All scenarios (Behave and Robot) are tagged `@tdd_bug @tdd_bug_1038 @tdd_expected_fail`. Because the bug is still present, the underlying assertions fail (proving the bug exists) and the `@tdd_expected_fail` tag inverts the result so CI passes. Once bug #1038 is fixed, the tag must be removed.

### Spec Contradiction Note

Rui Hu's investigation (issue #1038 comment #70755) found that the formal CLI reference (specification.md lines 9279–9290) does NOT include `--required`/`--informational` flags — they appear only in walkthrough examples and specification.md line 22334. Additionally, specification.md line 30761 states: "For entity registration commands (`actor add`, `skill add`, `tool add`, `validation add`, …), the YAML configuration file is the sole source of truth — the `--config` file fully defines the entity and no CLI override flags are accepted." The resolution may be to add the flags to the CLI OR to clean up the spec. This is documented in both the feature file header and step definition module docstring.

### Key Design Decisions

- **YAML configs use flat `mode` key** — matches what `Validation.from_config()` actually reads (`config.get("mode", "required")`). The specification's Validation Configuration schema (specification.md line 34109) nests mode under a `validation` key, but the current `from_config()` reads a flat top-level `mode` key. A comment in the step definitions documents this discrepancy.
- **Mock uses `side_effect = lambda v: v`** — the `register_tool` mock returns the actual Validation object passed by the CLI, so the CLI output reflects real CLI processing rather than hard-coded mock configuration. This prevents false positives where the mock masks an incomplete fix.
- **False-positive guard** — the Then step verifying mode checks both CLI output (via `f"mode: {expected_mode}"` prefix match) AND `register_tool` call args to prevent the mock from masking an incomplete fix.
- **Robot helper follows established pattern** — uses `typer.testing.CliRunner` with mocked `_get_tool_registry_service`, matching the pattern from `helper_tdd_plan_apply_yes_flag.py`. Reports real outcome (exit 0 + sentinel on bug fixed, exit 1 on bug present) and relies on `tdd_expected_fail_listener` for result inversion.
- **Symmetric override scenarios** — both directions are tested: `--required` overriding informational YAML config, and `--informational` overriding required YAML config.
- **Positional NAME omitted** — the bug report includes a positional NAME argument, but that is a separate spec inconsistency not under test here.
- **Mutual exclusivity deferred** — testing what happens when both `--required` and `--informational` are passed simultaneously is deferred to the bug-fix PR for #1038.

Closes #1102

Reviewed-on: cleveragents/cleveragents-core#1133
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-28 18:09:27 +00:00
brent.edwards bc6a41deb6 tdd(cli): prevent actor list from triggering database updates (#1151)
## Summary

- Fix bug #797: `agents actor list` no longer triggers database writes (`upsert_actor`, `set_default_actor`) by removing `ensure_built_in_actors()` from `ActorRegistry.list()` and `ActorRegistry.list_actors()`
- Include TDD regression tests from #841 with `@tdd_expected_fail` removed per Bug Fix Workflow
- Update three existing test suites that relied on the old behavior to call `ensure_built_in_actors()` explicitly

## Root Cause

`ActorRegistry.list_actors()` and `ActorRegistry.list()` both unconditionally called `self.ensure_built_in_actors()` before delegating to the actor service. `ensure_built_in_actors()` iterates all configured providers and calls `_actor_service.upsert_actor()` for each — a database WRITE operation. It may also call `_actor_service.set_default_actor()` if no default exists — another WRITE. This means every read-only `agents actor list` command triggered database writes and could prompt for pending migrations on fresh checkouts.

## Changes

### Bug Fix
- **`src/cleveragents/actor/registry.py`** — Removed `self.ensure_built_in_actors()` from `list()` and `list_actors()`. Both methods now delegate directly to the service layer without triggering writes. All write-heavy methods (`add`, `upsert_actor`, `get`, `get_actor`, `remove`, `remove_actor`, `set_default_actor`, `get_default_actor`) still call `ensure_built_in_actors()`.

### TDD Tests (from #841, `@tdd_expected_fail` removed)
- **`features/tdd_actor_list_no_db_update.feature`** — 2 Behave scenarios verifying `upsert_actor` and `set_default_actor` are not called during `actor list`
- **`features/steps/tdd_actor_list_no_db_update_steps.py`** — Step definitions
- **`robot/tdd_actor_list_no_db_update.robot`** — 2 Robot Framework integration tests
- **`robot/helper_tdd_actor_list_no_db_update.py`** — Robot helper script

### Test Adjustments
Three existing test suites relied on the old (buggy) behavior where `list_actors()` called `ensure_built_in_actors()`:
1. **`features/consolidated_actor.feature`** — Scenario updated to explicitly call `ensure_built_in_actors()` before `list_actors()`
2. **`features/steps/tdd_actor_list_validation_steps.py`** + **`robot/helper_tdd_actor_list_validation.py`** (bug #592) — Updated to call `ensure_built_in_actors()` explicitly before CLI invocation
3. **`features/steps/actor_list_empty_steps.py`** (bug #592) — Updated with explicit `ensure_built_in_actors()` call and capturing upsert pattern

## Quality Gates

| Gate | Result |
|------|--------|
| `nox -e lint` |  Pass |
| `nox -e typecheck` |  Pass (0 errors) |
| `nox -e unit_tests` |  Pass (468 features, 12367 scenarios, 0 failures) |
| `nox -e integration_tests` | ⚠️ 6 pre-existing failures (timeouts/OOM) |
| `nox -e coverage_report` |  98% (>= 97% threshold) |

The 6 integration test failures are pre-existing infrastructure issues (SIGTERM/SIGKILL timeouts) unrelated to this change: Container Resolve Crash (3), M3 E2E Verification (2), Resource CLI (1).

Closes #797

Reviewed-on: cleveragents/cleveragents-core#1151
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-28 05:27:14 +00:00
brent.edwards b59100cc6c test: add TDD bug-capture test for #1141 — session create persistence (#1144)
## Summary

Adds TDD bug-capture coverage for bug #1141 (`session create` -> `session list` lifecycle) on metadata branch `tdd/m3-session-create-persist`.

### What changed

- Added Behave bug-capture feature: `features/tdd_session_create_persist.feature`
  - Required tags: `@tdd_bug @tdd_bug_1141 @tdd_expected_fail`
  - Scenario path: `init --force --yes` -> `session list --format json` -> `session create` -> `session list --format json`
- Added Behave step definitions: `features/steps/tdd_session_create_persist_steps.py`
  - Uses root CLI app (`cleveragents.cli.main:app`) for realistic command routing
  - Asserts expected list totals; underlying assertion intentionally fails while bug is present
- Added Robot E2E bug-capture test: `robot/e2e/e2e_session_create_persist.robot`
  - Required tags: `E2E`, `tdd_expected_fail`, `tdd_bug`, `tdd_bug_1141`
  - Includes explicit in-file note to remove `tdd_expected_fail` when #1141 is fixed
- Updated changelog (`CHANGELOG.md`, Unreleased)

### Notes

- The issue subtask referenced `robot/e2e/e2e_session_lifecycle.robot`, which is not present on current `master`; equivalent E2E coverage is implemented in `robot/e2e/e2e_session_create_persist.robot`.
- The branch was force-updated to remove stale merge-based history and keep an atomic, rebase-clean commit for this issue.

## Quality gates

| Gate | Result |
|---|---|
| `nox -s lint` |  pass |
| `nox -s typecheck` |  pass |
| `nox -s unit_tests -- features/tdd_session_create_persist.feature` |  pass (TDD inversion active; underlying assert fails) |
| `nox -s integration_tests -- robot/e2e/e2e_session_create_persist.robot` |  pass |
| `nox -s e2e_tests` |  pass |
| `nox -s coverage_report` |  pass (97.66173849218832%) |
| `nox` (full default suite) |  pass |

## Related issue

Closes #1142

Reviewed-on: cleveragents/cleveragents-core#1144
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent Edwards <brent.edwards@cleverthis.com>
2026-03-28 04:34:47 +00:00
brent.edwards 10df485401 test: add TDD bug-capture test for #1039 — missing validation YAML (#1132)
## Summary

This PR adds a Behave TDD bug-capture test that proves bug #1039 exists: the specification references `validations/unit-tests.yaml` as an example validation configuration file (Workflow Examples, Example 1, Step 1) but the file does not exist in the project's `examples/validations/` directory.

## Changes

- **New feature file:** `features/tdd_missing_validation_unit_tests_yaml.feature` — four scenarios tagged `@tdd_expected_fail @tdd_bug @tdd_bug_1039` that verify:
  1. The file `examples/validations/unit-tests.yaml` exists
  2. The YAML config contains `name: local/unit-tests`
  3. The validation mode is `required`
  4. The description field is non-empty
- **New step file:** `features/steps/tdd_missing_validation_unit_tests_yaml_steps.py` — step implementations that locate the project root and check file existence/content

All four scenarios currently fail with `AssertionError` (the file doesn't exist, proving the bug exists). The `@tdd_expected_fail` tag inverts the results so the test suite passes CI.

## Robot Integration Test

N/A — this bug is purely about a missing example file referenced by the specification. There is no integration-level behavior to test.

## Quality Gate Results

| Gate | Result |
|------|--------|
| `nox -s lint` | passed |
| `nox -s typecheck` | passed (0 errors) |
| `nox -s unit_tests` | 462 features, 12234 scenarios, 0 failures |
| `nox -s integration_tests` | 1672 tests, 0 failures |
| `nox -s e2e_tests` | 37 tests, 0 failures |
| `nox -s coverage_report` | 98% (>= 97% threshold) |

Closes #1103

Reviewed-on: cleveragents/cleveragents-core#1132
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-28 04:15:49 +00:00
brent.edwards 26632f79e9 test: add TDD bug-capture test for #1079 — project context set missing flag (#1130)
## Summary

This PR adds TDD bug-capture tests for bug #1079 (`project context set` missing `--execution-env-priority` flag). Per the Bug Fix Workflow in CONTRIBUTING.md, the first step in fixing any bug is to write a test that proves the bug exists.

### What was done

- **Behave unit tests** (6 scenarios in `features/project_context_set_exec_env_priority.feature`):
  - `project context set --execution-env-priority override` with `--execution-environment` should succeed and persist
  - `project context set --execution-env-priority fallback` with `--execution-environment` should succeed and persist
  - `--execution-env-priority` without `--execution-environment` should be rejected
  - Default to `fallback` when only `--execution-environment` is specified
  - Invalid `--execution-env-priority` value should be rejected
  - `project context show` should reflect persisted `execution_env_priority` in JSON output

- **Robot integration tests** (3 test cases in `robot/project_context_set_exec_env_priority.robot`):
  - Override acceptance with persistence verification
  - Fallback acceptance with persistence verification
  - Full round-trip persistence check

All tests are tagged `@tdd_bug @tdd_bug_1079 @tdd_expected_fail`. The underlying assertions fail (confirming the bug exists — `--execution-env-priority` is "No such option"), and the `@tdd_expected_fail` tag inverts the result so CI passes.

### Bug confirmed

The `context_set()` function in `cleveragents.cli.commands.project_context` does not accept `execution_env_priority` as a parameter. The specification (§Execution Environment Routing, precedence table) requires this flag at precedence level 2 for project-level execution environment priority control.

### Quality gates

| Gate | Result |
|------|--------|
| `nox -s lint` | PASS |
| `nox -s typecheck` | PASS (0 errors) |
| `nox -s unit_tests` | PASS (12236 scenarios, 0 failed) |
| `nox -s integration_tests` | My 3 tests pass (13 pre-existing failures) |
| `nox -s coverage_report` | PASS (98%, threshold 97%) |

Closes #1100

Reviewed-on: cleveragents/cleveragents-core#1130
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-28 02:14:55 +00:00
brent.edwards b46a0e6102 feat(plan): implement real revert-mode re-execution from decision point (#1153)
## Summary

Implements the full revert-mode re-execution pipeline from the specification (§ Correction Flow, Revert Mode), enabling `agents plan correct --mode=revert` to truly re-execute from the targeted decision point rather than only logically marking decisions as superseded.

### Changes

- **Resource rollback**: `CorrectionService.execute_revert` now delegates to `CheckpointService.rollback_to_checkpoint` when a decision-aligned checkpoint exists, performing real `git reset --hard` in the sandbox. Gracefully degrades when checkpointing is unavailable.
- **Reasoning rollback**: Extracts `actor_state_ref` from the target decision's `ContextSnapshot` for downstream LangGraph actor restoration.
- **Guidance injection**: Generates a `user_intervention` decision ID for callers to record correction guidance in the decision tree.
- **Phase transition**: Signals that the plan should re-enter Strategize via `phase_transition_target="strategize"` on the result.
- **Model extension**: Added `checkpoint_restored`, `actor_state_ref`, `user_intervention_decision_id`, and `phase_transition_target` fields to `CorrectionResult` with backward-compatible defaults.
- **Dispatch update**: `execute_correction` now accepts and forwards a `decisions` parameter.

### Design Approach

Uses a signal-based architecture: rather than having `CorrectionService` directly manipulate plan state or LangGraph actors, the enhanced `execute_revert` returns signal fields on `CorrectionResult` that downstream consumers use for the full pipeline. This preserves separation of concerns.

### Tests

- **Behave**: 16 new BDD scenarios in `features/revert_re_execution.feature`
- **Robot**: 7 integration tests in `robot/revert_re_execution.robot`
- **Coverage**: 98% (above 97% threshold)

### Quality Gates

| Gate | Status |
|------|--------|
| `nox -s lint` | Pass |
| `nox -s typecheck` | Pass (0 errors) |
| `nox -s unit_tests` | Pass (12376 scenarios, 0 failed) |
| `nox -s integration_tests` | Pass (1631 passed, 3 pre-existing failures #647) |
| `nox -s e2e_tests` | Pass (37 passed) |
| `nox -s coverage_report` | Pass (98%) |

Closes #844

Reviewed-on: cleveragents/cleveragents-core#1153
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-28 00:57:51 +00:00
brent.edwards efacfd61c3 test: add TDD bug-capture test for #1023 — implicit init requirement (#1113)
## Summary

Adds TDD bug-capture tests proving that CLI commands (`resource add`, `project create`) fail in a fresh environment without explicit `agents init`, even when `CLEVERAGENTS_AUTO_APPLY_MIGRATIONS=true` is set.

Bug #1023 reports that `CLEVERAGENTS_AUTO_APPLY_MIGRATIONS=true` triggers schema migrations on an existing database but does NOT create the database file or its parent directory structure, causing `sqlite3.OperationalError: unable to open database file`.

## Changes

### Behave Tests
- `features/tdd_e2e_implicit_init.feature` — Two scenarios tagged `@tdd_expected_fail @tdd_bug @tdd_bug_1023`:
  - `resource add` in a fresh environment without prior init
  - `project create` in a fresh environment without prior init
  - Both scenarios include output content assertions
- `features/steps/tdd_e2e_implicit_init_steps.py` — Step definitions using Typer's CliRunner with real CLI invocation in a pristine temp directory.
  - Env isolation: saves/removes `CLEVERAGENTS_AUTO_APPLY_MIGRATIONS`, `CLEVERAGENTS_DATABASE_URL`, `CLEVERAGENTS_TEST_DATABASE_URL`, `CLEVERAGENTS_HOME`, `BEHAVE_TESTING`, `CLEVERAGENTS_TEMPLATE_DB`, `CLEVERAGENTS_TESTING_USE_MOCK_AI`
  - CliRunner invoked without `catch_exceptions=False` to ensure proper `@tdd_expected_fail` inversion

### Robot Tests
- `robot/tdd_e2e_implicit_init.robot` — Two test cases tagged `tdd_bug tdd_bug_1023 tdd_expected_fail` (listener inverts results while bug is unfixed).
- `robot/helper_tdd_e2e_implicit_init.py` — Helper script exercising CLI in a fresh environment.
  - Env isolation includes `ROBOT_TESTING`, `CLEVERAGENTS_TEMPLATE_DB`, `CLEVERAGENTS_TESTING_USE_MOCK_AI`
  - Env restoration in `finally` block ensures cleanup even on exceptions
  - CliRunner invoked without `catch_exceptions=False` for correct inversion

### Changelog
- Updated `CHANGELOG.md` with entry for this TDD test addition.

## Quality Gates

| Gate | Result |
|------|--------|
| `nox -s lint` | PASS |
| `nox -s typecheck` | PASS (0 errors) |
| `nox -s unit_tests` | PASS (462 features, 12232 scenarios, 0 failed) |
| `nox -s integration_tests` | PASS (pre-existing failures only; our suite correct) |
| `nox -s e2e_tests` | N/A (TDD test only, no E2E suite changes) |
| `nox -s coverage_report` | PASS (98% coverage, threshold 97%) |

Closes #1033

Reviewed-on: cleveragents/cleveragents-core#1113
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent Edwards <brent.edwards@cleverthis.com>
2026-03-28 00:30:26 +00:00
brent.edwards b51df2ee0f feat(server): add Kubernetes Helm chart for server deployment (#1085)
## Summary

This PR adds Kubernetes Helm deployment support for the CleverAgents server.

Closes #928

### What this PR includes

- Helm chart under `k8s/` with Deployment, Service, optional Ingress, ConfigMap,
  ServiceAccount, Secrets, NOTES, and optional Redis subchart configuration.
- Multi-stage `Dockerfile.server` for server runtime deployment.
- Deployment-focused docs in `k8s/README.md`.
- Behave + Robot + benchmark coverage for chart/deployment wiring.

### Review fixes applied (cycle 11 — hurui200320 review #2687)

**Critical fix:**

1. **CI SHA256 checksum verification fixed** — All 3 Helm install blocks (`unit_tests`, `integration_tests`, `helm` jobs) now save the tarball using its original filename (`helm-v3.16.4-linux-amd64.tar.gz`) instead of `helm.tgz`, so the `.sha256sum` file can correctly locate and verify it. This was causing all 3 CI Helm jobs to fail with "No such file or directory".

**Major fixes (test coverage gaps):**

2. **405 Allow header now tested** — The existing "POST to known path returns method not allowed" scenario now asserts that the `Allow: GET` header is present. Added a second 405 scenario testing `POST /live` for broader `_KNOWN_PATHS` coverage (review item #11).
3. **Security-hardening headers now tested** — New scenario "HTTP responses include security-hardening headers" verifies `content-length`, `x-content-type-options: nosniff`, and `cache-control: no-store` headers are present on HTTP responses.
4. **Lifespan warning logging now tested** — New scenario "Unrecognised lifespan message type logs warning and continues" queues `lifespan.startup` → `lifespan.bogus` → `lifespan.shutdown` and verifies: (a) the app completes the lifespan cycle cleanly, (b) a warning is logged mentioning the unrecognised type.

### Review fixes applied (cycle 10)

**Critical/Major fixes (from hurui200320's prior REQUEST_CHANGES):**

1. **Rebased branch onto master** — Removed merge commit per CONTRIBUTING.md rebase-only policy. Clean linear history restored.
2. **Fixed commit message body** — Replaced literal `\n` sequences with actual newlines. `ISSUES CLOSED: #928` footer is now on its own line after a blank separator.
3. **Moved uvicorn import to module top level** — `from uvicorn import run as uvicorn_run` is now at the top of `src/cleveragents/cli/commands/server.py` per Import Guidelines. Updated test mock target from `uvicorn.run` to `cleveragents.cli.commands.server.uvicorn_run`.
4. **Added SHA256 checksum verification** — All three Helm CLI install blocks in `.forgejo/workflows/ci.yml` now download and verify `helm.sha256sum` before extracting the binary.

**Minor fixes:**

5. **Added `Dockerfile.server` build to CI** — New "Build Docker image (Server)" step in the `docker` job validates the server Dockerfile.
6. **ASGI 405 Method Not Allowed** — Known paths (`/`, `/live`, `/ready`, `/health`) now return 405 with `Allow: GET` header for non-GET methods, per RFC 9110 §15.5.6. Added `_KNOWN_PATHS` frozenset.
7. **WebSocket close protocol fix** — App now calls `await receive()` to consume the `websocket.connect` event before closing. Changed close code from 1000 (Normal Closure) to 1008 (Policy Violation).
8. **Lifespan handler logging** — Unrecognised lifespan message types are now logged as warnings instead of silently consumed.
9. **Security-hardening headers** — `_send_response` now includes `content-length`, `x-content-type-options: nosniff`, and `cache-control: no-store` on all HTTP responses.
10. **`.dockerignore` credential patterns** — Added `*.pem`, `*.key`, `*.p12`, `*.pfx`, `credentials*.json`.
11. **`--log-level` validation** — Constrained to `click.Choice(["critical", "error", "warning", "info", "debug", "trace"])` for clean CLI validation errors.
12. **Reverted unrelated semgrep pre-commit change** — `pass_filenames` and `entry` restored to original values per atomic commit hygiene.
13. **Removed unused `ReceiveCallable` type alias** and `Callable`/`Awaitable` imports from `asgi_app_steps.py`.
14. **Fixed redundant `shutil.which("helm")` check** — `_skip_if_helm_missing` now returns `bool` to eliminate the duplicate check in `_render_chart`.
15. **Improved test deque error handling** — Lifespan test receive mock now raises descriptive `AssertionError` instead of opaque `IndexError`.
16. **Scope type dispatch** — Changed `if/if/if` to `if/elif/elif` for mutually exclusive ASGI scope types.
17. **Dockerfile.server base image** — Standardised to `python:3.13-slim` (floating minor) consistent with CLI Dockerfile.
18. **Dockerfile layer caching** — Split `uv pip install build` and `python -m build` into separate `RUN` instructions.
19. **Removed extraneous double blank line** in Dockerfile.server.

### Deferred items (acknowledged, not in scope)

- PodDisruptionBudget, HorizontalPodAutoscaler, NetworkPolicy — Follow-up for production hardening.
- `appVersion: "1.0.0"` placeholder — Needs tracking issue for release versioning alignment.
- Readiness probe with downstream dependency checks — Documented limitation.
- Cross-system test for probe paths matching ASGI routes — Test enhancement.
- Improved benchmarks (helm template timing vs PyYAML parsing) — Benchmark quality improvement.
- CI DRY violation (Helm install 3×) — Code quality improvement, consider composite action.
- File length limits exceeded (`k8s_helm_chart_steps.py` 551 lines, `helper_k8s_helm_chart.py` 678 lines) — Non-blocking, can be split in follow-up.
- `runAsGroup: 1000` in pod security context — Defense-in-depth improvement.
- HEAD method support on known paths — RFC compliance, does not affect K8s probes.
- `click.Choice` log-level validation via CLI runner test — Test gap.

### Scope note: status-check CI gate

The `status-check` job now includes `integration_tests`, `e2e_tests`, and `helm` in its `needs` list. The `helm` job is new in this PR. The `integration_tests` and `e2e_tests` additions fix previously-missing gate checks — included here since this PR modifies both of those jobs to install Helm.

### Quality gates

- `nox -e lint` 
- `nox -e typecheck` 
- `nox -e unit_tests`  (12,321 scenarios passed, 4 skipped)
- `nox -e integration_tests` — 3 pre-existing failures in unrelated areas (plan correction, resource types)
- `nox -e e2e_tests` — pre-existing failures (LLM API keys not available in local env)
- `nox -e coverage_report`  (**97.7%**)

Reviewed-on: cleveragents/cleveragents-core#1085
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-27 23:53:22 +00:00
brent.edwards 3fa07fb772 test: align #1080 expected-fail tags with TDD schema
Add @tdd_issue and @tdd_issue_1080 so the expected-fail scenario passes the global tag validator during unit_tests.
2026-03-27 22:53:33 +00:00
brent.edwards 5ac1e4eeb3 test: add TDD bug-capture test for #1080 — execution env resolution precedence
Write a Behave feature (features/tdd_exec_env_resolution_precedence.feature)
with three scenarios that capture bug #1080: the ExecutionEnvironmentResolver
does not honour the 6-level execution environment precedence chain defined in
the spec (§Execution Environment Routing).

The critical scenario ("Project-level override beats plan-level fallback")
demonstrates the bug by calling resolve() with plan_env="host" and
project_env="container", where the project has priority "override" and the
plan has priority "fallback".  Per the spec, project override (level 2)
should beat plan fallback (level 4), but the current flat resolver chain
returns "host" (plan always wins).  The @tdd_expected_fail tag inverts
this assertion failure to a CI pass.

Two regression-guard scenarios (without @tdd_expected_fail) verify:
(1) Plan-level override still beats project-level override (level 1 vs 2).
(2) Project-level override beats host default (level 2 vs 6).
Both pass both competing environments to the resolver to ensure the
resolver sees the full context and must choose correctly.

Files added:
- features/tdd_exec_env_resolution_precedence.feature — tagged with
  @tdd_bug @tdd_bug_1080 @mock_only and one @tdd_expected_fail scenario
- features/steps/tdd_exec_env_resolution_precedence_steps.py — step
  definitions with priority validation, TODO markers for #1080 cleanup,
  and module-level imports

Robot test: N/A — ExecutionEnvironmentResolver is a pure domain service
with no I/O or integration boundaries.

ISSUES CLOSED: #1101
2026-03-27 22:53:33 +00:00
brent.edwards 69e2d1f179 test(tdd): add required issue tags for expected-fail session leak test
Use @tdd_issue and @tdd_issue_987 so the TDD tag validation hook accepts this expected-fail bug-capture feature and unit tests can execute normally.
2026-03-27 22:01:03 +00:00
brent.edwards a5888a08b7 test: add TDD bug-capture test for #987 — AutomationProfileRepository session leak
Add a Behave feature with four scenarios that capture the session leak
bug in AutomationProfileRepository.  The upsert() and delete() methods
commit when auto_commit is True but never call session.close() in a
finally block, leaking database sessions.  SessionRepository correctly
uses finally: if self._auto_commit: db_session.close() in every method
— AutomationProfileRepository does not.

The test uses a _TrackingSession subclass of sqlalchemy.orm.Session that
records whether close() was called, providing direct assertion of the bug.
Four scenarios cover: (1) upsert success path, (2) delete success path,
(3) upsert error path, (4) delete error path.  All scenarios are tagged
@tdd_expected_fail @tdd_bug @tdd_bug_987 so CI passes while the bug
remains unfixed.

Fixes from review:
- M1: Added 4th scenario for delete() error path
- M2: Registered engine.dispose() in _cleanup_handlers for all scenarios
- M3: Registered session.close() in _cleanup_handlers; explicitly close
  setup session in step_given_persisted_profile
- M4: Replaced type: ignore annotations with proper subclass pattern
  (_TrackingSession uses **kwargs: Any; _FailingFlushTrackingSession
  overrides flush() instead of monkey-patching)
- M5: Moved OperationalError import from function body to module level

Robot integration tests are N/A — this is purely a unit-level session
lifecycle bug within a single repository class.

ISSUES CLOSED: #1092
2026-03-27 22:01:03 +00:00
brent.edwards 8b8942817c feat(wf03): add plan prompt test coverage and confidence-threshold pausing verification 2026-03-27 13:35:17 -07:00
hamza.khyari ebafe8cddd fix(cli): implement --mount flag on resource add container-instance
Add --mount flag to 'resource add container-instance' supporting both
resource-reference mounts (--mount local/api-repo:/workspace) and
host-path mounts (--mount /var/config:/config:ro).

- _parse_mount_spec() parses SOURCE:TARGET and SOURCE:TARGET:MODE formats
- Resource-reference mounts (containing /) get type='resource', mode='rw'
- Host-path mounts support explicit rw/ro mode
- Mounts stored as JSON string in resource properties['mounts']
- _format_properties() renders mounts in human-readable format for 'resource show'
- Added mount to container-instance cli_args metadata
- 6 Behave scenarios testing single mount, dual mount, persistence, and error paths
- Updated CHANGELOG

ISSUES CLOSED: #1078
2026-03-27 12:57:03 +00:00
hamza.khyari 6a8f724299 feat(lsp): add missing LspCapability enum values
Align LspCapability enum with the spec's 11-capability set
(docs/specification.md lines 20705-20717):

Renamed: TYPE_INFO -> HOVER, SYMBOLS -> DOCUMENT_SYMBOLS,
         FORMAT -> FORMATTING
Added:   DEFINITIONS, SIGNATURE_HELP, WORKSPACE_SYMBOLS

Updated LspToolAdapter._CAPABILITY_TOOL_MAP (11 entries):
- code_actions suffix -> code-actions (hyphen per spec)
- RENAME gets dedicated schema with new_name parameter
- workspace_symbols gets query-based schema

Updated _input_schema_for() with 4 schema categories extracted
to module-level constants: file-only, position-based, rename
(with new_name), and query-based. Added defensive ValueError for
unmapped capabilities. Added additionalProperties: false.

Extended LspClient.initialize() to advertise all 11 capabilities
in the initialize request (hover, definition, references, rename,
codeAction, formatting, signatureHelp, documentSymbol, workspace
symbol).

Fixed _make_runtime_handler() to handle workspace_symbols as
query-only input (no file_path required), resolving the
schema/handler contract mismatch.

Fixed stale references: type_info -> hover, symbols ->
document_symbols in test fixtures, feature files, CLI docstring.

Updated docs/reference/lsp.md and CHANGELOG.md.

Behave tests (38 scenarios): enum completeness, tool spec generation
with structural validation, input schema per category, all 11
stubbed provider keys, negative tests for invalid capability and
defensive ValueError branch.

ISSUES CLOSED: #834
2026-03-27 11:10:56 +00:00
hamza.khyari 7bb77aee5c feat(resource): implement ResourceHandler sandbox and checkpoint methods
Extend the ResourceHandler protocol with four sandbox/lifecycle methods:
create_sandbox (idempotent), create_checkpoint, rollback_to, project_access.

Fixes from self-review:
- Instance-level checkpoint dict instead of ClassVar (no shared state)
- Git rollback uses 'git reset --hard' + 'git clean -fd' (not just checkout)
- FsDirectory rollback skips .git to preserve git metadata
- project_access separates ImportError (local mode) from ValueError (reject)
- Added discard_checkpoints() cleanup method
- Git rollback test verifies file content restored

ISSUES CLOSED: #836
2026-03-27 10:50:15 +00:00
hurui200320 1878998b7a refactor(testing): rename tdd_bug/tdd_bug_N tags to tdd_issue/tdd_issue_N
Rename the TDD tag system from tdd_bug/tdd_bug_<N> to tdd_issue/tdd_issue_<N>
across the entire codebase. The tdd_expected_fail tag is unchanged.

The TDD expected-failure workflow is not limited to bug fixes — it applies
equally to any issue type (features, tasks, refactors). The _bug suffix was
misleading and narrowed the perceived scope. The new _issue suffix accurately
reflects that the TDD tagging system applies to any Forgejo issue.

Changes span 92 files:
- features/environment.py: validate_tdd_tags(), should_invert_result(), and
  apply_tdd_inversion() updated — regex, variables, error messages
- robot/tdd_expected_fail_listener.py: _validate_tdd_tags(), _should_invert_result(),
  start_test(), end_test() updated consistently
- 33 Behave .feature files: all @tdd_bug/@tdd_bug_<N> tags renamed
- 29 Robot .robot files: all tdd_bug/tdd_bug_<N> tags renamed
- 3 Robot fixture files renamed (tdd_bug_alone, tdd_missing_tdd_bug,
  tdd_expected_fail_missing_bug_n) with content and references updated
- Tag validation tests and helpers updated (function names, command dispatch
  keys, output strings, fixture references)
- CONTRIBUTING.md: section renamed from 'TDD Bug Test Tags' to
  'TDD Issue Test Tags', all tag references and examples updated
- noxfile.py: comment references updated
- Step definition files, mock helpers, and benchmark files: docstring
  references updated

ISSUES CLOSED: #965
2026-03-27 05:58:35 +00:00
Luis Mendes 01efa2922c fix(test): align behave BDD scenarios with updated lifecycle-apply behavior
Updated automation_levels.feature to expect "applied" processing state
instead of "queued" after completing execute on a full_automation plan,
matching the new auto_progress() behavior that drives Apply to terminal
state.

Updated plan_cli_coverage_boost_steps.py to properly mock the full apply
lifecycle (get_plan → apply_plan → start_apply → complete_apply) instead
of returning a single plan from apply_plan, matching the updated
lifecycle_apply_plan command handler that now drives plans through to the
terminal applied state.

Refs: #753
2026-03-26 17:35:34 +00:00
CoreRasurae e4c01492d5 refactor(cli): align actor run signature with spec positional args
Aligned the `agents actor run` command signature with the
specification by introducing positional NAME and PROMPT arguments.
The --config/-c option is preserved as an optional fallback for
direct YAML invocation. When NAME is provided without --config,
the actor is resolved from the Actor Registry.

Updated both actor_run.py and actor.py run commands. Added backward
compatibility: if --config is provided, it takes precedence over
name-based resolution.

Review fixes applied (code review round 1):
- P1-1: Narrowed bare `except Exception` to `except NotFoundError`
  in _resolve_config_files to avoid masking infrastructure errors.
- P1-2: Moved _resolve_config_files call inside the try block in
  run() so container/registry init errors get user-friendly messages.
  Added `except click.exceptions.Exit: raise` to let typer.Exit
  propagate through the broadened try scope.
- P1-3: Added atexit.register cleanup for temp files created by
  _resolve_config_files (resource leak fix).
- P1-4: Added CHANGELOG.md entry for the breaking CLI change.
- P2-1: Extracted duplicated _resolve_config_files to shared module
  `_resolve_actor.py`; both actor.py and actor_run.py now import it.
- P2-2: Added guard for actors with no configuration data
  (config_blob=None) to produce a clear error instead of invalid YAML.
- P2-3/P2-4: Added 5 BDD scenarios exercising the real
  resolve_config_files function (registry path, yaml_text path,
  config_blob fallback, no-config-data error, not-found error).
- P2-5: Added @coverage tags to all new BDD scenarios.
- P3-1: Added timeout=120s and on_timeout=kill to Robot tests.
- Fixed rxpy_route_validation.robot tests that used the removed
  --prompt/-p option (replaced with positional NAME + PROMPT args).

Review fixes applied (code review round 2):
- P2-1: Aligned actor_run.py exception handler from
  `CleverAgentsException` to `CleverAgentsError`, matching actor.py
  so infrastructure errors from registry resolution get user-friendly
  messages instead of falling through to the generic handler.
- P2-2: Changed `yaml.dump` to `yaml.safe_dump` in _resolve_actor.py
  for fail-fast behavior on unexpected types, consistent with the
  codebase's dominant pattern.
- P3-6: Replaced defensive `getattr(actor, ...)` calls with direct
  Pydantic model attribute access (`actor.yaml_text`, `actor.config_blob`)
  for type-checker coverage.
- P3-1: Switched BDD temp file cleanup from post-assertion
  `unlink()` to `context.add_cleanup()` for leak-proof teardown.
- P2-3/P3-2/P3-3/P3-4: Added 3 BDD edge-case scenarios
  (empty config_blob dict, infrastructure error propagation, empty
  string name) and 1 Robot test case (actor_app registry resolution).

Review fixes applied (code review round 3):
- P1-1: Migrated 48 remaining `-p` invocations across 9 Robot test
  files to the new positional `NAME PROMPT` pattern (context_delete_all_yes,
  load_context_test, scientific_paper_e2e_test, routing_prefix_stripping,
  scientific_paper_basic, scientific_paper_writer_test,
  context_management_test, initial_next_command_test,
  system_prompt_template_rendering).
- P2-1: Documented `--config/-c` as a spec deviation in
  `_resolve_actor.py` module docstring (spec lines 4562-4566 define
  `actor run` with no --config option; issue #901 AC accepts keeping
  it as optional).
- P2-2: Corrected `--config` help text from "fallback" to "overrides
  registry-based name resolution" — the option takes precedence, not
  the other way around.
- P3-1: Added `.strip()` to `yaml_text` emptiness check in
  `_resolve_actor.py` to handle whitespace-only values that would
  otherwise bypass the `config_blob` fallback.
- P3-2: Added `from None` to the no-configuration-data
  `typer.Exit(code=2)` for consistency with the not-found path.
- P3-3: Added BDD scenario testing `--config` precedence for
  `actor_run_app` (was only tested for `actor_app`).
- P3-4: Strengthened config_blob BDD scenario to verify generated
  YAML is parseable via `yaml.safe_load` round-trip.
- P3-7: Replaced hardcoded `/tmp/dummy.yaml` with
  `tempfile.gettempdir()` for portability.
- P3-8: Moved 5 inline imports to module level per CONTRIBUTING.md
  §1292-1294 (3x `import click`, 1x InfrastructureError in steps;
  1x `import typer` in robot helper).
- P3-9: Added `encoding="utf-8"` to `_write_yaml` in Robot helper
  for consistency with production code.

Review fixes applied (code review round 4):
- P2-1: Wrapped `yaml.safe_dump` in `_resolve_actor.py` with
  `try/except yaml.YAMLError` so non-serialisable config_blob values
  produce a user-friendly error message instead of a raw traceback.
- P2-2: Moved remaining inline `import yaml` to module level in
  `actor_run_signature_steps.py` per CONTRIBUTING.md §1292-1294.
- P2-3: Replaced 3 bare `assert` statements in Robot helper
  `helper_actor_run_signature.py` with diagnostic `if/print/sys.exit`
  pattern matching the rest of the file, improving failure diagnostics.

Review fixes applied (code review round 5):
- P3-4: Replaced per-call `atexit.register(lambda)` in
  `_resolve_actor.py` with a module-level `_temp_files` set and a
  single `atexit` handler (`_cleanup_temp_files`) to prevent
  unbounded handler accumulation in same-process usage (e.g. test
  suites running multiple CliRunner invocations).
- P3-1/P3-2/P3-3: Added 4 BDD scenarios: whitespace-only
  `yaml_text` fallback to config_blob, config-precedence
  registry-not-consulted assertion for both `actor_app` and
  `actor_run_app`, multiple `--config` files with positional NAME.
- P4-1: Strengthened error-path BDD assertions to verify error
  message content (not-found, no-config-data, serialisation-error)
  alongside exit codes via captured stderr.
- P4-2: Added Robot test case for `actor_app` unknown-name error
  path and corresponding helper function.

ISSUES CLOSED: #901
2026-03-26 13:41:31 +00:00
CoreRasurae 00881a3e5f feat(observability): implement Event System Domain Event Taxonomy (full EventType enum + DomainEvent model)
Verified and completed the Event System Domain Event Taxonomy implementation.

Gap analysis:
- EventType enum (38 types across 12 domains): ALREADY EXISTED, verified complete
- DomainEvent Pydantic model (all 9 fields): ALREADY EXISTED, verified complete
- EventBus Protocol (emit + subscribe): ALREADY EXISTED, verified complete
- ReactiveEventBus (RxPY Subject, stream property): ALREADY EXISTED, verified
- LoggingEventBus (structlog-based): ALREADY EXISTED, verified complete

Added:
- In-memory audit_log on ReactiveEventBus (list[DomainEvent] with defensive copy)
- Emit ordering aligned with specification: RxPY stream push, then audit append,
  then handler dispatch (§Event System)
- Clarified audit_log docstring: volatile in-memory log, not durable SQLite
  persistence (durable persistence wired separately via audit service layer)
- Behave feature: features/observability/event_system_taxonomy.feature (36 scenarios)
- Step definitions: features/steps/event_system_taxonomy_steps.py
- Robot integration test: robot/event_system_taxonomy_integration.robot (13 test cases)
- Robot helper: robot/helper_event_system_taxonomy.py (13 subcommands)
- ASV benchmark: benchmarks/bench_event_bus.py (5 benchmark suites)
- vulture_whitelist.py: added audit_log property

Code review fixes applied:
- Reordered emit() to match spec: on_next → audit_log → handlers (was: audit → on_next → handlers)
- Clarified audit_log docstring to distinguish volatile in-memory log from spec SQLite table
- Fixed benchmark TaxonomyAuditLogSuite: parameterized setup instead of inline construction
- Added teardown to TaxonomyLoggingEmitSuite to restore logging.disable(NOTSET)
- Clear _received list between benchmark iterations to reduce noise
- Catch specific pydantic.ValidationError in frozen model mutation tests
- Consolidated duplicate singular/plural step definitions
- Tightened EventType count threshold from >=30 to >=38

Quality gates:
- lint: PASSED
- typecheck: PASSED (0 errors)
- unit_tests: 36/36 new scenarios PASSED (pre-existing 12 flaky failures unchanged)
- integration_tests: 1483/1483 PASSED
- security_scan: PASSED
- dead_code: PASSED

ISSUES CLOSED: #587
2026-03-26 11:52:52 +00:00
aditya 37b6d27d0b feat(acms): implement builtin/context skill for CRP
Wire the three CRP tool handlers in context_ops.py to the ACMS pipeline
and ContextTierService, replacing NotImplementedError stubs with
functional implementations:

- request_context: Sources fragments from ContextTierService, filters
  by query/focus keywords, and delegates to ACMSPipeline.assemble()
  for budget-constrained context assembly. Accepts optional plan_id
  for actor-context invocation.

- query_history: Searches across hot/warm/cold tiers for fragments
  matching the query string via case-insensitive substring matching.
  Returns results sorted by last-accessed timestamp.

- get_context_budget: Returns current token budget state (max, reserved,
  available, used) computed from ContextBudget defaults and hot-tier
  fragment token counts.

Additionally:
- Register ACMSPipeline as a Singleton in the DI container
- Create build_context_skill_definition() factory for SkillRegistry
  auto-registration of the builtin/context skill
- Replace 3 obsolete NotImplementedError test scenarios with 8 new
  functional BDD scenarios covering all handler paths

Lint, typecheck, and coverage (98%) all pass.

ISSUES CLOSED: #873
2026-03-26 07:52:09 +00:00
hurui200320 414abb1396 fix(cli): add missing --yes flag to plan apply command (#1127)
## Summary

Adds the `--yes`/`-y` flag to the `lifecycle-apply` CLI command as required by the specification (`agents plan apply [--yes|-y] <PLAN_ID>`). Without `--yes`, a confirmation prompt now displays before proceeding with the destructive Apply phase. With `--yes`, the apply proceeds immediately without prompting.

Closes #932

## Changes

### Source Code
- **`src/cleveragents/cli/commands/plan.py`**: Added `yes: Annotated[bool, typer.Option("--yes", "-y", help="Skip confirmation prompt")] = False` parameter to `lifecycle_apply_plan`. Added `typer.confirm()` prompt before the apply operation, consistent with the pattern used by `rollback_plan`, `correct_plan`, and other destructive commands.
  - Confirmation prompt text matches spec exactly: `"Apply changes for plan {plan_id}?"` producing `Apply changes for plan <ID>? [y/N]:`.
  - Fixed redundant plan ID display when `pre_plan` is `None` — now shows `"Apply changes for plan X?"` instead of `"Apply plan X (X)?"`.
  - Added `except ValueError` handler consistent with sibling commands `lifecycle_execute_plan` and `_lifecycle_apply_with_id`.
  - Added `except Exception` catch-all handler with `isinstance(e, (typer.Abort, typer.Exit))` re-raise guard, consistent with `lifecycle_execute_plan`.
  - Moved `PlanPhase` and `ProcessingState` imports to module level per CONTRIBUTING.md §Import Guidelines.

### TDD Tag Removal (Bug Fix Workflow)
- **`features/tdd_plan_apply_yes_flag.feature`**: Removed `@tdd_expected_fail` tag (leaving `@tdd_bug` and `@tdd_bug_932` as permanent regression guards).
- **`robot/tdd_plan_apply_yes_flag.robot`**: Removed `tdd_expected_fail` tag (leaving `tdd_bug` and `tdd_bug_932`).

### Test Updates
Updated all existing `lifecycle-apply` invocations across 17 test/benchmark files to pass `--yes`, since the new confirmation prompt would otherwise abort in non-interactive test environments:
- 9 Behave step definition files
- 3 Robot Framework helper scripts
- 2 Robot Framework e2e acceptance tests
- 3 ASV benchmark files (4 invocations: `cli_robot_flow_bench.py` ×2, `m1_sourcecode_smoke_bench.py` ×1, `plan_cli_smoke_bench.py` ×1)

### Confirmation Prompt Tests (New + Strengthened)
- **`features/tdd_plan_apply_yes_flag.feature`**: 5 scenarios total:
  - `lifecycle-apply recognises the --yes long flag` — verifies flag acceptance, prompt suppression, exit code 0, and `apply_plan` was called
  - `lifecycle-apply recognises the -y short flag` — same as above for short flag
  - `lifecycle-apply without --yes prompts for confirmation and user declines` — verifies `"Apply cancelled."` message, `exit_code == 0`, and `apply_plan` was NOT called
  - `lifecycle-apply without --yes prompts for confirmation and user accepts` — verifies prompt appears, `exit_code == 0`, and `apply_plan` was called
  - `lifecycle-apply catches unexpected exceptions cleanly` — verifies `"Unexpected error"` output, no traceback leak, non-zero exit code (exercises the `except Exception` catch-all)
- **`features/steps/tdd_plan_apply_yes_flag_steps.py`**: Refactored step definitions:
  - `_make_mock_plan` uses `PlanPhase` and `ProcessingState` enum types instead of raw strings
  - `_make_mock_plan` uses `datetime.now(tz=UTC)` instead of timezone-naive `datetime.now()`
  - Unified prompt suppression step handles both `--yes` and `-y` via parameterised step pattern
  - Added `When` step for unexpected error scenario with `RuntimeError` side_effect
  - Added `Then` step for non-zero exit code assertion
- **Feature/Robot documentation**: Updated stale descriptions that said "implementation does not accept --yes" to reflect the flag is now implemented.

### Documentation
- **`docs/reference/plan_cli.md`**: Updated `lifecycle-apply` section with:
  - `### Synopsis` heading with code block
  - `### Options` table listing `--yes/-y` and `--format/-f` flags
  - `### Arguments` table listing `PLAN_ID`
  - Matches the style used by other command sections in the same file

## Review Fixes (Cycle 3 — Luis's review)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| M1 | Medium | `typer.Abort()` on user decline produces exit code 1 and redundant "Aborted." | Changed to `raise typer.Exit(0)` — consistent with `correct_decision` and legacy `apply` |
| M2 | Medium | Missing exit code assertion on decline scenario | Added `And the lifecycle-apply exit code should be 0` to the decline scenario |
| M3 | Medium | Spec compliance: "summary of pending changes" not implemented | Deferred — spec example shows summary *after* confirmation, not before; implementation matches spec. Ticket-vs-spec ambiguity noted. |
| L1 | Low | Missing `except Exception` catch-all handler | Added catch-all matching `lifecycle_execute_plan` pattern; re-raises `typer.Abort`/`typer.Exit` |
| L2 | Low | Documentation description not updated | Expanded description in `plan_cli.md` to explain confirmation prompt and `--yes` |
| L3 | Low | Dead code `is not None` guards | Removed both guards — `get_plan()` raises `NotFoundError`, never returns `None` |
| I1 | Info | Duplicate `PlanPhase` import | Hoisted import to top of `try` block, eliminating duplicate at old line 2087 |
| I2 | Info | `typer.confirm` without explicit `default=False` | Added `default=False` for consistency with sibling commands |
| L4 | Low | No test for `--yes` after positional arg | Not addressed — Typer/Click handles both orderings; low risk |
| L5 | Low | No test for auto-select + interactive prompt | Not addressed — separate concern outside ticket scope |
| I3 | Info | Robot helper only tests flag recognition | By design — noted as informational |

## Review Fixes (Cycle 4 — Self-QA)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| Major-1 | Major | No test for `except Exception` catch-all handler | Added new scenario `"lifecycle-apply catches unexpected exceptions cleanly"` with `RuntimeError` side_effect; asserts `"Unexpected error"` output, no traceback, non-zero exit |
| Minor-2 | Minor | Missing `ValueError` handler inconsistent with siblings | Added `except ValueError as e:` with `"[red]Execution Error:[/red]"` before catch-all, matching `lifecycle_execute_plan` and `_lifecycle_apply_with_id` |
| Minor-3 | Minor | Flag scenarios don't verify `apply_plan` called | Added `And the lifecycle-apply should have called apply` to both `--yes` and `-y` scenarios |
| Minor-4 | Minor | Stale docstring in Robot helper references `tdd_expected_fail` inversion | Updated to reflect bug is fixed and tests serve as regression guards |
| Minor-5 | Minor | `plan_cli.md` lacks Options table for `lifecycle-apply` | Added Synopsis, Options, and Arguments sections matching sibling command style |
| Nit-6 | Nit | Duplicated step defs for `--yes` vs `-y` prompt suppression | Unified into single parameterised step `"the lifecycle-apply {flag} output should not contain the confirmation prompt"` |
| Nit-7 | Nit | `datetime.now()` timezone-naive | Changed to `datetime.now(tz=UTC)` |
| Nit-8 | Nit | `_make_mock_plan` params use `str` instead of enum types | Changed to `PlanPhase` and `ProcessingState` enum types |

## Review Fixes (Cycle 5 — Jeff's approval note)

| ID | Severity | Issue | Resolution |
|----|----------|-------|------------|
| Import-1 | Minor | `PlanPhase`/`ProcessingState` imports inside function body instead of module level | Moved to module-level import per CONTRIBUTING.md §Import Guidelines |

## Known Limitations / Deferred Items

- **M3: Ticket AC mentions "summary of pending changes"** but the spec example only shows `"Apply changes for plan <ID>? [y/N]: y"` without a change summary. The implementation shows plan ID only (matching the spec), not a change summary. This is a ticket-vs-spec ambiguity; recommend discussing with ticket author.
- **Legacy `apply` command** accepts `--yes` but does not pass it to `_lifecycle_apply_with_id()`. This is a pre-existing issue outside the scope of this ticket.
- **`pre_plan is None` branch** has no explicit test. Pre-existing architectural issue; no action taken.

## Quality Gates

| Gate | Result |
|------|--------|
| `nox -s lint` |  passed |
| `nox -s typecheck` |  passed (0 errors) |
| `nox -s unit_tests` |  passed (471 features, 12,424 scenarios, 0 failures) |
| `nox -s integration_tests` |  passed (1,727 tests, 0 failures) |
| `nox -s e2e_tests` |  passed (41 tests, 0 failures) |
| `nox -s coverage_report` |  passed (≥97% coverage) |

Reviewed-on: cleveragents/cleveragents-core#1127
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Rui Hu <rui.hu@cleverthis.com>
Co-committed-by: Rui Hu <rui.hu@cleverthis.com>
2026-03-26 07:50:09 +00:00
aditya d8d08facde feat(acms): budget enforcement with max_file_size and max_total_size constraints
Implements byte-size budget enforcement for the ACMS context assembly
pipeline.  enforce_size_budget() filters context fragments against
max_file_size (per-fragment) and max_total_size (cumulative) limits
defined in a ContextView.  New domain models BudgetViolation and
BudgetEnforcementResult provide structured violation reporting.

Pipeline integration in ACMSPipeline.assemble() applies enforcement
as a pre-filter when a context_view is provided.

Review feedback addressed:
- Fixed PR milestone to v3.4.0 (matching ticket #847)
- Rebased branch onto latest master
- Added CHANGELOG.md entry
- Extracted duplicated enforcement logic into _apply_budget_enforcement()
  shared method on ACMSPipeline, called by both parent and subclass
- Fixed _make_fragment return type from object to ContextFragment
- Moved all imports to top of files (acms_pipeline.py, steps file)
- Added errors="replace" to encode("utf-8") for surrogate safety
- Documented thread-safety caveat on last_enforcement_result property
- Added short-circuit early return when both limits are None
- Added multi-byte Unicode content test scenario
- Added total_size assertions to key scenarios
- Added violation type verification to mixed limits scenario
- Merged duplicate singular/plural step definitions
- Added edge case scenarios (empty list, all exceed, exact boundary)
- Fixed misleading docstring in ContextAssemblyPipeline.assemble()
- Re-exported VALID_PHASES through core __init__.py public API

ISSUES CLOSED: #847
2026-03-26 07:14:43 +00:00
aditya 968f33c17d test(cli): TDD failing tests for actor list database update (bug #797)
Add TDD tests that prove bug #797 exists: the read-only `agents actor list`
command triggers database writes via `ActorRegistry.list_actors()` calling
`ensure_built_in_actors()`, which invokes `upsert_actor()` and potentially
`set_default_actor()` on every invocation.

Behave scenarios (features/tdd_actor_list_no_db_update.feature):
- Assert upsert_actor is not called during actor list (fails: bug present)
- Assert set_default_actor is not called during actor list (fails: bug present)

Robot Framework tests (robot/tdd_actor_list_no_db_update.robot):
- Integration smoke tests via helper script with same assertions
- Both test cases fail as expected, inverted by tdd_expected_fail_listener

All tests tagged @tdd_expected_fail @tdd_bug @tdd_bug_797 per the mandatory
TDD bug-fix workflow. The inversion mechanism converts the expected failures
to passes in CI. When bug #797 is fixed, the @tdd_expected_fail tag will be
removed and the tests will run normally.

Quality gates: lint pass, typecheck pass, coverage 98.22% (>97%).

ISSUES CLOSED: #841
2026-03-26 06:45:11 +00:00
aditya 42a32c2709 fix(acms): align budget allocation formula and protocol signatures with spec
Align ACMS pipeline protocol signatures and allocation formula with the
specification (docs/specification.md §42630-42937):

- BudgetAllocator.allocate() now accepts optional request: ContextRequest
  parameter per BudgetAllocatorProtocol spec (§44754-44766), enabling
  future request-aware allocation strategies.

- Allocation formula changed from proportional to confidence alone to
  proportional to confidence * quality_score per spec §45003. Both
  DefaultBudgetAllocator and ProportionalBudgetAllocator use the new
  weighted formula. With default quality_score=1.0, behavior is backward-
  compatible.

- DetailDepthResolver.resolve() now accepts budget: int parameter per
  DetailDepthResolverProtocol spec (§44803-44814), enabling future
  budget-aware depth resolution decisions.

- Default packer changed from no-op DefaultBudgetPacker to production
  GreedyKnapsackPacker per spec §45013. ACMSPipeline uses lazy import
  to avoid circular dependency; ContextAssemblyPipeline imports directly.

- Added quality_score field to service-layer StrategyCapabilities
  (already present in domain model).

- Updated 5 feature scenarios and 1 Robot helper where GreedyKnapsackPacker
  relevance-based ordering changed fragment output order vs. the old no-op
  packer. Fixed by adjusting test fragment relevance scores to preserve
  expected ordering.

ISSUES CLOSED: #924
2026-03-26 06:19:14 +00:00
brent.edwards 52511533b5 Merge remote-tracking branch 'origin/master' into merge-master-tmp 2026-03-26 02:58:55 +00:00
brent.edwards 3e704ff9c5 test: add TDD bug-capture test for #1076 — use_action automation_profile propagation (#1116)
## Summary

Add Behave BDD scenarios that capture the bug described in #1076 where `PlanLifecycleService.use_action()` does not resolve or propagate the `automation_profile` from the Action (or any other source in the spec's precedence chain) to the created Plan.

This is the TDD counterpart to bug #1076, following the project's [Bug Fix Workflow](CONTRIBUTING.md#bug-fix-workflow). The test proves the bug exists and will serve as a regression guard once the fix is merged.

### Changes

- **`features/tdd_use_action_automation_profile.feature`** — Three Behave scenarios tagged `@tdd_expected_fail @tdd_bug @tdd_bug_1076`:
  1. Action with `automation_profile="full-auto"` — Plan's `automation_profile` should be `AutomationProfileRef(profile_name="full-auto", provenance=ACTION)` but is `None`.
  2. Action without `automation_profile` but with project-scoped config `"trusted"` — Plan's `automation_profile` should be `AutomationProfileRef(profile_name="trusted", provenance=PROJECT)` but is `None`.
  3. Action without `automation_profile` — Plan's `automation_profile` should resolve to the global default `"supervised"` with `provenance=GLOBAL` but is `None`.

- **`features/steps/tdd_use_action_automation_profile_steps.py`** — Step definitions exercising `PlanLifecycleService.use_action()` and asserting the expected behavior per the specification (docs/specification.md lines 18919, 18967). Shared `_use_action_on_project()` helper eliminates duplicate When step bodies. Guard assertion on `Action.automation_profile` after `create_action()` ensures the Action itself stores the profile correctly.

- **`CHANGELOG.md`** — Entry added under `## Unreleased` describing the TDD test addition.

### How It Works

All three scenarios fail at the assertion level (confirming the bug exists), but the `@tdd_expected_fail` tag inverts the result so the test suite passes CI. When the bug is fixed in #1076, the `@tdd_expected_fail` tag will be removed and the tests will run normally.

### Quality Gates

| Gate | Result |
|------|--------|
| `nox -s lint` | PASS |
| `nox -s typecheck` | PASS |
| `nox -s unit_tests` | PASS (463 features, 12236 scenarios, 0 failures) |
| `nox -s integration_tests` | Pre-existing Pabot infrastructure failure (identical on master) |
| `nox -s e2e_tests` | PASS (37/37) |
| `nox -s coverage_report` | PASS (98%, threshold 97%) |

Closes #1098

Reviewed-on: cleveragents/cleveragents-core#1116
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-26 02:56:45 +00:00
brent.edwards 8ccda647a8 Merge remote-tracking branch 'origin/master' into merge-master-tmp 2026-03-26 02:09:54 +00:00
brent.edwards 0e407f7f19 test: add TDD bug-capture test for #1022 — InvariantService persistence (#1109)
## Summary

Add Behave BDD and Robot Framework integration tests that capture bug #1022 — `InvariantService` stores invariants in an in-memory dict only, losing them across CLI process invocations.

### Motivation

Per the project's TDD Bug Fix Workflow (CONTRIBUTING.md), every bug must first have a test written that captures the buggy behavior before the fix is implemented. This PR fulfills the TDD counterpart issue #1032 for bug #1022.

### What was done

**Behave tests** (`features/tdd_invariant_persistence.feature`): Four scenarios tagged `@tdd_expected_fail @tdd_bug @tdd_bug_1022 @mock_only` that simulate separate CLI invocations via fresh `InvariantService` instances and assert cross-instance data visibility:

1. Project invariant persistence across service instances
2. Global invariant persistence across service instances
3. CLI add/list across separate invocations
4. Cross-instance soft-delete by invariant ID

Step definitions in `features/steps/tdd_invariant_persistence_steps.py`.

**Robot integration tests** (`robot/tdd_invariant_persistence.robot`): Three integration test cases with a helper script (`robot/helper_tdd_invariant_persistence.py`) that exercises add-then-list and add-then-remove across fresh service instances.

### How it passes CI

All tests carry `@tdd_expected_fail`, which inverts the test result via the environment hooks (Behave) and listener (Robot). The underlying assertions fail (proving the bug exists), but are reported as passed to CI. The `@tdd_expected_fail` tag will be removed when bug #1022 is fixed, at which point the tests will run normally and must pass.

### Review fixes applied

- **M1**: Moved `InvariantScope` import from function body in `robot/helper_tdd_invariant_persistence.py` to module-level top imports.
- **m1**: Added `@mock_only` tag to `features/tdd_invariant_persistence.feature` — these tests use purely in-memory services and never touch the database.
- **m2**: Added `exit_code == 0` assertion to `step_invoke_list_cli` in `features/steps/tdd_invariant_persistence_steps.py`, matching the pattern used in `step_invoke_add_cli`.
- **m3**: Changed all `context: Any` parameter types to `context: Context` (from `behave.runner`), following the project-wide convention used in 335+ step files.
- **m4**: Narrowed `except Exception` to `except NotFoundError` in `robot/helper_tdd_invariant_persistence.py`, matching the specific error being tested.
- **n3**: Branch rebased onto latest `master`.

### Quality Gates

- **lint**:  passed
- **typecheck**:  passed (Pyright, 0 errors)
- **unit_tests**:  1 feature, 4 scenarios, 16 steps — all passed via `@tdd_expected_fail`
- **coverage**: ≥ 97% (test-only changes, no source modifications)

Closes #1032

Co-authored-by: Jeffrey Phillips Freeman <the@jeffreyfreeman.me>
Reviewed-on: cleveragents/cleveragents-core#1109
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
2026-03-26 02:06:28 +00:00
brent.edwards 498dca8137 Merge remote-tracking branch 'origin/master' into merge-master-tmp 2026-03-26 01:06:58 +00:00
brent.edwards 3d961fa0aa test: add TDD bug-capture test for #988 — ReactiveEventBus.emit swallows exceptions (#1106)
## Summary

Add a TDD bug-capture Behave scenario that proves bug #988 exists: `ReactiveEventBus.emit()` exception handler logs only `type(exc).__name__` (e.g., "ValueError") without the exception message (`str(exc)`) or traceback (`exc_info=True`). When a subscriber fails, the log contains zero diagnostic detail, making production debugging impossible.

## What was done

- **Feature file**: `features/tdd_event_bus_exception_swallow.feature` — tagged `@tdd_expected_fail @tdd_bug @tdd_bug_988`
- **Step definitions**: `features/steps/tdd_event_bus_exception_swallow_steps.py` — subscribes a handler that raises `ValueError("detailed error message for debugging")`, emits an event, captures the structlog warning via `structlog.testing.capture_logs()`, and asserts:
  1. The exception message text appears in the log entry (scenario 1)
  2. The `exc_info` key is present and truthy, confirming traceback logging (scenario 2)
- **Changelog**: Updated `CHANGELOG.md` with the new entry

## How the test works

1. A `ReactiveEventBus` is created with a subscriber that raises `ValueError` with a distinctive message
2. An event is emitted, triggering the failing handler
3. The `emit()` exception handler catches the error and logs a warning via structlog
4. **Scenario 1** asserts the exception **message** (not just the type name) appears in the log entry
5. **Scenario 2** asserts the log entry includes **`exc_info`** (traceback), per bug #988's acceptance criteria requiring `exc_info=True`
6. Both assertions **FAIL** because the current code only logs `type(exc).__name__` — confirming the bug
7. The `@tdd_expected_fail` tag inverts these failures to CI passes

## Test verification

- `nox -s unit_tests`  passes (462 features, 12,232 scenarios passed, 0 failed)
- Both underlying assertions correctly fail, proving the bug exists
- Tag validation rules pass: `@tdd_bug_988` has corresponding `@tdd_bug`, and `@tdd_expected_fail` has both
- `nox -s lint`  passes
- `nox -s typecheck`  passes (0 errors on changed files)

## Review fixes applied

- **C1 (Critical)**: Rebased onto latest `master` (`5f5ef891`) to eliminate unrelated `docs/timeline.md` regression that was overwriting Day 42 data with stale Day 39 content
- **m1**: Added docstrings to all four step functions
- **m2**: Renamed parameter `ctx` → `context` across all step functions to match project convention (97%+ of codebase uses `context`)
- **m3**: Added second scenario "Bug #988 — emit() logs traceback via exc_info when handler raises" with new `step_then_log_contains_traceback` step that verifies the `exc_info=True` requirement from bug #988 acceptance criteria
- **n1 (Informational)**: Feature-level tags are valid Gherkin; no change needed
- **n2 (Informational)**: `# type: ignore[import-untyped]` on behave imports is the established project convention (106+ files); no change needed

## Robot test

N/A — this is a purely unit-level concern (testing a single class's internal error handling, no external services or IPC involved).

Closes #1093

Reviewed-on: cleveragents/cleveragents-core#1106
Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com>
Co-authored-by: Brent Edwards <brent.edwards@cleverthis.com>
Co-committed-by: Brent Edwards <brent.edwards@cleverthis.com>
2026-03-26 01:02:10 +00:00
brent.edwards 2802af13d0 Merge branch 'master' into tdd/m6-server-connect-non-atomic 2026-03-26 00:13:28 +00:00
brent.edwards 3dee9441a1 Merge remote-tracking branch 'origin/master' into tdd/m6-server-connect-non-atomic 2026-03-25 20:45:08 +00:00
brent.edwards 3ec5114e01 Merge remote-tracking branch 'origin/master' into tdd/m4-correction-checkpoint-wiring 2026-03-25 20:44:56 +00:00
Luis Mendes 34c2acc354 fix(acms): implement context tier runtime promotion/demotion/eviction
Implement the missing runtime logic for the ACMS context tier service.
Previously only data models and manual promote()/demote()/evict_lru()
methods existed. This commit adds:

- Auto-promotion on access: get() now promotes fragments one tier up
  when access_count reaches the configurable promotion_threshold
  (default: 5 accesses).  The access counter resets after each
  successful promotion so fragments must accumulate fresh accesses
  before the next tier transition.
- Staleness enforcement: new enforce_staleness() method demotes hot
  fragments older than hot_ttl (default: 24h) to warm, and warm
  fragments older than warm_ttl to cold. A snapshot of existing
  warm-tier IDs prevents double-demotion in a single pass.
- Budget enforcement on store and promote: store() and promote()
  now enforce TierBudget.max_tokens_hot by evicting LRU hot-tier
  fragments until the token budget is met.  The eviction loop uses
  incremental token tracking to avoid recomputing the sum.
- Event emission: Added TIER_PROMOTED, TIER_DEMOTED, TIER_EVICTED
  event types to EventType enum. All tier transitions emit
  DomainEvent instances through the optional EventBus.
- Configuration: Added context_tier_promotion_threshold,
  context_tier_hot_ttl_hours, context_tier_warm_ttl_hours settings.
  The warm TTL setting also accepts the spec-defined env var
  CLEVERAGENTS_CTX_WARM_HOURS as an alias.
- DI wiring: container.py now injects event_bus into
  context_tier_service.

Review fixes applied (code review on PR #1150):
- C1: Reset access_count to 0 after each auto-promotion to prevent
  chain promotion that bypassed the warm tier.
- C2: Call _enforce_hot_budget() inside promote() warm-to-hot path
  so auto-promoted fragments respect the token budget.
- H1: Corrected _enforce_hot_budget() docstring: actual complexity
  is O(n + n*k) not O(n), since min() scans remaining entries on
  each eviction.
- M1: Added CLEVERAGENTS_CTX_WARM_HOURS as an additional env var
  alias for context_tier_warm_ttl_hours per specification line 30555.

Review fixes applied (second code review on PR #1150):
- B-CRIT-1: Fixed data loss in promote() warm-to-hot: emit
  TIER_PROMOTED before _enforce_hot_budget(), and if the promoted
  fragment is evicted by budget, restore it to the warm tier instead
  of silently losing it.
- B-HIGH-1: Fixed self-eviction on store(): fragments whose
  token_count exceeds the entire hot-tier budget are now redirected
  to the warm tier with a warning log.
- B-MED-1: Wrapped _emit_tier_event() in try/except so a failing
  event bus does not break tier operations (best-effort emission).
- B-MED-2: Fixed event ordering so TIER_PROMOTED fires before any
  budget-triggered TIER_EVICTED events.
- D-LOW-1: Fixed type hint in Robot helper (dict[str, Callable]).
- D-LOW-2: Added __all__ export to context_tiers.py.
- S-LOW-1: Added thread-safety docstring note to ContextTierService.

Review fixes applied (third code review on PR #1150):
- B-MED-1: Added TIER_DEMOTED event emission for oversized fragment
  redirect in store(), closing the observability gap where the only
  tier transition without event emission was the hot-to-warm redirect
  for fragments exceeding the entire hot-tier budget.
- S-LOW-1: Added CLEVERAGENTS_CTX_HOT_HOURS as an additional env var
  alias for context_tier_hot_ttl_hours, for consistency with the
  warm-tier alias CLEVERAGENTS_CTX_WARM_HOURS.
- S-LOW-2: Added docstring note to enforce_staleness() reconciling
  the hot-tier TTL with the specification statement that hot-tier
  retention is "Until resource removed" (TTL controls tier placement,
  not data retention).

Review fixes applied (fourth code review on PR #1150):
- B-HIGH-1: Reset access_count to 0 on demotion so that demoted
  fragments must accumulate fresh accesses before re-promotion.
  Without this reset, a previously popular fragment whose
  access_count already exceeded the promotion threshold would be
  re-promoted on the very next get() call, making staleness
  enforcement ineffective.

Review fixes applied (freemo APPROVED review on PR #1150):
- #1: Removed all # type: ignore annotations from test files.
  Fixed _EventCollector, _FailingBus, and _NullBus subscribe()
  signatures to use Callable[[DomainEvent], None] matching the
  EventBus protocol.  Replaced dict-spread TieredFragment construction
  with explicit keyword arguments and post-construction assignment.
- #2: Extracted runtime policy logic (enforce_staleness,
  _maybe_auto_promote, _re_fetch_after_promotion, _enforce_hot_budget,
  _emit_tier_event) into TierRuntimeMixin in tier_runtime.py to
  reduce context_tiers.py toward the 500-line guideline.
- #3: Added fragment_id non-empty validation guard to promote() and
  demote() per CONTRIBUTING.md argument validation policy.
- #8: Renamed _resolve to _re_fetch_after_promotion for clarity.

Removed @tdd_expected_fail from TDD tests (Behave + Robot) as the
bug is now fixed. All 3 TDD scenarios pass normally.

Tests: 27 Behave scenarios (24 feature + 3 TDD), 4 Robot integration
tests, 4 ASV benchmark suites.

ISSUES CLOSED: #821
2026-03-25 17:49:06 +00:00
freemo 06130212ed test(e2e): workflow example 5 — database schema migration with safety nets (review profile) (#816)
## Summary

E2E test for Workflow Example 5 — database schema migration with safety nets using the **review** automation profile. Exercises the full spec-aligned workflow:

- **Custom resource type registration** via `resource type add --config` (postgres-db type with `transaction_rollback` sandbox strategy, `--host`, `--port`, `--database`, `--schema` CLI args with flat `type`/`default` fields per `ResourceTypeArgument` schema)
- **Custom resource instantiation** — attempts `resource add` with the custom type to exercise mixed resource types, followed by `project link-resource` to link DB resource to the project
- **Custom skill creation** with spec-aligned database tools: `local/query_db` (read-only), `local/execute_migration` (writes, checkpointable), `local/backfill_column` (writes, checkpointable) — registered via `skill add --config`, with namespaced tool reference names per `SkillToolRefSchema` validation
- **Action creation** with `automation_profile: review`, `reusable: true`, `state: available`, spec invariants, and typed `arguments` section (`table_name`, `column_name`, `column_type`, `backfill_source` — all required per spec, using `arguments` field per `ActionConfigSchema`)
- **Plan use** with `--arg` flags exercising parameterized action invocation including `backfill_source=audit_log`, plus **explicit `--automation-profile review`** flag (action-to-plan profile propagation is not yet wired in `PlanLifecycleService.use_action`)
- **Phased child plan verification** via `plan tree --format json` with `decision_count >= 2` hard assertion on framework decisions plus WARN tiers for LLM decomposition quality (`< 3`, `< 5`)
- **Plan phase assertion** — hard assertion that phase is populated after execute
- **Checkpoint-based rollback** with hard assertions: `rc=0` on rollback success, `rc!=0` on fake checkpoint, None guard for JSON null checkpoint IDs, re-execute with Traceback/INTERNAL checks on success and explanatory comment on failure path
- **Plan diff** with hard `rc=0` assertion and content-signal verification
- **Migration content verification** — baseline SHA saved before apply, diff against baseline (not `HEAD~1`), WARN-level check on migration keywords (`last_login`, `schema`, `migration`, `column`, `alter`) — flexible per LLM non-determinism
- **Commit count** assertion `>= 2` (fixture baseline: Create Temp Git Repo + DB fixture commit), WARN if no additional commits from lifecycle-apply
- **Backfill evidence** WARN-level check in plan tree/execution output (`backfill`, `batch`, `populate`, `last_login`) with explanatory comment noting tree covers decomposition plan
- **Combined AC #6 gate** — if *both* migration content *and* backfill evidence are absent, explicit WARN visibility for CI debugging
- **Terminal state assertion** after `lifecycle-apply` — `plan status` call verifies phase/processing_state reflects terminal or apply-progress outcome
- **Automation profile fallback verification** — if `plan use` output omits `automation_profile`, falls back to `plan status` for secondary verification (hard assertion always runs)
- **Traceback and INTERNAL checks** on all CLI commands (resource add, project create, resource type add, skill add, action create, plan use, strategize, execute, plan tree, plan status, plan diff, plan rollback, re-execute after rollback, lifecycle-apply) including custom resource error paths
- **Dynamic actor selection** — detects available API keys (Anthropic/OpenAI) at suite setup
- **Skip If No LLM Keys** guard for graceful CI degradation
- **Test-level teardown** with diagnostic logging for both plan status and plan tree on failure
- **30-minute timeout** covering worst-case rollback+re-execute path
- **Force Tags** for consistency with `m6_acceptance.robot`
- **Timeout parameters** (`timeout=60s on_timeout=kill`) on all local `Run Process` git commands
- **Sequential section numbering** (1 through 15) for readability

Closes #751

ISSUES CLOSED: #751

## Approach

Follows the patterns established by `m6_acceptance.robot` and `m2_acceptance.robot`:
- `WF05 Suite Setup` initialises the workspace, generates a unique run suffix, and detects available LLM API keys
- `Safe Parse Json Field` from `common_e2e.resource` for JSON field extraction with None guards for JSON null values
- All CLI commands use `--format json` for predictable, parseable output
- `expected_rc=None` with explicit `Should Be Equal As Integers` for detailed failure messages
- Hard assertions on infrastructure/framework behavior (CLI commands, phase transitions, tool registration)
- WARN-level assertions on LLM-dependent output (decision decomposition, migration content, backfill evidence, commit count) — per ticket requirement "output validation is flexible"
- Traceback and INTERNAL checks on all CLI commands following `m2_acceptance.robot` pattern
- Baseline SHA approach for post-apply diff verification eliminates false positives from fixture commits

## Bug Fix: LifecyclePlanRepository.update() UNIQUE Constraint Violation

**Root cause**: `LifecyclePlanRepository.update()` called `clear()` on child relationship collections (project_links, arguments, invariants) followed by `append()` with new items, but only flushed at the end. SQLAlchemy's default operation ordering can emit INSERTs before DELETEs within the same flush, causing `UNIQUE constraint failed: plan_arguments.plan_id, plan_arguments.name` when plans have arguments.

**Fix**: Group all three `clear()` calls together and flush them before appending new rows. This ensures the DELETEs are committed before any INSERTs, preventing the UNIQUE constraint violation.

**Impact**: This was a latent bug affecting ALL plans with arguments when `update()` is called. Previously undetected because existing E2E tests (M1, M2, M5, M6) create plans without `--arg` flags.

## Review Fixes (addressing medium findings from @CoreRasurae review)

| # | Finding | Fix |
|---|---------|-----|
| **BUG-1** | No regression test for UNIQUE constraint fix | Added targeted BDD scenario in `repositories_coverage_boost.feature` — creates plan with argument `x=v1`, updates to `x=v2`, asserts no `IntegrityError` |
| **TEST-1** | AC #4 weakened — fragile string counting | Replaced raw `count('"decision_id"')` with proper JSON parsing via `json.loads()`, recursive tree walking for decision counting, structural `children_key_count` and `child_link_count` verification |
| **TEST-2** | AC #5 conditionally tested | Added explicit WARN log when no checkpoint_id is present ("AC #5 visibility"); fake checkpoint test now runs unconditionally (moved outside IF/ELSE) with Traceback/INTERNAL checks |
| **TEST-3** | No terminal state assertion after lifecycle-apply | Added `plan status` call after apply with phase/processing_state extraction; hard assertion on terminal state or apply-phase progress |
| **TEST-4** | AC #6 migration/backfill WARN-only | Added combined gate (`has_ac6_evidence`): if *both* migration and backfill evidence are absent, explicit WARN for CI visibility. WARN-only is intentional per ticket AC "output validation is flexible" |
| **TEST-5** | Automation profile silently skipped | Added fallback to `plan status --format json` when `plan use` output omits `automation_profile`; hard assertion (`Should Be Equal As Strings review`) now always executes |
| **TEST-8** | Missing Traceback/INTERNAL on custom resource error paths | Added Traceback/INTERNAL checks inside both `resource add` and `project link-resource` ELSE branches with `NoSuchOption` guard |

## Quality Gates

- `nox -e lint` 
- `nox -e typecheck`  (0 errors)
- `nox -e unit_tests`  (471 features, 12,422 scenarios, 0 failures)
- `nox -e integration_tests`  (1,727 tests, 0 failures)
- `nox -e e2e_tests`  (42 tests, 42 passed, 0 failed)
- `nox -e coverage_report`  (98%, meets threshold)

## Manual Verification

### Prerequisites
- `OPENAI_API_KEY` or `ANTHROPIC_API_KEY` environment variable set

### Commands

```bash
nox -e e2e_tests
# Or run just this suite:
python -m robot --outputdir build/reports/robot --include E2E robot/e2e/wf05_db_migration.robot
```

Reviewed-on: cleveragents/cleveragents-core#816
Reviewed-by: Luis Mendes <luis.mendes@cleverthis.com>
Co-authored-by: Jeffrey Phillips Freeman <the@jeffreyfreeman.me>
Co-committed-by: Jeffrey Phillips Freeman <the@jeffreyfreeman.me>
2026-03-25 13:05:04 +00:00
aditya 37d55f4321 feat(acms): context assembly CLI functional (context list/add/show/clear)
Verified the context assembly CLI commands (context list, add, show, clear)
are fully functional and integrated with ContextService and the DI container.
Both Robot E2E acceptance tests (ACMS Scoped Context Output Per Phase and
Context Policy Clear And Inheritance Fallback) pass, confirming phase-scoped
view resolution, size limit narrowing, inheritance fallback on clear, and
persistence round-trip correctness.

Added 5 new Behave scenarios to m5_acms_smoke.feature covering the required
unit test patterns: multiple resources management (list, add, show with
multiple files), clear-then-re-add flow, and clear on empty project. All
12,235 unit test scenarios pass. Coverage is 98.38% (threshold >97%).

ISSUES CLOSED: #848
2026-03-25 12:05:40 +00:00
Luis Mendes 3cf3f1f69e feat(validation): implement Fix-then-Revalidate orchestration loop for required validations
Implemented FixThenRevalidateOrchestrator with full diagnosis → self-fix →
re-validation → retry limit → strategy revision → user escalation → terminal
failure flow. Added retry counting per validation per plan, configurable retry
limit (default 3), auto_strategy_revision flag support, and validation_fix_history
recording in plan execution metadata.

Review fixes applied (round 1):
- Added try/except around fix_callback and revalidate_callback (spec: validation
  errors treated as required failure regardless of mode)
- Wired optional EventBus for VALIDATION_FIX_ATTEMPTED/SUCCEEDED/EXHAUSTED events
- Fixed total_attempts derivation from fix_history length instead of cumulative
  retry counts
- Updated module docstring to reflect implemented vs caller-responsible steps
- Added exhausted-retry logging on re-invocation
- Replaced bare assert in Robot helper with explicit _check() function
- Added 10 new Behave scenarios covering exception paths, multi-failure, reset,
  field validation, boundary values, event_bus property, and exhausted re-invocation

Review fixes applied (round 2 — PR #711):
- B1+B7: Fixed TOCTOU race in _fix_single_validation; atomic claim-per-iteration
  under RLock; eliminated defaultdict auto-vivification with .get() reads
- B2: Added validation_name mismatch check on revalidate_callback return
- B3: Added fix_description truncation to 2000 chars before FixAttemptRecord
- T1: Switched threading.Lock to threading.RLock for defensive reentrancy
- A1: Added model_validator preventing escalated+terminal_failure both True
- R1: Added max_length=255 to FixAttemptRecord.validation_name
- S7: Fixed misleading docstring about automation_profile integration
- D4: Fixed timestamp field description to specify UTC
- D1: Renamed misleading benchmark time_fix_after_two_retries
- Added 5 new Behave scenarios for truncation, model constraints, name mismatch

ISSUES CLOSED: #583
2026-03-24 22:01:48 +00:00
freemo e52c79e958 test: add TDD bug-capture test for #1024 — SQLite DB URL CWD resolution
Add Behave BDD scenarios (tagged @tdd_bug @tdd_bug_1024 @tdd_expected_fail)
that verify the default database_url resolves inside CLEVERAGENTS_HOME rather
than the current working directory.  Two scenarios exercise both the container
get_database_url() helper and the Settings model default.

Add Robot Framework integration tests with a helper script exercising the same
resolution paths via subprocess, verifying database URL resolution and CLI
database file placement relative to CLEVERAGENTS_HOME.

The tests currently fail as expected because the bug in #1024 is still present:
the relative SQLite path sqlite:///cleveragents.db resolves against CWD.  The
@tdd_expected_fail tag inverts the result so CI passes.

ISSUES CLOSED: #1034
2026-03-24 20:29:49 +00:00