Compare commits

..

2 Commits

Author SHA1 Message Date
HAL9000 1037e8e3ab fix(database/migration_runner): add check_same_thread=False to get_current_revision() SQLite engine
CI / benchmark-publish (pull_request) Has been skipped
CI / push-validation (pull_request) Successful in 34s
CI / helm (pull_request) Successful in 46s
CI / build (pull_request) Successful in 53s
CI / lint (pull_request) Successful in 1m0s
CI / benchmark-regression (pull_request) Failing after 1m29s
CI / quality (pull_request) Successful in 1m34s
CI / typecheck (pull_request) Successful in 1m42s
CI / security (pull_request) Successful in 1m42s
CI / integration_tests (pull_request) Successful in 3m42s
CI / unit_tests (pull_request) Successful in 4m35s
CI / e2e_tests (pull_request) Successful in 4m35s
CI / docker (pull_request) Successful in 1m29s
CI / coverage (pull_request) Successful in 13m7s
CI / status-check (pull_request) Successful in 4s
- Fix extra blank line in step file causing lint failure (ruff E302)
- Consolidate duplicate step_when_capture_engine_args and step_when_verify_thread_safe_args
  into a shared _run_get_current_revision_and_capture_kwargs() helper
- Remove duplicate Scenario 3 (identical assertion to Scenario 1, different step impl)
- Unify context attribute name to engine_creation_kwargs across all steps
- Add explicit type annotations to all step function signatures
2026-05-05 17:31:40 +00:00
HAL9000 023c0944a7 fix(database/migration_runner): add check_same_thread=False to get_current_revision() SQLite engine
MigrationRunner.get_current_revision() was creating a SQLAlchemy engine
with create_engine(self.database_url) without passing
connect_args={"check_same_thread": False} for SQLite databases.  When
called from a background thread (e.g. async startup flows), SQLite raised
ProgrammingError: SQLite objects created in a thread can only be used in
that same thread.

The sibling method init_or_upgrade() already passes check_same_thread=False
for SQLite, making this an inconsistency in the same class.  Because
get_pending_migrations() and check_migrations_needed() both delegate to
get_current_revision(), the threading bug propagated to all three methods.

This fix adds connect_args={"check_same_thread": False} to the
create_engine() call in get_current_revision() when the database URL
starts with "sqlite", consistent with the existing pattern in
init_or_upgrade().

ISSUES CLOSED: #10507
2026-05-05 17:31:40 +00:00
21 changed files with 14 additions and 1562 deletions
+2 -2
View File
@@ -41,7 +41,7 @@ jobs:
- name: Install uv and nox
run: |
pip install -q uv=${{ env.UV_VERSION }} nox
pip install -q uv==${{ env.UV_VERSION }} nox
- name: Cache uv packages
uses: actions/cache@v3
@@ -126,7 +126,7 @@ jobs:
- name: Install uv and nox
run: |
pip install -q uv=${{ env.UV_VERSION }} nox
pip install -q uv==${{ env.UV_VERSION }} nox
- name: Cache uv packages
uses: actions/cache@v3
+1 -1
View File
@@ -446,7 +446,7 @@ jobs:
needs: [lint, typecheck, security, quality, unit_tests]
runs-on: docker
container:
image: ${{vars.docker_prefix}}docker:dind
image: docker:dind
options: --privileged
steps:
- name: Start Docker daemon and install dependencies
+2 -2
View File
@@ -32,7 +32,7 @@ jobs:
- name: Install dependencies
run: |
python -m pip install -U pip
python -m pip install asv virtualenv uv=${{ env.UV_VERSION }} nox
python -m pip install asv virtualenv uv==${{ env.UV_VERSION }} nox
- name: Restore prior ASV benchmarks
env:
@@ -92,7 +92,7 @@ jobs:
- name: Install dependencies
run: |
python -m pip install -U pip
python -m pip install asv virtualenv uv=${{ env.UV_VERSION }} nox
python -m pip install asv virtualenv uv==${{ env.UV_VERSION }} nox
- name: Sync prior benchmark results from S3
env:
+1 -1
View File
@@ -80,7 +80,7 @@ jobs:
- name: Run quality gates script
run: |
python scripts/check-quality-gates.py --coverage-min 97 --complexity-max F || echo "Quality gates script not found or failed, skipping..."
python scripts/check-quality-gates.py --coverage-min 96.5 --complexity-max F || echo "Quality gates script not found or failed, skipping..."
- name: Generate quality trend data
run: |
+1 -1
View File
@@ -45,7 +45,7 @@ jobs:
build-docker:
runs-on: docker
container:
image: ${{vars.docker_prefix}}docker:dind
image: docker:dind
options: --privileged
needs: [build-wheel]
steps:
@@ -245,16 +245,6 @@ each work group's fetch algorithm:
The prompt body to pass to workers you spawn:
```
Implement or fix the indicated issue or pull request.
PR Compliance Checklist (MANDATORY — complete ALL items before creating a PR):
[ ] 1. CHANGELOG.md — add entry under [Unreleased] section
[ ] 2. CONTRIBUTORS.md — add or update contribution entry
[ ] 3. Commit footer — include `ISSUES CLOSED: #<issue-number>` in the commit message
[ ] 4. CI passes — all quality gates and tests green before requesting review
[ ] 5. BDD/Behave tests — added or updated for the changed behaviour
[ ] 6. Epic reference — PR description references the parent Epic issue number
[ ] 7. Labels — applied via forgejo-label-manager: State/In Review, Priority/<level>, MoSCoW/<level>, Type/<type>
[ ] 8. Milestone — PR assigned to the earliest open milestone matching the issue
```
```
-30
View File
@@ -7,8 +7,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
### Changed
- **[CI] Optimize test execution speed with caching and parallelism (#10889):** Fixed docker runner image references to use the `${{vars.docker_prefix}}` variable instead of hardcoded `docker:dind`, ensuring proper integration with the internal harbor registry. Coverage quality threshold tightened from 96.5% to 97%. Version pins for uv loosened from exact (`==`) to compatible release (`=`) in benchmark workflows for more resilient dependency resolution. Benchmark-regression restored as a PR-only informational job on master.yml, not blocking merges but providing performance regression detection.
- **`agents session list` now displays full 26-character session ULIDs** (#10970): The Rich table
and Summary panel ("Most Recent" / "Oldest") previously showed only the first 8 characters of
each session ULID. This made the output unusable for copy-paste into `session tell`,
@@ -26,16 +24,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
versions (<3.13.4) cannot be installed even if upstream transitive dependencies have loose
version constraints.
### Fixed
- **Implementation Supervisor PR Compliance Checklist** (#9824): Added a mandatory
8-item PR Compliance Checklist to the worker prompt body in `implementation-supervisor.md`
that every implementation worker must complete before creating a PR. Checklist covers:
CHANGELOG.md update, CONTRIBUTORS.md update, commit footer (`ISSUES CLOSED: #N`),
CI verification, BDD tests, Epic reference, label application via `forgejo-label-manager`,
and milestone assignment. This eliminates systemic PR merge blockers caused by workers
omitting required items.
### Changed
- Restored `benchmark-regression` CI job to `master.yml` with `pull_request` trigger guard
@@ -298,14 +286,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
### Added
- **ACMS Index Data Model and File Traversal Engine** (#9579): Implements the
foundational ACMS index data model with structured fields for file metadata
(path, size, last modified, type), tag system, and hot/warm/cold/archive
storage tier assignment. Introduces a timeout-safe large-project file traversal
engine capable of handling 10,000+ files without memory exhaustion through
chunked processing. Provides a complete index entry pipeline for creation,
storage, and retrieval with full queryability by path, tag, type, and recency.
- **ACMS Large-Project Indexing BDD Coverage** (#8726): Added 7 Behave scenarios
covering walk-based indexing of 10,000+ files without timeout, binary-file
skipping, oversized-file skipping, git-checkout indexing, fallback to walk when
@@ -314,7 +294,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
syscalls). Added `timeout=120` to git subprocess calls to prevent CI hangs.
Cached `get_scoped_view` results in `When` steps to avoid redundant re-queries
in `Then` steps.
- **Agent Evolution Pool Supervisor PR Metadata Assignment** (#7888): The
agent-evolution-pool-supervisor now automatically looks up the Type/Automation
label and the earliest open milestone from the repository before dispatching
@@ -621,15 +600,6 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
response format from the OpenCode API `/session/status` endpoint instead of an array.
Workers now dispatch and verify correctly, preventing incorrect session deletion.
---
### Fixed
- **CLI (`agents actor remove`)** (#6491): Restores output parity with the
other actor commands by honoring `--format`/`-f` for JSON/YAML/plain/Rich
envelopes. Adds a Robot Framework regression test to assert the JSON
envelope structure and updates the CLI synopsis in `docs/specification.md`
to document the option.
---
## [3.8.0] -- 2026-04-05
+2 -4
View File
@@ -7,6 +7,7 @@
* Jeffrey Phillips Freeman <jeffrey.freeman@syncleus.com>
* Luis Mendes <luis.p.mendes@gmail.com>
* Rui Hu <rui.hu@cleverthis.com>
* HAL 9000 <hal9000@cleverthis.com>
# Details
@@ -15,7 +16,6 @@ Below are some of the specific details of various contributions.
* Jeffrey Phillips Freeman has acted as Lead Developer, daily contributor, and Project Owner.
* Brent E. Edwards has contributed quality assurance, test coverage, and CI pipeline improvements.
* HAL 9000 has contributed automated implementation, bug fixes, and feature development as part of the CleverAgents automation pool.
* HAL 9000 has contributed the CI optimization and docker infrastructure fixes (#10889): corrected docker runner image references to use harbor prefix, tightened nightly quality coverage threshold, loosened version pins for resilient dependencies, and restored benchmark-regression job as non-blocking PR trigger.
* HAL 9000 has contributed concurrency safety improvements, including thread-safe context tier management (issue #7547) for parallel plan execution.
* HAL 9000 has contributed the plan concurrency race-condition fix (#7989): wired `LockService` into the plan lifecycle, guarding `execute_plan()` and `apply_plan()` with plan-level advisory locks and unique per-invocation owner identities to prevent silent concurrent state corruption.
* HAL 9000 has contributed the bug-hunt-pool-supervisor non-blocking tracking fix: updated step 5 to be best-effort and added rule 9 to prevent the automation-tracking-manager call from blocking the main supervisor loop.
@@ -25,10 +25,8 @@ Below are some of the specific details of various contributions.
* This project was made possible thanks to considerable donation of time, money, and resources by CleverThis, Inc.
* HAL 9000 has contributed automated bug fixes, CLI output formatting improvements, and ongoing maintenance as part of the CleverAgents automation system.
* HAL 9000 has contributed the file edit encoding parameter fix (PR #8258 / issue #7559).
<<* HAL 9000 has contributed the architecture-pool-supervisor milestone assignment feature (PR #8188 / issue #7521): added `forgejo_update_pull_request` permission and documented the PR workflow for major spec changes, enabling automatic milestone assignment for specification PRs.
* HAL 9000 has contributed the architecture-pool-supervisor milestone assignment feature (PR #8188 / issue #7521): added `forgejo_update_pull_request` permission and documented the PR workflow for major spec changes, enabling automatic milestone assignment for specification PRs.
* HAL 9000 has contributed the git worktree TOCTOU race condition fix (PR #8178 / issue #7507): replaced the unsafe mkdtemp() + rmdir() pattern with a parent-directory approach to eliminate the race window in concurrent git worktree operations.
* HAL 9000 has contributed the git_tools TOCTOU race condition fix (PR #8255 / issue #7619): eliminated the Time-Of-Check-To-Time-Of-Use race in `_get_base_env()` by adding double-checked locking with a module-level `threading.Lock`, preventing concurrent threads from writing conflicting environment snapshots.
* HAL 9000 has contributed the mandatory PR compliance checklist to `implementation-supervisor.md` (#9824): added an 8-item checklist to the worker prompt body with concrete items covering CHANGELOG.md, CONTRIBUTORS.md, commit footer, CI verification, BDD tests, Epic reference, labels, and milestone assignment to eliminate systemic PR merge blockers.
* HAL 9000 has contributed comprehensive milestone documentation for v3.6.0 (Advanced Concepts & Deferred Features) and v3.7.0 (TUI Implementation) (PR #9903): split into sub-documents covering context strategies, LLM backends, resource types, A2A rename, container tool execution, scope chain resolution, cost/safety budgets, E2E workflow tests, code review examples, plugin architecture, TUI layout, persona system, reference/command input, session management, configuration, and TuiMaterializer integration.
* HAL 9000 has contributed the LLMTraceRepository data-integrity fix (PR #8185 / issue #7505): replaced the unconditional `session.commit()` in `LLMTraceRepository.save()` with a dual-path implementation that respects the UnitOfWork pattern — flushing only when an external session is provided, and flushing + committing + closing when operating standalone. This eliminates premature transaction commits, loss of rollback capability, and a docstring/implementation mismatch.
* HAL 9000 has contributed the ACMS Index Data Model and File Traversal Engine (PR #9664 / issue #9579): foundational data structures for indexed context entries with hot/warm/cold/archive storage tier classification, tag system, and a timeout-safe chunked file traversal engine for large projects with 10,000+ files.
+1 -1
View File
@@ -277,7 +277,7 @@ The following standards are integrated into the architecture:
[(<span style="color: magenta;"><span style="color: cyan;">--temperature</span>|<span style="color: yellow;">-t</span></span>) <span style="color: #66cc66;">&lt;TEMP&gt;</span>] [<span style="color: cyan;">--allow-rxpy-in-run-mode</span>]
[<span style="color: cyan;">--skill</span> <span style="color: #66cc66;">&lt;SKILL&gt;</span>]... <span style="color: #66cc66;">&lt;NAME&gt;</span> <span style="color: #66cc66;">&lt;PROMPT&gt;</span>
<span style="color: cyan; font-weight: 600;">agents</span> actor add <span style="color: cyan;">--config</span>|<span style="color: yellow;">-c</span> <span style="color: #66cc66;">&lt;FILE&gt;</span> [<span style="color: cyan;">--update</span>]
<span style="color: cyan; font-weight: 600;">agents</span> actor remove [<span style="color: cyan;">--format</span> <span style="color: #66cc66;">&lt;FORMAT&gt;</span>] <span style="color: #66cc66;">&lt;NAME&gt;</span>
<span style="color: cyan; font-weight: 600;">agents</span> actor remove <span style="color: #66cc66;">&lt;NAME&gt;</span>
<span style="color: cyan; font-weight: 600;">agents</span> actor list
<span style="color: cyan; font-weight: 600;">agents</span> actor show <span style="color: #66cc66;">&lt;NAME&gt;</span>
<span style="color: cyan; font-weight: 600;">agents</span> actor context remove [<span style="color: cyan;">--yes</span>|<span style="color: yellow;">-y</span>] (<span style="color: magenta;"><span style="color: cyan;">--all</span>|<span style="color: yellow;">-a</span>|<span style="color: #66cc66;">&lt;NAME&gt;</span></span>)
@@ -1,138 +0,0 @@
Feature: ACMS Index Data Model and File Traversal Engine
As a developer
I want to index large projects with 10,000+ files
So that I can efficiently query and manage context entries at scale
Background:
Given I have an ACMS index
And I have a file traversal engine with chunk size 100
Scenario: Create an index entry with file metadata
When I create an index entry with:
| path | /project/src/main.py |
| file_type | python |
| size_bytes | 1024 |
Then the index entry should have path "/project/src/main.py"
And the index entry should have file type "python"
And the index entry should have size 1024 bytes
Scenario: Add tags to an index entry
Given I have an index entry with path "/project/src/main.py"
When I add tag "core" to the entry
And I add tag "important" to the entry
Then the entry should have tag "core"
And the entry should have tag "important"
And the entry should have 2 tags
Scenario: Set tier level for an index entry
Given I have an index entry with path "/project/src/main.py"
When I set the tier level to "hot"
Then the entry should have tier level "hot"
Scenario: Add entry to index
Given I have an index entry with path "/project/src/main.py"
When I add the entry to the index
Then the index should contain 1 entry
And I should be able to retrieve the entry by path "/project/src/main.py"
Scenario: Query index by path pattern
Given I have an index with entries:
| path | file_type |
| /project/src/main.py | python |
| /project/src/utils.py | python |
| /project/tests/test_main.py | python |
| /project/docs/readme.md | markdown |
When I query the index by path pattern "src"
Then I should get 2 results
And the results should include "/project/src/main.py"
And the results should include "/project/src/utils.py"
Scenario: Query index by file type
Given I have an index with entries:
| path | file_type |
| /project/src/main.py | python |
| /project/src/utils.py | python |
| /project/src/app.js | javascript |
| /project/docs/readme.md | markdown |
When I query the index by file type "python"
Then I should get 2 results
And all results should have file type "python"
Scenario: Query index by tag
Given I have an index with entries:
| path | tags |
| /project/src/main.py | core,important |
| /project/src/utils.py | supporting |
| /project/tests/test_main.py | test,important |
When I query the index by tag "important"
Then I should get 2 results
And the results should include "/project/src/main.py"
And the results should include "/project/tests/test_main.py"
Scenario: Query index by tier level
Given I have an index with entries:
| path | tier |
| /project/src/main.py | hot |
| /project/src/utils.py | warm |
| /project/tests/test_main.py | cold |
When I query the index by tier level "hot"
Then I should get 1 result
And the result should have path "/project/src/main.py"
Scenario: Query index by recency
Given I have an index with entries from different dates
When I query the index for entries modified after "2026-04-01"
Then I should get entries modified after that date
Scenario: Traverse and index a directory with multiple files
Given I have a test directory with 50 files
When I traverse and index the directory
Then the index should contain 50 entries
And all entries should have valid file paths
Scenario: Handle large project traversal with chunked processing
Given I have a test directory with 1000 files
When I traverse and index the directory with chunk size 100
Then the index should contain 1000 entries
And the traversal should complete without timeout
Scenario: Exclude patterns during traversal
Given I have a test directory with files including:
| path |
| /project/src/main.py |
| /project/.git/config |
| /project/__pycache__/main.cpython-39.pyc |
| /project/src/utils.py |
When I traverse and index the directory excluding ".git" and "__pycache__"
Then the index should contain 2 entries
And the index should not contain ".git" paths
And the index should not contain "__pycache__" paths
Scenario: Get all entries from index
Given I have an index with 5 entries
When I get all entries from the index
Then I should get 5 results
Scenario: Get entry count from index
Given I have an index with 10 entries
When I get the entry count
Then the count should be 10
Scenario: Remove entry from index
Given I have an index with 3 entries
When I remove an entry by path
Then the index should contain 2 entries
Scenario: Combined query with multiple filters
Given I have an index with entries:
| path | file_type | tags | tier |
| /project/src/main.py | python | core,important | hot |
| /project/src/utils.py | python | supporting | warm |
| /project/tests/test_main.py | python | test | cold |
| /project/docs/readme.md | markdown | docs | cold |
When I query the index with filters:
| path_pattern | src |
| file_type | python |
| tier | hot |
Then I should get 1 result
And the result should have path "/project/src/main.py"
-5
View File
@@ -59,11 +59,6 @@ Feature: Actor CLI YAML-first alignment
When I run actor remove with namespaced name
Then the actor remove should succeed for namespaced name
Scenario: Actor remove outputs JSON format
Given an actor CLI runner
When I run actor remove with format json
Then the actor remove output should be valid JSON envelope
Scenario: Actor update outputs JSON format
Given an actor CLI runner
When I run actor update with format json
-6
View File
@@ -686,12 +686,6 @@ def after_scenario(context, scenario):
pass # Ignore cleanup errors
context.test_dir = None
# Clean up TemporaryDirectory objects created by ACMS index traversal tests
if hasattr(context, "temp_dir") and context.temp_dir is not None:
with contextlib.suppress(Exception):
context.temp_dir.cleanup()
context.temp_dir = None
# Clean up environment variables set during tests
if hasattr(context, "env_vars_to_clean"):
for key in context.env_vars_to_clean:
-58
View File
@@ -1,58 +0,0 @@
@mock_only
Feature: PR Compliance Checklist in Implementation Supervisor
As an implementation supervisor
I want to pass a mandatory PR compliance checklist to every worker prompt
So that workers complete all required items before creating a PR and avoid systemic merge blockers
Background:
Given the implementation-supervisor.md agent definition exists
Scenario: Supervisor worker prompt includes the PR compliance checklist
When I read the implementation supervisor agent definition
Then the worker prompt body includes the PR compliance checklist section
And the checklist is marked as MANDATORY
Scenario: Checklist item 1 — CHANGELOG.md update required
When I read the implementation supervisor agent definition
Then the worker prompt body includes a CHANGELOG.md checklist item
And the item instructs workers to add an entry under the Unreleased section
Scenario: Checklist item 2 — CONTRIBUTORS.md update required
When I read the implementation supervisor agent definition
Then the worker prompt body includes a CONTRIBUTORS.md checklist item
And the item instructs workers to add or update their contribution entry
Scenario: Checklist item 3 — commit footer required
When I read the implementation supervisor agent definition
Then the worker prompt body includes a commit footer checklist item
And the item specifies the ISSUES CLOSED footer format
Scenario: Checklist item 4 — CI must pass before PR creation
When I read the implementation supervisor agent definition
Then the worker prompt body includes a CI passes checklist item
And the item instructs workers to verify all quality gates are green
Scenario: Checklist item 5 — BDD/Behave tests required
When I read the implementation supervisor agent definition
Then the worker prompt body includes a BDD tests checklist item
And the item instructs workers to add or update Behave feature files
Scenario: Checklist item 6 — Epic reference required in PR description
When I read the implementation supervisor agent definition
Then the worker prompt body includes an Epic reference checklist item
And the item instructs workers to reference the parent Epic issue number
Scenario: Checklist item 7 — Labels must be applied
When I read the implementation supervisor agent definition
Then the worker prompt body includes a labels checklist item
And the item instructs workers to apply labels via forgejo-label-manager
Scenario: Checklist item 8 — Milestone must be assigned
When I read the implementation supervisor agent definition
Then the worker prompt body includes a milestone checklist item
And the item instructs workers to assign the earliest open milestone
Scenario: All 8 checklist items are present in the worker prompt
When I read the implementation supervisor agent definition
Then the worker prompt body contains all 8 mandatory checklist items
@@ -1,395 +0,0 @@
"""Step definitions for ACMS Index Data Model and File Traversal Engine tests."""
from __future__ import annotations
import tempfile
from datetime import datetime, timedelta
from pathlib import Path
from behave import given, then, when
from cleveragents.acms.index import (
ACMSIndex,
FileTraversalEngine,
FileType,
IndexEntry,
TierLevel,
)
@given("I have an ACMS index")
def step_create_index(context):
"""Create a new ACMS index."""
context.index = ACMSIndex()
@given("I have a file traversal engine with chunk size {chunk_size:d}")
def step_create_traversal_engine(context, chunk_size):
"""Create a file traversal engine with specified chunk size."""
context.engine = FileTraversalEngine(chunk_size=chunk_size)
@when("I create an index entry with:")
def step_create_index_entry(context):
"""Create an index entry from table data."""
data = {row["key"]: row["value"] for row in context.table}
file_type = FileType(data.get("file_type", "other"))
size_bytes = int(data.get("size_bytes", "0"))
context.entry = IndexEntry(
path=data["path"],
file_type=file_type,
size_bytes=size_bytes,
created_at=datetime.now(),
modified_at=datetime.now(),
)
@then('the index entry should have path "{path}"')
def step_check_entry_path(context, path):
"""Verify the index entry has the expected path."""
assert context.entry.path == path
@then('the index entry should have file type "{file_type}"')
def step_check_entry_file_type(context, file_type):
"""Verify the index entry has the expected file type."""
assert context.entry.file_type == FileType(file_type)
@then("the index entry should have size {size:d} bytes")
def step_check_entry_size(context, size):
"""Verify the index entry has the expected size."""
assert context.entry.size_bytes == size
@given('I have an index entry with path "{path}"')
def step_create_entry_with_path(context, path):
"""Create an index entry with a specific path."""
context.entry = IndexEntry(
path=path,
file_type=FileType.PYTHON,
size_bytes=1024,
created_at=datetime.now(),
modified_at=datetime.now(),
)
@when('I add tag "{tag}" to the entry')
def step_add_tag_to_entry(context, tag):
"""Add a tag to the current entry."""
context.entry.add_tag(tag)
@then('the entry should have tag "{tag}"')
def step_check_entry_has_tag(context, tag):
"""Verify the entry has a specific tag."""
assert context.entry.has_tag(tag)
@then("the entry should have {count:d} tags")
def step_check_entry_tag_count(context, count):
"""Verify the entry has the expected number of tags."""
assert len(context.entry.tags) == count
@when('I set the tier level to "{tier}"')
def step_set_entry_tier(context, tier):
"""Set the tier level for the entry."""
context.entry.set_tier(TierLevel(tier))
@then('the entry should have tier level "{tier}"')
def step_check_entry_tier(context, tier):
"""Verify the entry has the expected tier level."""
assert context.entry.tier == TierLevel(tier)
@when("I add the entry to the index")
def step_add_entry_to_index(context):
"""Add the current entry to the index."""
context.index.add_entry(context.entry)
@then("the index should contain {count:d} entry")
def step_check_index_entry_count_singular(context, count):
"""Verify the index has the expected number of entries."""
assert context.index.get_entry_count() == count
@then("the index should contain {count:d} entries")
def step_check_index_entry_count(context, count):
"""Verify the index has the expected number of entries."""
assert context.index.get_entry_count() == count
@then('I should be able to retrieve the entry by path "{path}"')
def step_retrieve_entry_by_path(context, path):
"""Verify we can retrieve an entry by path."""
entry = context.index.get_entry(path)
assert entry is not None
assert entry.path == path
@given("I have an index with entries:")
def step_create_index_with_entries(context):
"""Create an index with multiple entries from table data."""
context.index = ACMSIndex()
for row in context.table:
path = row["path"]
file_type = FileType(row.get("file_type", "other"))
entry = IndexEntry(
path=path,
file_type=file_type,
size_bytes=1024,
created_at=datetime.now(),
modified_at=datetime.now(),
)
# Add tags if present
if "tags" in row:
for tag in row["tags"].split(","):
entry.add_tag(tag.strip())
# Set tier if present
if "tier" in row:
entry.set_tier(TierLevel(row["tier"]))
context.index.add_entry(entry)
@given("I have an index with {count:d} entries")
def step_create_index_with_n_entries(context, count):
"""Create an index with a specified number of entries."""
context.index = ACMSIndex()
for i in range(count):
entry = IndexEntry(
path=f"/project/file{i}.py",
file_type=FileType.PYTHON,
size_bytes=1024,
created_at=datetime.now(),
modified_at=datetime.now(),
)
context.index.add_entry(entry)
@when('I query the index by path pattern "{pattern}"')
def step_query_by_path_pattern(context, pattern):
"""Query the index by path pattern."""
context.query_results = context.index.query_by_path(pattern)
@then("I should get {count:d} results")
def step_check_query_result_count(context, count):
"""Verify the query returned the expected number of results."""
assert len(context.query_results) == count, (
f"Expected {count} results, got {len(context.query_results)}"
)
@then("I should get {count:d} result")
def step_check_query_result_count_singular(context, count):
"""Verify the query returned the expected number of results (singular)."""
assert len(context.query_results) == count, (
f"Expected {count} result, got {len(context.query_results)}"
)
@then('the results should include "{path}"')
def step_check_result_includes_path(context, path):
"""Verify the query results include a specific path."""
paths = [entry.path for entry in context.query_results]
assert path in paths
@when('I query the index by file type "{file_type}"')
def step_query_by_file_type(context, file_type):
"""Query the index by file type."""
context.query_results = context.index.query_by_type(FileType(file_type))
@then('all results should have file type "{file_type}"')
def step_check_all_results_file_type(context, file_type):
"""Verify all results have the expected file type."""
expected_type = FileType(file_type)
for entry in context.query_results:
assert entry.file_type == expected_type
@when('I query the index by tag "{tag}"')
def step_query_by_tag(context, tag):
"""Query the index by tag."""
context.query_results = context.index.query_by_tag(tag)
@when('I query the index by tier level "{tier}"')
def step_query_by_tier(context, tier):
"""Query the index by tier level."""
context.query_results = context.index.query_by_tier(TierLevel(tier))
@then('the result should have path "{path}"')
def step_check_single_result_path(context, path):
"""Verify the single result has the expected path."""
assert len(context.query_results) == 1
assert context.query_results[0].path == path
@given("I have an index with entries from different dates")
def step_create_index_with_dated_entries(context):
"""Create an index with entries from different dates."""
context.index = ACMSIndex()
now = datetime.now()
dates = [
now - timedelta(days=10),
now - timedelta(days=5),
now - timedelta(days=1),
now,
]
for i, date in enumerate(dates):
entry = IndexEntry(
path=f"/project/file{i}.py",
file_type=FileType.PYTHON,
size_bytes=1024,
created_at=date,
modified_at=date,
)
context.index.add_entry(entry)
@when('I query the index for entries modified after "{date_str}"')
def step_query_by_recency(context, date_str):
"""Query the index for entries modified after a specific date."""
# Parse date string (format: YYYY-MM-DD)
date = datetime.strptime(date_str, "%Y-%m-%d")
context.query_results = context.index.query_by_recency(date)
@then("I should get entries modified after that date")
def step_check_recency_results(context):
"""Verify the recency query returned valid results."""
assert len(context.query_results) > 0
@given("I have a test directory with {count:d} files")
def step_create_test_directory(context, count):
"""Create a temporary directory with test files."""
context.temp_dir = tempfile.TemporaryDirectory()
temp_path = Path(context.temp_dir.name)
# Create subdirectories and files
for i in range(count):
subdir = temp_path / f"subdir{i % 10}"
subdir.mkdir(exist_ok=True)
file_path = subdir / f"file{i}.py"
file_path.write_text(f"# Test file {i}\nprint('Hello {i}')\n")
@when("I traverse and index the directory")
def step_traverse_and_index(context):
"""Traverse and index the test directory."""
context.engine.reset_index()
context.index = context.engine.traverse_and_index(context.temp_dir.name)
@when("I traverse and index the directory with chunk size {chunk_size:d}")
def step_traverse_and_index_with_chunk_size(context, chunk_size):
"""Traverse and index the test directory with a specific chunk size."""
engine = FileTraversalEngine(chunk_size=chunk_size)
context.index = engine.traverse_and_index(context.temp_dir.name)
@then("all entries should have valid file paths")
def step_check_valid_file_paths(context):
"""Verify all entries have valid file paths."""
for entry in context.index.get_all_entries():
assert entry.path
assert len(entry.path) > 0
@then("the traversal should complete without timeout")
def step_check_no_timeout(context):
"""Verify the traversal completed without timeout."""
# This step passes if we got here without timing out
assert True
@given("I have a test directory with files including:")
def step_create_test_directory_with_specific_files(context):
"""Create a test directory with specific files."""
context.temp_dir = tempfile.TemporaryDirectory()
temp_path = Path(context.temp_dir.name)
for row in context.table:
file_path = temp_path / row["path"].lstrip("/")
file_path.parent.mkdir(parents=True, exist_ok=True)
file_path.write_text("test content")
@when('I traverse and index the directory excluding "{exclude1}" and "{exclude2}"')
def step_traverse_with_exclusions(context, exclude1, exclude2):
"""Traverse and index with exclusion patterns."""
context.engine.reset_index()
context.index = context.engine.traverse_and_index(
context.temp_dir.name,
exclude_patterns=[exclude1, exclude2],
)
@then('the index should not contain "{pattern}" paths')
def step_check_no_excluded_paths(context, pattern):
"""Verify the index doesn't contain paths matching the pattern."""
for entry in context.index.get_all_entries():
assert pattern not in entry.path
@when("I get all entries from the index")
def step_get_all_entries(context):
"""Get all entries from the index."""
context.query_results = context.index.get_all_entries()
@when("I get the entry count")
def step_get_entry_count(context):
"""Get the entry count from the index."""
context.entry_count = context.index.get_entry_count()
@then("the count should be {count:d}")
def step_check_entry_count_value(context, count):
"""Verify the entry count matches the expected value."""
assert context.entry_count == count
@when("I remove an entry by path")
def step_remove_entry(context):
"""Remove an entry from the index."""
# Get the first entry's path
entries = context.index.get_all_entries()
if entries:
context.index.remove_entry(entries[0].path)
@when("I query the index with filters:")
def step_query_with_combined_filters(context):
"""Query the index with multiple filters."""
filters = {row["filter"]: row["value"] for row in context.table}
path_pattern = filters.get("path_pattern")
file_type_str = filters.get("file_type")
tier_str = filters.get("tier")
file_type = FileType(file_type_str) if file_type_str else None
tier = TierLevel(tier_str) if tier_str else None
context.query_results = context.index.query_combined(
path_pattern=path_pattern,
file_type=file_type,
tier=tier,
)
-57
View File
@@ -337,63 +337,6 @@ def step_remove_namespaced_ok(context: Any) -> None:
)
@when("I run actor remove with format json")
def step_remove_format_json(context: Any) -> None:
with (
patch("cleveragents.cli.commands.actor._get_services") as mock_svc,
patch("cleveragents.cli.commands.actor._compute_actor_impact") as mock_impact,
):
mock_registry = MagicMock()
mock_service = MagicMock()
actor = _make_actor(
name="local/remove-json",
provider="json-provider",
model="gpt-json",
)
mock_registry.get_actor.return_value = actor
mock_impact.return_value = (2, 1, 3)
mock_svc.return_value = (mock_service, mock_registry)
context.result = context.runner.invoke(
actor_app,
["remove", actor.name, "--format", "json"],
)
context.mock_actor_registry = mock_registry
context.actor = actor
context.impact_counts = (2, 1, 3)
@then("the actor remove output should be valid JSON envelope")
def step_remove_json_valid(context: Any) -> None:
assert context.result.exit_code == 0
parsed = json.loads(context.result.output.strip())
assert _ENVELOPE_KEYS.issubset(parsed.keys())
assert parsed["command"] == f"agents actor remove {context.actor.name}"
assert parsed["status"] == "ok"
assert parsed["exit_code"] == 0
data = _unwrap_envelope(parsed)
assert isinstance(data, dict)
actor_data = data.get("actor_removed", {})
assert actor_data.get("name") == context.actor.name
assert actor_data.get("provider") == context.actor.provider
assert actor_data.get("model") == context.actor.model
impact = data.get("impact", {})
expected_sessions, expected_plans, expected_actions = context.impact_counts
assert impact.get("sessions") == expected_sessions
assert impact.get("active_plans") == expected_plans
assert impact.get("actions_referencing") == expected_actions
cleanup = data.get("cleanup", {})
assert cleanup.get("config") == "kept on disk"
assert cleanup.get("contexts") == "0 orphaned"
messages = parsed.get("messages", [])
assert messages, "expected messages in envelope"
first_message = messages[0]
assert first_message.get("level") == "ok"
assert "Actor removed" in first_message.get("text", "")
context.mock_actor_registry.remove_actor.assert_called_once_with(context.actor.name)
# ------------------------------------------------------------------
# Update with --format
# ------------------------------------------------------------------
@@ -1,185 +0,0 @@
"""Step definitions for PR compliance checklist in implementation supervisor."""
from pathlib import Path
from typing import Any
from behave import given, then, when
PROJECT_ROOT = Path(__file__).resolve().parents[3]
AGENT_DEF_PATH = PROJECT_ROOT / ".opencode" / "agents" / "implementation-supervisor.md"
@given("the implementation-supervisor.md agent definition exists")
def step_agent_def_exists(context: Any) -> None:
"""Verify the implementation supervisor agent definition file exists."""
assert AGENT_DEF_PATH.exists(), f"Agent definition not found at {AGENT_DEF_PATH}"
context.agent_def_path = AGENT_DEF_PATH
@when("I read the implementation supervisor agent definition")
def step_read_agent_def(context: Any) -> None:
"""Read the implementation supervisor agent definition."""
context.agent_def_content = AGENT_DEF_PATH.read_text(encoding="utf-8")
@then("the worker prompt body includes the PR compliance checklist section")
def step_prompt_includes_checklist(context: Any) -> None:
"""Verify the worker prompt body includes the PR compliance checklist."""
assert "PR Compliance Checklist" in context.agent_def_content, (
"Worker prompt body does not include 'PR Compliance Checklist'"
)
@then("the checklist is marked as MANDATORY")
def step_checklist_is_mandatory(context: Any) -> None:
"""Verify the checklist is marked as MANDATORY."""
assert "MANDATORY" in context.agent_def_content, (
"PR Compliance Checklist is not marked as MANDATORY"
)
@then("the worker prompt body includes a CHANGELOG.md checklist item")
def step_prompt_includes_changelog_item(context: Any) -> None:
"""Verify the worker prompt body includes a CHANGELOG.md checklist item."""
assert "CHANGELOG.md" in context.agent_def_content, (
"Worker prompt body does not include a CHANGELOG.md checklist item"
)
@then("the item instructs workers to add an entry under the Unreleased section")
def step_changelog_item_unreleased(context: Any) -> None:
"""Verify the CHANGELOG.md item mentions the Unreleased section."""
assert "[Unreleased]" in context.agent_def_content, (
"CHANGELOG.md checklist item does not mention the [Unreleased] section"
)
@then("the worker prompt body includes a CONTRIBUTORS.md checklist item")
def step_prompt_includes_contributors_item(context: Any) -> None:
"""Verify the worker prompt body includes a CONTRIBUTORS.md checklist item."""
assert "CONTRIBUTORS.md" in context.agent_def_content, (
"Worker prompt body does not include a CONTRIBUTORS.md checklist item"
)
@then("the item instructs workers to add or update their contribution entry")
def step_contributors_item_add_update(context: Any) -> None:
"""Verify the CONTRIBUTORS.md item instructs workers to add or update."""
assert "add or update" in context.agent_def_content, (
"CONTRIBUTORS.md checklist item does not instruct workers to add or update"
)
@then("the worker prompt body includes a commit footer checklist item")
def step_prompt_includes_commit_footer_item(context: Any) -> None:
"""Verify the worker prompt body includes a commit footer checklist item."""
assert "Commit footer" in context.agent_def_content, (
"Worker prompt body does not include a commit footer checklist item"
)
@then("the item specifies the ISSUES CLOSED footer format")
def step_commit_footer_issues_closed(context: Any) -> None:
"""Verify the commit footer item specifies the ISSUES CLOSED format."""
assert "ISSUES CLOSED" in context.agent_def_content, (
"Commit footer checklist item does not specify the ISSUES CLOSED format"
)
@then("the worker prompt body includes a CI passes checklist item")
def step_prompt_includes_ci_item(context: Any) -> None:
"""Verify the worker prompt body includes a CI passes checklist item."""
assert "CI passes" in context.agent_def_content, (
"Worker prompt body does not include a CI passes checklist item"
)
@then("the item instructs workers to verify all quality gates are green")
def step_ci_item_quality_gates(context: Any) -> None:
"""Verify the CI item instructs workers to verify quality gates."""
assert "quality gates" in context.agent_def_content, (
"CI checklist item does not mention quality gates"
)
@then("the worker prompt body includes a BDD tests checklist item")
def step_prompt_includes_bdd_item(context: Any) -> None:
"""Verify the worker prompt body includes a BDD/Behave tests checklist item."""
assert "BDD/Behave tests" in context.agent_def_content, (
"Worker prompt body does not include a BDD/Behave tests checklist item"
)
@then("the item instructs workers to add or update Behave feature files")
def step_bdd_item_feature_files(context: Any) -> None:
"""Verify the BDD item instructs workers to add or update feature files."""
assert "added or updated" in context.agent_def_content, (
"BDD checklist item does not instruct workers to add or update feature files"
)
@then("the worker prompt body includes an Epic reference checklist item")
def step_prompt_includes_epic_item(context: Any) -> None:
"""Verify the worker prompt body includes an Epic reference checklist item."""
assert "Epic reference" in context.agent_def_content, (
"Worker prompt body does not include an Epic reference checklist item"
)
@then("the item instructs workers to reference the parent Epic issue number")
def step_epic_item_parent_reference(context: Any) -> None:
"""Verify the Epic item instructs workers to reference the parent Epic."""
assert "parent Epic" in context.agent_def_content, (
"Epic checklist item does not instruct workers to reference the parent Epic"
)
@then("the worker prompt body includes a labels checklist item")
def step_prompt_includes_labels_item(context: Any) -> None:
"""Verify the worker prompt body includes a labels checklist item."""
assert "Labels" in context.agent_def_content, (
"Worker prompt body does not include a labels checklist item"
)
@then("the item instructs workers to apply labels via forgejo-label-manager")
def step_labels_item_forgejo_label_manager(context: Any) -> None:
"""Verify the labels item instructs workers to use forgejo-label-manager."""
assert "forgejo-label-manager" in context.agent_def_content, (
"Labels checklist item does not mention forgejo-label-manager"
)
@then("the worker prompt body includes a milestone checklist item")
def step_prompt_includes_milestone_item(context: Any) -> None:
"""Verify the worker prompt body includes a milestone checklist item."""
assert "Milestone" in context.agent_def_content, (
"Worker prompt body does not include a milestone checklist item"
)
@then("the item instructs workers to assign the earliest open milestone")
def step_milestone_item_earliest(context: Any) -> None:
"""Verify the milestone item instructs workers to assign the earliest open milestone."""
assert "earliest open milestone" in context.agent_def_content, (
"Milestone checklist item does not mention the earliest open milestone"
)
@then("the worker prompt body contains all 8 mandatory checklist items")
def step_prompt_contains_all_8_items(context: Any) -> None:
"""Verify the worker prompt body contains all 8 mandatory checklist items."""
required_items = [
"CHANGELOG.md",
"CONTRIBUTORS.md",
"ISSUES CLOSED",
"CI passes",
"BDD/Behave tests",
"Epic reference",
"forgejo-label-manager",
"earliest open milestone",
]
missing = [item for item in required_items if item not in context.agent_def_content]
assert not missing, (
f"Worker prompt body is missing the following checklist items: {missing}"
)
-18
View File
@@ -1,18 +0,0 @@
*** Settings ***
Documentation Integration tests for actor remove CLI format output
Resource ${CURDIR}/common.resource
Suite Setup Setup Test Environment
Suite Teardown Cleanup Test Environment
*** Variables ***
${HELPER} ${CURDIR}/helper_actor_remove_cli.py
*** Test Cases ***
Actor Remove Format JSON Emits Spec Envelope
[Documentation] Verify that ``actor remove --format json`` emits a spec-compliant envelope
[Tags] actor_remove_cli_format
${result}= Run Process ${PYTHON} ${HELPER} remove-json cwd=${WORKSPACE}
Log ${result.stdout}
Log ${result.stderr}
Should Be Equal As Integers ${result.rc} 0
Should Contain ${result.stdout} actor-remove-json-format-ok
-164
View File
@@ -1,164 +0,0 @@
"""Helper script for Robot integration tests covering ``actor remove --format`` output.
Exercises the real ``agents`` CLI via subprocess no mocking of any kind.
A test actor is seeded via ``agents actor add``, then removed via
``agents actor remove --format json``, and the resulting JSON envelope is
validated against the spec.
Usage::
python helper_actor_remove_cli.py <command>
Where *command* is one of:
- ``remove-json`` seed an actor, remove it with ``--format json``, validate envelope
"""
from __future__ import annotations
import json
import os
import sys
from pathlib import Path
# Ensure src is importable when run from workspace root
_SRC = str(Path(__file__).resolve().parents[1] / "src")
if _SRC not in sys.path:
sys.path.insert(0, _SRC)
# Ensure robot/ is on the import path for helper_e2e_common.
_ROBOT = str(Path(__file__).resolve().parent)
if _ROBOT not in sys.path:
sys.path.insert(0, _ROBOT)
from helper_e2e_common import cleanup_workspace, run_cli, setup_workspace # noqa: E402
_ACTOR_NAME = "local/robot-remove-actor"
_ACTOR_CONFIG: dict[str, object] = {
"name": _ACTOR_NAME,
"provider": "openai",
"model": "gpt-4",
}
def _write_actor_config(workspace: str) -> str:
"""Write actor config JSON to a temp file and return its path."""
config_path = os.path.join(workspace, "robot_remove_actor.json")
with open(config_path, "w", encoding="utf-8") as fh:
json.dump(_ACTOR_CONFIG, fh)
return config_path
def test_remove_format_json() -> None:
"""Seed an actor via the real CLI, remove it with ``--format json``.
Validates the resulting JSON envelope against the spec.
"""
workspace = setup_workspace(prefix="robot_actor_remove_")
try:
config_path = _write_actor_config(workspace)
# Step 1: Add the actor using the real CLI.
add_result = run_cli(
"actor",
"add",
_ACTOR_NAME,
"--config",
config_path,
workspace=workspace,
)
assert add_result.returncode == 0, (
f"actor add failed (rc={add_result.returncode}):\n"
f"stdout: {add_result.stdout}\nstderr: {add_result.stderr}"
)
# Step 2: Remove the actor with --format json using the real CLI.
remove_result = run_cli(
"actor",
"remove",
_ACTOR_NAME,
"--format",
"json",
workspace=workspace,
)
assert remove_result.returncode == 0, (
f"actor remove --format json failed (rc={remove_result.returncode}):\n"
f"stdout: {remove_result.stdout}\nstderr: {remove_result.stderr}"
)
# Step 3: Parse and validate the JSON envelope.
output = remove_result.stdout.strip()
assert output, (
f"actor remove --format json produced no output.\n"
f"stderr: {remove_result.stderr}"
)
payload = json.loads(output)
assert payload["command"] == f"agents actor remove {_ACTOR_NAME}", (
f"Unexpected command field: {payload.get('command')!r}"
)
assert payload["status"] == "ok", (
f"Unexpected status: {payload.get('status')!r}"
)
assert payload["exit_code"] == 0, (
f"Unexpected exit_code: {payload.get('exit_code')!r}"
)
data = payload["data"]
removed = data.get("actor_removed", {})
assert removed.get("name") == _ACTOR_NAME, (
f"actor_removed.name mismatch: {removed.get('name')!r}"
)
assert removed.get("provider") == _ACTOR_CONFIG["provider"], (
f"actor_removed.provider mismatch: {removed.get('provider')!r}"
)
assert removed.get("model") == _ACTOR_CONFIG["model"], (
f"actor_removed.model mismatch: {removed.get('model')!r}"
)
impact = data.get("impact", {})
assert "sessions" in impact, f"Missing 'sessions' in impact: {impact}"
assert "active_plans" in impact, f"Missing 'active_plans' in impact: {impact}"
assert "actions_referencing" in impact, (
f"Missing 'actions_referencing' in impact: {impact}"
)
cleanup = data.get("cleanup", {})
assert cleanup.get("config") == "kept on disk", (
f"cleanup.config mismatch: {cleanup.get('config')!r}"
)
assert "contexts" in cleanup, f"Missing 'contexts' in cleanup: {cleanup}"
messages = payload.get("messages", [])
assert messages, f"Expected non-empty messages list, got: {messages}"
assert messages[0].get("level") == "ok", (
f"Unexpected message level: {messages[0].get('level')!r}"
)
assert "Actor removed" in messages[0].get("text", ""), (
f"Expected 'Actor removed' in message text: {messages[0].get('text')!r}"
)
print("actor-remove-json-format-ok")
finally:
cleanup_workspace(workspace)
def main() -> None:
command = sys.argv[1] if len(sys.argv) > 1 else "remove-json"
dispatch: dict[str, object] = {
"remove-json": test_remove_format_json,
}
handler = dispatch.get(command)
if handler is None:
print(f"Unknown command: {command}", file=sys.stderr)
sys.exit(1)
assert callable(handler)
handler()
if __name__ == "__main__":
main()
+3 -21
View File
@@ -5,22 +5,12 @@ technology-specific vocabulary extensions, and the DetailLevelMap
inheritance mechanism for resolving named detail levels across the
ontology hierarchy (Layer 3 -> Layer 2 -> Layer 1 -> Layer 0).
Also provides the ACMS index data model and file traversal engine for
indexing large projects.
Based on ``docs/specification.md`` ~lines 42333-42422, 44405-44420.
"""
from __future__ import annotations
from cleveragents.acms import uko as _uko
from cleveragents.acms.index import (
ACMSIndex,
FileTraversalEngine,
FileType,
IndexEntry,
TierLevel,
)
from cleveragents.acms.uko import (
CODE_DETAIL_LEVEL_MAP,
FUNC_DETAIL_LEVEL_MAP,
@@ -72,14 +62,6 @@ from cleveragents.acms.uko import (
resolve_detail_level,
)
# Combine exports from both uko and index modules
_uko_exports = list(_uko.__all__)
_index_exports = [
"ACMSIndex",
"FileTraversalEngine",
"FileType",
"IndexEntry",
"TierLevel",
]
__all__: list[str] = _uko_exports + _index_exports
# Re-export everything published by the ``uko`` sub-package so the two
# ``__all__`` lists stay in sync automatically.
__all__: list[str] = list(_uko.__all__)
-412
View File
@@ -1,412 +0,0 @@
"""ACMS Index Data Model and File Traversal Engine.
Provides the foundational data model for indexed context entries and a
file traversal engine that can handle 10,000+ files without timeout using
chunked processing.
Based on issue #9579 and ``docs/specification.md`` ~lines 44405-44420.
"""
from __future__ import annotations
from collections.abc import Iterator
from dataclasses import dataclass, field
from datetime import datetime
from enum import StrEnum
from pathlib import Path
class FileType(StrEnum):
"""File type enumeration for index entries."""
PYTHON = "python"
JAVASCRIPT = "javascript"
TYPESCRIPT = "typescript"
JAVA = "java"
RUST = "rust"
MARKDOWN = "markdown"
TEXT = "text"
JSON = "json"
YAML = "yaml"
XML = "xml"
OTHER = "other"
class TierLevel(StrEnum):
"""Tier assignment levels for context entries.
Aligns with the hot/warm/cold storage tier vocabulary used throughout
the ACMS specification and milestone v3.4.0 acceptance criteria.
"""
HOT = "hot" # Core/essential
WARM = "warm" # Important
COLD = "cold" # Supporting
ARCHIVE = "archive" # Reference
@dataclass
class IndexEntry:
"""Represents a single indexed context entry.
Attributes:
path: Absolute or relative file path
file_type: Type of file (from FileType enum)
size_bytes: File size in bytes
created_at: File creation timestamp
modified_at: File modification timestamp
tags: Set of tags for categorization
tier: Tier assignment level
metadata: Additional metadata dictionary
"""
path: str
file_type: FileType
size_bytes: int
created_at: datetime
modified_at: datetime
tags: set[str] = field(default_factory=set)
tier: TierLevel = TierLevel.COLD
metadata: dict[str, str] = field(default_factory=dict)
def add_tag(self, tag: str) -> None:
"""Add a tag to this entry.
Args:
tag: Tag string to add. Must be non-empty.
Raises:
ValueError: If tag is empty or whitespace-only.
"""
if not tag or not tag.strip():
raise ValueError("tag must be a non-empty string")
self.tags.add(tag)
def remove_tag(self, tag: str) -> None:
"""Remove a tag from this entry.
Args:
tag: Tag string to remove. Must be non-empty.
Raises:
ValueError: If tag is empty or whitespace-only.
"""
if not tag or not tag.strip():
raise ValueError("tag must be a non-empty string")
self.tags.discard(tag)
def has_tag(self, tag: str) -> bool:
"""Check if entry has a specific tag."""
return tag in self.tags
def set_tier(self, tier: TierLevel) -> None:
"""Set the tier level for this entry."""
self.tier = tier
@dataclass
class ACMSIndex:
"""ACMS Index for storing and querying indexed context entries.
Provides storage and query capabilities for indexed files with support
for filtering by path, tags, type, and recency.
Attributes:
entries: Dictionary mapping file paths to IndexEntry objects
"""
entries: dict[str, IndexEntry] = field(default_factory=dict)
def add_entry(self, entry: IndexEntry) -> None:
"""Add an index entry to the index.
Args:
entry: The IndexEntry to add. Must not be None.
Raises:
TypeError: If entry is not an IndexEntry instance.
"""
if not isinstance(entry, IndexEntry):
raise TypeError(
f"entry must be an IndexEntry instance, got {type(entry).__name__}"
)
self.entries[entry.path] = entry
def remove_entry(self, path: str) -> bool:
"""Remove an entry by path. Returns True if removed, False if not found."""
if path in self.entries:
del self.entries[path]
return True
return False
def get_entry(self, path: str) -> IndexEntry | None:
"""Get an entry by path."""
return self.entries.get(path)
def query_by_path(self, path_pattern: str) -> list[IndexEntry]:
"""Query entries by path pattern (substring match).
Args:
path_pattern: Substring pattern to match against entry paths.
Must be non-empty.
Raises:
ValueError: If path_pattern is empty or whitespace-only.
"""
if not path_pattern or not path_pattern.strip():
raise ValueError("path_pattern must be a non-empty string")
return [entry for entry in self.entries.values() if path_pattern in entry.path]
def query_by_tag(self, tag: str) -> list[IndexEntry]:
"""Query entries by tag."""
return [entry for entry in self.entries.values() if entry.has_tag(tag)]
def query_by_type(self, file_type: FileType) -> list[IndexEntry]:
"""Query entries by file type."""
return [
entry for entry in self.entries.values() if entry.file_type == file_type
]
def query_by_tier(self, tier: TierLevel) -> list[IndexEntry]:
"""Query entries by tier level."""
return [entry for entry in self.entries.values() if entry.tier == tier]
def query_by_recency(
self, after: datetime, before: datetime | None = None
) -> list[IndexEntry]:
"""Query entries by modification recency.
Args:
after: Return entries modified after this datetime.
before: Return entries modified before this datetime (optional).
When both are provided, before must be >= after.
Returns:
List of entries matching the recency criteria.
Raises:
ValueError: If before is earlier than after when both are provided.
"""
if before is not None and before < after:
raise ValueError(
f"before ({before}) must be >= after ({after}) when both are provided"
)
results = [
entry for entry in self.entries.values() if entry.modified_at >= after
]
if before:
results = [entry for entry in results if entry.modified_at <= before]
return results
def query_combined(
self,
path_pattern: str | None = None,
tags: set[str] | None = None,
file_type: FileType | None = None,
tier: TierLevel | None = None,
after: datetime | None = None,
) -> list[IndexEntry]:
"""Query entries with multiple filters (AND logic).
Args:
path_pattern: Filter by path pattern
tags: Filter by any of these tags
file_type: Filter by file type
tier: Filter by tier level
after: Filter by modification date
Returns:
List of entries matching all specified criteria
"""
results = list(self.entries.values())
if path_pattern:
results = [entry for entry in results if path_pattern in entry.path]
if tags:
results = [
entry for entry in results if any(tag in entry.tags for tag in tags)
]
if file_type:
results = [entry for entry in results if entry.file_type == file_type]
if tier:
results = [entry for entry in results if entry.tier == tier]
if after:
results = [entry for entry in results if entry.modified_at >= after]
return results
def get_all_entries(self) -> list[IndexEntry]:
"""Get all entries in the index."""
return list(self.entries.values())
def get_entry_count(self) -> int:
"""Get the total number of entries in the index."""
return len(self.entries)
class FileTraversalEngine:
"""Engine for traversing and indexing files in large projects.
Handles 10,000+ files without timeout using chunked processing to
prevent memory exhaustion.
Attributes:
chunk_size: Number of files to process in each chunk
index: The ACMS index to populate
"""
def __init__(self, chunk_size: int = 100, index: ACMSIndex | None = None) -> None:
"""Initialize the traversal engine.
Args:
chunk_size: Number of files to process per chunk (default: 100).
Must be a positive integer.
index: Optional pre-populated ACMSIndex to use. If None, a new
empty index is created (supports Dependency Inversion).
Raises:
ValueError: If chunk_size is not a positive integer.
"""
if chunk_size <= 0:
raise ValueError(f"chunk_size must be a positive integer, got {chunk_size}")
self.chunk_size = chunk_size
self.index = index if index is not None else ACMSIndex()
def _get_file_type(self, file_path: Path) -> FileType:
"""Determine file type from extension."""
suffix = file_path.suffix.lower()
type_map = {
".py": FileType.PYTHON,
".js": FileType.JAVASCRIPT,
".ts": FileType.TYPESCRIPT,
".tsx": FileType.TYPESCRIPT,
".java": FileType.JAVA,
".rs": FileType.RUST,
".md": FileType.MARKDOWN,
".txt": FileType.TEXT,
".json": FileType.JSON,
".yaml": FileType.YAML,
".yml": FileType.YAML,
".xml": FileType.XML,
}
return type_map.get(suffix, FileType.OTHER)
def _create_index_entry(self, file_path: Path) -> IndexEntry | None:
"""Create an index entry from a file path.
Args:
file_path: Path to the file
Returns:
IndexEntry if successful, None if file cannot be read
"""
try:
stat = file_path.stat()
return IndexEntry(
path=str(file_path),
file_type=self._get_file_type(file_path),
size_bytes=stat.st_size,
created_at=datetime.fromtimestamp(stat.st_ctime),
modified_at=datetime.fromtimestamp(stat.st_mtime),
)
except (OSError, ValueError):
return None
def _traverse_directory(
self,
root_path: Path,
) -> Iterator[Path]:
"""Recursively traverse directory and yield file paths.
Args:
root_path: Root directory to traverse
Yields:
Path objects for each file found
"""
try:
for item in root_path.iterdir():
if item.is_file():
yield item
elif item.is_dir():
# Recursively traverse subdirectories
yield from self._traverse_directory(item)
except (OSError, PermissionError):
# Skip directories we can't read
pass
def traverse_and_index(
self,
root_path: str | Path,
exclude_patterns: list[str] | None = None,
) -> ACMSIndex:
"""Traverse a directory and index all files.
Uses chunked processing to handle large projects without timeout
or memory exhaustion.
Args:
root_path: Root directory to traverse
exclude_patterns: List of patterns to exclude
(e.g., ['.git', '__pycache__'])
Returns:
Populated ACMSIndex
"""
root = Path(root_path)
if not root.exists():
raise ValueError(f"Path does not exist: {root_path}")
exclude_patterns = exclude_patterns or []
chunk: list[IndexEntry] = []
for file_path in self._traverse_directory(root):
# Check if file matches any exclude pattern
if any(pattern in str(file_path) for pattern in exclude_patterns):
continue
# Create index entry
entry = self._create_index_entry(file_path)
if entry:
chunk.append(entry)
# Process chunk when it reaches the size limit
if len(chunk) >= self.chunk_size:
self._process_chunk(chunk)
chunk = []
# Process remaining entries
if chunk:
self._process_chunk(chunk)
return self.index
def _process_chunk(self, chunk: list[IndexEntry]) -> None:
"""Process a chunk of index entries.
Args:
chunk: List of IndexEntry objects to add to the index
"""
for entry in chunk:
self.index.add_entry(entry)
def get_index(self) -> ACMSIndex:
"""Get the current index."""
return self.index
def reset_index(self) -> None:
"""Reset the index to empty state."""
self.index = ACMSIndex()
__all__ = [
"ACMSIndex",
"FileTraversalEngine",
"FileType",
"IndexEntry",
"TierLevel",
]
+1 -51
View File
@@ -816,28 +816,12 @@ def update(
@app.command()
def remove(
name: Annotated[str, typer.Argument(help="Actor name to remove")],
fmt: Annotated[
str,
typer.Option("--format", "-f", help=_FORMAT_HELP),
] = OutputFormat.RICH.value,
) -> None:
def remove(name: Annotated[str, typer.Argument(help="Actor name to remove")]) -> None:
"""Remove a custom actor.
Specify the namespaced name (e.g. ``local/my-actor``).
"""
# Validate --format argument first; fail fast before any side effects.
fmt_value = fmt.lower()
_valid_formats = {f.value for f in OutputFormat}
if fmt_value not in _valid_formats:
raise typer.BadParameter(
f"Invalid format {fmt!r}. "
f"Supported values: {', '.join(sorted(_valid_formats))}",
param_hint="'--format'",
)
service, registry = _get_services()
try:
# Get actor details before removal for display
@@ -860,40 +844,6 @@ def remove(
else:
service.remove_actor(name)
command_name = f"agents actor remove {name}"
payload = {
"actor_removed": {
"name": name,
"provider": actor_provider,
"model": actor_model,
},
"impact": {
"sessions": session_count,
"active_plans": active_plan_count,
"actions_referencing": action_count,
},
"cleanup": {
"config": "kept on disk",
# NOTE: context-cleanup count is deferred; always 0 for now.
# Follow-up: implement dynamic orphaned-context detection.
"contexts": "0 orphaned",
},
}
messages = [{"level": "ok", "text": "Actor removed"}]
if fmt_value != OutputFormat.RICH.value:
rendered = format_output(
payload,
fmt_value,
command=command_name,
status="ok",
exit_code=0,
messages=messages,
)
if rendered:
console.print(rendered)
return
# Display Actor Removed panel
actor_info = (
f"[cyan bold]Name:[/cyan bold] {name}\n"