Commit Graph

74 Commits

Author SHA1 Message Date
brent.edwards aa5d5eeaf5 test(session): add TDD failing tests for session create DI error
Implement TDD bug-capture tests for bug #570 where `agents session create`
fails because `_get_session_service()` calls `container.db()` which does
not exist on the DI Container class (AttributeError). Same root cause as
bug #554.

Behave BDD scenarios tagged @tdd_bug @tdd_bug_570 @tdd_expected_fail
exercise the real DI path (no mocks). Includes Robot Framework integration
smoke tests with self-inverting helper and ASV benchmark baseline.

ISSUES CLOSED: #631
2026-03-11 02:42:13 +00:00
Brent E. Edwards 6bce5479f3 Merge branch 'master' into tdd/session-list-di-error
# Conflicts:
#	features/environment.py
#	noxfile.py
2026-03-10 23:34:36 +00:00
Brent E. Edwards d0689573e0 test(cli): add failing tests for session create DI container error
Add TDD regression tests for bug #570 where `_get_session_service()`
calls `container.db()` but the DI `Container` class has no `db`
provider, raising `AttributeError`.  Same root cause as bug #554.

Includes 4 Behave BDD scenarios tagged `@tdd_bug @tdd_bug_570
@tdd_expected_fail`, Robot Framework integration smoke tests with
`--format plain`, and ASV service-layer benchmarks.  Tests exercise the
real DI path by resetting `_service = None` and using a file-based
SQLite database.

Implements the `@tdd_expected_fail` inversion infrastructure:
- Behave: `after_scenario` hook in `features/environment.py` inverts
  pass/fail for scenarios tagged `@tdd_expected_fail`
- Robot: `robot/tdd_expected_fail_listener.py` listener (API v3)
  performs the same inversion for Robot test cases
- `noxfile.py`: registers the listener via `--listener` in both the
  `integration_tests` and `slow_integration_tests` sessions

Migrates 18 existing TDD scenarios across 5 feature files from the old
`@tdd @bugNNN` convention to the standardised `@tdd_bug @tdd_bug_NNN`
tags per CONTRIBUTING.md § TDD Bug Test Tags.

Refs: #570
2026-03-10 21:39:31 +00:00
brent.edwards 06bbe48a9c test(session): add TDD failing tests for session list DI error
Implement TDD bug-capture tests for bug #554 where `agents session list`
fails because `_get_session_service()` calls `container.db()` which does
not exist on the DI Container class (AttributeError).

Behave BDD scenarios tagged @tdd_bug @tdd_bug_554 @tdd_expected_fail
exercise the real DI path (no mocks) and assert correct behavior. The
@tdd_expected_fail handler in environment.py inverts failed→passed while
the bug is present, keeping CI green.

Also adds:
- @tdd_expected_fail infrastructure in features/environment.py
  (tag validation + status inversion in after_scenario hook)
- behave-parallel exit logic fix to use summary-based failure
  detection (compatible with TDD status inversion)
- Robot Framework integration smoke tests with self-inverting helper
- ASV benchmark baseline for session list command throughput

ISSUES CLOSED: #630
2026-03-10 19:05:12 +00:00
brent.edwards 4e3bf7d3ad test(cli): add failing tests for agents init --yes missing option
Add TDD-style Behave BDD tests for the missing agents init --yes flag
(bug #522). Five Gherkin scenarios cover: exit code validation, prompt
suppression, -y alias, output summary fields, and interactive-mode
regression guard. Includes Robot Framework smoke tests (tagged @wip)
and ASV benchmarks.

Configure behave.ini to exclude @wip scenarios globally and noxfile.py
to exclude wip-tagged Robot suites, so TDD-failing tests do not break CI.

Review feedback addressed:
- Remove unnecessary # type: ignore from benchmark (outside Pyright scope)
- Fix Then...Then to Then...And in Gherkin (L1)
- Fix CHANGELOG 'three scenarios' to 'five scenarios' (L2)
- Add behave.ini documentation for @wip workaround (Aditya F1)
- Rename Scenario 2 title to 'suppresses interactive prompts' (Aditya F2)

Closes #536
2026-03-06 20:28:37 +00:00
freemo f26fcfc44e perf(tests): replace behave-parallel subprocess model with in-process parallelism
Replace the subprocess-per-feature execution model (342 Python interpreter
startups) with direct use of behave's Runner API for in-process
execution.

Sequential mode (--processes 1 or BEHAVE_PARALLEL_COVERAGE=1): All
features run in a single Runner.run() call. Steps and hooks load once.

Parallel mode (--processes N, N>1): Features split into N chunks,
dispatched via multiprocessing.Pool with fork. Heavy modules shared
copy-on-write.

Proper format defaulting (mirrors behave.__main__.run_behave() logic
for -q flag). Summary extracted from runner.features status attributes
instead of regex-parsing stdout.

Simplified coverage pipeline: single slipcover invocation wraps the
entire behave-parallel process. No per-worker UUID files, no --merge
step needed. Coverage data produced in one build/coverage.json file.

Removed: behave-parallel tarball download from PyPI, tarfile and
urllib.request imports, per-worker subprocess.run() calls,
__SLIPCOVER_OUT__ placeholder mechanism, _build_base_args(),
_parse_summary(), regex-based summary parsing.

Results: nox -s unit_tests 24m21s -> 2m05s (91%); nox -s coverage_report
75m20s -> 3m00s (96%). Coverage: 98% (above 97% threshold).

ISSUES CLOSED: #481
2026-03-02 02:01:27 +00:00
freemo a8f7ed57cb perf(tests): reduce per-feature startup cost with shared fixtures and lazy imports
Created scripts/create_template_db.py that builds a pre-migrated SQLite
template database using Base.metadata.create_all() + alembic stamp
(~5ms for 34 tables, vs ~0.5-3s x 25 Alembic migrations per scenario).

Nox unit_tests and coverage_report sessions generate the template before
test execution and propagate CLEVERAGENTS_TEMPLATE_DB env var to all
workers.

features/environment.py before_all() installs a monkey-patch on
MigrationRunner.init_or_upgrade that copies the template for fresh
scenario temp DBs, falling through to real migrations for :memory:,
existing files, and migration-runner unit tests.

Quick wins: sleep(0.5) -> sleep(0.05) in cli_streaming wait step;
removed redundant Background re-declaration in cli_streaming.feature
scenario 7.

ISSUES CLOSED: #483
2026-03-02 02:01:27 +00:00
freemo 74772280e6 perf(tests): optimize coverage instrumentation and reporting pipeline
Replace coverage.py (sys.settrace-based) with slipcover (bytecode-based
instrumentation) for significantly faster coverage collection:

- Each behave-parallel worker runs under slipcover, producing per-feature
  JSON coverage files with unique UUIDs to avoid write contention.
- After all workers finish, slipcover --merge combines per-worker data
  into a single build/coverage.json report.
- XML report generated via slipcover --merge --xml for CI tooling.
- Terminal report with --fail-under=97 threshold enforcement.
- Robust JSON key-fallback logic handles both slipcover and coverage.py
  output formats.
- CI workflow (ci.yml, nightly-quality.yml) updated with defensive key
  lookup instead of hardcoded coverage.py format.
- Documentation updated to reflect slipcover as the coverage tool.
- CHANGELOG.md updated.

ISSUES CLOSED: #482
2026-03-02 02:01:27 +00:00
mngrif 4e750b9b87 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
mngrif 800835f6a0 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
mngrif 57ff467321 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
mngrif 5f76637b21 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
mngrif ce4bbd4303 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
mngrif 8105117c10 asv runners have the same machine name now 2026-02-24 20:48:41 +00:00
CoreRasurae c5990904fb test(robot): Disable tests that depend on code blocks
Code blocks exec()/eval()/compile() where removed as part of m4-security-eval
2026-02-21 16:32:11 +00:00
brent.edwards 56c38a04ce fix(ci): remove stale AutomationLevel refs from benchmarks and prevent ANSI in JSON output
- Remove AutomationLevel imports from cli_robot_flow_bench.py and
  persistence_robot_bench.py (enum was removed by master's automation
  refactor); replace with AutomationProfileRef where needed.
- Use typer.echo() instead of console.print() for machine-readable
  output (JSON/YAML/plain) in config.py and session.py to prevent
  Rich from injecting ANSI escape codes that corrupt json.loads().
- Set NO_COLOR=1 in noxfile unit_tests, integration_tests, and
  coverage_report sessions as a belt-and-suspenders safeguard for
  all CLI commands that route format_output through Rich.
2026-02-20 21:46:46 +00:00
brent.edwards cb96af7286 fix(ci): pre-compile bytecode before parallel integration tests
Add python -m compileall -q src/ step to the integration_tests nox
session before launching pabot.  On CI runners with high core counts
(e.g. 32 processes) the first Robot test to spawn a Python subprocess
could fail because 30+ processes simultaneously cold-compile the
entire source tree from scratch.  Pre-compiling eliminates the
thundering-herd race and lets all workers read cached .pyc files.
2026-02-20 21:09:45 +00:00
mngrif 9a4e55cab0 testing airspeedvelocity 2026-02-19 12:48:31 -05:00
mngrif b19ac44573 testing airspeedvelocity 2026-02-19 12:28:49 -05:00
mngrif d5af40eefd Merge branch 'master' into mngrif-asv-test 2026-02-19 00:03:54 +00:00
mngrif c2d6e71563 testing airspeedvelocity 2026-02-18 18:25:28 -05:00
mngrif f9f4503c6c testing airspeedvelocity 2026-02-18 17:45:07 -05:00
mngrif 745228e7c6 testing airspeedvelocity 2026-02-18 17:30:18 -05:00
mngrif 341495769d testing airspeedvelocity 2026-02-18 17:24:04 -05:00
mngrif 8ea70953e9 testing airspeedvelocity 2026-02-18 15:02:15 -05:00
mngrif 8f6b2a29b8 testing airspeedvelocity 2026-02-18 13:00:15 -05:00
freemo 4cda00abf6 Build: configured coverage report to run in parallal, removed cap on number of parallel processes, limited by number of cores of env variable 2026-02-17 21:00:02 -05:00
mngrif fb055193c5 testing airspeedvelocity 2026-02-17 15:43:46 -05:00
mngrif 033bfd04a7 testing airspeedvelocity 2026-02-17 15:30:07 -05:00
Jeffrey Phillips Freeman 4dc05051dd feat(cli): add project commands (core) 2026-02-16 23:56:55 -05:00
Jeffrey Phillips Freeman 0c6ed1c709 feat(repo): add project repositories 2026-02-16 23:50:59 -05:00
Jeffrey Phillips Freeman ee5f0376d7 Replaced hand written Reference section with one generated from docstrings 2026-02-16 23:40:33 -05:00
brent.edwards 7ddd99b07d chore(merge): integrate master into feature/q0-min-coverage
Resolve three merge conflicts from master integration:

- benchmarks/coverage_report_bench.py: remove duplicate PYPROJECT_PATH
  constant (already defined below the conflict region)
- noxfile.py: combine HEAD's os.makedirs('build') guard with master's
  resolved PYTHONPATH (Path('src').resolve())
- implementation_plan.md: take master for completed Q0-min-ci block,
  preserve HEAD's in-progress Q0-min-coverage state, adopt master's
  Q0-Advanced consolidation (no standalone commits)
2026-02-17 00:56:43 +00:00
Jeffrey Phillips Freeman d0f265ef62 feat(service): persist plan lifecycle via repositories 2026-02-17 00:37:53 +00:00
Jeffrey Phillips Freeman e273b52769 feat(db): add projects and project links tables 2026-02-16 12:18:51 -05:00
Jeffrey Phillips Freeman dbca60c98e feat(db): add projects and project links tables 2026-02-15 11:04:18 -05:00
Jeffrey Phillips Freeman 1f06bb2f76 feat(resource): add resource type model + schema loader 2026-02-15 15:19:09 +00:00
brent.edwards 75dc57845a feat(qa): align security scans with nox 2026-02-13 19:24:24 +00:00
brent.edwards 35c74d1a52 feat(qa): enforce coverage >=97% in CI 2026-02-13 18:54:15 +00:00
brent.edwards bb09434d5e chore(ci): pin ruff version in nox and nightly workflow
Ensures linting uses the same ruff version range in CI and nightly
jobs to avoid rule drift from latest releases.
2026-02-13 04:57:01 +00:00
brent.edwards da3d0dfd06 chore(ci): re-enable pabot, add discovery resource stub
- Restore pabot-based integration_tests with conservative parallelism
  (<=2 processes by default) and support PABOT_PROCESSES/--processes
- Remove the temporary CI debug dump from noxfile.py
- Add robot/discovery_common.resource to silence non-fatal warning
- Document the change in implementation_plan.md
2026-02-13 04:32:45 +00:00
brent.edwards 51a6e76e9e fix(ci): merge duplicate Settings blocks, inject venv Python, add CI debug
Three fixes targeting CI integration_tests failures (18 failures on e8aa5ac):

1. Merge duplicate *** Settings *** blocks in 14 robot files into single
   blocks. Multiple Settings sections are non-standard RF practice and
   may cause resource import failures in certain Robot Framework versions
   or CI environments.

2. Replace bare 'python' with ${PYTHON} variable in all Run Process
   calls (14 files). Noxfile now passes --variable PYTHON:<venv-path>
   to robot so tests use the venv interpreter regardless of PATH. This
   fixes '/usr/local/bin/python: No module named cleveragents' on CI.

3. Add comprehensive CI debug output in noxfile.py: file existence
   checks for .resource files, PATH/Python resolution, fixture dir
   checks, and RF version. This will diagnose any remaining resource
   import issues.

Also: remove hardcoded '/app/src' sys.path.insert in
system_prompt_template_rendering.robot (not portable to CI), and
add trailing newline to common.resource.

All 204 tests pass locally (4 excluded: 2 slow, 2 discovery).
2026-02-13 03:43:44 +00:00
brent.edwards e8aa5ac268 fix(ci): restore venv PATH, use absolute resource paths, and add timeouts in robot tests
- Restore session.env["PATH"] in integration_tests nox session to ensure
  Run Process uses venv Python instead of system Python
- Convert bare Resource references to ${CURDIR}/ absolute paths across
  30 robot files to fix CI resolution failures
- Add timeout=30s to all Run Process calls in rxpy_route_validation.robot
  to prevent hanging tests
- Tag 2 rxpy tests as slow (require running actors unavailable on CI)
- Fix LangGraph test to use correct config file (LANGGRAPH_CONFIG)
- All 204 tests pass (4 excluded: 2 slow + 2 discovery)
2026-02-13 02:42:57 +00:00
brent.edwards 0059050836 fix(ci): replace pabot with robot for reliable CI execution
- integration_tests: replace pabot with sequential robot execution to
  eliminate FileNotFoundError caused by subprocess/FD exhaustion in
  constrained CI containers (pabot spawns a robot subprocess per suite;
  after ~24 suites the container cannot execve new processes)

- integration_tests: add resource debug output (open FD count, ulimit
  values) to help diagnose future CI container issues

- integration_tests: remove _pabot_parallel_args (no longer needed);
  slow_integration_tests session still available for parallel runs

- load_context_test: add env:TERM=dumb alongside NO_COLOR=1 to disable
  all Rich terminal styling (NO_COLOR only disables color, not bold/
  reset ANSI codes that may split substrings)

- load_context_test: add repr() debug logging around the --load-context
  match to reveal any invisible characters on CI
2026-02-13 01:37:32 +00:00
brent.edwards 25e6f95eb9 fix(ci): resolve security_scan and integration_tests CI failures
- security_scan: create build/ directory before bandit writes its JSON
  report (fails on fresh CI checkout where directory does not exist)

- integration_tests: cap pabot parallelism to 2 processes and explicitly
  propagate venv bin/ to PATH, preventing FileNotFoundError for the
  robot binary under CI resource constraints

- integration_tests: add env:NO_COLOR=1 to load_context_test.robot help
  text assertions so Rich ANSI escape codes do not break substring
  matching on CI

- integration_tests: tag initial_next_command_test as slow (requires
  OPENAI_API_KEY for LLM agent invocation, unavailable on CI)
2026-02-13 01:04:22 +00:00
brent.edwards 321ab18a37 fix(ci): resolve unit_tests and integration_tests CI failures
- Fix Rich Console line-wrapping breaking assertions in
  context_unit_tests_steps.py: collapse newlines before checking for
  filenames and overflow summaries (CI temp paths exceed 80 columns)
- Fix features.mocks import failure in database_integration.robot:
  replace hardcoded sys.path '/app' with portable ${CURDIR}/..
- Fix --load-context help text assertions in load_context_test.robot:
  merge stderr into stdout via stderr=STDOUT and remove duplicate test
- Add standalone dead_code nox session running vulture directly
- Rename security nox session to security_scan to match CI references
- Restore --exclude discovery to integration_tests nox session (lost
  during merge conflict resolution)
2026-02-13 00:15:58 +00:00
brent.edwards 3e3530de48 ci(git-merge): Merging from master 2026-02-12 23:01:35 +00:00
freemo 89184689dc build: Coverage was broken, now fixed. 2026-02-12 17:02:21 -05:00
brent.edwards e801eb1ee8 feat(ci): add nox-based PR validation workflow
- Rewrite .forgejo/workflows/ci.yml to route all jobs through nox sessions
- Fix coverage_report nox session: serial behave mode replaces broken parallel
  mode (22% -> 97% accuracy), raise fail-under from 85% to 97%
- Pass posargs through format nox session for CI --check support
- Add 11 CI workflow validation scenarios (Behave) + Robot smoke test + ASV bench
- Add 108 new Behave scenarios covering 6 largest coverage gaps to reach 97%:
  yaml_template_engine, actor/config, actor/registry, message_router,
  context_analysis, context_service
- Update docs/development/ci-cd.md with nox-based CI docs and 97% threshold
- Restore implementation_plan.md verbose style, check off completed CI tasks

Verified: 1673 scenarios pass, 97% coverage, lint clean, typecheck clean
2026-02-12 22:01:51 +00:00
CoreRasurae e7ad541b71 build(env): Update security Vulture check exceptions 2026-02-12 20:19:57 +00:00