5.4 KiB
Scale Testing Runbook
Overview
Scale testing validates that CleverAgents performs within acceptable bounds when indexing and decomposing repositories of varying sizes (1K, 5K, and 10K files). The test suite uses pre-defined fixtures rather than live repositories to ensure repeatable, deterministic results.
Approach
- Fixture-based: Scale profiles are defined in JSON metadata files that describe repository characteristics (file count, size, language mix) without requiring actual large repositories.
- Threshold matrices: Baseline performance thresholds are maintained in a separate JSON file with p50/p95/p99 percentile targets for indexing and decomposition operations.
- Simulation: File distribution generation simulates how files would be allocated across languages for each profile, validating the distribution logic without creating actual files.
Environment Requirements
- Python 3.13+
- At least 8 GB RAM (for large profile validation)
- Local SSD storage recommended
- All dependencies installed via
uv sync
Fixture Files
All scale fixtures live in features/fixtures/scale/:
| File | Purpose |
|---|---|
scale_metadata.json |
Scale profile definitions (file count, size, language mix) |
baseline_thresholds.json |
Performance threshold matrices (p50/p95/p99 per operation) |
generator_instructions.md |
Manual instructions for generating synthetic test repos |
Running Scale Tests
Behave (Unit Tests)
# Run scale test scenarios only
nox -s unit_tests -- features/scale_test.feature
# Run all unit tests including scale tests
nox -s unit_tests
Robot Framework (Integration Tests)
# Run scale test integration suite only
nox -s integration_tests -- --suite robot/scale_test.robot
# Run all integration tests including scale tests
nox -s integration_tests
ASV Benchmarks
# Run all benchmarks including scale fixture benchmarks
nox -s benchmark
Interpreting Results
Behave Results
All 20 scenarios should pass. Failures indicate:
- Fixture loading failures: Missing or malformed JSON files
- Threshold validation failures: Percentile ordering violations (p50 >= p95)
- Language mix failures: Distribution ratios don't sum to 1.0
- Distribution generation failures: File count allocation doesn't match profile
Robot Results
All 6 test cases should pass. Each test prints a sentinel string on success:
| Sentinel | Meaning |
|---|---|
scale-metadata-ok |
Metadata fixture loaded and has valid structure |
scale-profiles-ok |
All profiles have required fields and valid language mixes |
scale-thresholds-ok |
Threshold matrix has all required sections and percentiles |
scale-monotonicity-ok |
All thresholds satisfy p50 < p95 < p99 |
scale-distribution-ok |
File distributions match language mix ratios |
scale-generator-docs-ok |
Generator instructions document is present |
ASV Benchmarks
The ScaleFixtureSuite measures:
| Benchmark | Description | Expected Range |
|---|---|---|
time_metadata_loading |
JSON parse time for metadata | < 1 ms |
time_threshold_loading |
JSON parse time for thresholds | < 1 ms |
time_threshold_validation |
Threshold matrix validation | < 0.1 ms |
time_language_mix_validation |
Language mix sum checks | < 0.1 ms |
time_file_distribution_generation |
Distribution calculation | < 0.5 ms |
time_full_fixture_roundtrip |
Complete load + validate + distribute | < 2 ms |
Threshold Adjustment Guidelines
When to Adjust Thresholds
- After significant indexing algorithm changes
- When targeting new hardware profiles (e.g., CI runners with different specs)
- When adding support for new file types that affect parsing performance
- After profiling reveals consistent threshold violations
How to Adjust
-
Run benchmarks on the target environment:
nox -s benchmark -
Collect actual p50/p95/p99 values from multiple runs (minimum 10 runs recommended).
-
Update
features/fixtures/scale/baseline_thresholds.jsonwith new values. -
Ensure the monotonicity invariant holds:
p50 < p95 < p99for all entries. -
Update the
scaling_expectationssection if the algorithmic complexity has changed. -
Run the full test suite to verify:
nox -s unit_tests -- features/scale_test.feature nox -s integration_tests -- --suite robot/scale_test.robot
Threshold Safety Margins
Current thresholds include a 2x safety margin over measured values to account for:
- CI runner variability
- Concurrent workload interference
- OS-level caching effects
- GC pauses in the Python runtime
When adjusting, maintain at least a 1.5x margin over observed p99 values.
Adding New Scale Profiles
To add a new scale profile (e.g., "extra-large" for 50K files):
-
Add the profile entry to
scale_metadata.json:{ "name": "extra-large", "file_count": 50000, "total_size_mb": 2500, "language_mix": { ... }, "expected_index_time_s": 250.0, "expected_decomposition_time_s": 500.0, "description": "Extra-large enterprise monorepo" } -
Add threshold entries to
baseline_thresholds.jsonfor"50000_files". -
Update Behave scenarios in
features/scale_test.featureif exact profile count assertions need updating. -
Run tests to verify the new profile integrates correctly.