Files
cleveragents-core/docs/reference/scaling_baselines.md
T
brent.edwards e77cf8bb1c
CI / lint (pull_request) Successful in 21s
CI / quality (pull_request) Successful in 47s
CI / typecheck (pull_request) Successful in 50s
CI / security (pull_request) Successful in 1m6s
CI / benchmark-publish (pull_request) Has been skipped
CI / build (pull_request) Successful in 22s
CI / unit_tests (pull_request) Successful in 3m44s
CI / integration_tests (pull_request) Successful in 3m50s
CI / docker (pull_request) Successful in 1m0s
CI / e2e_tests (pull_request) Successful in 6m26s
CI / coverage (pull_request) Successful in 8m0s
CI / benchmark-regression (pull_request) Failing after 59m59s
feat(perf): large project scaling tests
Add ASV benchmark suites for large project scaling: indexing at
10K/50K/100K files with walk_and_index throughput and incremental
refresh metrics, context assembly at 1K/5K/10K fragments, and
execution throughput for sequential/concurrent plans. Update scale
fixture thresholds for 50K/100K profiles. Add Behave and Robot
tests for baseline validation. Document scaling baselines.

Also fixes pre-existing test failures and integrates recent master
changes discovered during full nox suite verification.

ISSUES CLOSED: #859
2026-03-19 22:25:29 +00:00

4.0 KiB

Scaling Baselines

Performance baselines and expectations for CleverAgents at production scale.

Overview

CleverAgents is designed to handle repositories ranging from small microservices (1K files) to large enterprise monorepos (100K+ files). This document captures the baseline performance thresholds and scaling expectations used by the CI benchmark suite.

Scale Profiles

Profile Files Size (MB) Index Time (s) Decompose Time (s)
small 1,000 50 5 10
medium 5,000 250 25 50
large 10,000 500 50 100
xlarge 50,000 2,500 200 400
xxlarge 100,000 5,000 450 900

Indexing Thresholds

Latency percentiles for walk_and_index on local SSD with 8+ GB RAM:

File Count p50 (ms) p95 (ms) p99 (ms) Max (ms)
1,000 2,000 5,000 8,000 12,000
5,000 10,000 25,000 40,000 60,000
10,000 18,000 45,000 72,000 110,000
50,000 75,000 180,000 290,000 450,000
100,000 140,000 340,000 550,000 850,000

Decomposition Thresholds

Decomposition typically takes 2x indexing time:

File Count p50 (ms) p95 (ms) p99 (ms) Max (ms)
1,000 4,000 10,000 16,000 24,000
5,000 20,000 50,000 80,000 120,000
10,000 40,000 100,000 160,000 240,000
50,000 160,000 400,000 640,000 960,000
100,000 300,000 750,000 1,200,000 1,800,000

Context Assembly Thresholds

ACMS pipeline latency for varying fragment counts:

Fragments p50 (ms) p95 (ms) p99 (ms)
100 50 150 300
1,000 500 1,500 3,000
5,000 2,500 7,500 15,000
10,000 5,000 15,000 30,000

Execution Throughput Thresholds

Plan execution dispatch latency:

Plans p50 (ms) p95 (ms) p99 (ms)
10 200 500 1,000
50 1,000 2,500 5,000
100 2,000 5,000 10,000

Memory Usage

File Count Peak (MB) Steady State (MB)
1,000 256 128
5,000 768 384
10,000 1,536 768
50,000 6,144 3,072
100,000 12,288 6,144

Scaling Expectations

  • Indexing: O(n log n) — sub-linear scaling expected
  • Decomposition: O(n log n) — sub-linear scaling expected
  • Context Assembly: O(n log n) — sub-linear for large fragment sets
  • Execution: O(n) — linear scaling expected
  • Memory: O(n) — linear scaling expected

Sub-linear Scaling Invariants

The benchmark suite enforces these invariants:

  • 10K indexing p50 < 10x the 1K p50
  • 50K indexing p50 < 5x the 10K p50
  • 100K indexing p50 < 2.5x the 50K p50

ASV Benchmark Suites

Suite Params Timeout
IndexingScalingSuite 1K, 10K, 50K, 100K 600s
ContextAssemblyScalingSuite 100, 1K, 5K, 10K 300s
ExecutionThroughputSuite 10, 50, 100 300s

Fixture Files

  • features/fixtures/scale/scale_metadata.json — profile definitions
  • features/fixtures/scale/baseline_thresholds.json — p50/p95/p99 thresholds
  • features/fixtures/scale/generator_instructions.md — fixture generation guide