Feature: Large-project hierarchical decomposition As a plan orchestrator I want to decompose large file sets into bounded subplans So that each subplan can be executed independently within token budgets Background: Given a decomposition service # --- decomposition depth ----------------------------------------------- Scenario: Decomposition produces multiple levels for a deep project Given a project with 2000 files across 5 directory levels When I decompose with max_depth 4 Then the decomposition result should have max_depth_reached >= 1 And every leaf node should have at most 50 files Scenario: Decomposition respects max_depth limit Given a project with 2000 files across 5 directory levels When I decompose with max_depth 2 Then the decomposition result should have max_depth_reached <= 2 # NOTE: The test fixture ``make_files`` places all files under a single # tmpdir root, so ``_directory_key(depth=2)`` yields one bucket and # meaningful multi-level splits depend on bucket-size chunking. # In production, real directory trees with distinct top-level # subtrees will decompose to 4+ levels. A separate Robot E2E suite # (m6_e2e_verification.robot) validates the full 4+ level requirement # against a realistic project layout. Scenario: Decomposition produces meaningful hierarchy for large file sets Given a project with 5000 files across 8 directory levels When I decompose with max_depth 4 Then the decomposition result should have max_depth_reached >= 1 Scenario: Small project decomposition depth does not increase unexpectedly Given a project with 40 files across 3 directory levels When I decompose with max_depth 4 Then the decomposition result should have max_depth_reached <= 2 # --- bounded context --------------------------------------------------- Scenario: Leaf nodes respect max_files_per_subplan Given a project with 1500 files in a single directory When I decompose with max_files_per_subplan 500 Then every leaf node should have at most 500 files Scenario: Leaf nodes respect max_tokens_per_subplan Given a project with files totalling 500000 estimated tokens When I decompose with max_tokens_per_subplan 100000 Then every leaf node should have estimated tokens at most 100000 # --- language clustering ----------------------------------------------- Scenario: Files are clustered by language Given a project with 100 Python files and 100 TypeScript files When I decompose with language clustering Then at least two clusters should have different dominant languages # --- directory clustering ---------------------------------------------- Scenario: Files are clustered by directory Given a project with files in src/api and src/web directories When I decompose with directory clustering Then at least two clusters should have different directory prefixes # --- dependency closure ------------------------------------------------ Scenario: Compute transitive closure of root files Given a dependency graph with chain A -> B -> C -> D When I compute closure from A with no cutoff Then the closure should contain A, B, C, D Scenario: Bounded closure respects cutoff Given a dependency graph with chain A -> B -> C -> D When I compute closure from A with cutoff 2 Then the closure should contain at most 2 files # --- DAG ordering ------------------------------------------------------ Scenario: Topological sort produces valid execution order Given a dependency graph with edges A->B, A->C, B->D, C->D When I compute the topological sort Then A should appear before B and C And B and C should appear before D # --- cycle detection --------------------------------------------------- Scenario: Detect cycles in dependency graph Given a dependency graph with cycle A -> B -> C -> A When I check for cycles Then at least one cycle should be detected Scenario: No cycles in acyclic graph Given a dependency graph with chain A -> B -> C -> D When I check for cycles Then no cycles should be detected # --- fallback strategy ------------------------------------------------- Scenario: Fallback when heuristics produce a single cluster Given a project with 100 files all in the same directory with same extension When I decompose with max_files_per_subplan 200 Then the decomposition should produce a single root node # --- threshold guard --------------------------------------------------- Scenario: Skip decomposition when below threshold Given a project with 5 files When I decompose with min_files_per_subplan 10 Then the decomposition should produce a single root node And the metrics should contain skipped equal to 1 # --- config keys ------------------------------------------------------- Scenario: Default config has expected values Given default decomposition config Then max_depth should be 4 And max_child_depth should be 5 And max_files_per_subplan should be 500 And max_tokens_per_subplan should be 100000 And min_files_per_subplan should be 10 Scenario: Invalid config values are rejected When I create a config with max_depth 0 Then a decomposition ValueError should be raised # --- metrics ----------------------------------------------------------- Scenario: Decomposition result contains metrics Given a project with 2000 files across 5 directory levels When I decompose with default config Then metrics should contain total_nodes And metrics should contain leaf_nodes And metrics should contain max_depth # --- decision recording ------------------------------------------------ Scenario: Decisions are recorded in DecisionService Given a decomposition service with a mock decision service And a project with 2000 files across 5 directory levels When I decompose and record decisions for plan "01HXBBBBBBBBBBBBBBBBBBBBBB" Then at least one strategy_choice decision should be recorded And at least one subplan_spawn decision should be recorded # --- deterministic ordering -------------------------------------------- Scenario: Deterministic sort produces stable ordering Given two identical file sets presented in different order When I decompose both Then both decomposition results should have identical node ordering # --- memoization ------------------------------------------------------- Scenario: Closure computation is memoised for repeated queries Given a dependency graph with chain A -> B -> C -> D When I compute closure from A twice Then the second computation should return the same result # --- additional config validation ---------------------------------------- Scenario: Config rejects max_files_per_subplan of zero When I create a config with max_files_per_subplan 0 Then a decomposition ValueError should be raised Scenario: Config rejects max_tokens_per_subplan of zero When I create a config with max_tokens_per_subplan 0 Then a decomposition ValueError should be raised Scenario: Config rejects min_files_per_subplan of zero When I create a config with min_files_per_subplan 0 Then a decomposition ValueError should be raised # --- estimate_tokens edge cases ------------------------------------------ Scenario: Estimate tokens for nonexistent file returns zero When I estimate tokens for a nonexistent file Then the token estimate should be 0 # --- cluster_by_size with token_map -------------------------------------- Scenario: Cluster by size uses provided token_map Given a set of files with known token counts When I cluster by size with a token_map and max_tokens 100 Then the clusters should respect the token_map values # --- closure trimming with multiple roots -------------------------------- Scenario: Closure trims deterministically when result exceeds cutoff Given a dependency graph with two roots each reaching 3 nodes When I compute closure from both roots with cutoff 3 Then the closure should contain at most 3 files # --- clear cache --------------------------------------------------------- Scenario: Clear cache drops memoised results Given a dependency graph with chain A -> B -> C -> D When I compute closure from A and then clear the cache Then computing closure from A again should still work # --- service wrapper methods --------------------------------------------- Scenario: Service compute_dependency_order delegates to closure computer Given a dependency graph with edges A->B, A->C, B->D, C->D When I compute dependency order through the service Then A should appear before B and C in the service order And B and C should appear before D in the service order Scenario: Service compute_closure delegates to closure computer Given a dependency graph with chain A -> B -> C -> D When I compute closure through the service from A Then the service closure should contain A, B, C, D # --- decision_service property ------------------------------------------- Scenario: Decision service property returns injected service Given a decomposition service with a mock decision service Then the decision service property should return the mock Scenario: Decision service property returns None when not injected Then the decision service property should return None # --- empty file edge cases ----------------------------------------------- Scenario: Dominant extension of empty list returns empty string When I compute dominant extension of an empty list Then the dominant extension should be empty Scenario: Common prefix of empty list returns empty string When I compute common prefix of an empty list Then the common prefix should be empty # --- directory key short path -------------------------------------------- Scenario: Directory key handles short paths When I compute directory key for a single-component path Then the directory key should be empty string # --- absolute path clustering ------------------------------------------ Scenario: Directory clustering works correctly with absolute file paths Given a project with absolute paths in distinct subdirectories When I decompose with directory clustering Then at least two clusters should have different directory prefixes Scenario: _directory_key returns correct key for absolute path with root When I compute directory key for an absolute path with a root Then the directory key should reflect the relative path structure Scenario: Directory clustering does not collapse all absolute paths into one bucket Given a project with absolute paths spanning multiple top-level directories When I decompose with max_files_per_subplan 50 Then the decomposition result should have max_depth_reached >= 1