# Database Schema Reference ## Overview CleverAgents uses SQLAlchemy ORM models with SQLite as the default backend. All tables are created via `Base.metadata.create_all` through the `init_database()` helper in `src/cleveragents/infrastructure/database/models.py`. ## Resource Registry Tables The resource registry consists of three core tables introduced in Stage B1: | Table | Model | Description | |--------------------|----------------------|------------------------------------| | `resource_types` | `ResourceTypeModel` | Schema-level resource type defs | | `resources` | `ResourceModel` | Registered resource instances | | `resource_edges` | `ResourceEdgeModel` | Parent-child DAG edges | ### resource_types Primary key: `name` (namespaced, e.g. `builtin/git-checkout`). Stores kind (`physical`/`virtual`), handler references, JSON argument schemas, allowed parent/child types, auto-discovery config, and capabilities. ### resources Primary key: `resource_id` (26-char ULID). FK to `resource_types.name`. Stores optional namespaced name, location, JSON properties/metadata, and content hash for equivalence tracking. ### resource_edges Composite primary key: `(parent_id, child_id)`. Both columns FK to `resources.resource_id` with `CASCADE` delete. Stores `link_type` (`contains`, `references`, `derived_from`) and a self-loop check constraint. ## Robot Migration Smoke Suite A Robot Framework smoke suite validates that schema creation produces the expected resource registry tables with correct columns. ### Running the suite ```bash # Via nox (recommended — runs all Robot integration tests): nox -s integration_tests # Directly with robot: robot --outputdir build/reports/robot robot/resource_registry_migration.robot ``` ### Test cases | Test Case | What it checks | |--------------------------------------------------|--------------------------------------------------| | Schema Creation Produces Resource Registry Tables | `resource_types`, `resources`, `resource_edges` exist | | Resource Registry Tables Have Expected Columns | Core columns present on each table | | Migration Is Idempotent | Calling `init_database` twice is safe | The helper script (`robot/helper_resource_registry_migration.py`) uses `init_database()` with a temporary SQLite file and validates via `sqlalchemy.inspect`. ## Behave BDD Tests The Behave feature file `features/resource_registry_tables.feature` contains comprehensive scenarios covering: - Table existence after schema creation - CRUD operations for `ResourceTypeModel`, `ResourceModel`, `ResourceEdgeModel` - Constraint enforcement (uniqueness, FK, check constraints) - ORM relationship navigation - Migration smoke verification Run Behave tests via: ```bash nox -s unit_tests ``` ## ASV Benchmarks Performance benchmarks for resource registry operations live in `benchmarks/resource_registry_migration_bench.py` and measure: - Schema creation time - Insert throughput for types, resources, and edges - DAG query performance Run benchmarks via: ```bash nox -s benchmark ``` ## Tool and Validation Registry Tables ## Tables ### `tools` Stores registered tool and validation definitions. | Column | Type | Constraints | Description | |-----------------|---------|-------------------------------------|------------------------------------------| | `tool_id` | TEXT | PRIMARY KEY | ULID identifier | | `name` | TEXT | NOT NULL, UNIQUE | Namespaced name (`namespace/short_name`) | | `description` | TEXT | | Human-readable description | | `tool_type` | TEXT | NOT NULL, CHECK (`tool`/`validation`) | Discriminator for registry queries | | `source_type` | TEXT | NOT NULL, CHECK (see below) | Implementation source | | `input_schema` | TEXT | | JSON Schema for tool inputs | | `output_schema` | TEXT | | JSON Schema for tool outputs | | `read_only` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool only reads, never writes | | `writes` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool can write to resources | | `checkpointable`| BOOLEAN | NOT NULL, DEFAULT FALSE | Tool supports checkpoint/rollback | | `side_effects` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool has known side effects | | `config_yaml` | TEXT | | Raw YAML config for reconstruction | | `created_at` | TEXT | NOT NULL | ISO-8601 timestamp | | `updated_at` | TEXT | NOT NULL | ISO-8601 timestamp | **Check constraints:** - `ck_tools_tool_type`: `tool_type IN ('tool', 'validation')` - `ck_tools_source_type`: `source_type IN ('mcp', 'agent_skill', 'builtin', 'custom', 'wrapped')` **Indexes:** - `ix_tools_name` on `(name)` - `ix_tools_tool_type` on `(tool_type)` --- ### `tool_bindings` Stores resource slot bindings for tools. | Column | Type | Constraints | Description | |----------------|---------|--------------------------------------|--------------------------------------| | `binding_id` | TEXT | PRIMARY KEY | ULID identifier | | `tool_id` | TEXT | NOT NULL, FK → `tools.tool_id` CASCADE | Parent tool reference | | `slot_name` | TEXT | NOT NULL | Named resource slot | | `resource_type`| TEXT | NOT NULL | Expected resource type | | `binding_mode` | TEXT | NOT NULL, CHECK (see below) | How the slot is resolved | | `access_level` | TEXT | NOT NULL, DEFAULT `read_only` | Access level on bound resource | | `created_at` | TEXT | NOT NULL | ISO-8601 timestamp | **Check constraints:** - `ck_tool_bindings_mode`: `binding_mode IN ('context', 'static', 'parameter')` - `ck_tool_bindings_access`: `access_level IN ('read_only', 'read_write')` **Unique constraints:** - `uq_tool_bindings_tool_slot`: `UNIQUE(tool_id, slot_name)` --- ### `validation_attachments` Links validations to resources with mode semantics. | Column | Type | Constraints | Description | |-------------------|---------|-----------------------------------------------|------------------------------------| | `attachment_id` | TEXT | PRIMARY KEY | ULID identifier | | `resource_id` | TEXT | NOT NULL, FK → `resources.resource_id` CASCADE | Target resource | | `validation_name` | TEXT | NOT NULL | References `tools.name` (validation)| | `mode` | TEXT | NOT NULL, DEFAULT `required` | `required` or `informational` | | `created_at` | TEXT | NOT NULL | ISO-8601 timestamp | **Check constraints:** - `ck_validation_attachments_mode`: `mode IN ('required', 'informational')` **Unique constraints:** - `uq_validation_attachments_resource_validation`: `UNIQUE(resource_id, validation_name)` **Indexes:** - `ix_validation_attachments_resource_id` on `(resource_id)` --- ## Relationships ``` tools 1 ──< tool_bindings (tool_id FK, CASCADE delete) resources 1 ──< validation_attachments (resource_id FK, CASCADE delete) validation_attachments.validation_name → tools.name (logical, not enforced by FK) ``` --- ### `actors` Stores actor configurations with YAML text retention, schema version tracking, and optional compiled metadata. | Column | Type | Constraints | Description | |---------------------|----------|-----------------------------|------------------------------------------------| | `id` | INTEGER | PRIMARY KEY, AUTOINCREMENT | Auto-generated row ID | | `name` | TEXT(255)| NOT NULL, UNIQUE | Namespaced name (`namespace/identifier`) | | `provider` | TEXT(255)| NOT NULL | Provider identifier (e.g. `openai`) | | `model` | TEXT(255)| NOT NULL | Model identifier (e.g. `gpt-4`) | | `config_blob` | JSON | NOT NULL, DEFAULT `{}` | Canonical actor configuration blob | | `config_hash` | TEXT(128)| NOT NULL | SHA-256 hash of `config_blob` | | `graph_descriptor` | JSON | | Adapter-produced graph descriptor | | `yaml_text` | TEXT | | Original YAML source text for the actor config | | `schema_version` | TEXT(20) | NOT NULL, DEFAULT `'1.0'` | Actor config schema version | | `compiled_metadata` | JSON | | Compiler-produced metadata (graph topology etc)| | `unsafe` | BOOLEAN | NOT NULL, DEFAULT FALSE | True when actor is marked unsafe | | `is_built_in` | BOOLEAN | NOT NULL, DEFAULT FALSE | Generated from provider registry | | `is_default` | BOOLEAN | NOT NULL, DEFAULT FALSE | Designated default actor | | `created_at` | DATETIME | NOT NULL | Row creation timestamp | | `updated_at` | DATETIME | NOT NULL | Last update timestamp | **Notes:** - `yaml_text` preserves the original YAML source so that actors can be reconstructed or exported without loss. - `schema_version` tracks which version of the actor YAML schema was used to create the configuration. Defaults to `1.0`. - `compiled_metadata` stores JSON output from the actor compiler (graph topology, resolved tools, etc.). - `name` follows the `namespace/identifier` convention enforced by the domain model validator (e.g. `local/my-actor`, `openai/gpt-4`). --- ## Decision Tree Tables The decision tree subsystem consists of two tables introduced in Stage M3.2: | Table | Model | Description | |--------------------------|--------------------------|----------------------------------| | `decisions` | `DecisionModel` | Decision tree nodes | | `decision_dependencies` | `DecisionDependencyModel` | Decision influence DAG edges | ### decisions Primary key: `decision_id` (26-char ULID). FK to `v3_plans.plan_id` (CASCADE delete). Self-referential FKs for `parent_decision_id`, `corrects_decision_id`, and `superseded_by` (SET NULL). | Column | Type | Constraints | Description | |--------------------------------|-------------|---------------------------------------|-----------------------------------------| | `decision_id` | String(26) | PRIMARY KEY | ULID identifier | | `plan_id` | String(26) | NOT NULL, FK -> v3_plans CASCADE | Parent plan reference | | `parent_decision_id` | String(26) | FK -> decisions SET NULL | Parent decision in tree | | `sequence_number` | Integer | NOT NULL | Monotonic order within plan | | `decision_type` | String(30) | NOT NULL, CHECK (11 enum values) | Classification of the decision | | `question` | Text | NOT NULL | What question was being answered | | `chosen_option` | Text | NOT NULL | The option that was chosen | | `alternatives_considered_json` | Text | | JSON array of alternative options | | `confidence_score` | Float | CHECK (0.0-1.0 or NULL) | Confidence in the decision | | `context_snapshot_json` | Text | NOT NULL | JSON blob of ContextSnapshot | | `rationale` | Text | | Human-readable rationale | | `actor_reasoning` | Text | | Raw LLM reasoning trace | | `downstream_decision_ids_json` | Text | | JSON array of dependent decision ULIDs | | `downstream_plan_ids_json` | Text | | JSON array of spawned plan ULIDs | | `artifacts_produced_json` | Text | | JSON array of ArtifactRef objects | | `created_at` | String(30) | NOT NULL | ISO-8601 timestamp | | `is_correction` | Boolean | NOT NULL, DEFAULT FALSE | Whether this is a correction | | `corrects_decision_id` | String(26) | FK -> decisions SET NULL | Decision being corrected | | `correction_reason` | Text | | Why the correction was made | | `superseded_by` | String(26) | FK -> decisions SET NULL | Decision that supersedes this one | **Check constraints:** - `ck_decisions_type`: `decision_type IN ('prompt_definition', 'invariant_enforced', 'strategy_choice', 'implementation_choice', 'resource_selection', 'subplan_spawn', 'subplan_parallel_spawn', 'tool_invocation', 'error_recovery', 'validation_response', 'user_intervention')` - `ck_decisions_confidence`: `confidence_score IS NULL OR (confidence_score >= 0.0 AND confidence_score <= 1.0)` **Indexes:** - `ix_decisions_plan` on `(plan_id)` - `ix_decisions_parent` on `(parent_decision_id)` - `ix_decisions_type` on `(decision_type)` - `ix_decisions_created` on `(created_at)` - `ix_decisions_plan_seq` on `(plan_id, sequence_number)` ### decision_dependencies Composite primary key: `(source_decision_id, target_decision_id)`. Both columns FK to `decisions.decision_id` with `CASCADE` delete. Stores `relationship_type` (default `influences`) and a self-loop check constraint. | Column | Type | Constraints | Description | |-----------------------|-------------|----------------------------------------|---------------------------------| | `source_decision_id` | String(26) | PK, FK -> decisions CASCADE | Decision that influences | | `target_decision_id` | String(26) | PK, FK -> decisions CASCADE | Decision that is influenced | | `relationship_type` | String(30) | NOT NULL, DEFAULT 'influences' | Type of influence relationship | | `created_at` | String(30) | NOT NULL | ISO-8601 timestamp | **Check constraints:** - `ck_decision_deps_no_self_loop`: `source_decision_id != target_decision_id` **Indexes:** - `ix_decision_deps_source` on `(source_decision_id)` - `ix_decision_deps_target` on `(target_decision_id)` --- ## Migration - **Revision**: `c1_001_tool_registry` - **Depends on**: `b0_001_projects` - **File**: `alembic/versions/c1_001_tool_registry.py`