Files
cleveragents-core/docs/reference/database_schema.md
T
freemo 7235d46ade
CI / quality (pull_request) Successful in 19s
CI / lint (pull_request) Successful in 21s
CI / benchmark-publish (pull_request) Has been skipped
CI / security (pull_request) Successful in 50s
CI / typecheck (pull_request) Successful in 58s
CI / build (pull_request) Successful in 29s
CI / integration_tests (pull_request) Successful in 4m16s
CI / unit_tests (pull_request) Successful in 12m18s
CI / docker (pull_request) Successful in 1m30s
CI / benchmark-regression (pull_request) Successful in 25m15s
CI / coverage (pull_request) Successful in 1h21m51s
CI / build (push) Successful in 15s
CI / quality (push) Successful in 17s
CI / lint (push) Successful in 21s
CI / security (push) Successful in 28s
CI / typecheck (push) Successful in 31s
CI / benchmark-regression (push) Has been skipped
CI / integration_tests (push) Successful in 2m47s
CI / unit_tests (push) Successful in 11m0s
CI / docker (push) Successful in 40s
CI / benchmark-publish (push) Successful in 12m22s
CI / coverage (push) Successful in 44m42s
feat(skill): persist flattened tool sets
Add Alembic migration m4_002_skill_flattened_tools to extend the skills
table with five new columns: flattened_tools_json, includes_json,
capability_summary_json, yaml_text, and flattening_hash (SHA-256).  A
defence-in-depth uniqueness constraint (uq_skills_name) is also added.

Update SkillModel with the new column definitions and extend
SkillRepository with update_flattened_tools(), get_flattened_tools(),
needs_refresh(), recompute_flattening_hash(), and
invalidate_cached_summaries() methods.  The existing update() method
now nulls all cached fields on mutation (hash-based invalidation).

All new repository methods follow the session-factory pattern with
@database_retry and flush-but-don-t-commit semantics.  Structured
logging via structlog records cache updates and invalidations.

Database schema docs updated with the new skills table columns and a
persistence-field-to-domain-model mapping table.

Tests:
- 6 Behave scenarios covering create, invalidation, hash staleness,
  refresh recomputation, uniqueness constraint, and namespace filtering
- 2 Robot Framework smoke tests (round-trip and invalidation)
- 3 ASV benchmarks (persist, refresh check, namespace list)

ISSUES CLOSED: #166
2026-02-27 20:09:26 +00:00

383 lines
20 KiB
Markdown

# Database Schema Reference
## Overview
CleverAgents uses SQLAlchemy ORM models with SQLite as the default backend.
All tables are created via `Base.metadata.create_all` through the
`init_database()` helper in
`src/cleveragents/infrastructure/database/models.py`.
## Resource Registry Tables
The resource registry consists of three core tables introduced in Stage B1:
| Table | Model | Description |
|--------------------|----------------------|------------------------------------|
| `resource_types` | `ResourceTypeModel` | Schema-level resource type defs |
| `resources` | `ResourceModel` | Registered resource instances |
| `resource_edges` | `ResourceEdgeModel` | Parent-child DAG edges |
### resource_types
Primary key: `name` (namespaced, e.g. `builtin/git-checkout`).
Stores kind (`physical`/`virtual`), handler references, JSON argument
schemas, allowed parent/child types, auto-discovery config, and
capabilities.
### resources
Primary key: `resource_id` (26-char ULID).
FK to `resource_types.name`. Stores optional namespaced name, location,
JSON properties/metadata, and content hash for equivalence tracking.
### resource_edges
Composite primary key: `(parent_id, child_id)`.
Both columns FK to `resources.resource_id` with `CASCADE` delete.
Stores `link_type` (`contains`, `references`, `derived_from`) and a
self-loop check constraint.
## Robot Migration Smoke Suite
A Robot Framework smoke suite validates that schema creation produces the
expected resource registry tables with correct columns.
### Running the suite
```bash
# Via nox (recommended — runs all Robot integration tests):
nox -s integration_tests
# Directly with robot:
robot --outputdir build/reports/robot robot/resource_registry_migration.robot
```
### Test cases
| Test Case | What it checks |
|--------------------------------------------------|--------------------------------------------------|
| Schema Creation Produces Resource Registry Tables | `resource_types`, `resources`, `resource_edges` exist |
| Resource Registry Tables Have Expected Columns | Core columns present on each table |
| Migration Is Idempotent | Calling `init_database` twice is safe |
The helper script (`robot/helper_resource_registry_migration.py`) uses
`init_database()` with a temporary SQLite file and validates via
`sqlalchemy.inspect`.
## Behave BDD Tests
The Behave feature file `features/resource_registry_tables.feature`
contains comprehensive scenarios covering:
- Table existence after schema creation
- CRUD operations for `ResourceTypeModel`, `ResourceModel`, `ResourceEdgeModel`
- Constraint enforcement (uniqueness, FK, check constraints)
- ORM relationship navigation
- Migration smoke verification
Run Behave tests via:
```bash
nox -s unit_tests
```
## ASV Benchmarks
Performance benchmarks for resource registry operations live in
`benchmarks/resource_registry_migration_bench.py` and measure:
- Schema creation time
- Insert throughput for types, resources, and edges
- DAG query performance
Run benchmarks via:
```bash
nox -s benchmark
```
## Tool and Validation Registry Tables
## Tables
### `tools`
Stores registered tool and validation definitions.
| Column | Type | Constraints | Description |
|-----------------|---------|-------------------------------------|------------------------------------------|
| `tool_id` | TEXT | PRIMARY KEY | ULID identifier |
| `name` | TEXT | NOT NULL, UNIQUE | Namespaced name (`namespace/short_name`) |
| `description` | TEXT | | Human-readable description |
| `tool_type` | TEXT | NOT NULL, CHECK (`tool`/`validation`) | Discriminator for registry queries |
| `source_type` | TEXT | NOT NULL, CHECK (see below) | Implementation source |
| `input_schema` | TEXT | | JSON Schema for tool inputs |
| `output_schema` | TEXT | | JSON Schema for tool outputs |
| `read_only` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool only reads, never writes |
| `writes` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool can write to resources |
| `checkpointable`| BOOLEAN | NOT NULL, DEFAULT FALSE | Tool supports checkpoint/rollback |
| `side_effects` | BOOLEAN | NOT NULL, DEFAULT FALSE | Tool has known side effects |
| `config_yaml` | TEXT | | Raw YAML config for reconstruction |
| `created_at` | TEXT | NOT NULL | ISO-8601 timestamp |
| `updated_at` | TEXT | NOT NULL | ISO-8601 timestamp |
**Check constraints:**
- `ck_tools_tool_type`: `tool_type IN ('tool', 'validation')`
- `ck_tools_source_type`: `source_type IN ('mcp', 'agent_skill', 'builtin', 'custom', 'wrapped')`
**Indexes:**
- `ix_tools_name` on `(name)`
- `ix_tools_tool_type` on `(tool_type)`
---
### `tool_bindings`
Stores resource slot bindings for tools.
| Column | Type | Constraints | Description |
|----------------|---------|--------------------------------------|--------------------------------------|
| `binding_id` | TEXT | PRIMARY KEY | ULID identifier |
| `tool_id` | TEXT | NOT NULL, FK → `tools.tool_id` CASCADE | Parent tool reference |
| `slot_name` | TEXT | NOT NULL | Named resource slot |
| `resource_type`| TEXT | NOT NULL | Expected resource type |
| `binding_mode` | TEXT | NOT NULL, CHECK (see below) | How the slot is resolved |
| `access_level` | TEXT | NOT NULL, DEFAULT `read_only` | Access level on bound resource |
| `created_at` | TEXT | NOT NULL | ISO-8601 timestamp |
**Check constraints:**
- `ck_tool_bindings_mode`: `binding_mode IN ('context', 'static', 'parameter')`
- `ck_tool_bindings_access`: `access_level IN ('read_only', 'read_write')`
**Unique constraints:**
- `uq_tool_bindings_tool_slot`: `UNIQUE(tool_id, slot_name)`
---
### `validation_attachments`
Links validations to resources with mode semantics.
| Column | Type | Constraints | Description |
|-------------------|---------|-----------------------------------------------|------------------------------------|
| `attachment_id` | TEXT | PRIMARY KEY | ULID identifier |
| `resource_id` | TEXT | NOT NULL, FK → `resources.resource_id` CASCADE | Target resource |
| `validation_name` | TEXT | NOT NULL | References `tools.name` (validation)|
| `mode` | TEXT | NOT NULL, DEFAULT `required` | `required` or `informational` |
| `created_at` | TEXT | NOT NULL | ISO-8601 timestamp |
**Check constraints:**
- `ck_validation_attachments_mode`: `mode IN ('required', 'informational')`
**Unique constraints:**
- `uq_validation_attachments_resource_validation`: `UNIQUE(resource_id, validation_name)`
**Indexes:**
- `ix_validation_attachments_resource_id` on `(resource_id)`
---
## Relationships
```
tools 1 ──< tool_bindings (tool_id FK, CASCADE delete)
resources 1 ──< validation_attachments (resource_id FK, CASCADE delete)
validation_attachments.validation_name → tools.name (logical, not enforced by FK)
```
---
### `actors`
Stores actor configurations with YAML text retention, schema version
tracking, and optional compiled metadata.
| Column | Type | Constraints | Description |
|---------------------|----------|-----------------------------|------------------------------------------------|
| `id` | INTEGER | PRIMARY KEY, AUTOINCREMENT | Auto-generated row ID |
| `name` | TEXT(255)| NOT NULL, UNIQUE | Namespaced name (`namespace/identifier`) |
| `provider` | TEXT(255)| NOT NULL | Provider identifier (e.g. `openai`) |
| `model` | TEXT(255)| NOT NULL | Model identifier (e.g. `gpt-4`) |
| `config_blob` | JSON | NOT NULL, DEFAULT `{}` | Canonical actor configuration blob |
| `config_hash` | TEXT(128)| NOT NULL | SHA-256 hash of `config_blob` |
| `graph_descriptor` | JSON | | Adapter-produced graph descriptor |
| `yaml_text` | TEXT | | Original YAML source text for the actor config |
| `schema_version` | TEXT(20) | NOT NULL, DEFAULT `'1.0'` | Actor config schema version |
| `compiled_metadata` | JSON | | Compiler-produced metadata (graph topology etc)|
| `unsafe` | BOOLEAN | NOT NULL, DEFAULT FALSE | True when actor is marked unsafe |
| `is_built_in` | BOOLEAN | NOT NULL, DEFAULT FALSE | Generated from provider registry |
| `is_default` | BOOLEAN | NOT NULL, DEFAULT FALSE | Designated default actor |
| `created_at` | DATETIME | NOT NULL | Row creation timestamp |
| `updated_at` | DATETIME | NOT NULL | Last update timestamp |
**Notes:**
- `yaml_text` preserves the original YAML source so that actors can be
reconstructed or exported without loss.
- `schema_version` tracks which version of the actor YAML schema was
used to create the configuration. Defaults to `1.0`.
- `compiled_metadata` stores JSON output from the actor compiler
(graph topology, resolved tools, etc.).
- `name` follows the `namespace/identifier` convention enforced by the
domain model validator (e.g. `local/my-actor`, `openai/gpt-4`).
---
## Skill Registry Tables
The skill registry stores namespaced skill definitions with cached
flattened tool sets for fast resolution.
| Table | Model | Description |
|----------------|--------------------|------------------------------------|
| `skills` | `SkillModel` | Registered skill definitions |
| `skill_items` | `SkillItemModel` | Child items (tools, includes, etc) |
### `skills`
Primary key: `name` (namespaced, e.g. `local/code-tools`).
| Column | Type | Constraints | Description |
|--------------------------|--------------|----------------------|------------------------------------------------|
| `name` | String(255) | PK, UNIQUE | Namespaced name (`namespace/short_name`) |
| `namespace` | String(100) | NOT NULL | Namespace extracted from name |
| `short_name` | String(150) | NOT NULL | Short name extracted from name |
| `description` | Text | NOT NULL | Human-readable description |
| `version` | String(50) | | Optional version string |
| `metadata_json` | Text | | JSON overrides metadata |
| `flattened_tools_json` | Text | | JSON-serialized flattened tool descriptors |
| `includes_json` | Text | | JSON-serialized includes list |
| `capability_summary_json`| Text | | JSON-serialized capability summary |
| `yaml_text` | Text | | Original YAML text of the skill definition |
| `flattening_hash` | String(64) | | SHA-256 hash for refresh invalidation |
| `created_at` | String(30) | NOT NULL | ISO-8601 timestamp |
| `updated_at` | String(30) | NOT NULL | ISO-8601 timestamp |
**Unique constraints:**
- `uq_skills_name`: `UNIQUE(name)` (defence-in-depth — PK already unique)
**Indexes:**
- `ix_skills_namespace` on `(namespace)` — used for filtered skill listing
### `skill_items`
Child table linking skill components to their parent skill.
| Column | Type | Constraints | Description |
|--------------|--------------|------------------------------------------|--------------------------------|
| `id` | Integer | PK, AUTOINCREMENT | Auto-generated row ID |
| `skill_name` | String(255) | NOT NULL, FK -> skills.name CASCADE | Parent skill reference |
| `item_type` | String(30) | NOT NULL | Item kind (tool_ref, include, etc) |
| `item_name` | String(500) | NOT NULL | Item identifier |
| `item_config`| Text | | JSON configuration blob |
| `item_order` | Integer | NOT NULL | Stable ordering index |
| `created_at` | String(30) | NOT NULL | ISO-8601 timestamp |
**Indexes:**
- `ix_skill_items_skill_name` on `(skill_name)`
- `ix_skill_items_item_type` on `(item_type)`
### Persistence field mapping
| Database Column | Domain Model Property |
|----------------------------|------------------------------------------|
| `flattened_tools_json` | `ResolvedToolEntry` list (JSON) |
| `includes_json` | `SkillInclude` list (JSON) |
| `capability_summary_json` | `SkillCapabilitySummary` (JSON) |
| `yaml_text` | Original YAML source text |
| `flattening_hash` | SHA-256 digest of `yaml_text` |
---
## Decision Tree Tables
The decision tree subsystem consists of two tables introduced in Stage M3.2:
| Table | Model | Description |
|--------------------------|--------------------------|----------------------------------|
| `decisions` | `DecisionModel` | Decision tree nodes |
| `decision_dependencies` | `DecisionDependencyModel` | Decision influence DAG edges |
### decisions
Primary key: `decision_id` (26-char ULID).
FK to `v3_plans.plan_id` (CASCADE delete). Self-referential FKs for
`parent_decision_id`, `corrects_decision_id`, and `superseded_by` (SET NULL).
| Column | Type | Constraints | Description |
|--------------------------------|-------------|---------------------------------------|-----------------------------------------|
| `decision_id` | String(26) | PRIMARY KEY | ULID identifier |
| `plan_id` | String(26) | NOT NULL, FK -> v3_plans CASCADE | Parent plan reference |
| `parent_decision_id` | String(26) | FK -> decisions SET NULL | Parent decision in tree |
| `sequence_number` | Integer | NOT NULL | Monotonic order within plan |
| `decision_type` | String(30) | NOT NULL, CHECK (11 enum values) | Classification of the decision |
| `question` | Text | NOT NULL | What question was being answered |
| `chosen_option` | Text | NOT NULL | The option that was chosen |
| `alternatives_considered_json` | Text | | JSON array of alternative options |
| `confidence_score` | Float | CHECK (0.0-1.0 or NULL) | Confidence in the decision |
| `context_snapshot_json` | Text | NOT NULL | JSON blob of ContextSnapshot |
| `rationale` | Text | | Human-readable rationale |
| `actor_reasoning` | Text | | Raw LLM reasoning trace |
| `downstream_decision_ids_json` | Text | | JSON array of dependent decision ULIDs |
| `downstream_plan_ids_json` | Text | | JSON array of spawned plan ULIDs |
| `artifacts_produced_json` | Text | | JSON array of ArtifactRef objects |
| `created_at` | String(30) | NOT NULL | ISO-8601 timestamp |
| `is_correction` | Boolean | NOT NULL, DEFAULT FALSE | Whether this is a correction |
| `corrects_decision_id` | String(26) | FK -> decisions SET NULL | Decision being corrected |
| `correction_reason` | Text | | Why the correction was made |
| `superseded_by` | String(26) | FK -> decisions SET NULL | Decision that supersedes this one |
**Check constraints:**
- `ck_decisions_type`: `decision_type IN ('prompt_definition', 'invariant_enforced', 'strategy_choice', 'implementation_choice', 'resource_selection', 'subplan_spawn', 'subplan_parallel_spawn', 'tool_invocation', 'error_recovery', 'validation_response', 'user_intervention')`
- `ck_decisions_confidence`: `confidence_score IS NULL OR (confidence_score >= 0.0 AND confidence_score <= 1.0)`
**Indexes:**
- `ix_decisions_plan` on `(plan_id)`
- `ix_decisions_parent` on `(parent_decision_id)`
- `ix_decisions_type` on `(decision_type)`
- `ix_decisions_created` on `(created_at)`
- `ix_decisions_plan_seq` on `(plan_id, sequence_number)`
### decision_dependencies
Composite primary key: `(source_decision_id, target_decision_id)`.
Both columns FK to `decisions.decision_id` with `CASCADE` delete.
Stores `relationship_type` (default `influences`) and a self-loop
check constraint.
| Column | Type | Constraints | Description |
|-----------------------|-------------|----------------------------------------|---------------------------------|
| `source_decision_id` | String(26) | PK, FK -> decisions CASCADE | Decision that influences |
| `target_decision_id` | String(26) | PK, FK -> decisions CASCADE | Decision that is influenced |
| `relationship_type` | String(30) | NOT NULL, DEFAULT 'influences' | Type of influence relationship |
| `created_at` | String(30) | NOT NULL | ISO-8601 timestamp |
**Check constraints:**
- `ck_decision_deps_no_self_loop`: `source_decision_id != target_decision_id`
**Indexes:**
- `ix_decision_deps_source` on `(source_decision_id)`
- `ix_decision_deps_target` on `(target_decision_id)`
---
## Migration
- **Revision**: `c1_001_tool_registry`
- **Depends on**: `b0_001_projects`
- **File**: `alembic/versions/c1_001_tool_registry.py`