Files

3.9 KiB

Sandbox Infrastructure

Overview

The sandbox infrastructure provides resource isolation during plan execution. Each sandbox creates an isolated environment where a plan can read and write to a resource without affecting the original until changes are explicitly committed.

Sandbox Strategies

Strategy Class Description
none NoSandbox No isolation; writes go directly to the original
copy_on_write CopyOnWriteSandbox Filesystem copy for isolation
git_worktree GitWorktreeSandbox Git worktree for git repositories

Sandbox Lifecycle

PENDING -> CREATED -> ACTIVE -> COMMITTED -> CLEANED_UP
                        |           |
                        |           +-> CLEANED_UP
                        |
                        +-> ROLLED_BACK -> ACTIVE
                        |
                        +-> ERRORED -> CLEANED_UP

Checkpoint and Rollback Hooks

Purpose

Checkpoint hooks preserve sandbox state at key points during plan execute/apply flows. When a failure occurs, the CheckpointManager can restore a sandbox to a previously captured checkpoint.

SandboxCheckpoint Model

A frozen Pydantic model capturing a point-in-time snapshot:

Field Type Description
checkpoint_id str ULID identifier
sandbox_id str Sandbox that was checkpointed
plan_id str Plan owning the sandbox
phase str Lifecycle phase (pre_execute, post_execute, pre_apply)
created_at datetime When captured
metadata dict[str, str] Key-value metadata (status, reason, etc.)
snapshot_path str Path to the snapshot directory

CheckpointManager API

mgr = CheckpointManager()

# Create a snapshot before execute
cp = mgr.create_checkpoint(sandbox, plan_id, "pre_execute", {})

# Create a snapshot after successful execute
cp2 = mgr.create_checkpoint(sandbox, plan_id, "post_execute", {"status": "success"})

# Rollback on failure
success = mgr.rollback_to(cp)

# List all checkpoints for a sandbox
checkpoints = mgr.list_checkpoints(sandbox.sandbox_id)

# Delete a checkpoint
mgr.delete_checkpoint(cp.checkpoint_id)

Checkpoint Lifecycle

  1. Pre-execute checkpoint: Created before the execute phase starts. Captures the sandbox state so that a failed execution can be rolled back.

  2. Post-execute checkpoint: Created after a successful execute phase. Preserves the post-execution state before apply begins.

  3. Pre-apply checkpoint: Created before the apply phase starts. Allows rollback if the apply fails or encounters merge conflicts.

  4. Rollback on failure: When execute or apply fails, the system attempts to restore the sandbox to the most recent checkpoint.

Integration with Plan Executor

The PlanExecutor accepts an optional checkpoint_manager parameter. When provided, checkpoint hooks are automatically invoked:

  • Before run_execute: create_checkpoint(sandbox, plan_id, "pre_execute")
  • After successful execute: create_checkpoint(sandbox, plan_id, "post_execute")
  • On execute failure: rollback_to(last_checkpoint)

When no CheckpointManager is injected, all hooks are silently skipped.

Integration with Plan Apply Service

The PlanApplyService also accepts an optional checkpoint_manager:

  • Before apply: create_checkpoint(sandbox, plan_id, "pre_apply")
  • On apply failure: rollback_to(last_checkpoint)

Thread Safety

The CheckpointManager is thread-safe. All mutable state is protected by a reentrant lock (threading.RLock).

Snapshot Storage

Snapshots are stored in temporary directories under the system temp folder. Each snapshot is a full copy of the sandbox working directory at the time of checkpoint creation. Snapshots are cleaned up when:

  • A checkpoint is explicitly deleted via delete_checkpoint()
  • The checkpoint manager goes out of scope (manual cleanup recommended)