## Summary This PR adds Kubernetes Helm deployment support for the CleverAgents server. Closes #928 ### What this PR includes - Helm chart under `k8s/` with Deployment, Service, optional Ingress, ConfigMap, ServiceAccount, Secrets, NOTES, and optional Redis subchart configuration. - Multi-stage `Dockerfile.server` for server runtime deployment. - Deployment-focused docs in `k8s/README.md`. - Behave + Robot + benchmark coverage for chart/deployment wiring. ### Review fixes applied (cycle 11 — hurui200320 review #2687) **Critical fix:** 1. **CI SHA256 checksum verification fixed** — All 3 Helm install blocks (`unit_tests`, `integration_tests`, `helm` jobs) now save the tarball using its original filename (`helm-v3.16.4-linux-amd64.tar.gz`) instead of `helm.tgz`, so the `.sha256sum` file can correctly locate and verify it. This was causing all 3 CI Helm jobs to fail with "No such file or directory". **Major fixes (test coverage gaps):** 2. **405 Allow header now tested** — The existing "POST to known path returns method not allowed" scenario now asserts that the `Allow: GET` header is present. Added a second 405 scenario testing `POST /live` for broader `_KNOWN_PATHS` coverage (review item #11). 3. **Security-hardening headers now tested** — New scenario "HTTP responses include security-hardening headers" verifies `content-length`, `x-content-type-options: nosniff`, and `cache-control: no-store` headers are present on HTTP responses. 4. **Lifespan warning logging now tested** — New scenario "Unrecognised lifespan message type logs warning and continues" queues `lifespan.startup` → `lifespan.bogus` → `lifespan.shutdown` and verifies: (a) the app completes the lifespan cycle cleanly, (b) a warning is logged mentioning the unrecognised type. ### Review fixes applied (cycle 10) **Critical/Major fixes (from hurui200320's prior REQUEST_CHANGES):** 1. **Rebased branch onto master** — Removed merge commit per CONTRIBUTING.md rebase-only policy. Clean linear history restored. 2. **Fixed commit message body** — Replaced literal `\n` sequences with actual newlines. `ISSUES CLOSED: #928` footer is now on its own line after a blank separator. 3. **Moved uvicorn import to module top level** — `from uvicorn import run as uvicorn_run` is now at the top of `src/cleveragents/cli/commands/server.py` per Import Guidelines. Updated test mock target from `uvicorn.run` to `cleveragents.cli.commands.server.uvicorn_run`. 4. **Added SHA256 checksum verification** — All three Helm CLI install blocks in `.forgejo/workflows/ci.yml` now download and verify `helm.sha256sum` before extracting the binary. **Minor fixes:** 5. **Added `Dockerfile.server` build to CI** — New "Build Docker image (Server)" step in the `docker` job validates the server Dockerfile. 6. **ASGI 405 Method Not Allowed** — Known paths (`/`, `/live`, `/ready`, `/health`) now return 405 with `Allow: GET` header for non-GET methods, per RFC 9110 §15.5.6. Added `_KNOWN_PATHS` frozenset. 7. **WebSocket close protocol fix** — App now calls `await receive()` to consume the `websocket.connect` event before closing. Changed close code from 1000 (Normal Closure) to 1008 (Policy Violation). 8. **Lifespan handler logging** — Unrecognised lifespan message types are now logged as warnings instead of silently consumed. 9. **Security-hardening headers** — `_send_response` now includes `content-length`, `x-content-type-options: nosniff`, and `cache-control: no-store` on all HTTP responses. 10. **`.dockerignore` credential patterns** — Added `*.pem`, `*.key`, `*.p12`, `*.pfx`, `credentials*.json`. 11. **`--log-level` validation** — Constrained to `click.Choice(["critical", "error", "warning", "info", "debug", "trace"])` for clean CLI validation errors. 12. **Reverted unrelated semgrep pre-commit change** — `pass_filenames` and `entry` restored to original values per atomic commit hygiene. 13. **Removed unused `ReceiveCallable` type alias** and `Callable`/`Awaitable` imports from `asgi_app_steps.py`. 14. **Fixed redundant `shutil.which("helm")` check** — `_skip_if_helm_missing` now returns `bool` to eliminate the duplicate check in `_render_chart`. 15. **Improved test deque error handling** — Lifespan test receive mock now raises descriptive `AssertionError` instead of opaque `IndexError`. 16. **Scope type dispatch** — Changed `if/if/if` to `if/elif/elif` for mutually exclusive ASGI scope types. 17. **Dockerfile.server base image** — Standardised to `python:3.13-slim` (floating minor) consistent with CLI Dockerfile. 18. **Dockerfile layer caching** — Split `uv pip install build` and `python -m build` into separate `RUN` instructions. 19. **Removed extraneous double blank line** in Dockerfile.server. ### Deferred items (acknowledged, not in scope) - PodDisruptionBudget, HorizontalPodAutoscaler, NetworkPolicy — Follow-up for production hardening. - `appVersion: "1.0.0"` placeholder — Needs tracking issue for release versioning alignment. - Readiness probe with downstream dependency checks — Documented limitation. - Cross-system test for probe paths matching ASGI routes — Test enhancement. - Improved benchmarks (helm template timing vs PyYAML parsing) — Benchmark quality improvement. - CI DRY violation (Helm install 3×) — Code quality improvement, consider composite action. - File length limits exceeded (`k8s_helm_chart_steps.py` 551 lines, `helper_k8s_helm_chart.py` 678 lines) — Non-blocking, can be split in follow-up. - `runAsGroup: 1000` in pod security context — Defense-in-depth improvement. - HEAD method support on known paths — RFC compliance, does not affect K8s probes. - `click.Choice` log-level validation via CLI runner test — Test gap. ### Scope note: status-check CI gate The `status-check` job now includes `integration_tests`, `e2e_tests`, and `helm` in its `needs` list. The `helm` job is new in this PR. The `integration_tests` and `e2e_tests` additions fix previously-missing gate checks — included here since this PR modifies both of those jobs to install Helm. ### Quality gates - `nox -e lint` ✅ - `nox -e typecheck` ✅ - `nox -e unit_tests` ✅ (12,321 scenarios passed, 4 skipped) - `nox -e integration_tests` — 3 pre-existing failures in unrelated areas (plan correction, resource types) - `nox -e e2e_tests` — pre-existing failures (LLM API keys not available in local env) - `nox -e coverage_report` ✅ (**97.7%**) Reviewed-on: #1085 Reviewed-by: Jeffrey Phillips Freeman <jeffrey.freeman@cleverthis.com> Co-authored-by: Brent E. Edwards <brent.edwards@cleverthis.com> Co-committed-by: Brent E. Edwards <brent.edwards@cleverthis.com>
CleverAgents Kubernetes Deployment
This directory contains the Helm chart for deploying the CleverAgents server to Kubernetes. The chart provisions a Deployment, Service, and optional Ingress with TLS termination, following the architecture described in the project specification.
Prerequisites
- Kubernetes cluster (v1.24+)
- Helm 3.x
- Container image built from
Dockerfile.serverat the repository root - PostgreSQL database (managed or self-hosted) for server persistence
- (Optional) Redis for multi-instance session affinity and rate limiting
Helm Dependency Setup
This chart declares Redis as an optional Helm dependency. Before running
helm lint, helm template, helm install, or helm upgrade, prepare chart
dependencies:
# Rebuild from Chart.lock/charts if present
helm dependency build ./k8s
# Or refresh dependency versions from upstream repositories
helm dependency update ./k8s
Quick Start
1. Build the Server Image
docker build -f Dockerfile.server -t cleveragents/server:1.0.0 .
Then configure the chart to use the same tag:
helm upgrade --install cleveragents ./k8s \
--set image.repository=cleveragents/server \
--set image.tag=1.0.0
2. Install the Chart
Build/update dependencies first:
helm dependency build ./k8s
Create a Kubernetes Secret for the database URL (recommended):
kubectl create secret generic cleveragents-db \
--from-literal=database-url="postgresql+asyncpg://user:pass@db-host:5432/cleveragents"
Install the chart referencing that secret:
helm install cleveragents ./k8s \
--set database.existingSecret=cleveragents-db \
--set database.existingSecretKey=database-url
3. Verify the Deployment
kubectl get pods -l app.kubernetes.io/name=cleveragents
kubectl get svc cleveragents
Configuration
All configuration is managed through values.yaml. Override values at install
time using --set flags or a custom values file (-f custom-values.yaml).
Key Configuration Options
| Parameter | Description | Default |
|---|---|---|
replicaCount |
Number of server replicas | 1 |
image.repository |
Container image repository | cleveragents/server |
image.tag |
Image tag (defaults to appVersion) | "" |
resources.limits.cpu |
CPU limit | 500m |
resources.limits.memory |
Memory limit | 512Mi |
resources.requests.cpu |
CPU request | 100m |
resources.requests.memory |
Memory request | 128Mi |
ingress.enabled |
Enable Ingress | false |
ingress.allowInsecure |
Allow HTTP-only ingress (dev-only) | false |
ingress.className |
Ingress class name | "" |
ingress.tls |
TLS configuration | [] |
redis.enabled |
Enable Redis subchart | false |
database.url |
PostgreSQL connection URL | "" |
database.existingSecret |
Secret name for DB URL | "" |
server.workers |
Uvicorn worker count | 1 |
server.logLevel |
Server log level | info |
Health Probes
The chart uses separate probe endpoints by default:
- Liveness (
livenessProbe.httpGet.path):/live - Readiness (
readinessProbe.httpGet.path):/ready
This split lets operators configure independent probe timing and failure thresholds. In the current implementation, readiness is process-level and does not yet perform dependency checks (for example live DB/Redis connectivity).
Scaling
Scale the deployment by adjusting replicaCount:
helm dependency build ./k8s
helm upgrade cleveragents ./k8s --set replicaCount=3
When running multiple replicas, enable Redis for session affinity:
helm dependency build ./k8s
helm upgrade cleveragents ./k8s \
--set replicaCount=3 \
--set redis.enabled=true \
--set redis.auth.existingSecret=cleveragents-redis \
--set redis.auth.existingSecretPasswordKey=redis-password
TLS Configuration
Enable Ingress with TLS termination for production deployments:
helm dependency build ./k8s
helm upgrade cleveragents ./k8s \
--set ingress.enabled=true \
--set ingress.className=nginx \
--set ingress.hosts[0].host=cleveragents.example.com \
--set ingress.hosts[0].paths[0].path=/ \
--set ingress.hosts[0].paths[0].pathType=Prefix \
--set ingress.tls[0].secretName=cleveragents-tls \
--set ingress.tls[0].hosts[0]=cleveragents.example.com \
--set ingress.annotations."cert-manager\.io/cluster-issuer"=letsencrypt-prod
TLS termination occurs at the Kubernetes Ingress controller level. The CleverAgents server itself communicates over plain HTTP within the cluster.
By default, the chart fails rendering if ingress.enabled=true and
ingress.tls is not configured. For local-only development, you can explicitly
opt out by setting ingress.allowInsecure=true.
helm dependency build ./k8s
helm upgrade cleveragents ./k8s \
--set ingress.enabled=true \
--set ingress.allowInsecure=true
Database Configuration
The server requires a database connection. For production, use a Kubernetes
Secret and set database.existingSecret.
The chart requires one of:
database.existingSecret(recommended)database.url(allowed for local/dev convenience)
Examples:
# Direct URL
database:
url: "postgresql+asyncpg://user:password@postgres-host:5432/cleveragents"
# Using an existing Secret
database:
existingSecret: "cleveragents-db-credentials"
existingSecretKey: "database-url"
Redis (Optional)
Redis support in this chart is currently infrastructure scaffolding:
- The chart can provision Redis (Bitnami subchart) and inject Redis-related environment variables into the server Deployment.
- The current server runtime does not yet implement Redis-backed session affinity/rate-limiting behavior end-to-end.
Use this section to provision Redis infrastructure now; runtime integration is a follow-up capability.
Redis is deployed as a subchart from the Bitnami Helm repository:
redis:
enabled: true
architecture: standalone
auth:
enabled: true
existingSecret: "redis-credentials"
existingSecretPasswordKey: "redis-password"
master:
persistence:
enabled: true
size: 2Gi
Resource Limits
Default resource limits are conservative. Adjust based on expected load:
resources:
limits:
cpu: "2"
memory: 2Gi
requests:
cpu: 500m
memory: 512Mi
Architecture
The Helm chart deploys the following Kubernetes resources:
- Deployment: Runs the CleverAgents ASGI server with configurable replicas
and writable volumes for
/tmpand/app/data - Service: ClusterIP service exposing the server port (8000)
- Ingress (optional): Routes external traffic with TLS termination
- ConfigMap: Server configuration environment variables
- ServiceAccount: Dedicated service account for pod identity
- Secrets: Auto-generated secrets for database URL and Redis password
(when
existingSecretis not configured) - Redis (optional): Bitnami Redis subchart for session affinity
The server container runs as a non-root user (appuser, uid 1000) with a
read-only root filesystem for security hardening. Writable emptyDir volumes
are mounted at /tmp and /app/data to support runtime operations.
Uninstalling
helm uninstall cleveragents