Deploy to Kubernetes¶
One command, graph-agents-cli deploy --env <env>, takes the
agent to any Kubernetes cluster with Helm. It checks everything it can before it changes
anything, rolls out with --wait, and undoes only its own failed revision.
This page covers the three environments, what deploy does in each CD mode, local clusters,
the image, the rules a deploy follows, rollouts and rollback, the Helm chart and an external
database. How CI builds and promotes images is on CI/CD; the app Secret is on
Secrets.
What you need¶
| Tool | Needed for |
|---|---|
helm, kubectl |
every deploy, secrets, infra check |
docker (or a Docker-compatible CLI) with BuildKit |
building the image (build, direct deploy); the Dockerfile's RUN --mount needs BuildKit (the buildx plugin, the default in Docker Desktop) |
git |
image tags (the short commit sha), argocd mode |
gh (or GITHUB_TOKEN) |
argocd mode: opening the pull request |
A tool missing from PATH makes deploy exit 2. The project needs the kubernetes
deployment target (the default of create; a -p/--prototype project has none) and a kube
context for the cluster.
Environments¶
Every Kubernetes project has three environments. Each one has a values file, a namespace and
an app Secret, and under argocd an Argo CD Application in deployment/argocd/.
dev |
staging |
prod |
|
|---|---|---|---|
| Values file | values-dev.yaml |
values-staging. |
values-prod.yaml |
| Namespace | <name>-dev |
<name>-staging |
<name>-prod |
APP_ENV |
dev |
staging |
prod |
| Database | bundled Postgres (and Redis for langgraph-server) |
external, from the Secret | external, from the Secret |
| Traffic | none (Gateway off) | Gateway API HTTPRoute (or an Ingress) |
HTTPRoute (or an Ingress) |
App Secret <name>-app |
optional (secretOptional: true) |
required | required |
| Replicas | 1 | 1 | 2, with a PodDisruptionBudget and topology spread |
| Kube context | the current one is fine | recorded or confirmed | recorded or confirmed |
<name> is the project name, which is also the Helm release name. The manifest records each
environment's namespace and, once you set it, its kube context:
environments:
dev: { context: "", namespace: my-agent-dev }
staging: { context: "staging-cluster", namespace: my-agent-staging }
prod: { context: "prod-cluster", namespace: my-agent-prod }
deploy passes values.yaml and then values-<env>.yaml, so an environment file lists only
what differs.
What deploy does in each CD mode¶
The --cd choice at create time (or scaffold enhance --cd later) fixes how changes reach a
cluster. CI/CD covers what the generated workflows do in each mode.
| Mode | What deploy --env <env> does |
|---|---|
skip (default) |
Direct, for any environment: builds the image, side-loads it into a local cluster or pushes it, applies the app Secret from the env file, then helm upgrade --install. |
helm-push |
Runs Helm with an image (--image, or one it builds and pushes). dev is deployed directly; staging and prod are refused outside CI, even with --image, unless you pass --force-direct. It checks the live Secret but never applies it, and refuses --env-file and --rotate-api-key: the Secret owner runs secrets apply. |
argocd |
Never runs Helm and contacts no cluster: opens a pull request that changes only image.tag in values-<env>.yaml, and Argo CD applies it after the merge. Refuses --env-file and --rotate-api-key like helm-push. --status and --restart do use the cluster. For a GitHub Enterprise Server remote, set GH_ (and GH_ or GITHUB_TOKEN). |
"Outside CI" means the GITHUB_ACTIONS variable is not true. The helm-push refusal reads:
Error: Refusing to deploy staging from outside CI in helm-push mode (with or without --image).
The staging and promote-to-prod workflows run `deploy --env staging --image <ref>` on the self-hosted runner (GITHUB_ACTIONS=true); pass --force-direct to deploy from here anyway.
Deploy to a local cluster¶
kind, k3d, minikube, k3s, Docker Desktop, Rancher Desktop and OrbStack all work. A side-loaded image never leaves your machine, so any valid registry name will do.
kind create cluster --name agents # or k3d, minikube, Docker Desktop
graph-agents-cli create my-agent --registry localhost/dev
cd my-agent
cp .env.example .env
graph-agents-cli login --write-env # API_KEY and your provider key
graph-agents-cli deploy --env dev --dry-run
graph-agents-cli deploy --env dev
graph-agents-cli deploy --env dev --status
In dev, deploy reads .env.dev, else .env, and applies its allow-listed keys to the
Secret my-agent-app. The pods use the model provider in the chart values, so the Secret needs
that provider's key. To try a cluster without one, add MODEL_PROVIDER: fake under env: in
values-dev.yaml.
How the cluster is recognised
deploy identifies a local cluster from its nodes (a kind:// provider id, k3d node
names, the minikube label, Docker Desktop or OrbStack node names), then confirms it with
the tool's own listing: kind get clusters, k3d cluster list or minikube profile list.
The context name decides only when the nodes cannot be read, and a kind-* name is still
confirmed with kind. Any other cluster gets the image pushed to the registry.
How the image reaches each kind of local cluster:
| Cluster | How the image is loaded |
|---|---|
| kind | kind load docker-image <image> --name <cluster> |
| k3d | k3d image import <image> -c <cluster> |
| minikube | minikube image load <image> (with -p <profile> for a named profile) |
| k3s | docker save, then k3s ctr images import (usually needs root) |
| Docker Desktop, Rancher Desktop, OrbStack | nothing to load: they share the Docker daemon |
The start of a dry run shows the plan, before the rendered manifests. Here the cluster could
not be reached, so the Secret could not be read and the mode falls back to pushing to the
registry; on a kind cluster the mode line reads direct, local-load:
Environment: dev namespace: my-agent-dev
Kube context: docs-unreachable (the kubeconfig's current context; server https://127.0.0.1:21249)
Mode: direct, registry (build, push, and helm upgrade) (its nodes cannot be read and the context name is not a dev cluster)
▸ kubectl get secret my-agent-app -o json -n my-agent-dev --context docs-unreachable
[dry-run] could not read Secret my-agent-app (Command failed (exit code 1): kubectl get secret my-agent-app -o json -n my-agent-dev --context docs-unreachable); the real run refuses unless it holds: OPENAI_API_KEY, API_KEY
[dry-run] docker build -t localhost/dev/my-agent:07e43d3 -f Dockerfile .
[dry-run] docker push localhost/dev/my-agent:07e43d3
No env file found (.env.dev or .env); the Secret my-agent-app is left as is.
[dry-run] helm dependency build deployment/helm/my-agent
[dry-run] fetching subchart(s) postgresql, redis into deployment/helm/my-agent/charts so the render below can run (local only).
[dry-run] helm upgrade --install my-agent deployment/helm/my-agent -f deployment/helm/my-agent/values.yaml -f deployment/helm/my-agent/values-dev.yaml --set image.repository=localhost/dev/my-agent --set image.tag=07e43d3 --set existingSecret=my-agent-app --create-namespace --wait --timeout 5m -n my-agent-dev --kube-context docs-unreachable
[dry-run] a failed rollout prints pod diagnostics and is rolled back (--no-atomic keeps it).
[dry-run] rendering with `helm template` instead:
The image¶
The registry¶
The image is <registry>/<name>:<tag>. create takes the registry from --registry, else
from the git origin remote (ghcr.io/<owner>), else it records the placeholder
ghcr.io/CHANGE-ME, which build and deploy refuse with exit 3:
Error: The image registry is still the placeholder 'ghcr.io/CHANGE-ME'.
Run `graph-agents-cli scaffold enhance --registry <host>/<org>` (it sets create_params.registry in graph-agents-cli-manifest.yaml, image.repository in the chart's values.yaml and IMAGE_REPOSITORY in .github/agent.env), ...
Set it everywhere it is read with one command:
Building¶
build runs docker build with the runtime's
Dockerfile. Its default tag is latest; --registry overrides the manifest, --push pushes
and --dry-run prints the commands:
$ graph-agents-cli build --tag 1.0.0 --push --dry-run
Dry run; would execute:
docker build -t localhost/dev/my-agent:1.0.0 -f Dockerfile .
docker push localhost/dev/my-agent:1.0.0
Tags¶
deploy never deploys latest. The tag of a workstation build is:
- the short commit sha (
07e43d3); <sha>-dirty-<timestamp>when the project has uncommitted changes, with a warning, so the rollout picks them up. Changes underdeployment/,.github/,tests/anddocs/do not count: they never reach the image;- a timestamp outside git;
- whatever
--tagsays.
--image deploys an image that already exists instead of building one. It takes a tagged
reference; a digest (repo@sha256:...) is refused with exit 3, because the chart renders
repository:tag only. An invalid reference or a placeholder registry is refused before
docker runs.
Build once, deploy many
In direct mode every deploy builds again, so dev and staging can run two different
builds under one tag (KI-028).
Push one build and deploy the later environments with --image <ref>, or use a CD mode,
where CI builds each commit once.
The rules a deploy follows¶
deploy checks everything that needs only the project first, then the cluster, and changes
nothing until every check has passed:
- The project: the chart and the values file, the Gateway settings, the image reference,
the env file, placeholders in the chart
env, and thejwtsettings. - The context: the kube context and its API server are printed, and confirmed outside
dev. - The Secret: the Secret the deploy would produce (the live keys merged with the env file)
must hold every required key. When one is missing,
deployexits 1 before anything is built or changed. - The release: a Helm operation already in progress on the release stops the deploy (exit 2).
- The changes: the image is built and loaded or pushed, the namespace created when missing, the Secret applied and Helm run.
--dry-run makes the same checks (it reads the live Secret, read-only), so it refuses what
the real run would. It prints every command, runs helm template instead of helm upgrade and
never prompts. secrets apply follows the env-file and context rules below too.
The env file¶
deploy reads --env-file, else .env.<env>. Only dev falls back to .env: any other
environment without an env file exits 3, so your development keys never reach staging or prod.
Error: No env file for staging: pass --env-file or create .env.staging (the local .env is never used for staging: it holds your development keys) with the allow-listed keys: OPENAI_API_KEY, JUDGE_API_KEY, POSTGRES_DSN, API_KEY, LANGSMITH_API_KEY
An env file that sets none of the allow-listed keys (a keyless project: the fake model, a
keyless openai-compatible endpoint, no shared-bearer key) leaves the Secret as it is, as
a missing env file does in dev. The check runs before anything is built. Outside dev the
chart requires the Secret, so when it does not exist deploy exits 1 before building
(set secretOptional: true in that environment's values if it needs no Secret):
.env.dev sets none of the allow-listed keys (OPENAI_API_KEY, JUDGE_API_KEY, POSTGRES_DSN, API_KEY, LANGSMITH_API_KEY); the Secret my-agent-app is left as is.
Only skip mode reads the env file and applies the Secret. helm-push checks the live Secret
without applying it; argocd never touches the cluster.
The kube context¶
The context is --context, else environments.<env>.context in the manifest, else the
kubeconfig's current context. Outside dev the current context needs a confirmation: a prompt
at a terminal, --yes otherwise. Without either, deploy exits 1:
Error: Refusing to deploy to staging on the kubeconfig's current context 'docs-unreachable' without confirmation.
Record it as environments.staging.context in graph-agents-cli-manifest.yaml or pass --context <name>, or pass --yes to accept the current context.
A context that is named explicitly but missing from the kubeconfig is exit 3.
Settings that cannot work¶
Some settings would give pods that cannot start or cannot answer. deploy refuses them outside
dev (exit 3) and warns in dev:
- A
CHANGE-MEplaceholder in the chartenv, such as the base URLapi addwrites or theOPENAI_BASE_URLof anopenai-compatibleproject. This applies in every CD mode. - Incomplete
jwtsettings. Outsidedevthe chartenvor the Secret must provide a verification key (AUTH_JWT_JWKS_URLorAUTH_JWT_PUBLIC_KEY),AUTH_JWT_ISSUERandAUTH_JWT_AUDIENCE. Indeva missing key source is a warning (the pods would answer 503). - A Gateway without a parent.
values-staging.yamlandvalues-prod.yamlenable the Gateway with a blankgateway.parentRef.name, whichdeployrefuses (exit 3) inskipandhelm-pushmode before anything runs. Set it (andgateway.hostname), or disable the Gateway and enable the Ingress. - An agent it could not call. For the other agents in
api-policy.yaml(protocol: a2a), the rules the agent itself applies: a credential (bearer,forwardorexchange) sent over plainhttpto a host that is not loopback, a single-label name or a cluster-internal.svcname; anauth: exchangeAPI withoutTOKEN_EXCHANGE_URLorTOKEN_EXCHANGE_CLIENT_IDin the chartenv; a plain-httpTOKEN_EXCHANGE_URLto a host that is not loopback (unlessTOKEN_EXCHANGE_ALLOW_HTTPistrue, for a trusted in-cluster issuer). It warns about an unset peer URL, and aboutTOKEN_EXCHANGE_CLIENT_SECRETorPRINCIPAL_HASH_SALTmissing fromsecrets.keys. This applies in every CD mode.
Error: The agent could not call the agents in api-policy.yaml in prod:
- ORDERS_AGENT_URL (http://orders.example.com) must use https outside APP_ENV=dev to carry credentials (auth: exchange); plain http is for loopback and cluster-internal names only
Set them in values-prod.yaml (or values.yaml) under env:.
A value listed in the chart env, even an empty one, overrides the same key in the Secret.
Protected environments¶
deploy --env staging and --env prod are refused while the manifest says
auth_policy_implemented: false, which is what a custom auth policy starts with until you
implement it. See Authentication.
Chart dependencies¶
The chart declares the Bitnami postgresql and redis subcharts. deploy runs
helm dependency build when one is missing from charts/, even under --dry-run, which needs
registry-1.docker.io unless you vendor them under deployment/helm/<name>/charts/.
Rollouts and rollback¶
A rollout¶
deploy runs helm upgrade --install --wait --timeout <--timeout> (default 5m). helm's
output is printed when it returns.
When the rollout fails, deploy prints the pods, their container states, this release's
warning events since the deploy started and the logs. Then, with --atomic (the default), it
rolls back to the newest good revision, or uninstalls a first install that never succeeded.
--no-atomic leaves the failed revision in place for inspection. Either way deploy exits 2.
It only ever undoes the revision it created. When another Helm operation holds the release (a
pending-* revision), deploy refuses up front (exit 2) and prints the command that clears a
lock left by an interrupted helm.
The Secret after a failure¶
Once the release is back where it was (rolled back, failed before a new revision, or a first install uninstalled), the app Secret this deploy applied is put back to its previous values, and a Secret it created is deleted. A bad value never waits for the pods' next restart. A Secret someone changed in the meantime is left alone.
When the release stays on the failed revision (--no-atomic, or another deploy's revision),
the Secret keeps the new values and the error names the changed keys.
Redeploying the same image¶
Redeploying the image the release already runs is announced up front. Helm then records a new
revision but replaces no pod unless the chart values changed. When only the Secret changed,
deploy says so: running pods read the Secret only when they start, so run
deploy --restart.
Status and restart¶
graph-agents-cli deploy --env staging --status # waits up to 60s
graph-agents-cli deploy --env staging --restart # after a Secret rotation
graph-agents-cli deploy --env staging --status --timeout 3m
| Command | What it does | Exit |
|---|---|---|
deploy --status |
Waits up to --timeout (default 60s) for the rollout, then prints replicas, image, Helm revision and each pod's readiness and restarts. When the rollout is incomplete it adds the pods' states, warning events and logs. |
1 when not complete |
deploy --restart |
Restarts the Deployment and waits up to --timeout (default 5m) for the new pods. The old pods keep serving until new ones are ready. |
2 with diagnostics when they do not become ready |
In argocd mode --status runs argocd app get <name>-<env> when the argocd CLI is
installed, and --restart warns that Argo CD's self-heal may revert the restart annotation.
Add the context to printed recovery commands
The helm rollback, helm uninstall and kubectl rollout undo commands deploy suggests
after a failure name the release and namespace but not the kube context. Add
--kube-context <ctx> (helm) or --context <ctx> (kubectl) before running one
(KI-029).
Deploy a system of agents¶
When several projects call each other and a
graph-agents-system.yaml names
them, system deploy deploys them all:
graph-agents-cli system apply --env dev # each project's side of every edge
graph-agents-cli system deploy --env dev # --parallel 3 --only a,b --keep-going --skip-check --dry-run
system check --env ENV(the projects' files only): an error stops here, before any deploy (--skip-checkgoes on anyway).- Callees first. The agents deploy in waves: a wave holds the agents whose callees are all deployed, so a caller's first card check finds them up. A cycle is broken at its first agent in file order, with a warning.
- A few at a time. At most
--paralleldeploys run at once (defaultdeploy.parallelin the file, else 3): seven images built and loaded at once made one agent take 158 s instead of 27-38 s. A failed wave stops the run unless--keep-going. Projects inargocdmode go one at a time: each writes a commit and opens a pull request, and projects may share a repository. - Each is
graph-agents-cli deploy --env ENVin its project directory, with every rule on this page, no terminal input, and its output kept in a log file whose path is printed. Each agent's build, image load and rollout times are printed, read from the commands the deploy prints; a failed one prints the end of its log. system check --live --env ENV: every in-cluster Service the callers dial has a ready endpoint, the other URLs resolve and serve their card, the token URL answers, and each caller's Secret holds the keys its edges need (names only). Skipped with--skip-checkor--dry-run.
Outside dev every project must record its kube context (environments.<env>.context in
its manifest): system deploy refuses (exit 3) before any deploy otherwise, since a deploy
would ask to confirm the kubeconfig's current context and nothing answers it. --live uses
the same contexts, and in dev the current one when none is recorded.
system apply sets, per environment, what makes the agents find each other: the URL each
caller dials (<PEER>_AGENT_URL: the called agent's in-cluster Service,
http://<release>.<namespace>.svc.cluster.local, or the environment's url template) and the
called agent's appUrl, which its card advertises (the two must match, SC05), the audience
(AUTH_JWT_AUDIENCE when empty) and the allowed actors, and the NetworkPolicy rules below.
Set AUTH_JWT_ISSUER yourself: system check compares it with the file's issuer.
The chart¶
The chart lives in deployment/helm/<name>/. It renders a Deployment, a Service, a ConfigMap
from env, a ServiceAccount, and optionally an HTTPRoute or Ingress, a cert-manager
Certificate, an HPA, a PodDisruptionBudget, a NetworkPolicy and a ServiceMonitor.
Pod security¶
The pods meet the restricted Pod Security Standard:
- uid and gid 1000,
runAsNonRoot, no privilege escalation, every capability dropped; - seccomp
RuntimeDefault; - a read-only root filesystem; a
/tmpemptyDir (256Mi) is the only writable path, andHOMEpoints there; - no service-account token (
automountServiceAccountToken: false): the agent never calls the Kubernetes API.
Probes and resources¶
| Probe | Path | Meaning |
|---|---|---|
| Startup | /health |
the process answers (up to 30 × 5 s for the first boot) |
| Liveness | /health |
the process answers; failing restarts the pod |
| Readiness | /ready |
the database answers within 2 s; failing takes the pod out of the Service without a restart |
Requests default to 100m CPU and 256Mi memory, with a 1Gi memory limit and no CPU limit. Prod
requests 250m and 512Mi. /health, /ready and /metrics are never published by the route.
Observability covers the probes and metrics in detail.
Values worth knowing¶
| Value | Default | Purpose |
|---|---|---|
image.tag |
"" |
The chart refuses to render without a tag and never defaults to latest. Quote it (tag: "0123456"): an unquoted number is refused. |
env |
model, auth, A2A_* settings |
Non-secret settings, rendered into a ConfigMap. A key listed here overrides the Secret. |
existingSecret |
"" (<release>-app) |
The app Secret; deploy passes it. |
secretOptional |
true in dev, false elsewhere |
Without the Secret, pods outside dev do not start. |
route. |
/chat (Exact), /threads, /a2a/ (PathPrefix) |
What the HTTPRoute or Ingress publishes. The chart refuses an empty list. |
route.devPaths |
/playground, /docs, /openapi.json |
Published only while env.APP_ENV is exactly dev. |
gateway.* |
enabled, blank parent | Set gateway. and gateway.hostname for staging and prod. |
ingress.* |
disabled | An Ingress instead of the Gateway (className, hostname, annotations). |
tls.* |
none | tls., or tls. with an issuerRef. |
appUrl |
"" |
The public URL the A2A card advertises; derived from the hostname when empty. |
metrics., metrics. |
off | Prometheus scraping; see Observability. |
tracing.* |
off, capture: metadata |
Sets TRACING_ENABLED, TRACE_CAPTURE, OTEL_, LANGSMITH_. |
networkPolicy.* |
off | See NetworkPolicy below. |
shutdown. |
5 |
A stopping pod keeps serving while its endpoints drain. |
shutdown. |
20 |
Then it gets SIGTERM and finishes in-flight requests and streams. |
terminationGracePeriodSeconds |
30 |
Must be longer than the two above together; the chart refuses it otherwise. Raise all three for long runs. |
hpa.*, pdb.*, topologySpread.* |
off (PDB and spread on in prod) | Autoscaling (needs resources. and metrics-server), disruption budget, spreading over nodes. |
probes.*, resources |
see above | Probe timings; requests and limits. |
extraVolumes, extraVolumeMounts |
[] |
Extra mounts, such as a database CA. |
postgresql.*, redis.* |
on in dev only | The Bitnami subcharts, pinned to exact chart versions and image digests. |
The bundled dev Postgres keeps its password in a chart-managed Secret that survives upgrades, and a restart of its pod shuts Postgres down in fast mode (seconds, no crash recovery). Use an external database, or another chart, in staging and prod.
NetworkPolicy¶
networkPolicy.enabled is off in every environment
(KI-033):
it needs a CNI that enforces NetworkPolicy and addresses only you know. Turned on, it admits
only the http port, from the ingressFrom sources when listed. With restrictEgress, egress
is limited to DNS and egressTo.
deployment/helm/<name>/examples/networkpolicy.yaml is a worked staging and prod example: in
from the Gateway's and Prometheus's namespaces; out to DNS, the database, the model endpoint and
HTTPS on public addresses only, with every private range and the metadata endpoint excluded.
Copy its networkPolicy block into values-<env>.yaml and replace every address marked
CHANGE. An abridged view:
networkPolicy:
enabled: true
ingressFrom:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: gateway-system # CHANGE
restrictEgress: true
egressTo:
- to:
- ipBlock:
cidr: 192.0.2.10/32 # CHANGE: the database
ports:
- port: 5432
protocol: TCP
deploy --env <env> --dry-run renders the policy. Once it is deployed, a pod in another
namespace must not reach the agent, and /ready must still answer 200.
Agents calling agents. system apply writes, for each cluster environment of a
system file, a caller's
egressTo rule to each agent it calls (its pods, by the chart's selector labels, in its
namespace, on the pods' port service.targetPort, not the Service's: a NetworkPolicy sees the connection after
the Service has translated it) and a called agent's ingressFrom entry for each caller's
pods, plus the Gateway's namespace (gateway.parentRef.namespace) while its route publishes
paths, since the Gateway must still reach /chat. The rules take effect once
networkPolicy.enabled (and restrictEgress for egress) is on; with restrictEgress, also
allow the token issuer's address. Entries you wrote stay: system apply adds and removes only
the rules for agents of the file.
networkPolicy:
ingressFrom:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: concierge-agent-dev}
podSelector:
matchLabels: {app.kubernetes.io/name: concierge-agent, app.kubernetes.io/instance: concierge-agent}
External database¶
In staging and prod, POSTGRES_DSN (or DATABASE_URI for langgraph-server) in the app
Secret points at a database you run. Give the agent a least-privileged role that owns its own
database. It creates its tables at startup and needs nothing else, and never a superuser:
CREATE ROLE agent LOGIN PASSWORD '...' NOSUPERUSER NOCREATEDB NOCREATEROLE;
CREATE DATABASE agent OWNER agent;
REVOKE ALL ON DATABASE agent FROM PUBLIC;
Require TLS and check the server's certificate:
POSTGRES_DSN=postgresql://agent:<password>@db.example.com:5432/agent?sslmode=verify-full&sslrootcert=/etc/db-ca/ca.crt
The DSN reaches psycopg unchanged, so every libpq parameter works. Mount the CA with
extraVolumes and extraVolumeMounts:
extraVolumes:
- name: db-ca
secret:
secretName: db-ca # a Secret (or ConfigMap) holding ca.crt
extraVolumeMounts:
- name: db-ca
mountPath: /etc/db-ca
readOnly: true
For a publicly trusted certificate use sslrootcert=system, or set PGSSLMODE and
PGSSLROOTCERT in the chart env. deploy, secrets apply and infra check warn outside
dev when the DSN does not require TLS; the value is never printed. On the server, hostssl
entries in pg_hba.conf refuse clear-text connections.
How the app uses the database:
- One connection pool per process (
DB_POOL_MIN_SIZE1,DB_POOL_MAX_SIZE10), health-checked on checkout. Sizemax_connectionsfor replicas × (DB_POOL_MAX_SIZE+ 1); the extra connection holds the run leases. Do not put a transaction-mode PgBouncer in front. - Every connection gets
connect_timeout=5and TCP keepalives unless the DSN sets them, so a dead database is noticed in seconds and the app is ready again seconds after Postgres is. - A replica that starts while Postgres is unreachable stays up, answers
/ready(and requests) with 503, and sets up its schema once Postgres answers, instead of crash-looping. - Schema changes run under an advisory lock, so replicas can start together.
- Backups are yours: pending approvals live in the same database as the threads.
Check the cluster¶
infra check --env <env> is a read-only
report of what the environment needs; it creates nothing and exits 1 when a required item is
missing:
- the tools, the cluster and its Kubernetes version;
- the Gateway API when
gateway.enabled, the ingress class, cert-manager, metrics-server when the HPA is on, and Argo CD inargocdmode; - the namespace, image pull secrets, the app Secret and its required keys, the metrics token
Secret, the
jwtsettings and whether an external DSN requires TLS; - every unreplaced
CHANGE-ME(registry, chart image, chartenv, CODEOWNERS, Argo CDrepoURL); - where pending approvals are kept, and the GitHub settings of the CD modes (see CI/CD).
A check the environment does not need is a skip row. Every command it prints runs as is.
infra check: mode helm-push, env prod
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃ Check ┃ Status ┃ Required ┃ Detail ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩
│ tool: kubectl │ ok │ yes │ found │
│ tool: helm │ ok │ yes │ found │
│ tool: docker │ ok │ yes │ found │
│ tool: gh │ ok │ yes │ found │
│ cluster reachable │ missing │ yes │ The connection to the │
│ │ │ │ server 127.0.0.1:21249 was │
│ │ │ │ refused - did you specify │
│ │ │ │ the right host or port? │
│ placeholder: registry │ ok │ yes │ registry.example.com/team │
│ placeholder: chart env │ ok │ yes │ none │
│ placeholder: CODEOWNERS │ missing │ yes │ 8 rule(s) name a CHANGE-ME │
│ │ │ │ owner │
...
Missing required prerequisites: cluster reachable, placeholder: CODEOWNERS
Limitations¶
| Limitation | What to do |
|---|---|
Concurrent deploys to one release. deploy refuses while another Helm operation holds the release, but two narrow races remain: a failed revision of another deploy can be attributed to this run (and rolled back), and a deploy can apply its Secret just before helm refuses it (KI-030). |
Serialize deploys to one environment: one CI concurrency group, one operator at a time. |
A database that stops answering without closing its connections (a paused host, a proxy that holds traffic) is found by /ready and TCP timeouts, not in seconds: requests on open connections can wait about a minute (KI-017). |
Set keepalives_* and tcp_user_timeout in the DSN; alert on /ready. |
The Bitnami subcharts come from registry-1., which rate-limits anonymous pulls; their images are pinned by digest, and a pin must be refreshed if the digest is withdrawn (KI-082). |
Authenticate pulls, or vendor the charts under charts/. |
| Rolling upgrades from an older build can run one thread on an old and a new pod at once, and the chart has no rollout strategy value (KI-032). | See Upgrading projects. |
A failed reinstall after helm uninstall --keep-history rolls back to the uninstalled release (KI-031). |
Avoid --keep-history. |
| A Secret-only change does not restart pods, and a failed first install leaves its namespace (KI-075). | deploy --restart; delete the namespace if unwanted. |
helm's failure reason is printed on stdout (KI-072); the tag is passed with --set (KI-074); some chart values are validated late (KI-081). |
Keep stdout in CI logs; keep route and metrics as maps. |
The langgraph-server licence is not checked before a deploy (KI-076). |
Add the licence variable to secrets.keys. |
| Rollback was verified with helm 4.3, and k3d and minikube detection against fakes (KI-077). | Run deploy --dry-run first there. |
Next steps¶
-
Build once per commit and promote through
helm-pushor Argo CD pull requests. -
Apply, check and rotate the app Secret.
-
The checklist before production traffic.
-
Every flag of
deploy,build,secretsandinfra check.