Environment variables¶
The variables the CLI reads, and every setting of the service it generates, with defaults and a link to the guide that explains each group.
Two sets of variables are on this page. The CLI's own change how graph-agents-cli
behaves on your machine or in CI. The service's configure the agent a project runs:
locally from .env, in a cluster from the chart values and the app Secret.
The CLI¶
GRAPH_AGENTS_CLI_API_KEY-
The bearer credential
run,approvalsandevalsend, locally and with--url, when noAuthorizationheader is given: anAPI_KEYor a JWT. Read from the process environment only, never from.env; it keeps the credential out of the process list and your shell history. Locally, ashared-bearerproject'sAPI_KEYfrom.envis used when it is unset.Default: unset.
GRAPH_AGENTS_CLI_APPROVER_API_KEY-
The bearer credential
eval generatedecidesrole:approval gates with. A gate that listsrequesteris decided as the eval identity instead. See Evaluation.Default: unset.
GRAPH_AGENTS_CLI_RUN_PORT-
Port of the local server
run,approvalsandevalstart on demand; used exactly (exit 3 when it is taken or not a port).run --portwins over it.Default: first free of
18080-18089. GRAPH_AGENTS_CLI_INSTALL_SPEC-
Where
setup,update, thescaffold upgradebaseline and generated projects' CI install the CLI from: a private mirror, a wheel, or a package index. See Install source override.Default: the release tag of this repository.
GRAPH_AGENTS_CLI_NO_UPDATE_CHECK-
1turns off the check for a newer release on GitHub (at most every 12 hours), the skills version check and the skills listing ofinfo(npx skills list, which may download theskillspackage;infothen says the skills were not listed). Set it for disconnected installs.Default: unset.
GRAPH_AGENTS_CLI_DEBUG-
1shows the traceback behind a one-line network, file or parse error. See Exit codes.Default: unset.
GRAPH_AGENTS_CLI_DISABLE_OVERRIDES-
1ignores extension overrides and additions; every generated CI and CD job sets it. See Extensions.Default: unset.
GRAPH_AGENTS_CLI_SKIP_VERSION_LOCK-
1letsscaffold enhancerun with the running build instead of the version the project records.Default: unset.
Install source override¶
GRAPH_AGENTS_CLI_INSTALL_SPEC takes any spec uv tool install accepts. Write {version}
where the release number goes, so the upgrade baseline can install an older release:
{version}is a release number (thecli_versiona project records), so the override names releases only. A build between two releases is named withscaffold upgrade --baseline-ref(see Upgrading projects).- An override without
{version}installs one fixed build: the upgrade baseline refuses it (exit 3), since it cannot install the release a project was made with. - An override with a control character or whitespace is refused (exit 3), except the two
spaces of a PEP 508
graph-agents-cli @ <url>reference.
The value ends up in each generated project as GRAPH_AGENTS_CLI_SPEC in
.github/agent.env (see Project manifest).
Read from the project¶
Some commands also read the project's .env, below the process environment (a variable set
in your shell wins):
| Command | What it reads |
|---|---|
login |
The provider key or OPENAI_BASE_URL, API_KEY, the AUTH_JWT_* key settings, LANGSMITH_ when TRACING_ENABLED is on, MODEL_PROVIDER and AUTH_POLICY |
run, approvals, eval |
The whole file, passed to the local server they start; AUTH_POLICY, CHECKPOINTER and API_KEY to decide how to call it |
eval grade, eval analyze --judge |
JUDGE_ and JUDGE_MODEL_NAME (then the agent's MODEL_PROVIDER and MODEL_NAME), unless eval_config.yaml names the judge |
eval submit |
LANGSMITH_, LANGSMITH_ |
auth dev-token |
AUTH_POLICY, APP_ENV, the AUTH_JWT_* key settings (and fills the blank ones) |
Read by the tools the CLI runs¶
The CLI passes its environment to helm, kubectl, docker, git, gh and uv, so their
own variables apply (KUBECONFIG, DOCKER_HOST, UV_INDEX_URL, ...). A few are also read
by the CLI itself:
| Variable | Used for |
|---|---|
GH_HOST, GITHUB_HOST, GITHUB_ |
Declare a GitHub Enterprise Server host, so argocd-mode deploy opens its pull request there. See CI/CD. |
GITHUB_TOKEN, GH_TOKEN, GH_ |
The token for that pull request (the first one set). infra check reports the repository's GitHub settings only when gh is logged in or GITHUB_TOKEN is set. |
GITHUB_ACTIONS |
true marks a GitHub Actions job: a helm-push project deploys to staging and prod directly only there (or with --force-direct). |
CI, BUILD_ID, GITHUB_ACTIONS, GITLAB_CI |
Any of them skips the skills version check and the skills listing of info. |
HTTP_PROXY, HTTPS_PROXY, ALL_PROXY, NO_PROXY |
The CLI's own requests to another machine (a --url agent, the login model probe, the GitHub API) go through them, SOCKS (socks5://, socks5h://) included. Requests to this machine (the local server of run, approvals and eval, the playground) never use a proxy. A proxy the CLI cannot use (another scheme) is a one-line error naming the variable, exit 3. |
The generated service¶
Every setting has a default in code and is documented in the project's .env.example, the
full contract. Locally the app reads .env when it starts; in a cluster the chart sets the
non-secret values from values-<env>.yaml and the Secret carries the allow-listed secret
ones (see Secrets).
How settings are parsed
.envsits below the process environment. The app reads.envfirst, as it is imported, so settings fixed at import (the A2A card and its auth scheme,/docs, CORS, the auth policy's startup check) follow it too.PYTHON_DOTENV_DISABLED=1skips the file; the project's own tests set it.- A value that does not parse stops the app at startup, naming every bad variable
at once: the guardrails and limits, the pool sizes,
LOG_LEVEL,LOG_FORMAT,METRICS_ENABLED,MODEL_TIMEOUT_S,MODEL_MAX_RETRIES,MODEL_REASONING_EFFORT,MODEL_USE_RESPONSES_API(and theirJUDGE_forms),TRACE_CAPTURE,A2A_TASK_TTL_S,MAX_MESSAGE_CHARS,RESPONSE_FORMAT_STRATEGYand the response schema. Nothing silently falls back to a default. APP_ENVcounts as dev only when it is exactlydev.DEV,dev,developmentor unset are a deployed environment.TRACING_ENABLEDis on only fortrue,yesor1(any case); any other value leaves it off.- The auth settings follow the policy's startup rule: an unknown
AUTH_POLICYnever starts, and a misconfigured policy stops the process outsideAPP_ENV=dev(under dev requests get 503). See Authentication.
Application and runtime¶
| Variable | Default | Meaning |
|---|---|---|
APP_ENV |
unset (dev in .env.example) |
Exactly dev enables /playground, /docs, /openapi.json and the dev relaxations. |
HOST |
127.0.0.1 (0.0.0.0 in the image) |
Address the server binds. |
PORT |
8000 |
Port the server listens on. |
RUNTIME |
detected | fastapi or langgraph-server; overrides the detection (the server image sets LANGGRAPH_). |
AGENT_VERSION |
0.1.0 |
The version the A2A card, the OpenAPI document and traces report. |
Model¶
| Variable | Default | Meaning |
|---|---|---|
MODEL_PROVIDER |
openai (the create choice in .env.example) |
openai, anthropic, gemini or openai-compatible; fake is the deterministic test model. |
MODEL_NAME |
the create choice |
The model. Required: a run fails with a message naming it when it is unset. |
OPENAI_API_KEY, ANTHROPIC_, GOOGLE_API_KEY, MODEL_API_KEY |
The provider key (a secret): one per provider, MODEL_API_KEY for openai-compatible. |
|
OPENAI_BASE_URL |
The endpoint of an openai-compatible server (vLLM, TGI, Ollama). |
|
MODEL_TIMEOUT_S |
60 |
Timeout of one model request, in seconds; 0 keeps the provider SDK's default. |
MODEL_ |
2 |
Retries of a failed model request. |
MODEL_ |
unset (the model's default) | openai and openai-compatible only: none, minimal, low, medium, high or xhigh (each model accepts some). Sent as reasoning_effort on Chat Completions and as reasoning.effort on the Responses API. |
MODEL_ |
unset (langchain-openai chooses) | openai and openai-compatible only: true sends every request to the Responses API (/v1/responses), false to Chat Completions. Unset: Chat Completions, except for the models langchain-openai knows need the Responses API. A model that refuses function tools with a reasoning effort on Chat Completions ("use /v1/responses") needs true, or MODEL_; a server without /v1/responses needs false or unset. |
JUDGE_, JUDGE_MODEL_NAME, JUDGE_BASE_URL, JUDGE_API_KEY |
the agent's provider, model and key | The eval judge. See Evaluation. |
JUDGE_, JUDGE_ |
the agent's values, when the judge is an OpenAI-API model too | The judge's own reasoning effort and API. |
Either of the two OpenAI settings set for another provider stops startup, since it would be
ignored (fake ignores them). Neither has a create flag: like the timeouts, they are
settings of each environment (.env, .env.<env>, the chart values), not of the project.
| Variable | Default | Meaning |
|---|---|---|
RESPONSE_ |
auto |
Only with a response schema (structured final answers): auto uses the provider's own structured output when the model has it (strict on OpenAI) and the model's client can send the schema (Anthropic's refuses a type list and a schema with no type), else a final_answer tool the model must call; provider or tool forces one (provider with a schema the client cannot send stops startup). |
RESPONSE_ |
<agent directory>/ |
Another response schema file. Set, it must exist, except none: no response schema (text answers) whatever file the project has. The file is part of the agent, so the default fits nearly every project; the project's tests set none (tests/), and their own schema tests set a file. |
Persistence¶
| Variable | Default | Meaning |
|---|---|---|
CHECKPOINTER |
memory (the chart sets postgres) |
fastapi runtime: memory keeps state in the process (lost on restart), postgres in POSTGRES_DSN. |
POSTGRES_DSN |
The database (a secret). Passed to psycopg unchanged, so every libpq parameter works. | |
DB_POOL_MIN_SIZE, DB_POOL_MAX_SIZE |
1, 10 |
The connection pool of each process, shared by the checkpointer and the app's tables. |
PGCONNECT_ |
When set, replaces the app's own connect_ default. |
|
DATABASE_URI, REDIS_URI |
langgraph-server runtime: the server's persistence (secrets). |
|
RETENTION_DAYS |
0 |
Delete threads idle for more than N days, hourly, on every replica; 0 keeps everything. |
The database behaviour (keepalives, /ready, schema setup) is in
Deploy to Kubernetes.
Authentication¶
| Variable | Default | Meaning |
|---|---|---|
AUTH_POLICY |
shared-bearer (the create choice) |
shared-bearer, jwt or custom. |
API_KEY |
shared-bearer: the key clients send as Authorization: Bearer <API_KEY> (a secret). Unset answers 503, never "no auth". |
|
AUTH_ |
empty | Comma list of roles that may read, never continue or delete, other principals' threads (and list their approvals). |
AUTH_ADMIN_ROLES |
empty (nobody) | langgraph-server: roles that may manage assistants, crons and the store. |
AUTH_ |
empty (no agent) | Comma list of the agents that may call this one for a user (*: any); any other delegated request gets 403. See Agents calling agents. |
AUTH_ |
empty | The roles a delegated request keeps; it never reads across, administers or decides as a role: approver whatever it keeps. |
AUTH_ |
3 |
How many agents may stand between the user and this one (1-8); a longer chain gets 401. |
A2A_ |
origin |
How require_ treats a request another agent presents: origin (the id must be in the user's own words the agent forwarded too), refuse, or request (the agent's request counts, as in 0.2). See When another agent asks for the user. |
A2A_CALLER_NOTE |
on |
off drops the system note that tells the model another agent wrote the request (its request stays fenced). |
AUTH_ |
authorization,cookie |
langgraph-server with LANGGRAPH_: the request headers passed on to the server's auth handler; empty forwards nothing. |
PRINCIPAL_ |
When set (a secret), principal ids in logs, traces and run records are HMAC-SHA256 with it instead of a plain hash. Keep it stable. | |
AUTH_JWT_* |
The jwt policy: AUTH_, AUTH_, AUTH_JWT_ISSUER, AUTH_, AUTH_, AUTH_, AUTH_JWT_SECRET, AUTH_, AUTH_, AUTH_, AUTH_, AUTH_, AUTH_ (act; empty reads every token as the user's own), AUTH_ (azp), AUTH_. Defaults and rules: Authentication. |
Outbound APIs¶
| Variable | Default | Meaning |
|---|---|---|
API_POLICY_PATH |
api-policy.yaml |
Path of the outbound API policy. See api-policy.yaml. |
each API's base_url_env |
The API's base URL, per environment in values-<env>.yaml. The URL may carry a path prefix. |
|
each auth: bearer API's token_env |
The API's token (a secret; api add adds it to secrets.keys). |
Token exchange¶
For auth: exchange APIs (RFC 8693).
Outside APP_ENV=dev the app refuses to start when such an API exists and the URL, the client
id or the secret is missing; a malformed value stops startup in every environment.
| Variable | Default | Meaning |
|---|---|---|
TOKEN_ |
unset | The issuer's token endpoint. https outside APP_ENV=dev, unless the host is loopback or TOKEN_. |
TOKEN_ |
unset | This agent's client at the issuer (the chart's values.yaml holds the project name as a placeholder). |
TOKEN_ |
unset | Its secret. api add --auth exchange adds it to secrets.keys; it never goes in a values file. |
TOKEN_ |
client_ |
Or client_ (the id and secret in the form). |
TOKEN_ |
urn:ietf:params:oauth:token-type:access_ |
Or urn:ietf:params:oauth:token-type:jwt. |
TOKEN_ |
2000 |
The whole exchange's deadline (connecting takes at most 1 s of it); 100-60000. |
TOKEN_ |
300 |
How long an exchanged token is reused at most (never past the caller's own token's expiry, less 30 s); 1-300. |
TOKEN_ |
10 |
How long an issuer's refusal is remembered, and how long its circuit breaker stays open after three failures in a row; 1-300. |
TOKEN_ |
10000 |
Exchanged tokens kept per process (least recently used first out). |
TOKEN_ |
false |
A plain-http TOKEN_ outside dev, for a trusted in-cluster issuer only. |
Guardrails and limits¶
| Variable | Default | Meaning |
|---|---|---|
RUN_TIMEOUT_S |
300 |
Wall-clock limit of one run; the run is cancelled with status timeout. |
RECURSION_LIMIT |
50 |
Graph steps per run: two to answer and two per sequential tool call (24 calls). The run then ends with status step_limit. |
MAX_ |
1048576 |
Request body cap (413 above it). |
MAX_ |
32000 |
Longest user message on /chat (422) and A2A (invalid params): one cap for every surface. |
MAX_ |
16 |
Keys in a /chat metadata object (422 above it). |
MAX_ |
256 |
Characters per metadata key and string value (422 above it). |
SSE_HEARTBEAT_S |
15 |
Idle seconds before the stream sends a : keep-alive comment. |
What each limit does to a request is in the HTTP API.
Logging, metrics and CORS¶
| Variable | Default | Meaning |
|---|---|---|
LOG_LEVEL |
INFO |
DEBUG, INFO, WARNING, ERROR or CRITICAL. DEBUG also turns on third-party debug output, which can hold message content. |
LOG_FORMAT |
text under APP_ENV=dev, else json |
fastapi runtime: json or text. The LangGraph Server formats its own lines (LOG_JSON). |
METRICS_ENABLED |
true |
Serve Prometheus text at GET /metrics (true or false). |
METRICS_TOKEN |
When set (a secret), /metrics answers only Authorization: Bearer <METRICS_. |
|
CORS_ |
empty (no CORS) | Comma list of browser origins allowed to call the API. |
Tracing¶
| Variable | Default | Meaning |
|---|---|---|
TRACING_ENABLED |
false |
On only for true, yes or 1. |
TRACE_CAPTURE |
metadata |
metadata: structure, timing, token counts, tool names, error types and hashed ids. full: also prompts, completions, tool I/O and the client's /chat metadata. |
LANGSMITH_ |
With tracing on, traces go to LangSmith (a secret). | |
LANGSMITH_ |
the project name | The LangSmith project. |
LANGSMITH_ |
LangSmith's default | A self-hosted LangSmith. |
OTEL_ |
Without a LangSmith key, spans go over OTLP/HTTP here. | |
OTEL_ |
the project name | The OTLP service name. |
PROPAGATE_ |
peers |
How far the request id and trace context go. peers: calls of other agents (protocol: a2a) and of auth: forward and auth: exchange APIs carry the request's X-Request-ID and, under OTLP, its W3C trace context (other APIs never receive them), and an incoming traceparent continues the caller's trace on the A2A routes (/a2a/...) only. all: every API receives them, and a traceparent is continued on every path (behind a tracing gateway). off (also false, 0, no): neither. true, 1, yes and on read as peers (logged once). Any other value stops startup. |
See Observability for what each mode sends where.
A2A¶
| Variable | Default | Meaning |
|---|---|---|
APP_URL |
http://HOST:PORT |
The public base URL the agent card advertises; the chart sets it from appUrl or the route hostname. |
A2A_NAME |
the agent directory (app) |
Mount name: the card at /a2a/, JSON-RPC at /a2a/. |
A2A_DESCRIPTION |
a generic description | What the card (and its chat skill) says the agent does. |
A2A_TASK_TTL_S |
3600 |
Seconds a task is kept after its last update; 0 keeps a task until its thread is deleted (in process memory, until restart). |
A2A_ |
4000 |
The most of the user's own words an agent forwards to the agents it calls, and reads from one calling it (the origin extension). |
A2A_ |
auto |
Whether the A2A client sends the user's own words to the agents it calls: auto (to a peer whose card declares the origin extension) or off. |
A2A_CARD_TTL_S |
300 |
Seconds a peer's checked agent card is reused; 0 reads it before every call. |
A2A_ |
6000 |
The most of a peer's reply the model reads (the rest is cut, marked [truncated]). |
See HTTP API.
langgraph-server runtime¶
| Variable | Default | Meaning |
|---|---|---|
LANGGRAPH_ |
in-process loopback | Reach the LangGraph Server over HTTP instead (its auth handler then sees AUTH_). |
LANGGRAPH_SERVER |
1 in the server image |
Marks the server runtime for detection. |
The server's own variables (licence, LOG_JSON, ...) are LangChain's; see
Develop your agent for the runtime.
Offline¶
For a disconnected install, set GRAPH_AGENTS_CLI_NO_UPDATE_CHECK=1 and point
GRAPH_AGENTS_CLI_INSTALL_SPEC at a mirror; the service needs MODEL_PROVIDER=openai-compatible
with OPENAI_BASE_URL on your network and tracing off or at an in-cluster OTLP collector.
The whole profile is in Offline profile.
-
The three policies and every
AUTH_JWT_*setting. -
What the limits and timeouts do to requests and runs.
-
Logging, metrics and tracing in practice.