Skip to content

Environment variables

The variables the CLI reads, and every setting of the service it generates, with defaults and a link to the guide that explains each group.

Two sets of variables are on this page. The CLI's own change how graph-agents-cli behaves on your machine or in CI. The service's configure the agent a project runs: locally from .env, in a cluster from the chart values and the app Secret.

The CLI

GRAPH_AGENTS_CLI_API_KEY

The bearer credential run, approvals and eval send, locally and with --url, when no Authorization header is given: an API_KEY or a JWT. Read from the process environment only, never from .env; it keeps the credential out of the process list and your shell history. Locally, a shared-bearer project's API_KEY from .env is used when it is unset.

Default: unset.

GRAPH_AGENTS_CLI_APPROVER_API_KEY

The bearer credential eval generate decides role: approval gates with. A gate that lists requester is decided as the eval identity instead. See Evaluation.

Default: unset.

GRAPH_AGENTS_CLI_RUN_PORT

Port of the local server run, approvals and eval start on demand; used exactly (exit 3 when it is taken or not a port). run --port wins over it.

Default: first free of 18080-18089.

GRAPH_AGENTS_CLI_INSTALL_SPEC

Where setup, update, the scaffold upgrade baseline and generated projects' CI install the CLI from: a private mirror, a wheel, or a package index. See Install source override.

Default: the release tag of this repository.

GRAPH_AGENTS_CLI_NO_UPDATE_CHECK

1 turns off the check for a newer release on GitHub (at most every 12 hours), the skills version check and the skills listing of info (npx skills list, which may download the skills package; info then says the skills were not listed). Set it for disconnected installs.

Default: unset.

GRAPH_AGENTS_CLI_DEBUG

1 shows the traceback behind a one-line network, file or parse error. See Exit codes.

Default: unset.

GRAPH_AGENTS_CLI_DISABLE_OVERRIDES

1 ignores extension overrides and additions; every generated CI and CD job sets it. See Extensions.

Default: unset.

GRAPH_AGENTS_CLI_SKIP_VERSION_LOCK

1 lets scaffold enhance run with the running build instead of the version the project records.

Default: unset.

Install source override

GRAPH_AGENTS_CLI_INSTALL_SPEC takes any spec uv tool install accepts. Write {version} where the release number goes, so the upgrade baseline can install an older release:

export GRAPH_AGENTS_CLI_INSTALL_SPEC='git+https://git.example.com/graph-agents-cli@v{version}'
  • {version} is a release number (the cli_version a project records), so the override names releases only. A build between two releases is named with scaffold upgrade --baseline-ref (see Upgrading projects).
  • An override without {version} installs one fixed build: the upgrade baseline refuses it (exit 3), since it cannot install the release a project was made with.
  • An override with a control character or whitespace is refused (exit 3), except the two spaces of a PEP 508 graph-agents-cli @ <url> reference.

The value ends up in each generated project as GRAPH_AGENTS_CLI_SPEC in .github/agent.env (see Project manifest).

Read from the project

Some commands also read the project's .env, below the process environment (a variable set in your shell wins):

Command What it reads
login The provider key or OPENAI_BASE_URL, API_KEY, the AUTH_JWT_* key settings, LANGSMITH_API_KEY when TRACING_ENABLED is on, MODEL_PROVIDER and AUTH_POLICY
run, approvals, eval The whole file, passed to the local server they start; AUTH_POLICY, CHECKPOINTER and API_KEY to decide how to call it
eval grade, eval analyze --judge JUDGE_MODEL_PROVIDER and JUDGE_MODEL_NAME (then the agent's MODEL_PROVIDER and MODEL_NAME), unless eval_config.yaml names the judge
eval submit LANGSMITH_API_KEY, LANGSMITH_ENDPOINT
auth dev-token AUTH_POLICY, APP_ENV, the AUTH_JWT_* key settings (and fills the blank ones)

Read by the tools the CLI runs

The CLI passes its environment to helm, kubectl, docker, git, gh and uv, so their own variables apply (KUBECONFIG, DOCKER_HOST, UV_INDEX_URL, ...). A few are also read by the CLI itself:

Variable Used for
GH_HOST, GITHUB_HOST, GITHUB_SERVER_URL Declare a GitHub Enterprise Server host, so argocd-mode deploy opens its pull request there. See CI/CD.
GITHUB_TOKEN, GH_TOKEN, GH_ENTERPRISE_TOKEN The token for that pull request (the first one set). infra check reports the repository's GitHub settings only when gh is logged in or GITHUB_TOKEN is set.
GITHUB_ACTIONS true marks a GitHub Actions job: a helm-push project deploys to staging and prod directly only there (or with --force-direct).
CI, BUILD_ID, GITHUB_ACTIONS, GITLAB_CI Any of them skips the skills version check and the skills listing of info.
HTTP_PROXY, HTTPS_PROXY, ALL_PROXY, NO_PROXY The CLI's own requests to another machine (a --url agent, the login model probe, the GitHub API) go through them, SOCKS (socks5://, socks5h://) included. Requests to this machine (the local server of run, approvals and eval, the playground) never use a proxy. A proxy the CLI cannot use (another scheme) is a one-line error naming the variable, exit 3.

The generated service

Every setting has a default in code and is documented in the project's .env.example, the full contract. Locally the app reads .env when it starts; in a cluster the chart sets the non-secret values from values-<env>.yaml and the Secret carries the allow-listed secret ones (see Secrets).

How settings are parsed

  • .env sits below the process environment. The app reads .env first, as it is imported, so settings fixed at import (the A2A card and its auth scheme, /docs, CORS, the auth policy's startup check) follow it too. PYTHON_DOTENV_DISABLED=1 skips the file; the project's own tests set it.
  • A value that does not parse stops the app at startup, naming every bad variable at once: the guardrails and limits, the pool sizes, LOG_LEVEL, LOG_FORMAT, METRICS_ENABLED, MODEL_TIMEOUT_S, MODEL_MAX_RETRIES, MODEL_REASONING_EFFORT, MODEL_USE_RESPONSES_API (and their JUDGE_ forms), TRACE_CAPTURE, A2A_TASK_TTL_S, MAX_MESSAGE_CHARS, RESPONSE_FORMAT_STRATEGY and the response schema. Nothing silently falls back to a default.
  • APP_ENV counts as dev only when it is exactly dev. DEV, dev, development or unset are a deployed environment.
  • TRACING_ENABLED is on only for true, yes or 1 (any case); any other value leaves it off.
  • The auth settings follow the policy's startup rule: an unknown AUTH_POLICY never starts, and a misconfigured policy stops the process outside APP_ENV=dev (under dev requests get 503). See Authentication.

Application and runtime

Variable Default Meaning
APP_ENV unset (dev in .env.example) Exactly dev enables /playground, /docs, /openapi.json and the dev relaxations.
HOST 127.0.0.1 (0.0.0.0 in the image) Address the server binds.
PORT 8000 Port the server listens on.
RUNTIME detected fastapi or langgraph-server; overrides the detection (the server image sets LANGGRAPH_SERVER=1).
AGENT_VERSION 0.1.0 The version the A2A card, the OpenAPI document and traces report.

Model

Variable Default Meaning
MODEL_PROVIDER openai (the create choice in .env.example) openai, anthropic, gemini or openai-compatible; fake is the deterministic test model.
MODEL_NAME the create choice The model. Required: a run fails with a message naming it when it is unset.
OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, MODEL_API_KEY The provider key (a secret): one per provider, MODEL_API_KEY for openai-compatible.
OPENAI_BASE_URL The endpoint of an openai-compatible server (vLLM, TGI, Ollama).
MODEL_TIMEOUT_S 60 Timeout of one model request, in seconds; 0 keeps the provider SDK's default.
MODEL_MAX_RETRIES 2 Retries of a failed model request.
MODEL_REASONING_EFFORT unset (the model's default) openai and openai-compatible only: none, minimal, low, medium, high or xhigh (each model accepts some). Sent as reasoning_effort on Chat Completions and as reasoning.effort on the Responses API.
MODEL_USE_RESPONSES_API unset (langchain-openai chooses) openai and openai-compatible only: true sends every request to the Responses API (/v1/responses), false to Chat Completions. Unset: Chat Completions, except for the models langchain-openai knows need the Responses API. A model that refuses function tools with a reasoning effort on Chat Completions ("use /v1/responses") needs true, or MODEL_REASONING_EFFORT=none; a server without /v1/responses needs false or unset.
JUDGE_MODEL_PROVIDER, JUDGE_MODEL_NAME, JUDGE_BASE_URL, JUDGE_API_KEY the agent's provider, model and key The eval judge. See Evaluation.
JUDGE_REASONING_EFFORT, JUDGE_USE_RESPONSES_API the agent's values, when the judge is an OpenAI-API model too The judge's own reasoning effort and API.

Either of the two OpenAI settings set for another provider stops startup, since it would be ignored (fake ignores them). Neither has a create flag: like the timeouts, they are settings of each environment (.env, .env.<env>, the chart values), not of the project.

Variable Default Meaning
RESPONSE_FORMAT_STRATEGY auto Only with a response schema (structured final answers): auto uses the provider's own structured output when the model has it (strict on OpenAI) and the model's client can send the schema (Anthropic's refuses a type list and a schema with no type), else a final_answer tool the model must call; provider or tool forces one (provider with a schema the client cannot send stops startup).
RESPONSE_SCHEMA_PATH <agent directory>/response_schema.json Another response schema file. Set, it must exist, except none: no response schema (text answers) whatever file the project has. The file is part of the agent, so the default fits nearly every project; the project's tests set none (tests/conftest.py), and their own schema tests set a file.

Persistence

Variable Default Meaning
CHECKPOINTER memory (the chart sets postgres) fastapi runtime: memory keeps state in the process (lost on restart), postgres in POSTGRES_DSN.
POSTGRES_DSN The database (a secret). Passed to psycopg unchanged, so every libpq parameter works.
DB_POOL_MIN_SIZE, DB_POOL_MAX_SIZE 1, 10 The connection pool of each process, shared by the checkpointer and the app's tables.
PGCONNECT_TIMEOUT When set, replaces the app's own connect_timeout=5 default.
DATABASE_URI, REDIS_URI langgraph-server runtime: the server's persistence (secrets).
RETENTION_DAYS 0 Delete threads idle for more than N days, hourly, on every replica; 0 keeps everything.

The database behaviour (keepalives, /ready, schema setup) is in Deploy to Kubernetes.

Authentication

Variable Default Meaning
AUTH_POLICY shared-bearer (the create choice) shared-bearer, jwt or custom.
API_KEY shared-bearer: the key clients send as Authorization: Bearer <API_KEY> (a secret). Unset answers 503, never "no auth".
AUTH_READ_ACROSS_ROLES empty Comma list of roles that may read, never continue or delete, other principals' threads (and list their approvals).
AUTH_ADMIN_ROLES empty (nobody) langgraph-server: roles that may manage assistants, crons and the store.
AUTH_ALLOWED_ACTORS empty (no agent) Comma list of the agents that may call this one for a user (*: any); any other delegated request gets 403. See Agents calling agents.
AUTH_DELEGATED_ROLES empty The roles a delegated request keeps; it never reads across, administers or decides as a role: approver whatever it keeps.
AUTH_MAX_DELEGATION_DEPTH 3 How many agents may stand between the user and this one (1-8); a longer chain gets 401.
A2A_DELEGATED_MENTIONS origin How require_user_mentioned treats a request another agent presents: origin (the id must be in the user's own words the agent forwarded too), refuse, or request (the agent's request counts, as in 0.2). See When another agent asks for the user.
A2A_CALLER_NOTE on off drops the system note that tells the model another agent wrote the request (its request stays fenced).
AUTH_FORWARD_HEADERS authorization,cookie langgraph-server with LANGGRAPH_SERVER_URL: the request headers passed on to the server's auth handler; empty forwards nothing.
PRINCIPAL_HASH_SALT When set (a secret), principal ids in logs, traces and run records are HMAC-SHA256 with it instead of a plain hash. Keep it stable.
AUTH_JWT_* The jwt policy: AUTH_JWT_JWKS_URL, AUTH_JWT_PUBLIC_KEY, AUTH_JWT_ISSUER, AUTH_JWT_AUDIENCE, AUTH_JWT_ALGORITHMS, AUTH_JWT_ALLOW_HS, AUTH_JWT_SECRET, AUTH_JWT_PRINCIPAL_CLAIM, AUTH_JWT_ROLES_CLAIM, AUTH_JWT_LEEWAY_S, AUTH_JWT_JWKS_CACHE_S, AUTH_JWT_JWKS_ALLOW_HTTP, AUTH_JWT_ACTOR_CLAIM (act; empty reads every token as the user's own), AUTH_JWT_CLIENT_CLAIM (azp), AUTH_JWT_DIRECT_CLIENTS. Defaults and rules: Authentication.

Outbound APIs

Variable Default Meaning
API_POLICY_PATH api-policy.yaml Path of the outbound API policy. See api-policy.yaml.
each API's base_url_env The API's base URL, per environment in values-<env>.yaml. The URL may carry a path prefix.
each auth: bearer API's token_env The API's token (a secret; api add adds it to secrets.keys).

Token exchange

For auth: exchange APIs (RFC 8693). Outside APP_ENV=dev the app refuses to start when such an API exists and the URL, the client id or the secret is missing; a malformed value stops startup in every environment.

Variable Default Meaning
TOKEN_EXCHANGE_URL unset The issuer's token endpoint. https outside APP_ENV=dev, unless the host is loopback or TOKEN_EXCHANGE_ALLOW_HTTP=true.
TOKEN_EXCHANGE_CLIENT_ID unset This agent's client at the issuer (the chart's values.yaml holds the project name as a placeholder).
TOKEN_EXCHANGE_CLIENT_SECRET unset Its secret. api add --auth exchange adds it to secrets.keys; it never goes in a values file.
TOKEN_EXCHANGE_CLIENT_AUTH client_secret_basic Or client_secret_post (the id and secret in the form).
TOKEN_EXCHANGE_SUBJECT_TOKEN_TYPE urn:ietf:params:oauth:token-type:access_token Or urn:ietf:params:oauth:token-type:jwt.
TOKEN_EXCHANGE_TIMEOUT_MS 2000 The whole exchange's deadline (connecting takes at most 1 s of it); 100-60000.
TOKEN_EXCHANGE_MAX_TTL_S 300 How long an exchanged token is reused at most (never past the caller's own token's expiry, less 30 s); 1-300.
TOKEN_EXCHANGE_FAILURE_TTL_S 10 How long an issuer's refusal is remembered, and how long its circuit breaker stays open after three failures in a row; 1-300.
TOKEN_EXCHANGE_CACHE_MAX 10000 Exchanged tokens kept per process (least recently used first out).
TOKEN_EXCHANGE_ALLOW_HTTP false A plain-http TOKEN_EXCHANGE_URL outside dev, for a trusted in-cluster issuer only.

Guardrails and limits

Variable Default Meaning
RUN_TIMEOUT_S 300 Wall-clock limit of one run; the run is cancelled with status timeout.
RECURSION_LIMIT 50 Graph steps per run: two to answer and two per sequential tool call (24 calls). The run then ends with status step_limit.
MAX_REQUEST_BYTES 1048576 Request body cap (413 above it).
MAX_MESSAGE_CHARS 32000 Longest user message on /chat (422) and A2A (invalid params): one cap for every surface.
MAX_METADATA_KEYS 16 Keys in a /chat metadata object (422 above it).
MAX_METADATA_VALUE_CHARS 256 Characters per metadata key and string value (422 above it).
SSE_HEARTBEAT_S 15 Idle seconds before the stream sends a : keep-alive comment.

What each limit does to a request is in the HTTP API.

Logging, metrics and CORS

Variable Default Meaning
LOG_LEVEL INFO DEBUG, INFO, WARNING, ERROR or CRITICAL. DEBUG also turns on third-party debug output, which can hold message content.
LOG_FORMAT text under APP_ENV=dev, else json fastapi runtime: json or text. The LangGraph Server formats its own lines (LOG_JSON).
METRICS_ENABLED true Serve Prometheus text at GET /metrics (true or false).
METRICS_TOKEN When set (a secret), /metrics answers only Authorization: Bearer <METRICS_TOKEN>.
CORS_ALLOW_ORIGINS empty (no CORS) Comma list of browser origins allowed to call the API.

Tracing

Variable Default Meaning
TRACING_ENABLED false On only for true, yes or 1.
TRACE_CAPTURE metadata metadata: structure, timing, token counts, tool names, error types and hashed ids. full: also prompts, completions, tool I/O and the client's /chat metadata.
LANGSMITH_API_KEY With tracing on, traces go to LangSmith (a secret).
LANGSMITH_PROJECT the project name The LangSmith project.
LANGSMITH_ENDPOINT LangSmith's default A self-hosted LangSmith.
OTEL_EXPORTER_OTLP_ENDPOINT Without a LangSmith key, spans go over OTLP/HTTP here.
OTEL_SERVICE_NAME the project name The OTLP service name.
PROPAGATE_TRACE_HEADERS peers How far the request id and trace context go. peers: calls of other agents (protocol: a2a) and of auth: forward and auth: exchange APIs carry the request's X-Request-ID and, under OTLP, its W3C trace context (other APIs never receive them), and an incoming traceparent continues the caller's trace on the A2A routes (/a2a/...) only. all: every API receives them, and a traceparent is continued on every path (behind a tracing gateway). off (also false, 0, no): neither. true, 1, yes and on read as peers (logged once). Any other value stops startup.

See Observability for what each mode sends where.

A2A

Variable Default Meaning
APP_URL http://HOST:PORT The public base URL the agent card advertises; the chart sets it from appUrl or the route hostname.
A2A_NAME the agent directory (app) Mount name: the card at /a2a/<name>/.well-known/agent-card.json, JSON-RPC at /a2a/<name>.
A2A_DESCRIPTION a generic description What the card (and its chat skill) says the agent does.
A2A_TASK_TTL_S 3600 Seconds a task is kept after its last update; 0 keeps a task until its thread is deleted (in process memory, until restart).
A2A_ORIGIN_MAX_CHARS 4000 The most of the user's own words an agent forwards to the agents it calls, and reads from one calling it (the origin extension).
A2A_FORWARD_ORIGIN auto Whether the A2A client sends the user's own words to the agents it calls: auto (to a peer whose card declares the origin extension) or off.
A2A_CARD_TTL_S 300 Seconds a peer's checked agent card is reused; 0 reads it before every call.
A2A_REPLY_MAX_CHARS 6000 The most of a peer's reply the model reads (the rest is cut, marked [truncated]).

See HTTP API.

langgraph-server runtime

Variable Default Meaning
LANGGRAPH_SERVER_URL in-process loopback Reach the LangGraph Server over HTTP instead (its auth handler then sees AUTH_FORWARD_HEADERS).
LANGGRAPH_SERVER 1 in the server image Marks the server runtime for detection.

The server's own variables (licence, LOG_JSON, ...) are LangChain's; see Develop your agent for the runtime.

Offline

For a disconnected install, set GRAPH_AGENTS_CLI_NO_UPDATE_CHECK=1 and point GRAPH_AGENTS_CLI_INSTALL_SPEC at a mirror; the service needs MODEL_PROVIDER=openai-compatible with OPENAI_BASE_URL on your network and tracing off or at an in-cluster OTLP collector. The whole profile is in Offline profile.

  • Authentication

    The three policies and every AUTH_JWT_* setting.

  • HTTP API

    What the limits and timeouts do to requests and runs.

  • Observability

    Logging, metrics and tracing in practice.