HTTP API¶
The HTTP surface of a generated agent: /chat and its
server-sent events, threads, approvals, health, readiness, metrics and A2A, with their status
codes and limits.
Both runtimes, fastapi and langgraph-server, serve the same routes with the same auth
policy (see Develop your agent for the runtimes). The examples on this
page were captured from a fresh project on the fake model (MODEL_PROVIDER=fake) under the
shared-bearer policy.
Routes¶
| Route | What it does | Auth action |
|---|---|---|
POST /chat |
Run the agent on one message; the answer streams as server-sent events. | chat.send |
GET /threads |
The caller's threads, most recent first (threads). | thread.list |
GET /threads/ |
A thread's messages. | thread.read |
DELETE /threads/ |
Delete a thread and everything attached to it. | thread.delete |
GET /threads/ |
A thread's approvals. | approval.read |
GET /approvals |
Approvals across threads that the caller may see. | approval.read |
POST /threads/ |
Approve or reject a paused call; the resumed run streams. | approval.decide |
GET /health |
Liveness (probes). | none |
GET /ready |
Readiness. | none |
GET /metrics |
Prometheus text. | none, or METRICS_TOKEN |
GET /a2a/ |
The A2A agent card. | card.read |
POST /a2a/ |
A2A JSON-RPC. | a2a.invoke |
GET /playground, /docs, /openapi.json |
Dev chat page and API docs, only under APP_ENV=dev (404 otherwise). |
none |
<agent> is the agent directory (app by default; A2A_NAME overrides it).
Authentication¶
Every route except the probes, /metrics and the dev-only pages goes through the project's
auth policy (AUTH_POLICY; see Authentication). Under
shared-bearer, clients send the key as a bearer token:
curl -N http://127.0.0.1:8000/chat \
-H "Authorization: Bearer $API_KEY" \
-H 'Content-Type: application/json' \
-H 'Accept: text/event-stream' \
-d '{"message": "What'\''s the weather in San Francisco?"}'
| Status | When |
|---|---|
401 |
No credential, or an invalid one, with a WWW-Authenticate: Bearer challenge ({"detail": "Missing or invalid bearer token."} under shared-bearer). |
403 |
The principal may not take the action, or the thread belongs to another principal. |
503 |
The policy is not configured on the server (an unset API_KEY, incomplete jwt settings): never "no auth". |
Every response carries X-Request-ID; a valid one sent by the caller is echoed, and it is on
every log line of the request. The agent passes it on, with the W3C trace context under OTLP
tracing, to the other agents (protocol: a2a) and the auth: forward and auth: exchange
APIs its tools call (never to other APIs), and a traceparent sent to the A2A routes
continues the caller's trace (under the default PROPAGATE_TRACE_HEADERS=peers; all
continues it on every route)
(Observability).
POST /chat¶
The request body:
| Field | Type | Rules |
|---|---|---|
message |
string | Required, 1 to MAX_ (32 000) characters, valid Unicode. |
thread_id |
string | Optional. Omit it to start a thread (the server generates a random id); send it to continue one you own. 1-128 characters of [A-Za-z0-9_.:-]. |
metadata |
object | Optional, flat: at most MAX_ (16) keys, string, number, boolean or null values, keys and strings at most MAX_ (256) characters. Kept in the run record, never in checkpoints. |
The response is text/event-stream. A real stream, one tool call on the fake model:
event: message.start
data: {"thread_id": "94afcb0f-1a65-4dad-b1d2-5e97c370205d", "run_id": "bc32135d-e8f0-4595-a4f6-76382ec18779"}
event: tool.call
data: {"id": "call_get_weather", "name": "get_weather", "args": {"query": "San Francisco"}}
event: tool.result
data: {"id": "call_get_weather", "name": "get_weather", "result": "It's 60 degrees and foggy.", "is_error": false}
event: message.delta
data: {"text": "Here "}
event: message.delta
data: {"text": "is "}
...
event: message.delta
data: {"text": "foggy."}
event: message.end
data: {"thread_id": "94afcb0f-1a65-4dad-b1d2-5e97c370205d", "run_id": "bc32135d-e8f0-4595-a4f6-76382ec18779", "usage": {"input_tokens": 11, "output_tokens": 11}, "latency_ms": 12, "status": "ok"}
Events¶
| Event | Fields | Notes |
|---|---|---|
message.start |
thread_id, run_id |
First event of every run. A run resumed by a decision adds approval_id and decision. |
message.delta |
text |
A piece of the answer, in order. With a response schema, one event: the answer's JSON text. |
tool.call |
id, name, args |
The model called a tool. args is {} when the model's arguments were not valid JSON. |
tool.result |
id, name, result, is_error |
The tool's result. A failed call also has error_id, and outside APP_ENV=dev its result reads The tool call did not succeed. Reference: <error_ |
message.end |
thread_id, run_id, usage (input_tokens, output_tokens), latency_ms, status |
Last event of a run that ended normally. Paused runs add approval and approvals; a completed run of a project with a response schema adds structured_. |
error |
code, message, error_id, run_id |
Last event of a run that failed. Under APP_ENV=dev it also has detail. |
Idle streams get a : keep-alive comment line every SSE_HEARTBEAT_S (15) seconds; SSE
clients ignore it.
message.end status¶
status |
Meaning |
|---|---|
ok |
The agent answered. |
step_limit |
The run used its RECURSION_LIMIT graph steps. The last message.delta says so, and everything the run did stays in the thread: "continue" picks up with a fresh budget. |
awaiting_ |
The run paused before a gated API call. approval is the call waiting for a decision (approvals lists every one the run waits for). See Human approval. |
Error codes¶
The error event's message is generic and names a reference; the detail is in the server
log under error_id.
code |
When |
|---|---|
run_failed |
An unexpected error during the run. |
timeout |
The run passed RUN_TIMEOUT_S and was cancelled. |
recursion_limit |
The step limit was reached and the closing reply could not be written. |
thread_busy |
Another run holds the thread. |
approval_pending |
The thread waits for a decision (with approvals). |
unavailable |
The database is unreachable, or the run could no longer confirm it was the only run on the thread. |
forbidden |
The thread is no longer the caller's. |
unsupported_ |
The graph paused for input this server cannot collect (an interrupt() of your own). |
invalid_ |
The project has a response schema, and no try of the model's answer fitted it (3 tries), or the graph gave no answer (an agent.py not built with response_), or its answer does not fit and was never checked (an agent.py without StructuredAnswer() in its middleware: it is not delivered on /chat or A2A, but it stays in the thread, KI-172). With StructuredAnswer() the thread keeps nothing of a failed try. |
Structured answers¶
A project with a response schema (app/response_schema.json,
Develop) answers in JSON of that shape. A
completed run sends the answer twice: its JSON text as the run's only message.delta, and the
object as message.end's structured_response. Nothing else the model writes is streamed, and
the answer tool of the tool strategy (final_answer) never appears as a tool.call:
event: message.start
data: {"thread_id": "7f3c...", "run_id": "0b9e..."}
event: tool.call
data: {"id": "call_lookup_order", "name": "lookup_order", "args": {"order_id": "ORD-10442"}}
event: tool.result
data: {"id": "call_lookup_order", "name": "lookup_order", "result": "{\"status\": \"shipped\"}", "is_error": false}
event: message.delta
data: {"text": "{\"category\": \"billing\", \"priority\": \"high\", \"order_id\": \"ORD-10442\", \"summary\": \"Charged twice for a shipped order.\"}"}
event: message.end
data: {"thread_id": "7f3c...", "run_id": "0b9e...", "usage": {...}, "latency_ms": 2140, "status": "ok", "structured_response": {"category": "billing", "priority": "high", "order_id": "ORD-10442", "summary": "Charged twice for a shipped order."}}
A run that pauses for an approval has no answer yet (awaiting_approval); the run the decision
resumes ends with it. A step_limit end has none either (its reply says why). A run whose
answer never fits ends with the error code invalid_structured_response.
On /chat, a busy thread, a pending approval and a request that breaks a limit are refused
before the stream starts, with an HTTP status instead:
| Status | Body | When |
|---|---|---|
409 |
{"code": "thread_busy", "detail": ...} |
A run is in progress on the thread. |
409 |
{"code": "approval_ |
The thread waits for an approval: decide it, or wait until it expires. |
413 |
{"detail": "Request body exceeds MAX_ |
The body is over MAX_. |
422 |
{"detail": [{"type": ..., "loc": [...], "msg": ..., "ctx": {...}}]} |
A field breaks a rule. The error names the field and the rule, never the submitted value. |
500 |
{"detail": "Internal server error. Reference: <id>.", "error_id": ...} |
An unhandled error; the detail is only in the log. |
503 |
{"detail": ..., "error_id": ...} |
The database is unreachable: answered within a few seconds (2 s once the app knows it is down), logged as one warning line. |
Guardrails¶
- One run per thread
- A second
/chaton a thread with a run in progress gets 409thread_busy. The lock is in the process and, underCHECKPOINTER=postgres, a lease row shared by every replica. The holder renews it every 5 s; a replica lost without closing its connections frees its threads 30 s later. A lease is a row, not a database session, so a Postgres restart or failover keeps it and the run goes on. A run whose lease cannot be renewed is stopped (run statusinterrupted) before it writes, so two replicas never run one thread at once. - Limits
- Bodies over
MAX_REQUEST_BYTES(1 MiB) get 413. A message overMAX_MESSAGE_CHARS(32 000) gets 422 on/chatand an invalid-params error over A2A. Metadata over its caps, with nested values, or text that is not valid Unicode gets 422. - Thread ids
- One namespace shared by every caller: an id another principal sent first is theirs (403),
so a predictable id can be claimed ahead of its user, and a 403 reveals that an id is
taken. Omit
thread_idon the first turn, or generate unguessable ids (UUID4) in the client. Underlanggraph-serverthread ids are UUIDs. - Timeouts
- A run is cancelled after
RUN_TIMEOUT_S(300 s; statustimeout). Each model request hasMODEL_TIMEOUT_S(60 s) andMODEL_MAX_RETRIES(2). A client that disconnects cancels its run (statuscancelled). - Step limit
RECURSION_LIMIT(50) graph steps: two to answer and two per tool call made after the previous one returned, so 24 sequential tool calls. The run ends with a reply saying so and statusstep_limit, not an error. The app warns at startup when an API'slimits.max_calls_per_runcannot be reached within the limit.- A valid history after any stop
- Model providers reject a tool call without its result. A timeout, a disconnect, a crash or a database outage can leave one, so every run first answers its thread's open tool calls with an error result placed right after the call (and moves misplaced results back).
- Tool arguments that are not valid JSON
- No tool runs. The agent answers the call with an error result saying so and asks the
model again in the same step, at most twice. The client sees a
tool.callwithargs: {}and an errortool.result. - Tool output that is not valid Unicode
- A tool's result can hold a lone surrogate (an upstream JSON
"\ud800"escape decodes to one), which UTF-8 cannot encode. It is replaced with U+FFFD as the result leaves the tool (UntrustedToolResults), so the thread, the model's next request,tool.result, the answer and A2A replies hold U+FFFD in its place, and the run goes on. - Failed tool calls
- Outside
APP_ENV=deva failed call'stool.result, and its message in the thread history, carry only the generic text and itserror_id. The error text (policy rule, limit, upstream status) goes to the model, which may still paraphrase it in its answer.
Every run is recorded when it starts (running) and updated when it ends: ok,
step_limit, awaiting_approval, error, timeout, cancelled or interrupted. Records
left running by a dead process are marked interrupted within about a minute of its lease
expiring. See Observability for the metrics they feed.
Threads¶
GET /threads?limit=&offset=&scope=- The caller's threads, most recent first:
limit1-100 (default 20),offsetfrom 0. Each row is{thread_id, owner, created_at, updated_at},ownerbeing the hashed principal id.scope=alllists every principal's threads, for a role inAUTH_READ_ACROSS_ROLESonly (403 otherwise); the defaultscope=ownlists only the caller's, read-across roles included. A user's own threads include those an agent started for them; an agent calling for a user (a delegated request, see Authentication) lists, reads, continues and deletes only the threads it started for that user. GET /threads/{thread_id}/messages- The thread's messages in order, for its owner or a read-across role:
{id, role, content}, plustool_calls(id,name,args) on an assistant message andtool_call_id,name,is_erroron a tool result. A failed tool result reads as in the stream: anerror_idand, outside dev, the generic text. DELETE /threads/{thread_id}- Deletes the thread, its checkpoints, run records, approvals and A2A tasks, for its owner only: 204, or 403, 404 (unknown), 409 (a run in progress), 422 (not a valid id).
$ curl -s http://127.0.0.1:8000/threads -H "Authorization: Bearer $API_KEY"
[{"thread_id":"94afcb0f-1a65-4dad-b1d2-5e97c370205d","owner":"a4d26868017c0ccf","created_at":"2026-09-25T02:45:43.215203+00:00","updated_at":"2026-09-25T02:45:43.215230+00:00"}]
Approvals¶
A run that reaches a call the API policy gates pauses and
ends with status awaiting_approval. The concepts and the CLI commands are in
Human approval; this is the wire format.
The approval object¶
| Field | Meaning |
|---|---|
approval_id, thread_id, run_id |
Which approval, on which thread, paused by which run. |
status |
pending, approved, rejected or expired. |
api, method, path, operation_id |
The call: the full path with ids filled in. |
rpc_method, a2a_operation |
Only for a call to a JSON-RPC API (protocol: jsonrpc\|a2a): the request's JSON-RPC method and, for an A2A message, whether it approves or rejects one of the other agent's approvals, read from the body. |
query, body |
The call's query and JSON body, the fields the tool named in redact= masked. Shown to the owner and the deciders while the approval is pending (to read-across roles only under TRACE_); dropped once it is decided unless TRACE_. |
tool, reason |
The tool that made the call and the reason the model gave. |
approvers |
requester and/or role:<name> entries of the rule that gated the call. |
requester, decided_by |
Hashed principal ids. |
requester_actor |
The agent the requester's run acted through (a delegated request), or null. |
decide_with |
How the requester decides: direct, or relayed by the agents the rule lists. |
decided_via |
The agent that relayed the decision, or null (decided directly). |
digest |
sha256:<hex> of the call as shown (api, method, path, operation id, JSON-RPC method and A2A decision, query and body, masked fields masked): a relayed decision must name it. |
nested |
Only for a decision this agent relays to another agent (an A2A message that approves): the other agent's approval it decides, as that agent reported it (agent, approval_id, call with api, method, path, operation_id, query, body, reason, expires_at, digest, reported_by, decide_with), and in its own nested the approval that one relays in turn. Copied from the message the approval binds, so it is exactly what is sent. |
effect |
With nested: the call that will actually happen once approved (the innermost nested call, the agent that makes it, and via, the agents between, the called one first), shown first. The approval expires 5 s before the approval it decides, at the latest. Its query and body (and every nested call's) are dropped once decided, as the call's own are. |
created_at, expires_at, decided_at |
ISO 8601 times in UTC (+00:00), whatever the database's time zone. |
comment |
The decider's comment. |
Routes¶
GET /threads/{thread_id}/approvals- The thread's approvals, newest first. The owner and read-across roles see them all, a decider the ones it may decide; anyone else gets 403.
GET /approvals?status=&limit=&offset=- Across threads, newest first: the caller's own approvals, the ones naming one of its
roles (it may decide them), and every one for a read-across role.
statusfilters by status,limit1-100 (default 20). Each row carries itsthread_id. POST /threads/{thread_id}/approvals/{approval_id}- Body
{"decision": "approve" | "reject", "comment": "...", "digest": "sha256:..."}(the comment at most 1000 characters; the approval'sdigest, required of an agent relaying the person's decision and checked when a person sends it). The run resumes and streams the rest with the/chatevents. An approval is decided once:
| Status | code |
When |
|---|---|---|
403 |
not_an_approver |
The caller is not one of the approvers (a requester decides their own call only when requester is listed). |
403 |
approval_ |
An agent (a delegated request) tried to decide a decide_with: direct approval: the person decides, with their own credentials. |
409 |
approval_ |
The decision named no digest (a relayed one must) or another one: it was taken on a different view of the call. |
404 |
approval_ |
No such approval on this thread. |
409 |
approval_ |
Decided already (the body names its status). |
409 |
thread_busy |
Another run holds the thread. |
410 |
approval_expired |
It expired (approval.), which counts as rejected. |
Health, readiness and metrics¶
| Route | Answer |
|---|---|
GET /health |
Liveness, the process only: {"status": "ok", "runtime": "fastapi", "checkpointer": "memory"}. |
GET /ready |
200 {"status": "ready"} when the database (and the run store) is set up and answers within 2 s, else 503 {"status": "not_ready"}. |
GET /metrics |
Prometheus text when METRICS_ENABLED (default true; 404 otherwise). With METRICS_TOKEN set, only Authorization: Bearer <METRICS_ is answered (401 otherwise). |
The metrics: http_requests_total, http_request_duration_seconds, agent_runs_total (by
status), agent_active_runs, agent_run_duration_seconds, agent_tokens_total,
agent_approvals_total and agent_database_up. The chart probes /ready and /health and
never publishes these three routes; scraping and alerts are in
Observability.
A2A¶
The agent speaks the A2A protocol at /a2a/<agent>. The card
names the URL to call: APP_URL when set, else http://HOST:PORT, so a deployed agent needs
APP_URL (the chart sets it from appUrl or the route hostname).
$ curl -s http://127.0.0.1:8000/a2a/app/.well-known/agent-card.json -H "Authorization: Bearer $API_KEY"
{"name": "app", "description": "my-agent: a LangGraph agent served over the A2A protocol.",
"supportedInterfaces": [{"url": "http://127.0.0.1:8000/a2a/app", "protocolBinding": "JSONRPC", "protocolVersion": "1.0"}],
"version": "0.1.0", "capabilities": {"streaming": true, "extensions": [{"uri": "https://ss7172.github.io/graph-agents-cli/a2a/ext/origin/v1", ...}]},
"securitySchemes": {"bearer": {"httpAuthSecurityScheme": {"description": "Shared bearer key (API_KEY).", "scheme": "bearer"}}}, ...}
- Versions
- A2A 1.0 requests carry the
A2A-Version: 1.0header and use the 1.0 method names (SendMessage,SendStreamingMessage,GetTask,ListTasks,CancelTask,SubscribeToTask). A request without the header is served as A2A 0.3 (message/send, ...) on the same URL, with the same error codes. - Card
- Its description (and its one
chatskill's) isA2A_DESCRIPTION, its versionAGENT_VERSION, and its security scheme follows the auth policy. - Replies
- The A2A
contextIdis the chat thread id.SendMessagereturns the reply as one text part of aresponseartifact;SendStreamingMessagestreams it in chunks, the last markedlastChunk. With a response schema the text part is the answer's exact JSON text, and the artifact (its last chunk, streamed) adds a data part holding the answer,mediaTypeapplication/json; the card listsapplication/jsonamong its output modes. A protobufValueholds every number as a double (1reads1.0), so read the text part where exact integers matter. - Errors
- A message with no text, an empty text part, or over
MAX_MESSAGE_CHARSis an invalid-params error (-32602) before a task is created. An unknown task is-32001under both versions. - Tasks
- A task belongs to the principal that created it: another principal's task id reads as
not found. For an agent calling for a user, the owner is the user and the agent together,
so another agent acting for the same user cannot read, list, continue or cancel it. The
user, calling directly (not delegated), reads (
GetTask), lists (ListTasks) and cancels (CancelTask) the tasks their agents started for them as well as their own; continuing one (a message naming itstaskId) andSubscribeToTaskstay with the agent that started it. UnderCHECKPOINTER=postgres(and underlanggraph-serverwith a PostgresDATABASE_URI) tasks are kept in the database, in tablea2a_tasks(agent_a2a_tasks): every replica sees them and they survive restarts and rollouts, soGetTask,ListTasksand a decision naming ataskIdwork whichever pod they reach. UnderCHECKPOINTER=memorythey are kept in process memory. A task is droppedA2A_TASK_TTL_S(3600 s) after its last update (0: when its thread is deleted), and deleting a thread deletes its tasks. A task whose run ended with its process (a crash, an OOM kill) turnsfailedinstead of stayingworking. Postgres cannot store the character U+0000, so a stored task holds U+FFFD in its place (in a message, a tool's output the reply repeats, or an id); the answer to the request itself, and the memory store, keep it as sent. - Approvals
- A gated run moves the task to
input-required, with a data part{"type": "approval_request", "approval": {...}, "approvals": [...], "approval_json": "..."}. AStructholds every number as a double (1reads1.0, a large integer rounds), soapproval_jsonrepeats the approvals as exact JSON text: read that. The text part names each waiting call, its body (at most 2,000 characters of JSON) and, for a relayed decision, itseffect. The client decides with a message whose data part is{"approval_id": "...", "decision": "approve" | "reject", "comment": "...", "digest": "..."}, under the same checks as the HTTP route: on the task (taskId), or on its context alone (contextId, the task named inreferenceTaskIds), which runs as a new task. A task belongs to its requester, so only the requester decides over A2A;role:approvers use the HTTP routes. An agent calling for a user decides only an approval whose rule relays through it (decide_with: relayed, its actor inrelayers, the approval'sdigestnamed); otherwise its decision leaves the taskinput-requiredwith the note (approval_direct_only: ...), and the person decides over HTTP with their own credentials. - Tasks follow their approval
- A task waiting on an approval ends as the approval does, whichever way it is decided:
once the resumed run ends, every
input-requiredtask of the approval's requester on that thread that lists it (or that the decision named inreferenceTaskIds) takes the run's outcome (completed,failed, orinput-requiredwith the approvals it waits on now). Its status says where the run continued:Continued in task <id>.(a decision sent on the context),Approval <id> was approved outside this task; the run continued there.(over HTTP),Approval <id> was rejected. ..., followed by the run's reply (its first 2,000 characters). With a response schema a completed run's answer is not in that text: the task adds it whole as its lastresponseartifact, the JSON text and the data part, as a decision sent on the task gives. An approval that expires fails it (Approval <id> expired before anyone decided.). Another principal's task on the same thread is left as it is.role:gates are decided over HTTP, not over A2A (a design choice). - Error parts
- A failed task (and a decision refused while approvals still wait) carries a data part
{"type": "error", "code": "..."}:thread_busy(send the message again),forbidden,delegation_too_deep, the approval codes (approval_direct_only,approval_digest_mismatch,approval_expired, ...),approval_pending(a message sent while an approval waits), or the run's error code. Branch on the code, not the text. - The origin extension
- The card lists
https://ss7172.github.io/graph-agents-cli/a2a/ext/origin/v1incapabilities.extensions(optional). An agent calling for a user (a delegated principal) may put the user's own words in the message metadata under that URI:{"origin": {"text": "...", "truncated": false, "hops": 1}}(and, in a decision it relays,approving: the approval it decides). They reach this agent's run with the request's private credentials, whererequire_user_mentionedchecks record ids against them and the model's note quotes them. They are read for delegated principals only (a person's own message is their words), capped atA2A_ORIGIN_MAX_CHARS, never stored with the task (every task is saved without them) and never traced.hopsoverAUTH_MAX_DELEGATION_DEPTHfails the task:delegation chain too deep. The run a decision resumes acts on the words of the request that paused it, not on those the decision carries: the approval keeps them while it waits (never shown; dropped once it is decided or expired). This agent trusts the calling agent's code to relay the words faithfully: it guards against an injected model, not a compromised agent.
A 1.0 call and its answer (trimmed):
$ curl -s http://127.0.0.1:8000/a2a/app -H "Authorization: Bearer $API_KEY" \
-H 'Content-Type: application/json' -H 'A2A-Version: 1.0' \
-d '{"jsonrpc": "2.0", "id": 1, "method": "SendMessage",
"params": {"message": {"messageId": "m1", "role": "ROLE_USER", "parts": [{"text": "weather in Paris?"}]}}}'
{"result": {"task": {"id": "d5d9d335-...", "contextId": "ad0855c8-...",
"status": {"state": "TASK_STATE_COMPLETED", ...},
"artifacts": [{"name": "response", "parts": [{"text": "Here is what I found: It's 90 degrees and sunny."}]}], ...}},
"id": 1, "jsonrpc": "2.0"}
graph-agents-cli run --mode a2a is a ready-made client (see the
CLI reference).
Under langgraph-server¶
The LangGraph Server serves the graph and mounts the same app as custom routes, so every route above behaves the same, with these differences:
DELETE /threads/{thread_id}is the server's own route, with the same owner rule; the app then drops the thread's run records, approvals and A2A tasks./threadson the public route also exposes the server's native thread and run routes. Native run creation skips/chat's guardrails (run timeout, one run per thread, run records); the auth handler still limits it to the caller's threads, and its tools act for the caller: a run context the request sends (context, orconfig.configurable) is replaced with the caller's own id, roles and public attributes,@actorincluded.- The native state routes (
GET /threads/{id}/state,POST /threads/{id}/history,GET /threads/{id}, search and run joins) return the stored state as it is, a failed tool call's error text included; only/chat,/threads/{id}/messagesand A2A replace it with an error id. - A native run cannot resume a paused run (decide through the approval routes). On a thread that has approvals or waits on a gated call, a native run without input or from a checkpoint is refused (403), and such a thread is not copied.
Limitations
- A running A2A task streams from its own replica. While its run goes on,
SubscribeToTaskandCancelTaskreaching another pod are refused (-32004,-32002): retry, or follow the task withGetTask(KI-024). - The run lease is checked in the process, not by the database in the same
transaction: a write already sent when a network partition starts can land after
another replica took the thread (normal reads still follow the newer run's
checkpoints). Keep
tcp_user_timeoutin the DSN below the 30 s lease (KI-018). A thread whose run was on a replica that died answers 409thread_busyfor up to 30 s. - Native runs under
langgraph-servercarry the caller's raw id in checkpoint metadata, and store reads are open to every authenticated principal: namespace per-user data by principal, and do not publish the native routes you do not need.
Each is listed with its impact and workaround in Known issues (for example KI-001 and KI-034).
-
Gate calls, decide them from the CLI, and what binds a decision to its call.
-
The policies behind every route on this page.
-
Every limit, timeout and switch mentioned here, with its default.