Skip to main content

HTTP API

The gateway binary serves two listeners:

  • the API, on SYNAPSE_ADDR (default 0.0.0.0:8080), with every endpoint on this page except metrics;
  • the metrics endpoint, on SYNAPSE_METRICS_ADDR (default 0.0.0.0:9090).

Both speak plain HTTP and neither authenticates callers; see Security.

Endpoints​

MethodPathPurpose
GET/healthLiveness check.
GET/v1/modelsList the chat route aliases.
POST/v1/chat/completionsOpenAI-compatible chat completions, streaming or not.
POST/v1/embeddingsOpenAI-compatible embeddings.
POST/v1beta/models/{model}:{action}Gemini-native passthrough to Vertex AI.
POST/v1/models/{model}:{action}Gemini-native passthrough to Vertex AI.
POST/google/models/{model}:{action}Gemini-native passthrough to Vertex AI.
POST/typesafe/v1/systemoneJev passthrough to TypeSafe System One.
POST/internal/a2a/agentsRegister an A2A agent.
DELETE/internal/a2a/agents/{id}Remove an A2A agent.
GET/.well-known/a2a-agent-catalog.jsonList the registered A2A agents.
GET/a2a/agents/{id}/.well-known/agent-card.jsonOne agent's A2A card.
GET/a2a/agents/{id}/resolveOne agent's endpoint and card.
GET/metrics and / (metrics port)Prometheus metrics.

Any other path returns 404 with an empty body, and a known path with the wrong method returns 405.

Request headers​

Every endpoint except health, models, A2A and metrics reads these optional attribution headers and records them on the request's cost ledger rows. Tenant attribution explains each one.

HeaderEffect
x-synapse-tenantTenant. Without it, SYNAPSE_DEFAULT_TENANT (default unattributed).
x-synapse-workspaceWorkspace within the tenant.
x-synapse-userEnd user.
x-synapse-threadConversation or agent thread.
x-synapse-messageMessage id. On chat completions and the passthroughs it also becomes the request id.
x-synapse-user-task-typeYour own label for the work, recorded as sent.
x-synapse-ai-task-typeOverrides the AI task type.

Response headers​

Chat completion responses carry the routing report. Jev router describes when each header appears.

HeaderValues
x-synapse-routingstatic, jev or static-override. Always present on chat completions.
x-synapse-tierThe tier whose leg served the request.
x-synapse-tier-decidedThe tier Jev chose, when a different tier served.
x-synapse-reasoning-effortThe serving leg's effort, or client when the request set its own.
x-synapse-routing-degradedtimeout, error, low_confidence or jev_unavailable.

Errors​

Gateway errors are JSON in the OpenAI shape:

{
"error": {
"type": "model_not_found",
"message": "unknown model alias 'chat-typo'",
"code": "model_not_found"
}
}
StatuscodeWhen
400invalid_request_errorThe request is invalid for its route, such as an unknown routing_strategy, a jev block on the wrong kind of route, a mismatched embedding dimensions, or a native Vertex 4xx that stops the chain. Also returned by a passthrough whose provider isn't configured.
400native_feature_unsupportedThe request uses native Vertex features and the route has no vertex leg.
400content_blockedA guardrail policy blocked the request. type is content_policy_violation and a scanners array names the scanners. See Block response.
404model_not_foundThe model is not a chat route alias (chat) or an embedding alias (embeddings).
502all_legs_failedEvery leg failed. A failures array lists each leg's provider, model and message. See When every leg fails.
502upstream_errorA provider failed where no further fallback was possible, such as the last native Vertex leg or an embedding leg.

The message is human-readable and can change between versions; match on code. A body that isn't valid JSON, or doesn't match the endpoint's schema, is rejected before it reaches the gateway, with a plain-text message instead of this shape: 400 for malformed JSON, 415 without a Content-Type: application/json header, and 422 for a missing or mistyped field. Bodies larger than 2 MB are rejected with 413.

The passthrough endpoints return the provider's status and body unchanged when the provider answers. See Gemini passthrough and Jev passthrough.

GET /health​

Returns 200 with the body ok. It doesn't check providers or the ledger; it only shows that the process is serving HTTP.

GET /v1/models​

Lists the chat route aliases from routes.toml, sorted by name:

{
"object": "list",
"data": [
{ "id": "auto", "object": "model", "owned_by": "synapse" },
{ "id": "chat", "object": "model", "owned_by": "synapse" }
]
}

Embedding aliases are not listed.

POST /v1/chat/completions​

An OpenAI chat completion request. model is a route alias; the route decides which providers and models serve it.

FieldTypeDescription
modelstringRequired. A route alias from routes.toml.
messagesarrayRequired. OpenAI messages: role, content (a string, an array of content parts, or null on an assistant tool-call turn), and optionally tool_calls, tool_call_id and name.
streambooleantrue for server-sent events. See Streaming.
temperaturenumberSampling temperature.
max_tokensintegerOutput token limit. Honoured on the native Vertex lane only.
response_formatobject{"type": "text" | "json_object" | "json_schema", "json_schema": {...}}. Standard lane only; on the native Vertex lane use vertex.response_schema.
toolsarrayOpenAI function tools. See Tool calling.
tool_choicestring or objectHonoured on the native Vertex lane only.
routing_strategystringstatic or jev. See Override the decision per request.
vertexobjectNative Vertex features: cached_content, media_uris, response_schema, thinking_config. Selects the native Vertex lane; see Lane detection and the native Vertex guide.
jevobjectTyped questions for the Jev lane: questions, optional state, optional extract. See the Jev lane guide.

The gateway accepts any other field without an error, but forwards very few of them:

  • The standard lane sends the provider messages, tools, temperature, response_format and an OpenAI reasoning_effort. It drops max_tokens, tool_choice and every other field, such as top_p, stop or seed.
  • The native Vertex lane sends messages, tools, tool_choice, temperature, max_tokens and the vertex block. See Other request fields.

Limitations lists the gaps.

Response​

A non-streaming request returns a chat.completion:

{
"id": "chatcmpl-3f1c9a52-7d4e-4b8a-9c1e-2a6b0f9d8e71",
"object": "chat.completion",
"created": 1790620000,
"model": "gemini-3.5-flash-lite",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Hello!" },
"finish_reason": "stop"
}],
"usage": { "prompt_tokens": 9, "completion_tokens": 3, "total_tokens": 12 }
}
  • id is chatcmpl- followed by the request id: the x-synapse-message header, or a generated UUID.
  • model is the model of the leg that served the request, not the route alias.
  • finish_reason is stop, length or tool_calls. A tool call has content: null and a tool_calls array.
  • A hybrid extraction response adds a jev object with answers, survivors and degraded.

Streaming response​

With "stream": true, the response is text/event-stream. Each event is a chat.completion.chunk on a data: line, and the stream ends with data: [DONE]:

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":0,"model":"gemini-3.5-flash-lite","choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":0,"model":"gemini-3.5-flash-lite","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Chunks carry no usage and created is always 0. A failure after the first chunk arrives as an event whose body is an error object, followed by [DONE]; see Fallback while streaming.

POST /v1/embeddings​

FieldTypeDescription
modelstringRequired. An embedding alias from the embeddings table of routes.toml.
inputstring or array of stringsRequired. The texts to embed.
dimensionsintegerOptional. Must equal the alias's dimensions.

The response is an OpenAI list of embedding objects, in input order, with usage.prompt_tokens and usage.total_tokens. model in the response is the alias. See Send a request.

Gemini passthrough​

POST /v1beta/models/{model}:{action}, POST /v1/models/{model}:{action} and POST /google/models/{model}:{action} behave identically; the three prefixes match the API versions Google's Gemini SDKs use. The body is a Gemini request, forwarded verbatim to Vertex AI with the gateway's credentials.

  • {model} is a Vertex AI model name, not a route alias, and {action} is the part after the last colon, such as generateContent, streamGenerateContent or countTokens.
  • streamGenerateContent with ?alt=sse streams the response through.
  • generateContent and streamGenerateContent write a ledger row with lane passthrough. On a 5xx, 429 or 408 response or a connection error, they try the following vertex legs of the route that lists the model; see Gemini SDK clients.
  • Other actions are forwarded once, without a ledger row.
  • A model that no route lists is still forwarded, as a single attempt.
  • Streamed calls (?alt=sse) can run for up to an hour. Every other call is bounded by SYNAPSE_REQUEST_TIMEOUT_SECS.
  • Vertex AI's status and body are returned unchanged. A connection error on the last attempt returns 502 with code upstream_error. If the gateway has no Vertex AI project configured, the endpoint returns 400.

Guardrails don't scan passthrough requests.

Jev passthrough​

POST /typesafe/v1/systemone forwards a TypeSafe System One {state, questions} body verbatim with the gateway's TYPESAFE_API_KEY. The body must be a JSON object; if it has no model, the gateway sets jev-latest. TypeSafe's status and body are returned unchanged, and usage is recorded in the ledger with lane passthrough. Without TYPESAFE_API_KEY, the endpoint returns 400. See Jev passthrough.

A2A agent registry​

The gateway binary serves an in-memory registry of A2A agents, seeded at startup from a2a.toml (see Configuration keys). Each gateway instance has its own registry.

POST /internal/a2a/agents​

Registers an agent. Every field except ttl_seconds is required:

{
"id": "invoice-agent",
"name": "Invoice agent",
"description": "Extracts line items from invoices",
"endpoint_url": "https://agents.example.com/invoice",
"card_url": "https://agents.example.com/invoice/.well-known/agent-card.json",
"tags": ["finance"],
"card": { "name": "Invoice agent", "skills": [] },
"ttl_seconds": 3600
}

Returns 204. If the id is already registered, the existing entry is kept and the request still returns 204. With ttl_seconds, the agent disappears from the catalogue that many seconds after registration.

DELETE /internal/a2a/agents/{id}​

Removes the agent. Returns 204, whether or not the id existed.

warning

The /internal/ endpoints have no authentication. Block them at your proxy; see The A2A admin endpoints are open.

GET /.well-known/a2a-agent-catalog.json​

{
"version": "1.0",
"agents": [{
"id": "invoice-agent",
"name": "Invoice agent",
"description": "Extracts line items from invoices",
"card_url": "https://agents.example.com/invoice/.well-known/agent-card.json",
"endpoint_url": "https://agents.example.com/invoice",
"tags": ["finance"]
}]
}

GET /a2a/agents/{id}/.well-known/agent-card.json​

Returns the agent's card as registered, or 404 with an empty body.

GET /a2a/agents/{id}/resolve​

Returns {"id", "endpoint_url", "card_url", "card"} for the agent, or 404 with an empty body.

Metrics​

GET /metrics and GET / on the metrics port return the Prometheus text format. See Metrics and the metrics catalogue.