Architecture
Synapse accepts OpenAI-compatible chat completion requests and serves each one through one of three backend lanes: the standard lane, the native Vertex lane or the Jev lane. The request body alone decides the lane, so a client opts into Vertex-only features or Jev decisions by adding an extension block, with no separate endpoint.
Request flow
client
│ POST /v1/chat/completions (model = route alias, x-synapse-tenant header)
▼
route lookup ─────────────► 404 model_not_found (unknown alias)
│
▼
guardrails (route policy) ─► 400 content_blocked
│
▼
route planning static route: its legs
│ strategy = "jev": the Jev router picks a tier
▼
lane detection
├─► standard lane (genai: OpenAI, Qwen, oai_compat, Vertex)
├─► native Vertex lane (Vertex REST: :streamGenerateContent)
└─► Jev lane (TypeSafe System One)
│
▼
fallback chain leg 1 → leg 2 → … until one succeeds
│
▼
provider ──► response to the client (JSON or server-sent events)
│
├─► cost ledger tokens and cost per tenant, written asynchronously
└─► metrics OpenTelemetry synapse_* metrics (Prometheus, OTLP)
Each route alias in routes.toml maps to an ordered list of legs (provider plus model).
Synapse tries the legs in order until one succeeds. On the standard lane, any failure moves
to the next leg: an error response, a first-chunk timeout or a broken stream. On the native
Vertex lane, only a 5xx, 429 or 408 response, a connection error or a timeout moves
on; any other 4xx stops the chain. A streaming response can fall back only until its
first chunk reaches the client. See the fallback chains guide.
The ledger write never blocks the response: if the ledger's queue is full, the event is
dropped and counted in synapse_ledger_dropped_total. See the
cost ledger guide and the
metrics catalogue. For a step-by-step walk through the
source, see Request pipeline.
Lanes
Standard lane
Requests without lane triggers use the standard lane, which calls providers through the
genai crate. Any provider with an OpenAI-compatible API
can appear in the chain: OpenAI, Qwen (DashScope) and self-hosted vLLM, Ollama or TGI
through the oai_compat provider. Vertex legs work here too, without the native-only
features. See Providers.
Native Vertex lane
Requests that use Vertex-only features go to the native Vertex lane, which calls the Vertex
AI :streamGenerateContent REST endpoint directly, for streaming and non-streaming clients
alike; for a non-streaming client, Synapse buffers the stream into one response. The OpenAI
message format is translated to Vertex's, and these fields of the request's vertex block
are preserved:
cached_content: acachedContentsresource name, for context caching.media_uris: Cloud Storage (gs://) URIs, attached as file parts with MIME typevideo/mp4.response_schema: a JSON schema sent asgenerationConfig.responseSchemafor constrained decoding.thinking_config: passed through verbatim asgenerationConfig.thinkingConfig.
Only the route's vertex legs take part; other providers cannot serve these features. If
the route has no vertex leg, Synapse returns 400 Bad Request with error code
native_feature_unsupported rather than silently dropping the features. See the
native Vertex guide.
Jev lane
Requests with a jev block carrying typed questions go to the route's typesafe legs.
TypeSafe System One (Jev) evaluates the questions against a state, which defaults to the
request's messages, and returns structured decisions as the message content. If every
typesafe leg fails with a retryable error, the route's remaining legs answer as a normal
chat completion, so clients must handle both response shapes. A route with typesafe legs
returns 400 Bad Request to requests without a jev block. See the
Jev lane guide.
Lane detection
Synapse checks the request body in this order:
- A
jevblock with a non-emptyquestionsmap selects the Jev lane. Avertexblock on the same request still applies if the chain falls back to a native Vertex leg. - A
vertexblock with any ofcached_content,response_schema,thinking_config, or amedia_urisentry that starts withgs://, selects the native Vertex lane. - Anything else uses the standard lane.
For example, this request uses the native Vertex lane:
{
"model": "gemini-flash",
"messages": [{ "role": "user", "content": "Summarise the video." }],
"vertex": {
"cached_content": "projects/my-gcp-project/locations/us/cachedContents/abc123",
"media_uris": ["gs://my-bucket/video.mp4"],
"response_schema": { "type": "object", "properties": { "summary": { "type": "string" } } }
}
}
Lane detection and routing strategy are independent. On a strategy = "jev" route, the Jev
router chooses which tier's legs form the chain, and the lane still follows the request
body: a native Vertex request is served by the nearest tier that has a vertex leg. See the
Jev router guide.