<?xml version="1.0" encoding="utf-8"?><?xml-stylesheet type="text/xsl" href="atom.xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <id>https://synapse-gateway.io/en/latest/blog/</id>
    <title>Synapse Blog</title>
    <updated>2026-09-28T10:00:00.000Z</updated>
    <generator>https://github.com/jpmonette/feed</generator>
    <link rel="alternate" href="https://synapse-gateway.io/en/latest/blog/"/>
    <subtitle>Synapse Blog</subtitle>
    <icon>https://synapse-gateway.io/en/latest/img/favicon.svg</icon>
    <entry>
        <title type="html"><![CDATA[Introducing Synapse: an LLM gateway that keeps native power]]></title>
        <id>https://synapse-gateway.io/en/latest/blog/introducing-synapse/</id>
        <link href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/"/>
        <updated>2026-09-28T10:00:00.000Z</updated>
        <summary type="html"><![CDATA[Synapse is an open-source Rust LLM gateway that speaks the OpenAI API to your clients and keeps Vertex AI's native features, Jev routing and per-tenant cost accounting behind it.]]></summary>
        <content type="html"><![CDATA[<p>Synapse is an open-source LLM gateway written in Rust. Your clients send standard OpenAI
<code>POST /v1/chat/completions</code> requests, and Synapse routes each one through a config-driven
fallback chain of providers, records what it cost, and hands back an OpenAI-shaped response.</p>
<p>What makes it different is what it refuses to throw away. Vertex AI's context caching,
Cloud Storage media and strict response schemas survive the trip, and a routing model can
decide how capable a model, and how much reasoning, each request deserves. This post explains why we built it, how it is put together
and where it is going.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="we-tried-not-to-write-one">We tried not to write one<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#we-tried-not-to-write-one" class="hash-link" aria-label="Direct link to We tried not to write one" title="Direct link to We tried not to write one" translate="no">​</a></h2>
<p>We started where most teams start: put a generic OpenAI-compatible proxy in front of every
provider and move on. That approach reaches Gemini through an OpenAI-shaped adapter, and the
adapter is where the features we depend on disappear. There is no OpenAI field for a Vertex
<code>cachedContents</code> resource or a <code>gs://</code> video, and an OpenAI <code>response_format</code> schema only
becomes a strict Vertex <code>responseSchema</code> if the adapter translates it, so a translation layer
either drops these features or never learns about them.</p>
<p>We wanted multi-provider routing and fallback, and we wanted Vertex's native features, not
one or the other. So Synapse keeps a dedicated native lane for Vertex, alongside the
OpenAI-compatible lane that serves everything else, and keeps the codebase small enough that
the routing, fallback, ledger and metrics code is yours to read, run and embed.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="three-lanes-one-endpoint">Three lanes, one endpoint<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#three-lanes-one-endpoint" class="hash-link" aria-label="Direct link to Three lanes, one endpoint" title="Direct link to Three lanes, one endpoint" translate="no">​</a></h2>
<p>Every chat request goes to the same endpoint. The request body alone decides which of three
lanes serves it, so a client opts into native features by adding an extension block, with no
second API to learn. The <a class="" href="https://synapse-gateway.io/en/latest/docs/overview/architecture/">architecture overview</a> walks through
the whole request flow.</p>
<ul>
<li class=""><strong>The standard lane</strong> calls providers through the <a href="https://crates.io/crates/genai" target="_blank" rel="noopener noreferrer" class=""><code>genai</code></a>
crate. OpenAI, Qwen (DashScope) and self-hosted vLLM, Ollama or TGI through the
<code>oai_compat</code> provider can all appear in one fallback chain, and so can Vertex, without its
native-only features.</li>
<li class=""><strong>The native Vertex lane</strong> takes any request whose <code>vertex</code> block carries
<code>cached_content</code>, <code>response_schema</code>, <code>thinking_config</code> or a <code>gs://</code> media URI. Synapse
translates the OpenAI messages to Vertex's format and calls Vertex AI's
<code>:streamGenerateContent</code> endpoint directly, buffering the stream for clients that did not
ask for one. Only the route's <code>vertex</code> legs take part; if a route has none, the request fails
with <code>native_feature_unsupported</code> instead of silently losing the features. See the
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/native-vertex/">native Vertex guide</a>.</li>
<li class=""><strong>The Jev lane</strong> sends typed questions to TypeSafe System One (Jev), which returns
structured decisions instead of free text. It can also judge and then extract in a single
call. See the <a class="" href="https://synapse-gateway.io/en/latest/docs/guides/jev-lane/">Jev lane guide</a>.</li>
</ul>
<p>A native request looks like any other chat completion, plus a <code>vertex</code> block:</p>
<div class="language-json codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-json codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token property">"model"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"gemini-flash"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token property">"messages"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token property">"role"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"user"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token property">"content"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"Summarise the video."</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token property">"vertex"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token property">"media_uris"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token string" style="color:rgb(255, 121, 198)">"gs://cloud-samples-data/video/animals.mp4"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token property">"response_schema"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token property">"type"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"object"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token property">"properties"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token property">"summary"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token property">"type"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"string"</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><br></div></code></pre></div></div>
<p>Here <code>gemini-flash</code> is a route alias for <code>gemini-3.5-flash-lite</code> on Vertex, as in the
<a class="" href="https://synapse-gateway.io/en/latest/docs/get-started/quickstart/">quickstart</a>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="routing-by-difficulty">Routing by difficulty<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#routing-by-difficulty" class="hash-link" aria-label="Direct link to Routing by difficulty" title="Direct link to Routing by difficulty" translate="no">​</a></h2>
<p>Lanes decide <em>how</em> a request reaches a provider. Routes decide <em>which</em> legs it tries. A
static route is an ordered list of legs, tried until one succeeds. A <code>strategy = "jev"</code> route
instead declares difficulty tiers, described by the kind of work they suit. For each request,
the Jev router asks Jev how demanding the conversation is and whether it needs step-by-step
reasoning, then serves it from the matching tier with a matching reasoning effort. The
response says what happened in <code>x-synapse-routing</code>, <code>x-synapse-tier</code> and related headers.</p>
<p>A companion post, <a class="" href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/">Jev routing explained</a>, covers the
decision in detail, and the <a class="" href="https://synapse-gateway.io/en/latest/docs/guides/jev-router/">Jev router guide</a> is the reference.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-baseline-not-an-add-on">The baseline, not an add-on<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#the-baseline-not-an-add-on" class="hash-link" aria-label="Direct link to The baseline, not an add-on" title="Direct link to The baseline, not an add-on" translate="no">​</a></h2>
<p>A few things we consider table stakes come with every deployment:</p>
<ul>
<li class=""><strong>Real streaming.</strong> On the standard and native Vertex lanes, Synapse always streams from
upstream, so <code>stream: true</code> clients get
token-by-token server-sent events. Non-streaming clients get the buffered result and, on
the standard lane, keep the full fallback chain. See
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/streaming-and-tools/">Streaming and tool calling</a>.</li>
<li class=""><strong>Tool calling on the standard and native Vertex lanes.</strong> The native lane also honours
<code>tool_choice</code> through Vertex <code>toolConfig</code>; the standard lane forwards tools but drops
<code>tool_choice</code>.</li>
<li class=""><strong>A cost ledger you own.</strong> Every request is attributed to a tenant from the
<code>x-synapse-tenant</code> header and priced from your <code>pricing.toml</code> into SQLite or Postgres, with
optional fan-out to Google Cloud Pub/Sub and AWS SNS. See the
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/cost-ledger/">cost ledger guide</a>.</li>
<li class=""><strong>Metrics.</strong> OpenTelemetry <code>synapse_*</code> metrics, served in Prometheus format on port 9090
and optionally pushed over OTLP. The
<a class="" href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/">OpenTelemetry metrics post</a> explains the pipeline.</li>
<li class=""><strong>Input guardrails.</strong> Named scanner policies that block or observe requests before they
reach a provider. See <a class="" href="https://synapse-gateway.io/en/latest/docs/configuration/guardrails-policy/">Guardrails policy</a>.</li>
</ul>
<p>Synapse is licensed under MPL-2.0, so you can build it into commercial products, and none of this sits behind a paid tier.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="run-it-or-embed-it">Run it, or embed it<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#run-it-or-embed-it" class="hash-link" aria-label="Direct link to Run it, or embed it" title="Direct link to Run it, or embed it" translate="no">​</a></h2>
<p>You can run Synapse as a single binary or as the <code>sustentabilitas/synapse-gateway</code> Docker
image, with the API on port 8080 and metrics on port 9090. The
<a class="" href="https://synapse-gateway.io/en/latest/docs/get-started/quickstart/">quickstart</a> gets a gateway answering in a few minutes.</p>
<p>If your service is written in Rust, you can skip the extra process entirely. Depend on the
<code>synapse-gateway</code> crate with <code>default-features = false</code>, build a <code>Gateway</code> in code, and call
<code>Gateway::chat()</code> in-process, with the same routing, fallback and ledger behaviour as the
binary and no HTTP hop. See
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/embedding-as-library/">Embedding Synapse as a library</a>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-family">The family<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#the-family" class="hash-link" aria-label="Direct link to The family" title="Direct link to The family" translate="no">​</a></h2>
<p>Synapse is a Cargo workspace, and the gateway has siblings that grew out of the same needs:</p>
<ul>
<li class=""><strong><a class="" href="https://synapse-gateway.io/en/latest/docs/synapse-family/proxy/overview/"><code>synapse-proxy</code></a></strong> is a config-driven reverse-proxy
sidecar. It routes by path prefix and stamps a bound identity, such as a tenant, on every
forwarded request, so a sandboxed workload cannot choose its own.</li>
<li class=""><strong><a class="" href="https://synapse-gateway.io/en/latest/docs/synapse-family/a2a/overview/"><code>synapse-a2a</code></a></strong> is an agent-to-agent (A2A) registry
with admin registration and public discovery, served by the gateway binary.</li>
<li class=""><strong><a class="" href="https://synapse-gateway.io/en/latest/docs/synapse-family/mcp/overview/"><code>synapse-mcp</code></a></strong> is an on-demand MCP gateway library
that routes Streamable HTTP tool calls per server and injects the current tenant's identity.</li>
<li class=""><strong><code>synapse-context</code></strong> is the shared context store behind the proxy and the MCP gateway.</li>
</ul>
<p><a class="" href="https://synapse-gateway.io/en/latest/docs/internals/workspace-crates/">Workspace crates</a> shows how they depend on each other.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="whats-next">What's next<a href="https://synapse-gateway.io/en/latest/blog/introducing-synapse/#whats-next" class="hash-link" aria-label="Direct link to What's next" title="Direct link to What's next" translate="no">​</a></h2>
<p>We would rather you learn the gaps here than in production. Inbound authentication, rate
limiting and dynamic route reloading are planned; today, run Synapse behind your own API
gateway, ingress or service mesh. The gateway records metrics but emits no trace spans, and
chat legs get one attempt each, with no retries or circuit breakers. The
<a class="" href="https://synapse-gateway.io/en/latest/docs/reference/limitations-roadmap/">limitations and roadmap</a> page lists every one of
these, with links to the details.</p>
<p>The code is on <a href="https://github.com/sustentabilitas/synapse-gateway" target="_blank" rel="noopener noreferrer" class="">GitHub</a>, and
<a class="" href="https://synapse-gateway.io/en/latest/docs/contributing/">contributions</a> are welcome. Try the quickstart, and tell us what breaks.</p>]]></content>
        <author>
            <name>Raj Wilkhu</name>
            <uri>https://github.com/rajwilkhu</uri>
        </author>
        <category label="Release" term="Release"/>
        <category label="Vertex AI" term="Vertex AI"/>
        <category label="Routing" term="Routing"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[OpenTelemetry metrics in Synapse]]></title>
        <id>https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/</id>
        <link href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/"/>
        <updated>2026-09-28T09:30:00.000Z</updated>
        <summary type="html"><![CDATA[Synapse 0.5.38 records its synapse_* metrics with opentelemetry-rust, serves them in Prometheus format, can push them over OTLP, and ships a Grafana dashboard to chart them.]]></summary>
        <content type="html"><![CDATA[<p>As of release 0.5.38, the Synapse gateway records its metrics with
<a href="https://github.com/open-telemetry/opentelemetry-rust" target="_blank" rel="noopener noreferrer" class="">opentelemetry-rust</a> instead of the
<code>metrics</code> crate. Your Prometheus scrape keeps working with the same series names and labels,
apart from two deliberate changes covered below, and setting one environment variable now pushes the same metrics to an OpenTelemetry
collector.</p>
<p>This post covers why we moved, what the exposition looks like, what the metrics cover and how
to get the example Grafana dashboard running.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-opentelemetry-rust">Why opentelemetry-rust<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#why-opentelemetry-rust" class="hash-link" aria-label="Direct link to Why opentelemetry-rust" title="Direct link to Why opentelemetry-rust" translate="no">​</a></h2>
<p>The gateway was the odd one out in its own workspace. <code>synapse-proxy</code> and <code>synapse-mcp</code>
already recorded through OpenTelemetry instruments, while the gateway used the <code>metrics</code>
crate and its own exporter. Moving the gateway onto opentelemetry-rust 0.32 puts all three
crates on one pipeline: instruments created from an OpenTelemetry <code>Meter</code>, exported by the
OpenTelemetry Prometheus exporter, with OTLP available beside it. (The proxy keeps the
exporter's default naming; see the <a class="" href="https://synapse-gateway.io/en/latest/docs/reference/metrics-catalogue/">metrics catalogue</a>.)</p>
<p>That matters most when you embed the gateway. A Rust service that builds a <code>Gateway</code> in code
now passes it a <code>GatewayMetrics</code> made from a meter of its own <code>MeterProvider</code>, and the
gateway's instruments land wherever that provider exports. The default is a no-op, so an
embedded gateway records nothing until you opt in; a global <code>metrics</code> recorder in the host
process no longer receives gateway series. With the <code>server</code> feature,
<code>synapse::telemetry::install</code> builds exactly the exporters the binary uses. See
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/embedding-as-library/#what-the-builder-leaves-out">Embedding Synapse as a library</a>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="prometheus-by-default-otlp-when-you-ask">Prometheus by default, OTLP when you ask<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#prometheus-by-default-otlp-when-you-ask" class="hash-link" aria-label="Direct link to Prometheus by default, OTLP when you ask" title="Direct link to Prometheus by default, OTLP when you ask" translate="no">​</a></h2>
<p>The binary always serves Prometheus text on its metrics listener, <code>SYNAPSE_METRICS_ADDR</code>,
which defaults to <code>0.0.0.0:9090</code>, at <code>GET /metrics</code>. The API port does not serve metrics.</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">curl -s http://localhost:9090/metrics | grep '^synapse_'</span><br></div></code></pre></div></div>
<p>Set <code>OTEL_EXPORTER_OTLP_ENDPOINT</code> to a collector's base URL, such as
<code>http://otel-collector:4318</code>, and the gateway also pushes the same metrics over OTLP/HTTP to
<code>&lt;endpoint&gt;/v1/metrics</code>, every 60 seconds by default, tagged with <code>service.name</code> from
<code>OTEL_SERVICE_NAME</code> (default <code>synapse-gateway</code>). Both paths carry the same names and labels.
The startup line <code>synapse-gateway metrics listening</code> tells you whether OTLP is on. The
<a class="" href="https://synapse-gateway.io/en/latest/docs/configuration/environment-variables/#telemetry">Telemetry</a> variables are listed with
the rest of the configuration.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="an-exposition-that-looks-hand-written">An exposition that looks hand-written<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#an-exposition-that-looks-hand-written" class="hash-link" aria-label="Direct link to An exposition that looks hand-written" title="Direct link to An exposition that looks hand-written" translate="no">​</a></h2>
<p>The OpenTelemetry Prometheus exporter adds things by default that would have broken existing
dashboards: a second <code>_total</code> on counters, unit suffixes, <code>otel_scope_*</code> labels and a
<code>target_info</code> series. The gateway turns all of those off, so names come out exactly as the
catalogue lists them, such as <code>synapse_requests_total</code> and <code>synapse_request_duration_seconds</code>.</p>
<p>Two things did change, both deliberately:</p>
<ul>
<li class=""><strong>Durations are histograms.</strong> The <code>*_duration_seconds</code> metrics now expose <code>_bucket</code>, <code>_sum</code>
and <code>_count</code> series, with buckets from 5 ms to 120 seconds, instead of summaries. Queries on
<code>{quantile=...}</code> move to <code>histogram_quantile(...)</code> over <code>_bucket</code>, and you can now aggregate
latency across instances, which summaries never allowed.</li>
<li class=""><strong>Cardinality is capped.</strong> Each metric keeps at most 2,000 label combinations, the
OpenTelemetry SDK default. Beyond that, new combinations fold into one series labelled
<code>otel_metric_overflow="true"</code>.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-the-metrics-cover">What the metrics cover<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#what-the-metrics-cover" class="hash-link" aria-label="Direct link to What the metrics cover" title="Direct link to What the metrics cover" translate="no">​</a></h2>
<p>The <a class="" href="https://synapse-gateway.io/en/latest/docs/reference/metrics-catalogue/">metrics catalogue</a> lists every instrument with its
type and labels. In short, the gateway counts:</p>
<ul>
<li class=""><strong>Chat requests:</strong> requests, latency, and input and output tokens, labelled by <code>route</code>,
serving <code>model</code>, <code>lane</code> and <code>system</code>, the provider family in OpenLLMetry's <code>gen_ai.system</code>
vocabulary (<code>vertexai</code>, <code>openai</code>, <code>dashscope</code>, <code>oai_compat</code>).</li>
<li class=""><strong>Embeddings and passthrough:</strong> embedding requests and latency, and passthrough calls by
action and status for the Gemini and Jev passthrough endpoints, plus Gemini passthrough
model fallbacks.</li>
<li class=""><strong>Jev:</strong> routing decisions by tier and outcome, decision latency, and hybrid extraction
responses.</li>
<li class=""><strong>Guardrails:</strong> scans by outcome, matches by scanner and severity, and scan latency.</li>
<li class=""><strong>The cost ledger:</strong> rows dropped because the queue was full, and rows a sink failed to
write.</li>
</ul>
<p>Two things are deliberately missing. Tenants are not labels, because their values come from
clients and are unbounded; per-tenant usage and cost live in the
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/cost-ledger/">cost ledger</a>, which you can query directly. And there is no
metric for failed chat completions: the request metrics count requests that produced a
response, so measure error rates at the load balancer or mesh in front of the gateway.
<a class="" href="https://synapse-gateway.io/en/latest/docs/operating/metrics/#what-the-request-metrics-count">What the request metrics count</a>
spells out the edge cases, such as abandoned streams.</p>
<p>The gateway records metrics only. It creates no OpenTelemetry trace spans; logs go through
<code>tracing</code>, filtered with <code>RUST_LOG</code>. See <a class="" href="https://synapse-gateway.io/en/latest/docs/operating/logging/#tracing">Logging</a>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-dashboard-to-start-from">A dashboard to start from<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#a-dashboard-to-start-from" class="hash-link" aria-label="Direct link to A dashboard to start from" title="Direct link to A dashboard to start from" translate="no">​</a></h2>
<p>The repository ships an example Grafana dashboard,
<a href="https://synapse-gateway.io/en/latest/dashboards/synapse-gateway-dashboard.json" target="_blank" rel="noopener noreferrer" class="">synapse-gateway-dashboard.json</a>. It
charts traffic by route, lane and model; latency percentiles; token rates; cost-ledger health;
and embeddings. The latency panels are the ones the move to histograms fixed: they query
<code>histogram_quantile</code> over <code>_bucket</code> series, which the old summaries never produced, so they
used to stay empty.</p>
<p>Two things to check before you import it: every panel names its Prometheus data source by UID
<code>prometheus</code>, and every query filters on a <code>$job</code> variable that defaults to <code>synapse-gateway</code>.
The <a class="" href="https://synapse-gateway.io/en/latest/docs/operating/grafana-dashboard/">Grafana dashboard</a> page covers importing it by hand,
provisioning it with Docker Compose, and loading it on Kubernetes through the
kube-prometheus-stack sidecar. Its Resilience row stays empty, because chat legs have no
retries or circuit breakers to report on.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="where-to-start">Where to start<a href="https://synapse-gateway.io/en/latest/blog/opentelemetry-metrics/#where-to-start" class="hash-link" aria-label="Direct link to Where to start" title="Direct link to Where to start" translate="no">​</a></h2>
<p>Point Prometheus at port 9090, import the dashboard, and add the starter alerts from
<a class="" href="https://synapse-gateway.io/en/latest/docs/operating/metrics/#alerts-to-start-with">Metrics</a>: lost ledger rows, a failing ledger
sink and degraded Jev routing are problems no client error will tell you about.</p>]]></content>
        <author>
            <name>Raj Wilkhu</name>
            <uri>https://github.com/rajwilkhu</uri>
        </author>
        <category label="Observability" term="Observability"/>
    </entry>
    <entry>
        <title type="html"><![CDATA[Jev routing explained: the right model and effort for every request]]></title>
        <id>https://synapse-gateway.io/en/latest/blog/jev-routing-explained/</id>
        <link href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/"/>
        <updated>2026-09-28T09:00:00.000Z</updated>
        <summary type="html"><![CDATA[How a strategy = "jev" route in Synapse asks TypeSafe Jev how demanding each request is, picks a tier and reasoning effort, reports the decision and degrades safely.]]></summary>
        <content type="html"><![CDATA[<p>Most applications send every request on a route to the same model. A chat assistant that
answers "thanks!" and debugs a race condition in the same afternoon pays for its strongest
model on both, or saves money on both and disappoints on the hard one.</p>
<p>Synapse's Jev router lets the request decide. A <code>strategy = "jev"</code> route declares difficulty
tiers, and for each request the gateway asks TypeSafe Jev how demanding it is, then serves it
from the matching tier with a matching reasoning effort. This post walks through how that
decision is made, what it costs, and what happens when it can't be made.</p>
<!-- -->
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="tiers-describe-work-not-models">Tiers describe work, not models<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#tiers-describe-work-not-models" class="hash-link" aria-label="Direct link to Tiers describe work, not models" title="Direct link to Tiers describe work, not models" translate="no">​</a></h2>
<p>A Jev route replaces a route's single list of legs with two to ten tiers, ordered from
easiest to hardest. Each tier has a name, a description of the work it suits, a reasoning
effort and its own legs:</p>
<div class="language-toml codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-toml codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token table class-name">routes."auto"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">strategy</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"jev"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token table class-name">routes."auto".jev_router</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">default_tier</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"moderate"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token table class-name">routes."auto".tiers</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">name</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"trivial"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">description</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"Greetings, chit-chat, one-line lookups or rewrites"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">effort</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"none"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">legs</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token key property">provider</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"vertex"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">model</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"gemini-3.5-flash-lite"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">region</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"us"</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token table class-name">routes."auto".tiers</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">name</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"moderate"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">description</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"Everyday Q&amp;A, summarising, simple extraction or code edits"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">effort</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"low"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">legs</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token key property">provider</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"vertex"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">model</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"gemini-3.6-flash"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">region</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"global"</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token table class-name">routes."auto".tiers</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">name</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"hard"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">description</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"Multi-step analysis, maths or proofs, non-trivial code or debugging"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">effort</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"medium"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token key property">legs</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">[</span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"> </span><span class="token key property">provider</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"vertex"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">model</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"gemini-3.1-pro-preview"</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token key property">region</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"global"</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><span class="token punctuation" style="color:rgb(248, 248, 242)">]</span><br></div></code></pre></div></div>
<p>A <code>jev</code> route needs <code>TYPESAFE_API_KEY</code>; without it, the default strict provider validation
refuses to start the gateway.</p>
<p>The descriptions matter more than anything else in the file, because they are what Jev
scores the request against. "Multi-step analysis, maths or proofs" gives Jev something to
judge; "Gemini Pro" does not. The order matters too: it is what the difficulty score indexes,
and it drives fallback between tiers. The <a class="" href="https://synapse-gateway.io/en/latest/docs/get-started/tutorials/jev-tiers/">Jev tiers tutorial</a>
builds this route step by step.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="two-questions-per-request">Two questions per request<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#two-questions-per-request" class="hash-link" aria-label="Direct link to Two questions per request" title="Direct link to Two questions per request" translate="no">​</a></h2>
<p>Before serving a request, Synapse sends Jev a bounded summary of the conversation: the
latest user message, the system prompt and as much recent history as fits, about 24,000
characters in all, plus whether the request carries tools or media. Jev answers two typed
questions about it.</p>
<p><strong>Difficulty</strong> is a score against your tier descriptions, in order. Synapse rounds it to the
nearest tier. Jev also reports its confidence in that score; below <code>min_confidence</code>
(default 0.5), the gateway ignores the score and serves the request from <code>default_tier</code>.</p>
<p><strong>Needs reasoning</strong> is the probability that a good reply needs careful step-by-step reasoning,
such as maths, logic, planning or debugging. At or above <code>reasoning_threshold</code> (default 0.7),
the tier's effort goes up one step, even when the difficulty score was too uncertain to use.</p>
<p>The decision call is bounded by <code>timeout_ms</code>, 400 ms by default, so it adds up to that much
latency to every request on the route. That is the price of the decision, and the reason a
static route is still the better choice when you already know which model each use case needs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="effort-translated-per-provider">Effort, translated per provider<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#effort-translated-per-provider" class="hash-link" aria-label="Direct link to Effort, translated per provider" title="Direct link to Effort, translated per provider" translate="no">​</a></h2>
<p>Each tier's <code>effort</code> is one of <code>none</code>, <code>minimal</code>, <code>low</code>, <code>medium</code>, <code>high</code>, <code>xhigh</code> or <code>max</code>,
and the reasoning bump moves one step along that list, stopping at <code>max</code>. Synapse translates
the effort for each leg:</p>
<ul>
<li class="">On the standard lane, OpenAI-style providers get <code>reasoning_effort</code>, with <code>max</code> sent as
<code>xhigh</code>, and Vertex Gemini 3 models get a <code>thinkingLevel</code>.</li>
<li class="">On the native Vertex lane, Synapse sets <code>thinkingBudget</code> itself: 512 tokens for <code>minimal</code>,
1,024 for <code>low</code>, 4,096 for <code>medium</code>, 8,192 for <code>high</code>, 16,384 for <code>xhigh</code> and 24,576 for
<code>max</code>.</li>
<li class=""><code>none</code> sends nothing, so the model's own default applies. On some Gemini models that default
is dynamic thinking, so use <code>minimal</code> when you want thinking kept small.</li>
</ul>
<p>The <a class="" href="https://synapse-gateway.io/en/latest/docs/configuration/routes/#effort">Effort</a> table in the routes reference has every
mapping. Thinking tokens are billed as output, which is why easy tiers usually run at <code>none</code>
or <code>minimal</code> and only hard tiers get <code>medium</code> or more.</p>
<p>A client's own effort always wins on its lane. A <code>reasoning_effort</code> in the request body
replaces the tier's effort on the standard lane, and a <code>vertex.thinking_config</code> replaces the
tier's <code>thinkingBudget</code> on the native lane. Jev still picks the tier, and the response
reports <code>x-synapse-reasoning-effort: client</code>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="every-response-says-what-happened">Every response says what happened<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#every-response-says-what-happened" class="hash-link" aria-label="Direct link to Every response says what happened" title="Direct link to Every response says what happened" translate="no">​</a></h2>
<p>Every chat completion carries <code>x-synapse-routing</code>: <code>jev</code> on a Jev route, <code>static-override</code>
when the client sent <code>"routing_strategy": "static"</code> to skip the decision, and <code>static</code> on
ordinary routes. When they apply, <code>x-synapse-tier</code> names the tier that served,
<code>x-synapse-tier-decided</code> the tier the decision picked (Jev's choice, or <code>default_tier</code> when
routing is degraded) if a different one served, and
<code>x-synapse-reasoning-effort</code> the effort the serving leg ran with. For streams, the headers
describe the leg that produced the first chunk. The
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/jev-router/#response-headers">response headers</a> section lists every value.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-jev-problem-never-fails-the-request">A Jev problem never fails the request<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#a-jev-problem-never-fails-the-request" class="hash-link" aria-label="Direct link to A Jev problem never fails the request" title="Direct link to A Jev problem never fails the request" translate="no">​</a></h2>
<p>The router is designed to degrade, not to break. If Jev times out, errors, answers with low
confidence, or the gateway has no Jev client at all, the request goes to <code>default_tier</code> and
the response carries <code>x-synapse-routing-degraded</code> with the reason: <code>timeout</code>, <code>error</code>,
<code>low_confidence</code> or <code>jev_unavailable</code>. Pick a <code>default_tier</code> that handles most requests
acceptably, usually a middle one.</p>
<p>If the serving tier's legs fail, Synapse tries each harder tier, nearest first, then each
easier tier. Requests with native Vertex features only use <code>vertex</code> legs, so they go to the
nearest tier that has one. Only when every eligible leg fails does the request fail, exactly
as it would on a static route. <a class="" href="https://synapse-gateway.io/en/latest/docs/guides/jev-router/#failure-behaviour">Failure behaviour</a>
has the details.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="decisions-you-can-audit">Decisions you can audit<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#decisions-you-can-audit" class="hash-link" aria-label="Direct link to Decisions you can audit" title="Direct link to Decisions you can audit" translate="no">​</a></h2>
<p>Each decision Jev answers writes its own row to the <a class="" href="https://synapse-gateway.io/en/latest/docs/guides/cost-ledger/">cost ledger</a>,
with provider <code>typesafe</code>, lane <code>jev</code> and the same <code>request_id</code> as the chat request it routed,
priced as <code>typesafe:&lt;model&gt;</code> from your <code>pricing.toml</code>. You see the cost of deciding next to
the cost of answering.</p>
<p>Two metrics track the router: <code>synapse_routing_decisions_total</code>, labelled by route, tier and
outcome, and <code>synapse_routing_decision_duration_seconds</code>. A rising share of <code>timeout</code> or
<code>error</code> outcomes means requests are landing on <code>default_tier</code> without a real decision, so
raise <code>timeout_ms</code> or check Jev's availability. <a class="" href="https://synapse-gateway.io/en/latest/docs/operating/metrics/#queries">Metrics</a>
has a ready-made query for that ratio.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="try-it">Try it<a href="https://synapse-gateway.io/en/latest/blog/jev-routing-explained/#try-it" class="hash-link" aria-label="Direct link to Try it" title="Direct link to Try it" translate="no">​</a></h2>
<p>Start with the <a class="" href="https://synapse-gateway.io/en/latest/docs/get-started/tutorials/jev-tiers/">Jev tiers tutorial</a>, then keep the
<a class="" href="https://synapse-gateway.io/en/latest/docs/guides/jev-router/">Jev router guide</a> at hand while you tune tier descriptions and
thresholds against your own traffic.</p>]]></content>
        <author>
            <name>Raj Wilkhu</name>
            <uri>https://github.com/rajwilkhu</uri>
        </author>
        <category label="Routing" term="Routing"/>
    </entry>
</feed>