Skip to main content

Introducing Synapse: an LLM gateway that keeps native power

· 7 min read
Raj Wilkhu
Maintainer, Synapse

Synapse is an open-source LLM gateway written in Rust. Your clients send standard OpenAI POST /v1/chat/completions requests, and Synapse routes each one through a config-driven fallback chain of providers, records what it cost, and hands back an OpenAI-shaped response.

What makes it different is what it refuses to throw away. Vertex AI's context caching, Cloud Storage media and strict response schemas survive the trip, and a routing model can decide how capable a model, and how much reasoning, each request deserves. This post explains why we built it, how it is put together and where it is going.

OpenTelemetry metrics in Synapse

· 5 min read
Raj Wilkhu
Maintainer, Synapse

As of release 0.5.38, the Synapse gateway records its metrics with opentelemetry-rust instead of the metrics crate. Your Prometheus scrape keeps working with the same series names and labels, apart from two deliberate changes covered below, and setting one environment variable now pushes the same metrics to an OpenTelemetry collector.

This post covers why we moved, what the exposition looks like, what the metrics cover and how to get the example Grafana dashboard running.

Jev routing explained: the right model and effort for every request

· 6 min read
Raj Wilkhu
Maintainer, Synapse

Most applications send every request on a route to the same model. A chat assistant that answers "thanks!" and debugs a race condition in the same afternoon pays for its strongest model on both, or saves money on both and disappoints on the hard one.

Synapse's Jev router lets the request decide. A strategy = "jev" route declares difficulty tiers, and for each request the gateway asks TypeSafe Jev how demanding it is, then serves it from the matching tier with a matching reasoning effort. This post walks through how that decision is made, what it costs, and what happens when it can't be made.