Skip to main content

SGLang Crossed 400,000 Production GPUs: Picking Your Self-Hosted PaaS's Default Model Server

10 min readDora NodaDora Noda
Share
On this page

For two years, picking the model server for a self-hosted GPU pool was a one-line decision: vLLM. That line expired in 2026. SGLang — the inference engine built around RadixAttention, a radix-tree KV-cache manager that automatically reuses shared prompt prefixes across requests — is now deployed across 400,000+ production GPUs at xAI, AMD, NVIDIA, Intel, LinkedIn, and Cursor. It is no longer the scrappy research alternative. It is a second production-grade default, and running a GPU node pool without re-examining the choice is now the lazy option.

Here is the verdict before the why: keep vLLM as the default for a general mix of tenant workloads, and switch a pool to SGLang when your traffic is prefix-heavy — multi-turn agent conversations, shared system prompts, structured tool-calling loops. On that workload shape SGLang is not marginally better; independent 2026 benchmarks show roughly 29% higher throughput, 4.5x faster time-to-first-token, and nearly double the KV-cache hit rate. On unique-prompt traffic the two engines tie, and vLLM's broader hardware support and deeper Kubernetes ecosystem keep it the safer universal pick. This post gives you the numbers, the mechanism behind them, and the measurement that settles the choice for your own pool.

The milestone that forced the re-pick​

SGLang comes from the LMSYS team (the group behind Chatbot Arena), and its headline feature is doing for the KV cache what a good CDN does for static assets: stop recomputing what you already have. When requests share a prompt prefix — a system prompt, few-shot examples, prior conversation turns — RadixAttention serves the shared part from cache and starts generation from the branch point. The engine now serves trillions of tokens daily across its 400,000-GPU footprint, with xAI running it behind Grok and NVIDIA, AMD, Intel, LinkedIn, and Cursor in production alongside.

That footprint matters for a reason beyond bragging rights. A model server becomes a platform default only when three things are true: it is fast, it runs everywhere you need it to, and someone else has already found the production bugs. The 400k number settles the third item. The rest of this post settles the first two — and they split by workload, which is why the answer is a decision matrix rather than a coronation.


PagedAttention vs RadixAttention in one paragraph each​

vLLM's PagedAttention manages the KV cache the way an OS manages virtual memory: fixed-size blocks allocated on demand, shared between requests when blocks match. Block-level sharing only triggers at multiples of the configured block size, so two requests sharing a 500-token system prompt but diverging mid-block still recompute part of what they share. It is a memory-efficiency innovation first — near-zero KV-cache waste, continuous batching of incoming requests — and prefix reuse (automatic prefix caching) arrived as a later addition on top.

SGLang's RadixAttention stores the KV cache in a radix tree keyed by token sequence, so sharing happens at any prefix length, not just block boundaries, and it is automatic: the scheduler matches each incoming request against the tree and skips recomputing whatever prefix already exists. The structural consequence is that prefix reuse is the engine's native mode rather than a bolt-on. SGLang pairs this with a zero-overhead CPU scheduler and a compressed finite-state machine for constrained decoding, which is why its structured-output story (guaranteed-valid JSON against a schema) is also a step ahead.

One paragraph of mechanism, one sentence of consequence: the more your requests share prefixes, the wider SGLang's lead; the more unique each prompt is, the closer the race. Every benchmark below is just that sentence wearing different numbers.


The numbers: head-to-head in 2026​

Start with the headline result. The PremAI 2026 benchmark on Llama 3.1 8B puts SGLang at roughly 16,200 tokens/s against vLLM's 12,500 — about 29% faster — with TTFT 23% faster (79ms vs 103ms) and inter-token latency 16% faster. That is prefix-heavy traffic, where RadixAttention does its best work.

The gap widens exactly where an agent platform lives. A multi-turn agent tool-calling benchmark (shared system prompt plus conversation history — the shape of a deploy-from-chat agent's own loop) measured SGLang at 85ms median TTFT versus vLLM's 380ms (4.5x faster), identical 4.5x advantage at P99, a 78.6% KV-cache hit rate against 41.2%, and 5,430 vs 3,120 output tokens/s. Other workload surveys agree on the hit-rate story: multi-turn chat lands at 75–90% on SGLang versus 10–20% on vLLM's prefix caching, and code analysis at 60–80% versus 5–15%. For structured generation, SGLang's constrained decoding runs JSON output roughly 3x faster than unconstrained generation with post-processing.

Now the honest counter-cases, because a benchmark section that only shows one engine winning is marketing. On unique-prompt traffic — no shared prefixes to reuse — the engines tie: Spheron's H100 test on Llama 3.3 70B FP8 at 50 concurrent requests measured SGLang at 1,920 tok/s against vLLM's 1,850, with TTFT within 20ms. A May 2026 comparison on Qwen2.5-7B on a single H100 had vLLM leading raw throughput outright (23,500 vs 16,800 tok/s), even as SGLang won TTFT by over 4x (1.8s vs 7.8s). One operator's standardized bench found them tied to within 0.3% on wall-clock for an identical 8-prompt workload, and a DGX Spark concurrency sweep documented a clean crossover: SGLang wins at 1–4 concurrent streams and prefill-heavy mixes, vLLM wins at 6+ streams. The sensitivity variable is always the same: prefix overlap and concurrency shape decide the winner, not the logo.

WorkloadvLLMSGLangWinner
Prefix-heavy throughput (Llama 3.1 8B)~12,500 tok/s~16,200 tok/sSGLang +29%
Multi-turn agent TTFT (P50 / P99)380ms / 1,250ms85ms / 280msSGLang ~4.5x
Multi-turn cache-hit rate~41%~79%SGLang ~2x
Unique-prompt throughput (Llama 3.3 70B)~1,850 tok/s~1,920 tok/sTie
Raw throughput, low-overlap (Qwen2.5-7B)~23,500 tok/s~16,800 tok/svLLM +40%
Constrained JSON decodingBaseline~3x vs post-processingSGLang

Read the table as a workload map, not a leaderboard. If your tenants are agents having long conversations with shared system prompts, the left column is leaving nearly half your TTFT and a third of your throughput on the floor. If they are firing unrelated one-shot prompts, the engines are interchangeable on speed and the decision moves to hardware and ecosystem.


Where vLLM still wins​

Speed ties go to the incumbent, and vLLM's moat is breadth. It has the broadest hardware support of any serving engine — NVIDIA, AMD, Intel, TPU, and Arm backends — and the AMD story specifically matured in 2026: a dedicated ROCm CI pipeline took vLLM's AMD pass rate from 37% in late 2025 to 93% by January 2026, and ROCm now delivers roughly 90–95% of H100-class throughput on current Instinct hardware for vLLM inference workloads. For a self-hosted fleet whose GPU mix is "whatever was available," that breadth is the whole game. SGLang covers NVIDIA, AMD, TPU, and Intel too, but its non-NVIDIA track record is thinner and younger.

The ecosystem gap is wider. vLLM sits at the center of the Kubernetes-native serving stack: llm-d (Red Hat and IBM's disaggregated prefill/decode serving, built on vLLM with P/D-aware routing via the Gateway API Inference Extension), Microsoft's KAITO autoscaling and multi-role inference, KServe integrations, and the vime reinforcement-learning framework extending into post-training. If your roadmap includes prefill/decode disaggregation — compute-bound prefill on one pool, memory-bound decode on another, KV transfer over fast interconnects — vLLM is where the production-hardened tooling lives today. Add the largest operator community, the most Stack Overflow answers, and the longest list of already-fixed production bugs, and vLLM remains the engine you pick when you cannot predict your tenants' workload mix. It is the safer universal default precisely because it is good at everything and best at generality.

Where SGLang wins — and its honest gotchas​

SGLang wins three workload shapes decisively. Prefix-heavy multi-turn traffic is the big one, per the numbers above — and that shape is exactly what AI-agent tenants generate: long shared system prompts, tool schemas repeated every turn, conversation history growing monotonically. Structured output is the second: guaranteed-valid JSON/XML at ~3x the speed of generate-then-validate matters the moment tenants use constrained decoding for tool calls rather than as a party trick. The third is model freshness: SGLang shipped day-one support for DeepSeek V3 and R1, a meaningful edge when tenants ask for a new open-weights release the week it drops.

The gotchas are real and worth pricing in. A production incident report describes SGLang's pure-Python router saturating a single core (127% CPU) at 150 concurrent requests on a small model, delivering 2.4x worse latency than vLLM until the operator switched to the Rust-based sglang-router — the fix exists, but the default bit someone. Separately, SGLang's cache_hit_rate Prometheus gauge has been observed stuck at 0.0 while prefix caching demonstrably works, so the metric you most want for capacity planning needs cross-checking against cached_tokens in server logs. Neither is disqualifying; both say "younger ops surface" the way vLLM's 2023 growing pains did. Budget a shakedown period, pin the Rust router from day one, and verify your dashboards against logs before trusting them.


The decision for your GPU pool​

For a Cluster-API-managed GPU node pool serving a general mix of tenant inference, the default stays vLLM: broadest hardware coverage, deepest Kubernetes serving ecosystem, and tied performance on unpredictable traffic. Switch a pool (or a deployment within it) to SGLang when any of these hold:

SignalThresholdAction
Prefix-cache hit rate on current engineAgent multi-turn consistently above ~60%, chat above ~30%Trial SGLang; expect the TTFT win to scale with hit rate
Tenant shapeMulti-turn agents with long shared system prompts / tool schemasSGLang is the native engine for this shape
Structured-output shareConstrained decoding on the hot path (tool calls, form filling)SGLang's ~3x constrained-decoding lead pays directly
HardwareHomogeneous NVIDIA poolSGLang's thinner non-NVIDIA record stops mattering
HardwareMixed-vendor or AMD-heavy poolStay on vLLM; breadth wins
RoadmapPrefill/decode disaggregation this yearStay on vLLM + llm-d until SGLang's story matures

The single most valuable measurement is the first row, and you can take it before committing to anything: vLLM exposes vllm:gpu_prefix_cache_hit_rate, and SGLang reports cached_tokens per request. Drive a load test that mimics your real prefix structure — shared system prompt plus multi-turn history — and scrape the hit rate. A synthetic all-unique-prompts load test will under-report SGLang's advantage and over-report vLLM's, which is exactly backwards from what your agent tenants will experience. Healthy targets from operator runbooks: chat above 0.3, agent multi-turn above 0.6. If your measured hit rate clears those bars and your pool is NVIDIA-homogeneous, the 400,000-GPU engine has earned the trial. If it doesn't, vLLM's generality remains the right default — and the existence of a credible second choice is what makes either answer defensible instead of habitual.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Run this on infrastructure you own

bex is the open-source, AI-native Render alternative — push a git repo and get a running HTTPS service on your own machines.

Get started with bex