~/LLM SERVING/sglang-v0-5-21-released-with-expanded-model-support-and-performance-upgrades

SGLang v0.5.21 Released with Expanded Model Support and Performance Upgrades

SGLang v0.5.21 has been released, incorporating 779 pull requests from 227 contributors. The release adds serving support for a wide array of new LLMs, VLMs, and diffusion models—including DeepSeek-V4.1 Flash, MiMo-V2.6, and DiffusionGemma—while moving its default prefix cache engine to a Rust core. This update strengthens SGLang's role as a leading open-source inference engine across unified language, multimodal, and image generation workloads. Improvements like dynamic prefill/decode instance switching without server restarts allow infrastructure engineers to maximize GPU hardware utilization and reduce latency in production. Key performance gains include a 22% faster time-to-first-token on long prompts for DeepSeek-V4.1 and a 20.6% prefill throughput boost for Kimi K3 in disaggregated serving. The update also introduces new `/v1/decisions` and `/v1/score` endpoints for low-latency classification and single-request multi-candidate scoring.

## BACKGROUND

SGLang is a high-performance open-source framework designed for serving large language models and multimodal models across setups ranging from single GPUs to large distributed clusters. In LLM serving, inference is divided into a compute-heavy prefill phase that processes the prompt and a memory-bandwidth-bound decode phase that generates tokens sequentially.

## REFERENCES

## KEYWORDS

#LLM Serving#Inference#SGLang#Open Source#AI Infrastructure

$ subscribe --daily

SGLang v0.5.21 Released with Expanded Model Support and Performance Upgrades | Daily News