OSS TanbouSign in with GitHub

Serve large language and multimodal models with low-latency inference

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
36,771
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 4, 2026
On this page

Overview

SGLang is an inference platform for serving large language and multimodal models over HTTP. Prefix caching, continuous batching, speculative decoding, and several parallelism strategies let operators tune latency and throughput from one GPU to distributed clusters.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Unify model execution behind compatible serving APIs

It lists language models such as Llama, Qwen, DeepSeek, Kimi, GLM, GPT, Gemma, and Mistral, plus embedding, reward, image, and video-generation models. RadixAttention prefix caching, paged attention, quantization, and multi-LoRA batching sit behind OpenAI-compatible and native generation APIs.

Sources: [1][2]

Best fit

Centralize inference policies for several applications

It fits systems that share models across applications and need batching, structured output, or multi-GPU partitioning in the serving layer. Start with a small model on one GPU, measure real prompt lengths and concurrency, and introduce tensor, pipeline, expert, or data parallelism only when the workload requires it.

Sources: [1]

Before adoption

Validate hardware, model compatibility, and exposure controls

v0.5.21 requires Python 3.10 or newer, and the standard NVIDIA installation targets CUDA 13. The quickstart recommends Linux and an NVIDIA sm80-or-newer GPU; AMD GPU, CPU, TPU, and NPU paths have separate instructions and support boundaries. Hugging Face compatibility does not guarantee every model or remote extension. The local example accepts requests without an authentication header, so external deployments need authentication and access controls. GitHub reports Apache-2.0.

Sources: [3][2][4][5][6]

Official sources

  1. [1]SGLang v0.5.21 README(2026-10-04)
  2. [2]SGLang v0.5.21 quickstart(2026-10-04)
  3. [3]SGLang v0.5.21 installation guide(2026-10-04)
  4. [4]SGLang v0.5.21 Python package metadata(2026-10-04)
  5. [5]sgl-project/sglang GitHub repository metadata(2026-10-04)
  6. [6]SGLang v0.5.21 LICENSE(2026-10-04)
Supplemental curator note

Fit the target model, quantization, and context length into real GPU memory, then measure latency and throughput under concurrency. The local example has no authentication; add authentication, transport security, and usage controls before exposure.

Try it in 3 steps

  1. 1

    Create an isolated Python environment

    On Linux with Python 3.10 or newer, CUDA 13, and an NVIDIA sm80-or-newer GPU, activate a disposable virtual environment. Continue in the same terminal and working directory.

    python3 -m venv sglang-demo-env && . sglang-demo-env/bin/activate
  2. 2

    Install SGLang 0.5.21

    Install the reviewed release. The next step downloads a small public Qwen model, so it needs network access and sufficient free storage.

    python -m pip install "sglang==0.5.21"
  3. 3

    Send a request to local inference

    Bind the public model to 127.0.0.1 and poll readiness up to 60 times at one-second intervals, with one-second connection and request limits. If it is not ready in roughly 120 seconds, print the log, fail, and do not call the chat API. After readiness, cap the chat request at 30 seconds and require non-empty content. The subshell stops the server on success or failure without leaving its trap in the caller. This does not validate production authentication.

    ( sglang serve --model-path qwen/qwen2.5-0.5b-instruct --host 127.0.0.1 --port 30000 > sglang.log 2>&1 & server_pid=$!; trap 'kill "$server_pid" 2>/dev/null || true; wait "$server_pid" 2>/dev/null || true' EXIT; ready=0; for attempt in 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60; do if curl -fsS --connect-timeout 1 --max-time 1 http://127.0.0.1:30000/health >/dev/null 2>&1; then ready=1; break; fi; kill -0 "$server_pid" 2>/dev/null || { cat sglang.log; exit 1; }; sleep 1; done; if test "$ready" != 1; then cat sglang.log; exit 1; fi; curl -fsS --connect-timeout 2 --max-time 30 http://127.0.0.1:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"qwen/qwen2.5-0.5b-instruct","messages":[{"role":"user","content":"Reply with OK"}],"temperature":0,"max_tokens":8}' > response.json && python -c 'import json; data=json.load(open("response.json")); assert data["choices"][0]["message"]["content"].strip()' )
Check the official README

Growth

Growth trends · Last 30 days

36,771 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
1,606
Open PRs
4,614

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • cuda
  • inference
  • llama
  • llm
  • moe
  • transformer
  • vlm
  • deepseek
  • blackwell
  • gpt-oss
  • diffusion
  • attention
Stars
36,771
Forks
9,311
Watchers
184
Open issues
932
Contributors
449
Owner type
Organization
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 4, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?