On this page
Overview
TensorRT-LLM is a library for optimizing large-language-model inference on supported NVIDIA GPUs and exposing it through Python or C++ runtimes and an OpenAI-compatible server. Quantization, continuous batching, and multi-GPU execution help teams tune latency and throughput.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Optimize models and expose an inference server
Teams can build TensorRT engines for supported models or use trtllm-serve to load a Hugging Face model and expose an OpenAI-compatible API. The runtime batches requests and manages the KV cache to coordinate GPU execution.
Best fit
Before adoption
Provision a supported GPU, CUDA stack, driver, and memory
The v1.2.1 documentation targets Linux x86-64 or aarch64 with NVIDIA Blackwell, Grace Hopper, Hopper, Ada Lovelace, or Ampere GPUs. Its release image is based on CUDA 13.1 and requires a compatible driver. Choose a GPU-memory budget that fits model weights, KV cache, and concurrent requests.
Review the container, Python environment, and model license
The official release container is the recommended path. Manual installation requires compatible Python 3.12, CUDA Toolkit 13.1, and PyTorch packages. TensorRT-LLM uses Apache License 2.0, while bundled components and downloaded models can carry separate terms that must be reviewed.
Official sources
- [1]NVIDIA/TensorRT-LLM README at v1.2.1(2026-10-04)
- [2]TensorRT-LLM quick start at v1.2.1(2026-10-04)
- [3]TensorRT-LLM Linux installation guide at v1.2.1(2026-10-04)
- [4]TensorRT-LLM support matrix at v1.2.1(2026-10-04)
- [5]NVIDIA/TensorRT-LLM license file at v1.2.1(2026-10-04)
- [6]TensorRT-LLM v1.2.1 release(2026-10-04)
Supplemental curator note
Compare serving stacks with a fixed model and GPU generation, measuring latency, concurrency, and GPU memory under the same requests. Start with the official container and a small public model to verify the path.
Try it in 3 steps
- 1
Pull the pinned release container
On Linux with a supported NVIDIA GPU, NVIDIA Container Toolkit, sufficient GPU memory, and storage, fetch the official v1.2.1 image.
docker pull nvcr.io/nvidia/tensorrt-llm/release:1.2.1 - 2
Create the request
Create a working directory and save one request for public model TinyLlama-1.1B-Chat-v1.0. Review the model license and terms separately.
mkdir trtllm-tinyllama-demo && cd trtllm-tinyllama-demo && printf '%s\n' '{' ' "model": "TinyLlama/TinyLlama-1.1B-Chat-v1.0",' ' "messages": [{"role": "user", "content": "Reply with hello."}],' ' "max_tokens": 16' '}' > request.json - 3
Start, wait, infer, and clean up in one subshell
Start without a fixed container name and remove only the captured container ID. Poll /health up to 60 times with a 2-second request bound and 5-second interval (about 7 minutes maximum), then make one inference request bounded to 120 seconds and save tinyllama-response.json. No inference runs after a readiness timeout.
( cid=''; cleanup() { if test -n "$cid"; then docker rm -f "$cid" >/dev/null 2>&1 || true; fi; }; trap cleanup EXIT INT TERM; cid=$(docker run -d --rm --ipc host --gpus all --ulimit memlock=-1 --ulimit stack=67108864 -p 127.0.0.1::8000 nvcr.io/nvidia/tensorrt-llm/release:1.2.1 trtllm-serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0") || exit 1; port=$(docker port "$cid" 8000/tcp | sed -n 's/.*://p') || exit 1; test -n "$port" || exit 1; ready=0; for attempt in 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60; do if curl -fsS --connect-timeout 1 --max-time 2 "http://127.0.0.1:$port/health" -o health.txt; then ready=1; break; fi; sleep 5; done; test "$ready" = 1 || exit 1; curl -fsS --connect-timeout 2 --max-time 120 -X POST "http://127.0.0.1:$port/v1/chat/completions" -H 'Content-Type: application/json' --data-binary @request.json -o tinyllama-response.json || exit 1; test -s tinyllama-response.json )
Growth
Growth trends · Last 30 days
14,763 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 688
- Open PRs
- 959
- Issues opened
- 206
- Issues closed
- 198
- PRs opened
- 3,557
- PRs merged
- 2,287
Issues
206 / 198
Pull requests
3,557 / 2,287
Maintenance
- Median first response
- 5.3 days
- Issue response rate
- 19.5% (17/87)
Based on up to the 100 newest issues opened by external users in the last 90 days. A first comment from an OWNER, MEMBER, or COLLABORATOR counts as a response; issues whose full comment history cannot be checked are excluded. The median and response rate update weekly.
Built with
Categories and tags
Categories
GitHub data
GitHub dataView detailed GitHub data
GitHub Topics
- blackwell
- cuda
- moe
- pytorch
- llm-serving
- Stars
- 14,763
- Forks
- 2,796
- Watchers
- 122
- Open issues
- 615
- Contributors
- 423
- Owner type
- Organization
- Primary language
- Python
- License
- Not determined
- Repository last updated
- Oct 5, 2026
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- Crawl4AI84,764 Stars
1 shared tag(s) · 2 shared category(s) · Same language
a Python crawler that turns browser-rendered web pages into LLM-ready Markdown and JSON
Python - AutoGen61,263 Stars
1 shared tag(s) · 2 shared category(s) · Same language
Build multi-stage AI workflows by coordinating specialized agents through conversations and events
Python - LiteLLM60,131 Stars
1 shared tag(s) · 2 shared category(s) · Same language
unify 100+ LLM providers behind OpenAI-compatible interfaces and centralize routing, cost controls, guardrails, and logging in an AI gateway
Python - CrewAI59,348 Stars
1 shared tag(s) · 2 shared category(s) · Same language
combine role-based Crews with event-driven Flows to build multi-agent automation
Python - Aider49,381 Stars
1 shared tag(s) · 2 shared category(s) · Same language
edit code from the terminal with repository-wide context
Python - smolagents29,676 Stars
1 shared tag(s) · 2 shared category(s) · Same language
Build code-writing AI agents with a compact Python library and pluggable models, tools, and execution backends
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.