OSS TanbouSign in with GitHub

Build and serve large-language-model inference on NVIDIA GPUs

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
14,763
Primary language
Python
License
Not determined
Repository last updated
Oct 5, 2026
On this page

Overview

TensorRT-LLM is a library for optimizing large-language-model inference on supported NVIDIA GPUs and exposing it through Python or C++ runtimes and an OpenAI-compatible server. Quantization, continuous batching, and multi-GPU execution help teams tune latency and throughput.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Optimize models and expose an inference server

Teams can build TensorRT engines for supported models or use trtllm-serve to load a Hugging Face model and expose an OpenAI-compatible API. The runtime batches requests and manages the KV cache to coordinate GPU execution.

Sources: [1][2]

Tune precision and scale across GPUs

Supported combinations include FP8, FP4, INT8, and INT4 precision, in-flight batching, and multi-GPU or multi-node execution. Availability varies by model, precision, and GPU generation.

Sources: [1][4]

Best fit

Use it to tune generative-AI serving on NVIDIA GPUs

TensorRT-LLM fits teams that need to tune latency, concurrency, and memory use from public-model experiments through dedicated NVIDIA GPU inference servers. It is not a CPU inference runtime or a path for non-NVIDIA GPUs.

Sources: [2][4]

Before adoption

Provision a supported GPU, CUDA stack, driver, and memory

The v1.2.1 documentation targets Linux x86-64 or aarch64 with NVIDIA Blackwell, Grace Hopper, Hopper, Ada Lovelace, or Ampere GPUs. Its release image is based on CUDA 13.1 and requires a compatible driver. Choose a GPU-memory budget that fits model weights, KV cache, and concurrent requests.

Sources: [3][4]

Review the container, Python environment, and model license

The official release container is the recommended path. Manual installation requires compatible Python 3.12, CUDA Toolkit 13.1, and PyTorch packages. TensorRT-LLM uses Apache License 2.0, while bundled components and downloaded models can carry separate terms that must be reviewed.

Sources: [3][5][2]

Official sources

  1. [1]NVIDIA/TensorRT-LLM README at v1.2.1(2026-10-04)
  2. [2]TensorRT-LLM quick start at v1.2.1(2026-10-04)
  3. [3]TensorRT-LLM Linux installation guide at v1.2.1(2026-10-04)
  4. [4]TensorRT-LLM support matrix at v1.2.1(2026-10-04)
  5. [5]NVIDIA/TensorRT-LLM license file at v1.2.1(2026-10-04)
  6. [6]TensorRT-LLM v1.2.1 release(2026-10-04)
Supplemental curator note

Compare serving stacks with a fixed model and GPU generation, measuring latency, concurrency, and GPU memory under the same requests. Start with the official container and a small public model to verify the path.

Try it in 3 steps

  1. 1

    Pull the pinned release container

    On Linux with a supported NVIDIA GPU, NVIDIA Container Toolkit, sufficient GPU memory, and storage, fetch the official v1.2.1 image.

    docker pull nvcr.io/nvidia/tensorrt-llm/release:1.2.1
  2. 2

    Create the request

    Create a working directory and save one request for public model TinyLlama-1.1B-Chat-v1.0. Review the model license and terms separately.

    mkdir trtllm-tinyllama-demo && cd trtllm-tinyllama-demo && printf '%s\n' '{' ' "model": "TinyLlama/TinyLlama-1.1B-Chat-v1.0",' ' "messages": [{"role": "user", "content": "Reply with hello."}],' ' "max_tokens": 16' '}' > request.json
  3. 3

    Start, wait, infer, and clean up in one subshell

    Start without a fixed container name and remove only the captured container ID. Poll /health up to 60 times with a 2-second request bound and 5-second interval (about 7 minutes maximum), then make one inference request bounded to 120 seconds and save tinyllama-response.json. No inference runs after a readiness timeout.

    ( cid=''; cleanup() { if test -n "$cid"; then docker rm -f "$cid" >/dev/null 2>&1 || true; fi; }; trap cleanup EXIT INT TERM; cid=$(docker run -d --rm --ipc host --gpus all --ulimit memlock=-1 --ulimit stack=67108864 -p 127.0.0.1::8000 nvcr.io/nvidia/tensorrt-llm/release:1.2.1 trtllm-serve "TinyLlama/TinyLlama-1.1B-Chat-v1.0") || exit 1; port=$(docker port "$cid" 8000/tcp | sed -n 's/.*://p') || exit 1; test -n "$port" || exit 1; ready=0; for attempt in 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60; do if curl -fsS --connect-timeout 1 --max-time 2 "http://127.0.0.1:$port/health" -o health.txt; then ready=1; break; fi; sleep 5; done; test "$ready" = 1 || exit 1; curl -fsS --connect-timeout 2 --max-time 120 -X POST "http://127.0.0.1:$port/v1/chat/completions" -H 'Content-Type: application/json' --data-binary @request.json -o tinyllama-response.json || exit 1; test -s tinyllama-response.json )
Check the official README

Growth

Growth trends · Last 30 days

14,763 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
688
Open PRs
959
Issues opened
206
Issues closed
198
PRs opened
3,557
PRs merged
2,287

Issues

206 / 198

Jul 8Oct 5
Issues openedIssues closed

Pull requests

3,557 / 2,287

Jul 8Oct 5
PRs openedPRs merged

Maintenance

Median first response
5.3 days
Issue response rate
19.5% (17/87)

Based on up to the 100 newest issues opened by external users in the last 90 days. A first comment from an OWNER, MEMBER, or COLLABORATOR counts as a response; issues whose full comment history cannot be checked are excluded. The median and response rate update weekly.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • blackwell
  • cuda
  • moe
  • pytorch
  • llm-serving
Stars
14,763
Forks
2,796
Watchers
122
Open issues
615
Contributors
423
Owner type
Organization
Primary language
Python
License
Not determined
Repository last updated
Oct 5, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?