OSS TanbouSign in with GitHub

Post-train foundation models with shared Python trainers and a CLI for SFT, preference optimization, and reward learning

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
19,445
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 4, 2026
On this page

Overview

TRL is a Python library for post-training language and multimodal foundation models. Built on the Transformers Trainer ecosystem, it provides dedicated trainers and CLI commands for supervised fine-tuning, GRPO, DPO, KTO, reward modeling, and related methods. The same workflow can grow from a single-GPU experiment to distributed training with DDP, DeepSpeed ZeRO, or FSDP.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Compose SFT, preference optimization, and reward learning with method-specific trainers

SFTTrainer, GRPOTrainer, DPOTrainer, KTOTrainer, RewardTrainer, and other classes accept models, datasets, and training configuration without requiring each algorithm to be implemented from scratch. They integrate with the Transformers training stack and PEFT adapters, making it easier to compare methods in a shared environment.

Sources: [2]

Choose Python trainers for control or CLI commands for a shorter path

Projects can use dedicated trainer classes when they need detailed control or start with commands such as trl sft and trl dpo for established workflows. Checkpoint resume behavior and vLLM-backed configurations continue to receive release maintenance, helping teams separate experiment definitions from execution infrastructure.

Sources: [2][3]

Best fit

Fits teams comparing several post-training methods in one Transformers environment

TRL is useful for research and engineering teams that want to evaluate instruction tuning, preference alignment, and reward-driven training with shared model assets and data pipelines. A configuration can be proven on a small model before it is expanded through DDP, DeepSpeed ZeRO, or FSDP.

Sources: [2]

Before adoption

Plan compute, data governance, and evaluation separately from the trainer API

Installation does not determine training quality or safety. GPU memory must match model size and sequence length, dataset rights and provenance need review, preference or reward bias needs measurement, and evaluation data should remain separate. Distributed and vLLM-backed runs also require compatible versions and tested checkpoint recovery. Python 3.10 or newer is required.

Sources: [2][4][3]

Official sources

  1. [1]huggingface/trl repository metadata(2026-10-04)
  2. [2]TRL v1.14.1 README(2026-10-04)
  3. [3]TRL v1.14.1 release(2026-10-04)
  4. [4]TRL v1.14.1 pyproject.toml(2026-10-04)
  5. [5]TRL v1.14.1 Apache-2.0 license(2026-10-04)
Supplemental curator note

When comparing post-training libraries, look beyond the number of supported algorithms and evaluate whether data contracts, evaluation, and checkpoint recovery fit one operating model. Reproducing the full workflow on a small model before scaling out is the practical adoption path.

Try it in 3 steps

  1. 1

    Create an isolated virtual environment

    Use Python 3.10 or newer and isolate the training dependencies.

    python -m venv .venv && source .venv/bin/activate
  2. 2

    Install the reviewed stable release

    Follow the README package installation while pinning the release used for this guide.

    python -m pip install "trl==1.14.1"
  3. 3

    Verify the CLI and trainer import

    Confirm the CLI and SFTTrainer before downloading a model or dataset and starting training.

    trl --help && python -c "from trl import SFTTrainer; print(SFTTrainer.__name__)"
Check the official README

Growth

Growth trends · Last 30 days

19,445 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
230
Open PRs
163

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data
Stars
19,445
Forks
3,042
Watchers
108
Open issues
114
Contributors
533
Owner type
Organization
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 4, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?