OSS Tanbou

quantize PyTorch models to 8-bit and 4-bit formats to reduce memory use for LLM inference and QLoRA training

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
8,510
Primary language
Python
License
MIT
Repository last updated
Sep 7, 2026
On this page

Overview

bitsandbytes is a k-bit quantization library for PyTorch. It provides LLM.int8 inference, 4-bit quantization for QLoRA workflows, and 8-bit optimizers to reduce memory requirements for model inference and training.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Reduce inference memory with LLM.int8

Its 8-bit method quantizes most features while handling outliers separately with higher-precision matrix multiplication, reducing the memory footprint of large-model inference.

Sources: [1]

Use 4-bit quantization for QLoRA-style fine-tuning

4-bit model weights combined with trainable LoRA parameters make large-model fine-tuning feasible under tighter memory budgets.

Sources: [1]

Integrate 8-bit optimizers and quantized linear layers into PyTorch

The library exposes optimizers plus layers such as Linear8bitLt and Linear4bit for integrating quantized operations into PyTorch workflows.

Sources: [1]

Best fit

Fits inference and training when VRAM or RAM limits are the primary constraint

It is useful on GPUs, workstations, and supported CPU backends when full-precision models exceed available memory.

Sources: [1]

Before adoption

Require Python 3.10+ and PyTorch 2.4+

The 0.50.2 README lists Python 3.10 or newer and PyTorch 2.4 or newer as minimum requirements across platforms.

Sources: [1]

Accelerator support and performance vary by backend

CPU, NVIDIA, AMD, Intel, Gaudi, and Apple Silicon backends have different feature and optimization levels, so the stable-release support matrix should be checked for target hardware.

Sources: [1][2]

Measure latency, quality, and kernel compatibility in addition to memory savings

Quantization effects vary by model and operation. Pin related Transformers or PEFT versions and regression-test representative data on actual hardware.

Sources: [1][2][3]

Official sources

  1. [1]bitsandbytes 0.50.2 README(2026-10-03)
  2. [2]bitsandbytes 0.50.2 release(2026-10-03)
  3. [3]bitsandbytes MIT license(2026-10-03)
Supplemental curator note

bitsandbytes is useful when model size exceeds comfortable VRAM or RAM limits, but speed, accuracy, and supported operations vary by quantization mode and hardware backend. Benchmark memory, latency, and model quality on the actual target hardware.

Try it in 3 steps

  1. 1

    Install bitsandbytes 0.50.2 in an isolated environment

    Pin the stable release in a Python 3.10+ environment. PyTorch 2.4+ is required.

    python3 -m venv .venv && . .venv/bin/activate && python -m pip install bitsandbytes==0.50.2
  2. 2

    Verify the installed version

    Read local package metadata without downloading a model.

    python -c "import importlib.metadata; print(importlib.metadata.version('bitsandbytes'))"
  3. 3

    Import the quantized layer APIs

    Confirm the 8-bit and 4-bit layer APIs load without a model or external service. Check the target accelerator support matrix before real workloads.

    python -c "import bitsandbytes as bnb; print(bnb.nn.Linear8bitLt, bnb.nn.Linear4bit)"
Check the official README

Growth

Growth trends · Last 30 days

8,510 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
0
Open PRs
46

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • llm
  • machine-learning
  • pytorch
  • qlora
  • quantization
Stars
8,510
Forks
937
Watchers
53
Open issues
48
Contributors
129
Owner type
Organization
Primary language
Python
License
MIT
Repository last updated
Sep 7, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?