OSS TanbouSign in with GitHub

load, transform, and stream ML datasets from the Hugging Face Hub or local data with an Apache Arrow-backed Python library

About these scores

OSS scale score is an unbounded metric that log-compresses and weights Stars, Watchers, Forks, and Contributors. Discovery score is the current OSS scale score minus the score at discovery. Update pace is commits in the last 30 days, growth momentum is the OSS scale score difference within the recent observation window, and OSS health is a 0–100 rating based on available recency, Community Health, and release data.

Stars
22,029
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 5, 2026
On this page

Overview

Hugging Face Datasets is a Python library for loading, preprocessing, and sharing AI/ML datasets. It provides one API for Hub datasets and local CSV, JSON, Parquet, and other formats while using Apache Arrow, caching, streaming, and multiprocessing for training and evaluation data pipelines.

Features and best fit

Based on official documentation; not hands-on tested · Content checked:

Key features

Load Hub datasets and local files through load_dataset

The library handles datasets from the Hugging Face Hub as well as CSV, JSON/JSONL, Parquet, Arrow, text, image, audio, video, PDF, and other local or remote sources.

Sources: [1]

Use Apache Arrow, caching, and multiprocessing for efficient preprocessing

Dataset objects use Arrow-backed storage and memory mapping, cache transformed results, and can parallelize preprocessing with multiprocessing.

Sources: [1]

Stream large datasets without downloading them in full

With streaming=True, IterableDataset can iterate data incrementally without storing the complete dataset locally, while interoperability covers NumPy, Pandas, Polars, PyTorch, TensorFlow, JAX, and other ecosystems.

Sources: [1]

Best fit

Fits reproducible Python pipelines for training and evaluation datasets

It is useful when dataset sources, splits, transforms, and output formats need to be managed as code and reused consistently across model training and evaluation.

Sources: [1]

Before adoption

Version 5.0.1 requires Python 3.10+ and PyArrow 21+

Package metadata for 5.0.1 requires Python >=3.10.0 and pyarrow>=21.0.0, with additional version constraints for fsspec, huggingface-hub, and other dependencies.

Sources: [3][4]

Version 5.0.1 contains archive-extraction and metadata-path security fixes

The release fixes a symlink-following arbitrary-file-write issue during archive extraction and a path traversal issue involving metadata file_name values in folder-based builders. Prefer 5.0.1 or later when processing untrusted archives or dataset metadata.

Sources: [2]

Official sources

  1. [1]Datasets 5.0.1 README(2026-10-04)
  2. [2]Datasets 5.0.1 release(2026-10-04)
  3. [3]Datasets 5.0.1 package metadata(2026-10-04)
  4. [4]Datasets 5.0.1 version metadata(2026-10-04)
  5. [5]Datasets Apache 2.0 license(2026-10-04)
Supplemental curator note

Hub datasets have their own content, licenses, and revisions. Do not treat the library's Apache-2.0 license as the dataset license, and pin dataset revisions when reproducibility matters.

Try it in 3 steps

  1. 1

    Create a virtual environment for Datasets 5.0.1

    Isolate the Python environment and pin the security-fixed 5.0.1 release.

    python3 -m venv datasets-demo && . datasets-demo/bin/activate && python -m pip install --upgrade pip && python -m pip install 'datasets==5.0.1'
  2. 2

    Create an in-memory Dataset without network access

    Create an Arrow-backed Dataset from a local Python object without contacting the Hub.

    . datasets-demo/bin/activate && python -c 'from datasets import Dataset; d=Dataset.from_dict({"text":["alpha","beta","gamma"]}); print(d)'
  3. 3

    Try a local transformation with map

    Inspect the transformation API and result without downloading data or training a model.

    . datasets-demo/bin/activate && python -c 'from datasets import Dataset; d=Dataset.from_dict({"text":["alpha","beta","gamma"]}); d=d.map(lambda x:{"length":len(x["text"])}); print(d.to_dict())'
Check the official README

Growth

Growth trends · Last 30 days

22,029 Stars

Trend data is still being collected.

Development activity

Last 90 days · weekly

Commits (last 30 days)
38
Open PRs
508

Development activity is still being collected.

Built with

Categories and tags

GitHub data

GitHub dataView detailed GitHub data

GitHub Topics

  • nlp
  • datasets
  • pytorch
  • tensorflow
  • pandas
  • numpy
  • natural-language-processing
  • computer-vision
  • machine-learning
  • deep-learning
  • speech
  • ai
Stars
22,029
Forks
3,493
Watchers
283
Open issues
943
Contributors
444
Owner type
Organization
Primary language
Python
License
Apache-2.0
Repository last updated
Oct 5, 2026
Write a related article

Share a guide or use case for this OSS in Markdown. Articles are published after administrator approval.

Report incorrect information

Tell us if any listing information is incorrect or outdated.

After reading this page, do you know what to do next?