On this page
Overview
Hugging Face Datasets is a Python library for loading, preprocessing, and sharing AI/ML datasets. It provides one API for Hub datasets and local CSV, JSON, Parquet, and other formats while using Apache Arrow, caching, streaming, and multiprocessing for training and evaluation data pipelines.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Load Hub datasets and local files through load_dataset
The library handles datasets from the Hugging Face Hub as well as CSV, JSON/JSONL, Parquet, Arrow, text, image, audio, video, PDF, and other local or remote sources.
Sources: [1]
Use Apache Arrow, caching, and multiprocessing for efficient preprocessing
Dataset objects use Arrow-backed storage and memory mapping, cache transformed results, and can parallelize preprocessing with multiprocessing.
Sources: [1]
Stream large datasets without downloading them in full
With streaming=True, IterableDataset can iterate data incrementally without storing the complete dataset locally, while interoperability covers NumPy, Pandas, Polars, PyTorch, TensorFlow, JAX, and other ecosystems.
Sources: [1]
Best fit
Fits reproducible Python pipelines for training and evaluation datasets
It is useful when dataset sources, splits, transforms, and output formats need to be managed as code and reused consistently across model training and evaluation.
Sources: [1]
Before adoption
Version 5.0.1 requires Python 3.10+ and PyArrow 21+
Package metadata for 5.0.1 requires Python >=3.10.0 and pyarrow>=21.0.0, with additional version constraints for fsspec, huggingface-hub, and other dependencies.
Version 5.0.1 contains archive-extraction and metadata-path security fixes
The release fixes a symlink-following arbitrary-file-write issue during archive extraction and a path traversal issue involving metadata file_name values in folder-based builders. Prefer 5.0.1 or later when processing untrusted archives or dataset metadata.
Sources: [2]
Official sources
- [1]Datasets 5.0.1 README(2026-10-04)
- [2]Datasets 5.0.1 release(2026-10-04)
- [3]Datasets 5.0.1 package metadata(2026-10-04)
- [4]Datasets 5.0.1 version metadata(2026-10-04)
- [5]Datasets Apache 2.0 license(2026-10-04)
Supplemental curator note
Hub datasets have their own content, licenses, and revisions. Do not treat the library's Apache-2.0 license as the dataset license, and pin dataset revisions when reproducibility matters.
Try it in 3 steps
- 1
Create a virtual environment for Datasets 5.0.1
Isolate the Python environment and pin the security-fixed 5.0.1 release.
python3 -m venv datasets-demo && . datasets-demo/bin/activate && python -m pip install --upgrade pip && python -m pip install 'datasets==5.0.1' - 2
Create an in-memory Dataset without network access
Create an Arrow-backed Dataset from a local Python object without contacting the Hub.
. datasets-demo/bin/activate && python -c 'from datasets import Dataset; d=Dataset.from_dict({"text":["alpha","beta","gamma"]}); print(d)' - 3
Try a local transformation with map
Inspect the transformation API and result without downloading data or training a model.
. datasets-demo/bin/activate && python -c 'from datasets import Dataset; d=Dataset.from_dict({"text":["alpha","beta","gamma"]}); d=d.map(lambda x:{"length":len(x["text"])}); print(d.to_dict())'
Growth
Growth trends · Last 30 days
22,029 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 38
- Open PRs
- 508
Development activity is still being collected.
Built with
Categories and tags
Categories
GitHub data
GitHub dataView detailed GitHub data
GitHub Topics
- nlp
- datasets
- pytorch
- tensorflow
- pandas
- numpy
- natural-language-processing
- computer-vision
- machine-learning
- deep-learning
- speech
- ai
- Stars
- 22,029
- Forks
- 3,493
- Watchers
- 283
- Open issues
- 943
- Contributors
- 444
- Owner type
- Organization
- Primary language
- Python
- License
- Apache-2.0
- Repository last updated
- Oct 5, 2026
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- AutoGen61,263 Stars
2 shared tag(s) · 1 shared category(s) · Same language
Build multi-stage AI workflows by coordinating specialized agents through conversations and events
Python - Apache Airflow47,054 Stars
2 shared tag(s) · 1 shared category(s) · Same language
orchestrate data, ML, and AI workflows with code-defined DAGs
Python - Ray43,970 Stars
2 shared tag(s) · 1 shared category(s) · Same language
Scale Python and AI workloads from a laptop to a cluster with distributed tasks, actors, and objects
Python - Gradio43,672 Stars
2 shared tag(s) · 1 shared category(s) · Same language
turn Python functions and ML models into web, chat, and shareable application interfaces
Python - JAX36,377 Stars
2 shared tag(s) · 1 shared category(s) · Same language
Compose automatic differentiation, JIT compilation, vectorization, and distributed execution for Python and NumPy-style programs
Python - spaCy33,939 Stars
2 shared tag(s) · 1 shared category(s) · Same language
build production NLP pipelines in Python and Cython for tokenization, NER, tagging, parsing, text classification, and training
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.