On this page
Overview
DVC is a command-line tool that connects machine-learning data, models, pipelines, and experiment results to Git history. Large payloads live in a cache or remote storage while lightweight metadata is versioned with the code. Teams can begin running and comparing experiments locally without first operating a separate tracking server.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Associate data and models with Git revisions
dvc add places data or models in the DVC cache and creates lightweight metadata that Git can track. This keeps large payloads out of Git while preserving the relationship among code, parameters, data, and models at each revision.
Sources: [1]
Reproduce pipelines and compare experiments
DVC pipelines describe input code and data, commands, and outputs as a computational graph so only affected stages need to run again. Experiment versioning records multiple trials in the local Git repository and compares parameters, metrics, and plots.
Sources: [1]
Best fit
Fits ML projects that need code and data changes to remain reproducible
DVC fits teams that need to reproduce training or preprocessing together with the exact input data, and projects that want to begin experiment tracking locally without another server. It also works with existing Git hosting plus cloud or on-premises storage.
Before adoption
Design remote-storage and Git operations separately
Git primarily stores code and DVC metadata. Sharing and backing up data payloads requires a configured remote such as S3, Azure, Google Cloud Storage, or SSH, and the selected backend may require an extra package such as dvc-s3.
DVC does not make a project reproducible automatically. Teams still need to register code, parameter, environment, and data changes consistently and align their Git and DVC push/pull, cache-retention, and access-control practices.
Sources: [1]
Official sources
- [1]DVC 3.67.1 README(2026-10-04)
- [2]DVC 3.67.1 release(2026-10-04)
Supplemental curator note
A key comparison point is keeping large payloads outside Git while versioning their relationships and reproduction steps alongside code. Team adoption should also design cache retention, remote-storage access, and cost controls.
Try it in 3 steps
- 1
Install DVC
Create an isolated virtual environment and install the reviewed official release only inside it.
mkdir -p dvc-demo && python3 -m venv dvc-demo/.venv && dvc-demo/.venv/bin/python -m pip install "dvc==3.67.1" - 2
Initialize a Git repository
Create Git and DVC metadata in an empty working directory.
cd dvc-demo && git init && .venv/bin/dvc init - 3
Track sample data
Use a small file to inspect the DVC metadata and the files that should be committed to Git.
printf "sample\n" > data.txt && .venv/bin/dvc add data.txt && git add data.txt.dvc .gitignore
Growth
Growth trends · Last 30 days
15,902 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 0
- Open PRs
- 33
Development activity is still being collected.
Built with
Categories and tags
GitHub data
GitHub dataView detailed GitHub data
GitHub Topics
- data-science
- machine-learning
- reproducibility
- data-version-control
- developer-tools
- ai
- unstructured-data
- Stars
- 15,902
- Forks
- 1,328
- Watchers
- 135
- Open issues
- 184
- Contributors
- 292
- Owner type
- Organization
- Primary language
- Python
- License
- Apache-2.0
- Repository last updated
- Sep 28, 2026
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- spaCy33,935 Stars
2 shared tag(s) · 2 shared category(s) · Same language
build production NLP pipelines in Python and Cython for tokenization, NER, tagging, parsing, text classification, and training
Python - TRL19,445 Stars
2 shared tag(s) · 2 shared category(s) · Same language
Post-train foundation models with shared Python trainers and a CLI for SFT, preference optimization, and reward learning
Python - bitsandbytes8,510 Stars
2 shared tag(s) · 2 shared category(s) · Same language
quantize PyTorch models to 8-bit and 4-bit formats to reduce memory use for LLM inference and QLoRA training
Python - Transformers166,931 Stars
2 shared tag(s) · 1 shared category(s) · Same language
Load, run, and train text, vision, audio, and multimodal models through shared AutoClass and Pipeline APIs
Python - Stable Diffusion WebUI165,189 Stars
2 shared tag(s) · 1 shared category(s) · Same language
run image generation, inpainting, LoRAs, and extensions from one Gradio browser UI
Python - ComfyUI136,026 Stars
2 shared tag(s) · 1 shared category(s) · Same language
save generative AI pipelines as visible graphs and rerun only what changed
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.