On this page
Overview
DeepSeek-V3 is a Mixture-of-Experts language model with 671B total parameters and 37B parameters activated per token. It uses Multi-head Latent Attention, DeepSeekMoE, auxiliary-loss-free load balancing, and multi-token prediction, while the repository provides FP8-oriented conversion and inference code plus guidance for several deployment runtimes.
Features and best fit
Based on official documentation; not hands-on tested · Content checked:
Key features
Use a 671B MoE architecture while activating 37B parameters per token
The README describes DeepSeek-V3 as a 671B-total-parameter MoE model with 37B active parameters per token and identifies MLA and DeepSeekMoE as core architecture elements.
Sources: [1]
Best fit
Fits evaluation of large-scale MoE architecture and deployment
It is useful for studying the model architecture, weight conversion, and tensor/pipeline-parallel inference on substantial GPU infrastructure.
Sources: [1]
Before adoption
The official demo assumes Linux, Python 3.10, and multi-node GPUs
The DeepSeek-Infer demo lists Linux and Python 3.10 only, and its example launches the 671B configuration with torchrun across two nodes and eight processes per node. It is not a lightweight Mac or Windows local-run path.
Distinguish the code license from the model license
Repository code is MIT-licensed, while Base/Chat model weights are governed by a separate DeepSeek License Agreement containing redistribution conditions and use-based restrictions. The README states commercial use is supported, but weight use still requires review of the model-license terms.
Official sources
- [1]DeepSeek-V3 v1.0.0 README(2026-10-04)
- [2]DeepSeek-V3 inference requirements(2026-10-04)
- [3]DeepSeek-V3 code MIT license(2026-10-04)
- [4]DeepSeek-V3 model license agreement(2026-10-04)
- [5]DeepSeek-V3 v1.0.0 archival release(2026-10-04)
Supplemental curator note
The repository code is MIT-licensed, but DeepSeek-V3 model weights are governed by a separate DeepSeek License Agreement with redistribution conditions and use-based restrictions. The official demo also assumes Linux, Python 3.10, and large multi-GPU infrastructure, so it should not be treated as a lightweight local model.
Try it in 3 steps
- 1
Clone the DeepSeek-V3 v1.0.0 repository
Pin the archival release and inspect code and configuration without downloading model weights.
git clone --depth 1 --branch v1.0.0 https://github.com/deepseek-ai/DeepSeek-V3.git && cd DeepSeek-V3 - 2
Preflight the 671B configuration
Inspect the 671B configuration without loading weights, including FP8 dtype, layer count, and routed/active expert counts.
python3 -c "import json; c=json.load(open('inference/configs/config_671B.json')); print({k:c[k] for k in ['dim','n_layers','n_routed_experts','n_activated_experts','dtype']})" - 3
Review runtime requirements and the model license
The official demo assumes Linux, Python 3.10, and substantial multi-GPU infrastructure. Review the pinned dependencies and the separate non-MIT model-license terms before downloading or running weights.
printf '%s ' '--- inference requirements ---' && cat inference/requirements.txt && printf '%s ' '--- model license header ---' && head -n 20 LICENSE-MODEL
Growth
Growth trends · Last 30 days
104,509 Stars
Trend data is still being collected.
Development activity
Last 90 days · weekly
- Commits (last 30 days)
- 0
- Open PRs
- 71
Development activity is still being collected.
Built with
Categories and tags
GitHub data
GitHub dataView detailed GitHub data
Related information
Write a related articleShare a guide or use case for this OSS in Markdown. Articles are published after administrator approval.
Explore next
- Khoj37,560 Stars
2 shared tag(s) · 1 shared category(s) · Same language
Combine documents, the web, local/cloud LLMs, agents, and automations in a self-hosted personal AI
Python - Onyx32,321 Stars
2 shared tag(s) · 1 shared category(s) · Same language
turn internal knowledge into a context layer for AI agents
Python - Haystack26,647 Stars
2 shared tag(s) · 1 shared category(s) · Same language
compose retrieval, routing, memory, and generation as component graphs for RAG and agent workflows
Python - Transformers166,909 Stars
5 shared tag(s) · 1 shared category(s) · Same language
Load, run, and train text, vision, audio, and multimodal models through shared AutoClass and Pipeline APIs
Python - Whisper109,887 Stars
3 shared tag(s) · 1 shared category(s) · Same language
run multilingual transcription, language identification, and English translation locally
Python - Stable Diffusion WebUI164,915 Stars
2 shared tag(s) · 1 shared category(s) · Same language
run image generation, inpainting, LoRAs, and extensions from one Gradio browser UI
Python
Report incorrect information
Tell us if any listing information is incorrect or outdated.