OSS探訪

671B parameterのMoE LLMとFP8/multi-node inference向けreference implementationを公開する

スコアの見方

OSS規模スコアはStars・Watchers・Forks・Contributorsを対数圧縮して重み付けした現在の規模指標(上限なし)です。発掘スコアは現在のOSS規模スコアから発掘時点のOSS規模スコアを引いた値、更新ペースは直近30日Commit数、成長モメンタムは直近の観測期間におけるOSS規模スコア差、OSS健全度は取得できた更新状況・Community Health・Releaseの0〜100評価です。

Stars
104,509
主要言語
Python
ライセンス
MIT
リポジトリ最終更新
2025/08/28
ページ内ナビ

概要

DeepSeek-V3は671B total parameterのMixture-of-Experts language modelで、tokenごとに37B parameterをactivateします。Multi-head Latent Attention、DeepSeekMoE、auxiliary-loss-free load balancing、multi-token prediction等を採用し、repositoryではFP8 weights向け変換・推論codeと複数runtimeへのdeploy案内を提供します。

特徴と向いている用途

公式資料に基づく紹介・実機未検証 · 内容確認日:

主な特徴

671B MoE architectureでtokenごとのactive parameterを抑える

READMEはDeepSeek-V3を671B total parameter、tokenごとに37B active parameterのMoE modelとして説明し、MLAとDeepSeekMoEを中核architectureとして挙げています。

出典:[1]

FP8 weightと複数のdistributed inference runtimeを扱う

公式repositoryはFP8 weightsを前提にreference demoを提供し、SGLang、LMDeploy、TensorRT-LLM、vLLM、LightLLM等のlocal/cloud deployment optionを案内しています。

出典:[1][2]

向いている用途

大規模MoE modelのarchitecture・deploymentを検証したい場合に向く

model architecture、weight conversion、tensor/pipeline parallelなinference構成を調査し、大規模GPU infrastructureでDeepSeek-V3をself-host評価したい用途に適します。

出典:[1]

導入前の確認

公式demoはLinux・Python 3.10・multi-node GPU前提である

DeepSeek-Infer demoはLinuxとPython 3.10のみをsystem requirementとし、exampleは2 node × 各8 processのtorchrunで671B configを起動します。Mac/Windows向けの簡易local runではありません。

出典:[1][2]

code licenseとmodel licenseを分けて確認する

repository codeはMIT Licenseですが、Base/Chat model weightsは別のDeepSeek License Agreementに従い、再配布条件とuse-based restrictionsを含みます。READMEはcommercial useをsupportするとしていますが、weight利用時はmodel license本文の条件確認が必要です。

出典:[3][4][1][5]

参考にした公式資料

  1. [1]DeepSeek-V3 v1.0.0 README(2026-10-04)
  2. [2]DeepSeek-V3 inference requirements(2026-10-04)
  3. [3]DeepSeek-V3 code MIT license(2026-10-04)
  4. [4]DeepSeek-V3 model license agreement(2026-10-04)
  5. [5]DeepSeek-V3 v1.0.0 archival release(2026-10-04)
編集部からの補足

repository codeはMITですが、DeepSeek-V3のmodel weightsは別のDeepSeek License Agreementに従い、利用・再配布条件とuse-based restrictionsがあります。さらに公式demoはLinux/Python 3.10と大規模multi-GPUを前提にするため、一般PC向けの軽量local modelとして扱わない方が安全です。

3ステップで試す

  1. 1

    DeepSeek-V3 v1.0.0 repositoryを取得する

    archival releaseへ固定し、model weightをdownloadせずcodeとconfigurationだけを取得します。

    git clone --depth 1 --branch v1.0.0 https://github.com/deepseek-ai/DeepSeek-V3.git && cd DeepSeek-V3
  2. 2

    671B configをpreflight確認する

    model weightを読み込まず、671B configurationがFP8・61 layers・256 routed expertsであること等を確認します。

    python3 -c "import json; c=json.load(open('inference/configs/config_671B.json')); print({k:c[k] for k in ['dim','n_layers','n_routed_experts','n_activated_experts','dtype']})"
  3. 3

    実行要件とmodel licenseを確認する

    公式demoはLinux/Python 3.10と大規模multi-GPUを前提にします。weight取得・実行前にrequirementsとMITではない別model licenseの条件を確認してください。

    printf '%s ' '--- inference requirements ---' && cat inference/requirements.txt && printf '%s ' '--- model license header ---' && head -n 20 LICENSE-MODEL
公式READMEで確認

成長

成長の推移 · 直近30日

104,509 Stars

推移データを蓄積中です。

開発アクティビティ

直近90日・週次

Commit(直近30日)
0
Open PR
71

開発アクティビティを蓄積中です。

Built with

カテゴリとタグ

GitHubデータ

GitHubのデータGitHubの詳細データを見る
Stars
104,509
Forks
16,715
Watchers
744
Open Issues
199
Contributors
23
所有者種別
Organization
主要言語
Python
ライセンス
MIT
リポジトリ最終更新
2025/08/28

このOSSの使い方や活用事例をMarkdownで投稿できます。管理者が承認した後に公開されます。

情報の誤りを報告

掲載内容に誤りや古い情報があればお知らせください。

このページを読んで、次に何をすればよいか分かりましたか?