DeepSeek V3 deepseek-v3
deepseek-ai · deepseek flagship
#42 overall #5 coding #19 agentic
- params
- 671B (A37B)
- arch
- moe
- context
- 160k
- license
- mit open weights
- released
- Dec 2024
- train compute
- 3.3e+24 FLOP
- downloads/30d
- 1.0M
# architecture
- layers
- 61
- d_model
- 7168
- heads
- 128q / 128kv
- head dim
- 56
- experts
- top-8 of 256 + 1 shared
- vocab
- 129280
- family
- deepseek_v3
signal path of one layer, generated from the registry's structured fields — dimension lines quote the model's real numbers; an MoE trace forks at the router, a dense trace runs straight through. Geometry fields come from the repo's config.json.
# why ranked
Overall: #42 — score 33.8 ● high — all signals present
| signal | weight | input (0–100) |
|---|---|---|
| aa_intelligence | 0.70 | 33.5 |
| bench_composite | 0.30 | 34.7 |
benchmark panel evidence:
| complete panel | score (0–100) |
|---|---|
| aider-polyglot-v1 | 41.8 |
| swe-bench-verified-v1 | 27.5 |
Coding: #5 — score 28.2 ● high — all signals present
| signal | weight | input (0–100) |
|---|---|---|
| aa_coding | 0.30 | 24.6 |
| aider_polyglot | 0.30 | 54.5 |
| swe_bench_verified | 0.40 | 11.1 |
Agentic: #19 — score 11.1 ● low — a single signal; treat with caution
| signal | weight | input (0–100) |
|---|---|---|
| swe_bench_verified | 0.60 | 11.1 |
| swe_rebench | 0.40 | — |
missing signals are dropped and the remaining weights renormalized — never imputed.
# trend
methodology changed during this history window; score movement across that boundary is not model movement. See methodology v5.
# benchmarks
| benchmark | score | source | date |
|---|---|---|---|
| AA Coding Index | 23.0 | Artificial Analysis | — |
| AA Intelligence Index | 9.2 | Artificial Analysis | — |
| AA Math Index | 41.0 | Artificial Analysis | — |
| Aider Polyglot | 55.1 | Aider | Mar 2025 |
| Aider-polyglot | 0.5 / 1 | LLM Stats | — |
| Aider-polyglot-edit | 0.8 / 1 | LLM Stats | — |
| AIME 2024 | 0.6 / 1 | LLM Stats | — |
| C-eval | 0.9 / 1 | LLM Stats | — |
| Cluewsc | 0.9 / 1 | LLM Stats | — |
| Cnmo-2024 | 0.4 / 1 | LLM Stats | — |
| Csimpleqa | 0.6 / 1 | LLM Stats | — |
| SWE-bench Verified | 42.0 | SWE-bench | Aug 2025 |
# where to run
| provider | quant | ctx | $/M in | $/M out | $/M cache | price src | tps | uptime | ✓ |
|---|---|---|---|---|---|---|---|---|---|
| SiliconFlow | — | — | — | — | siliconflow | — | — | ✓ | |
| Hyperbolic | 32k | $0.20 | $0.20 | — | via litellm | — | — | ||
| DeepSeek | 128k | $0.28 | $0.42 | — | via litellm | — | — | ||
| StreamLake | unknown | 125k | $0.26 | $1.03 | — | via openrouter | — | 99.4% | |
| DeepInfra | fp4 | 160k | $0.32 | $0.89 | — | deepinfra | — | 94.0% | ✓ |
| Novita | 64k | $0.40 | $1.30 | — | novita | — | — | ✓ | |
| Nebius | 125k | $0.50 | $1.50 | — | via litellm | — | — | ||
| Amazon Bedrock | 160k | $0.58 | $1.68 | — | via litellm | — | — | ||
| Together | 128k | $1.25 | $1.25 | — | via modelsdev | — | — | ||
| Replicate | 64k | $1.45 | $1.45 | — | via litellm | — | — | ||
| Azure | 125k | $1.14 | $4.56 | — | via litellm | — | — |
sorted by blended price ((3·input + output) / 4 per 1M) · ✓ = the provider's own catalog confirms the offer · "via …" prices are what the aggregator routing the offer charges, not the provider's own list price
# variants
source aliases
- aa
deepseek-v3,deepseek-v3-0324- aider
DeepSeek Chat V3 (prev),DeepSeek V3 (0324)- arena
DeepSeek-V3- epoch
DeepSeek-V3 (Mar 2025)- litellm
deepseek/deepseek-chat- llmstats
deepseek-v3-0324- openrouter
deepseek/deepseek-chat- swebench
DeepSeek-V3-0324