DeepSeek R1 deepseek-r1
deepseek-ai · deepseek reasoning flagship
#33 overall
- params
- 671B (A37B)
- arch
- moe
- context
- 64k
- license
- mit open weights
- released
- Jan 2025
- reasoning
- yes
- train compute
- 3.5e+24 FLOP
- downloads/30d
- 838.8k
# architecture
- layers
- 61
- d_model
- 7168
- heads
- 128q / 128kv
- head dim
- 56
- experts
- top-8 of 256 + 1 shared
- vocab
- 129280
- family
- deepseek_v3
signal path of one layer, generated from the registry's structured fields — dimension lines quote the model's real numbers; an MoE trace forks at the router, a dense trace runs straight through. Geometry fields come from the repo's config.json.
# why ranked
Overall: #33 — score 47.0 ● high — all signals present
| signal | weight | input (0–100) |
|---|---|---|
| aa_intelligence | 0.70 | 47.4 |
| bench_composite | 0.30 | 45.9 |
benchmark panel evidence:
| complete panel | score (0–100) |
|---|---|
| aider-polyglot-v1 | 45.9 |
Coding: provisional — score 63.6 ● low — a single signal; treat with caution
provisional: too few independent signals for a numbered position — the score is shown, but this model sorts after every ranked model.
| signal | weight | input (0–100) |
|---|---|---|
| aa_coding | 0.30 | — |
| aider_polyglot | 0.30 | 63.6 |
| swe_bench_verified | 0.40 | — |
missing signals are dropped and the remaining weights renormalized — never imputed.
# trend
methodology changed during this history window; score movement across that boundary is not model movement. See methodology v5.
# benchmarks
| benchmark | score | source | date |
|---|---|---|---|
| AA Intelligence Index | 13.9 | Artificial Analysis | — |
| AA Math Index | 76.0 | Artificial Analysis | — |
| Aider Polyglot | 56.9 | Aider | Jan 2025 |
# where to run
| provider | quant | ctx | $/M in | $/M out | $/M cache | price src | tps | uptime | ✓ |
|---|---|---|---|---|---|---|---|---|---|
| SiliconFlow | — | — | — | — | siliconflow | — | — | ✓ | |
| DeepSeek | 64k | $0.55 | $2.19 | — | via litellm | — | — | ||
| Nebius | 125k | $0.80 | $2.40 | — | via litellm | — | — | ||
| DeepInfra | 64k | $0.85 | $2.50 | $0.85 | via requesty | — | — | ||
| Hyperbolic | 160k | $2.00 | $2.00 | — | hyperbolic | — | — | ✓ | |
| Amazon Bedrock | 125k | $1.35 | $5.40 | — | via litellm | — | — | ||
| Azure | 160k | $1.35 | $5.40 | — | via modelsdev | — | — | ||
| Snowflake | 125k | $1.35 | $5.40 | — | via litellm | — | — | ||
| Novita | fp8 | 64k | $4.00 | $4.00 | — | novita | — | 100.0% | ✓ |
| Together | 163k | $3.00 | $7.00 | — | via modelsdev | — | — | ||
| Replicate | 64k | $3.75 | $10.00 | — | via litellm | — | — | ||
| SambaNova | 32k | $5.00 | $7.00 | — | via litellm | — | — |
sorted by blended price ((3·input + output) / 4 per 1M) · ✓ = the provider's own catalog confirms the offer · "via …" prices are what the aggregator routing the offer charges, not the provider's own list price
# variants
source aliases
- aa
deepseek-r1- aider
DeepSeek R1- arena
DeepSeek-R1- litellm
deepseek/deepseek-reasoner- openrouter
deepseek/deepseek-r1