Kimi K2 Thinking kimi-k2-thinking

moonshotai · kimi reasoning flagship agentic

#20 overall #11 agentic

params
1000B (A32B)
arch
moe
context
256k
license
modified-mit open weights
released
Nov 2025
reasoning
yes
train compute
4.2e+24 FLOP
downloads/30d
46.1k

# architecture

attn shared · reasoning MHA ×64 router top-8 of 384 +1 shared expert ×384 expert out ×61 layers ctx 262,144 1000B pool · A32B/token
layers
61
d_model
7168
heads
64q / 64kv
head dim
112
experts
top-8 of 384 + 1 shared
vocab
163840
family
kimi_k2

signal path of one layer, generated from the registry's structured fields — dimension lines quote the model's real numbers; an MoE trace forks at the router, a dense trace runs straight through. Geometry fields come from the repo's config.json.

# why ranked

Overall: #20 — score 66.8 ● high — all signals present

signalweightinput (0–100)
aa_intelligence 0.70 71.1
bench_composite 0.30 56.7

benchmark panel evidence:

complete panelscore (0–100)
swe-bench-verified-v156.7

Coding: provisional — score 50.0 ● low — a single signal; treat with caution

provisional: too few independent signals for a numbered position — the score is shown, but this model sorts after every ranked model.

signalweightinput (0–100)
aa_coding 0.30
aider_polyglot 0.30
swe_bench_verified 0.40 50.0

missing signals are dropped and the remaining weights renormalized — never imputed.

Agentic: #11 — score 50.0 ● low — a single signal; treat with caution

signalweightinput (0–100)
swe_bench_verified 0.60 50.0
swe_rebench 0.40

missing signals are dropped and the remaining weights renormalized — never imputed.

full methodology

# trend

OGM score · last 55 days
OGM score over 55 days
overall rank (up is better)
Overall rank over 55 days

methodology changed during this history window; score movement across that boundary is not model movement. See methodology v5.

# benchmarks

benchmarkscoresourcedate
AA Intelligence Index 25.9 Artificial Analysis
AA Math Index 94.7 Artificial Analysis
SWE-bench Verified 63.4 SWE-bench Dec 2025

# where to run

providerquantctx$/M in$/M out$/M cacheprice srctpsuptime
Amazon Bedrock $0.30$1.25 bedrock
Gmi 256k $0.80$1.20 via litellm
BaseTen $0.60$2.50 via litellm
Googleunknown 256k $0.60$2.50 via openrouter 100.0%
Moonshot AI 256k $0.60$2.50$0.15 via modelsdev
Novitabf16 256k $0.60$2.50 novita 99.8%
Crusoe 256k $2.50$2.50 via litellm

sorted by blended price ((3·input + output) / 4 per 1M) · ✓ = the provider's own catalog confirms the offer · "via …" prices are what the aggregator routing the offer charges, not the provider's own list price

source aliases
aa
kimi-k2-thinking
arena
Kimi-K2-Thinking
epoch
Kimi K2 Thinking
openrouter
moonshotai/kimi-k2-thinking