Qwen3-VL-8B-Thinking qwen3-vl-8b-thinking

qwen · qwen vision reasoning efficient

#46 overall

params
8.8B
arch
dense
context
128k
license
apache-2.0 open weights
reasoning
yes

# architecture

attn shared · reasoning GQA 32q/8kv feed-forward (dense) every parameter, every token out ×36 layers ctx 131,072 8.8B — no routing
layers
36
d_model
4096
heads
32q / 8kv
head dim
128
vocab
151936
family
qwen3_vl_text

signal path of one layer, generated from the registry's structured fields — dimension lines quote the model's real numbers; an MoE trace forks at the router, a dense trace runs straight through. Geometry fields come from the repo's config.json.

# why ranked

Overall: #46 — score 28.4 ● high — all signals present

signalweightinput (0–100)
aa_intelligence 0.70 17.5
bench_composite 0.30 53.7

benchmark panel evidence:

complete panelscore (0–100)
aime-2025-v153.7

full methodology

# trend

OGM score · last 54 days
OGM score over 54 days
overall rank (up is better)
Overall rank over 54 days

methodology changed during this history window; score movement across that boundary is not model movement. See methodology v5.

# benchmarks

benchmarkscoresourcedate
AA Intelligence Index 4.9 Artificial Analysis
AA Math Index 30.7 Artificial Analysis
Ai2d 0.8 / 1 LLM Stats
AIME 2025 0.8 / 1 LLM Stats
Arena-hard-v2 0.5 / 1 LLM Stats
Bfcl-v3 0.6 / 1 LLM Stats
Blink 0.7 / 1 LLM Stats
Cc-ocr 0.8 / 1 LLM Stats
Charadessta 0.6 / 1 LLM Stats
Charxiv-d 0.9 / 1 LLM Stats
Charxiv-r 0.5 / 1 LLM Stats
Creative-writing-v3 0.8 / 1 LLM Stats
Mm-mt-bench 8.0 LLM Stats

# where to run

providerquantctx$/M in$/M out$/M cacheprice srctpsuptime
Alibabaunknown 128k $0.18$2.10 via openrouter 100.0%

sorted by blended price ((3·input + output) / 4 per 1M) · ✓ = the provider's own catalog confirms the offer · "via …" prices are what the aggregator routing the offer charges, not the provider's own list price

source aliases
aa
qwen3-vl-8b-reasoning
llmstats
qwen3-vl-8b-thinking
openrouter
qwen/qwen3-vl-8b-thinking