LLM Index for Scientific Computing

Generated 2026-08-14 · 478 models scored from 579 measured by Artificial Analysis · raw cost data only

SoftMinZ vs cost per agentic task scatter with Pareto frontier

Each point is a model: vertical position is the SoftMinZ performance index, horizontal position the raw cost of one agentic task (log scale). The dashed line is the Pareto frontier — models for which no cheaper model is at least as good. Frontier members are colour- and shape-coded; everything else is grey.

Full ranking

Every scored model. Bold rows are Pareto frontier members. Use the two buttons to switch the ranking — they need no JavaScript. In cost view, models with no published cost data sort last.

Highest SoftMinZLowest cost per task Best model first (highest SoftMinZ). Cheapest model first (lowest cost per task).
Model / variantSoftMinZCost per task
Muse Spark 1.2 (xhigh)1.83$0.6126
Claude Opus 4.8 (max)1.80$3.0712
Gemini 3.1 Pro Preview1.77$0.4118
Grok 4.6 (high)1.76$1.4213
Kimi K3 (max)1.73$1.3209
Claude Fable 5 (with fallback)1.71$4.9043
Muse Spark 1.1 (xhigh)1.66$0.3738
Claude Opus 5 (xhigh)1.65$2.9689
GLM-5.2 (max)1.65$0.4537
Claude Sonnet 5 (max)1.64$2.6557
Claude Opus 5 (max)1.63$3.8996
Claude Opus 5 (high)1.61$1.9678
Claude Opus 4.7 (max)1.61$3.4450
Qwen3.8 2.4T A95B1.60$1.9605
Qwen3.8 Max1.59$1.9863
Qwen3.7 Max1.59$0.6966
Claude Opus 5 (medium)1.56$1.1109
Grok 4.5 (high)1.53$0.5272
Gemini 3.7 Flash (high)1.47$0.7208
Claude Opus 5 (low)1.42$0.6647
Gemini 3.6 Flash1.41$0.8421
Gemini 3.5 Flash1.39$1.1397
Kimi K2.61.35$0.4619
Gemini 3.5 Flash (medium)1.33
Claude Opus 4.6 (max)1.32
Qwen3.7 Plus1.29$0.2961
Gemini 3.7 Flash (medium)1.28$0.4684
Grok 4.3 (high)1.27$0.1823
Grok 4.20 0309 v21.17
MiniMax-M31.15$0.2209
MiMo-V2.5-Pro1.15$0.0448
Claude Opus 4.7 (Non-reasoning, high)1.15
Grok 4.20 03091.11
Motif 31.10
Claude Opus 4.51.08
GPT-5.1 (high)1.06$0.3826
Gemini 3.7 Flash (low)1.05$0.2807
GPT-5.2 (medium)1.05
Claude Sonnet 4.6 (max)1.05$1.4340
Grok 4.3 (medium)1.03
Solar Pro 41.03$0.3075
GPT-5.2 (xhigh)1.03
Qwen3.6 Max Preview1.03
Solar Open2 250B1.02
GLM-5.11.02$0.3716
GPT-5.2 Codex (xhigh)1.01
GPT-5.5 (xhigh)1.00$1.6981
Inkling Small0.99$0.0894
GPT-5.5 (high)0.99$1.1888
A.X-K20.97
GPT-5.6 Terra (max)0.96$0.7633
Motif 3 (Beta)0.95
GPT-5.6 Sol (max)0.92$2.0597
GPT-5.6 Sol (high)0.92$0.9466
GPT-5.5 (medium)0.92$0.7785
GPT-5.6 Sol (medium)0.91$0.6346
Muse Spark0.91
Qwen3.6 Plus0.91$0.5053
GPT-5.3 Codex (xhigh)0.90
GPT-5.6 Sol (xhigh)0.90$1.3706
GPT-5.6 Terra (xhigh)0.89$0.4965
MiMo-V2.50.89$0.0137
GPT-5.6 Sol (low)0.88$0.3943
Nemotron 3 Ultra0.88$0.5244
GLM-50.88
Inkling0.87$0.4421
GPT-5.4 (xhigh)0.86$1.6368
MiniMax-M2.70.84$0.1031
GPT-5.6 Terra (high)0.82$0.3574
Kimi K2.7 Code0.82$0.3134
GPT-5.4 (low)0.80
Gemini 3 Pro Preview (high)0.80
Kimi K2.50.80$0.1175
GPT-5.5 (low)0.79$0.4015
Grok 40.79
MiMo-V2-Pro0.78
GPT-5 Codex (high)0.75
GPT-5.5 Instant (May 2026)0.75
GPT-5.4 nano (xhigh)0.73$0.2076
Hy30.72$0.0438
GPT-5.1 Codex (high)0.72
Gemini 3 Flash0.70
GPT-5 (high)0.69$0.3388
Claude 4.5 Sonnet0.68$0.6070
GPT-5.6 Terra (medium)0.65$0.1936
DeepSeek V4 Pro (high)0.65$0.0557
GPT-5.6 Luna (max)0.64$0.0675
Kimi K3 (low)0.64$0.3735
Grok 4.3 (low)0.62
MiMo-V2-Flash (Feb 2026)0.62
MiMo-V2-Omni-03270.62
Ling 3.0 Flash0.62$0.0439
GPT-5.6 Terra (low)0.61$0.1484
DeepSeek V4 Flash 0731 (max)0.61$0.0435
GPT-5.4 mini (xhigh)0.61$0.6193
GPT-5.6 Luna (xhigh)0.60$0.0470
Claude Sonnet 5 (Non-reasoning)0.60$0.7065
Qwen3.6 27B0.60$0.4021
Kimi K2.60.60
Gemini 3.5 Flash (minimal)0.59
DeepSeek V4 Pro 0813 (max)0.59$0.3662
DeepSeek V4 Pro (max)0.59$0.0554
Gemini 3.5 Flash-Lite0.58$0.1292
GPT-5.6 Luna (high)0.57$0.0329
GPT-5 mini (high)0.57$0.0453
GPT-5.4 nano0.57
GPT-5 mini (medium)0.57
DeepSeek V3.2 Speciale0.57
Kimi K2 Thinking0.57
GPT-5.1 Codex mini (high)0.56
GLM-5-Turbo0.56
MiMo-V2-Omni0.55
Grok Build 0.1 06160.55$0.2791
K-EXAONE 2.00.54
Claude Opus 4.6 (high)0.54
Agnes 2.5 Pro Alpha0.51
DeepSeek V3.20.50
Grok 4.1 Fast0.49
JT-4.1 Flash 236B A21B0.48
Qwen3.6 35B A3B0.48$0.4079
Claude 4 Sonnet0.48
Grok 4 Fast0.48
KAT-Coder-Pro V20.47
Claude Sonnet 4.6 (Non-reasoning)0.47
Claude Sonnet 4.6 (Non-reasoning, Low Effort)0.46
Gemini 3 Pro Preview (low)0.45
Claude Opus 4.50.45
GPT-5.5 Instant (June 2026)0.44$0.6674
GPT-5 (medium)0.44
MiniMax-M2.10.44
DeepSeek V4 Flash (high)0.41$0.0662
DeepSeek V3.1 Terminus0.41
GLM-5.10.41
GLM 5V Turbo0.41
G9v3-39A5B0.40
GPT-5 (low)0.40
LongCat 2.00.39
Qwen3.5 Omni Plus0.39
GPT-5.6 Luna (medium)0.38$0.0173
o30.38
Claude 4.5 Haiku0.38$0.2658
DeepSeek V4 Flash (max)0.37$0.0845
Gemini 2.5 Pro0.35$0.2663
Grok 3 mini Reasoning (high)0.35
Qwen3.5 397B A17B0.34$0.4697
Hy3-preview0.34
Muse Glimmer (high)0.33$0.0959
Nex-N2-Pro0.33
Command A+0.32
Kimi K2.50.32
Claude 3.7 Sonnet0.31
DeepSeek V3.2 Exp0.31
Ring-2.6-1T0.29
GLM-4.70.29$0.5090
GPT-5.4 mini (medium)0.28
Gemini 3.1 Flash-Lite0.28$0.0461
Step 3.7 Flash0.27$0.0864
GLM-5.20.27
GPT-5.6 Luna (low)0.26$0.0135
Gemini 3 Flash0.26
Claude 4.5 Sonnet0.25
Qwen3.5 397B A17B0.24
DeepSeek V3.10.24
Qwen3.5 27B0.24
GPT-5.6 Sol (Non-reasoning)0.23$0.4037
Doubao Seed Code0.23
GPT-5.20.22
Gemma 4 31B0.21
MiniMax-M2.50.21
MiMo-V2-Flash0.21
GPT-5.5 (Non-reasoning)0.20$0.3376
o4-mini (high)0.20
GLM-4.50.20
Qwen3.5 122B A10B0.19$0.3213
DeepSeek R1 05280.19
GPT-5.4 (Non-reasoning)0.18
Gemini 2.5 Flash0.18
Nemotron 3.5 Lightning0.17$0.0841
GLM-50.17
Magistral Medium 1.20.16$1.4169
Nemotron 3 Super0.15$0.3404
Step 3.5 Flash0.14
Qwen3 Max Thinking0.13
Mistral Medium 3.50.12$0.5708
Step 3.5 Flash 26030.12
o10.11
Claude 4 Sonnet0.11
Qwen3.5 27B0.11
KAT-Coder-Pro V10.10
Qwen3.5 35B A3B0.10
Claude 4.5 Haiku0.09
JT-35B-Flash0.09
GPT-5 nano (medium)0.08
Mistral Small 40.06$0.1163
Gemini 2.5 Flash (Sep)0.06
Kimi K2 09050.06
gpt-oss-120b (high)0.05$0.1079
Ling 3.0 Tiny0.05
GPT-5 nano (high)0.05
MiniMax-M20.04
Nova 2.0 Pro Preview (low)0.02
GLM-4.60.02$0.3950
DeepSeek V4 Pro0.01
Qwen3 Coder 480B0.01
Qwen3 235B A22B 25070.01$0.2984
MiMo-V2.5-Pro0.00
Claude 3.7 Sonnet0.00
Kimi K2-0.00
Grok Code Fast 1-0.01
Qwen3.6 27B-0.02$0.5896
Hy3-preview-0.02
Qwen3 Max-0.03
GPT-5.6 Terra (Non-reasoning)-0.04$0.1680
Nova 2.0 Pro Preview (medium)-0.04
K-EXAONE-0.04
Gemma 4 12B-0.05
Cogito v2.1-0.05
Qwen3 VL 235B A22B-0.05
Qwen3 Max Thinking (Preview)-0.05
Gemma 4 31B-0.06$0.0609
Qwen3 Max (Preview)-0.06
Gemma 4 26B A4B-0.07$0.0668
Trinity Large Thinking-0.07
Qwen3.5 122B A10B-0.08$0.2630
DeepSeek V3.2-0.08
GPT-5.6 Luna (Non-reasoning)-0.09$0.0174
Qwen3 Next 80B A3B-0.09$0.2676
o3-mini (high)-0.09
DeepSeek V3.1-0.09
North Mini Code-0.10
Qwen3.5 9B-0.10$0.4006
Gemini 2.5 Flash (Sep)-0.10
DeepSeek V3.1 Terminus-0.10
GLM-4.6-0.10
Nova 2.0 Lite (high)-0.11
Mercury 2-0.11$0.1013
DeepSeek V3.2 Exp-0.11
Qwen3 235B 2507-0.11
EXAONE 4.5 33B-0.12
HyperNova 60B 2605-0.14
Nemotron 3 Nano-0.15$0.0359
K2 Think V2-0.15
INTELLECT-3-0.16
Nova 2.0 Lite (medium)-0.17
Nemotron Cascade 2 30B A3B-0.17
DeepSeek R1 (Jan)-0.17$0.5301
Grok 4.3 (Non-reasoning)-0.18$0.3780
Ring-1T-0.18
Seed-OSS-36B-Instruct-0.18
GPT-5.1-0.19
Gemini 2.5 Flash-Lite (Sep)-0.19
Qwen3 VL 32B-0.20
Llama Nemotron Super 49B v1.5-0.22
Grok 3-0.22
GLM-4.7-0.22
GLM-4.6V-0.22
DeepSeek V4 Flash-0.23
Nova 2.0 Omni (low)-0.23
ERNIE 5.0 Thinking Preview-0.23
GPT-5 (minimal)-0.23
Gemini 2.5 Flash-Lite (Sep)-0.23
Grok 4.20 0309-0.23
Ling-2.6-1T-0.23
MiMo-V2-Flash-0.24
GPT-5.4 nano-0.24
Qwen3 30B A3B 2507-0.24$0.1183
Mistral Large 3-0.25$0.1284
Apriel-v1.6-15B-Thinker-0.25
GPT-4.1-0.25
DeepSeek V3 0324-0.26$0.0628
Apriel-v1.5-15B-Thinker-0.26
MiniMax M1 80k-0.27
Gemma 4 26B A4B-0.30
Gemma 4 E4B-0.30
GPT-5 mini (minimal)-0.30
Gemini 2.5 Flash-0.30
GLM-4.5-Air-0.30
Grok 4.20 0309 v2-0.31
K2-V2 (high)-0.33
GPT-4o (Aug)-0.33
Mistral Medium 3-0.33
Ling-1T-0.33
gpt-oss-20b (high)-0.34$0.0442
gpt-oss-120b (low)-0.34$0.0412
Llama 4 Maverick-0.34$0.0648
Devstral 2-0.34
Nova 2.0 Omni (medium)-0.34
Nova 2.0 Pro Preview-0.34
Hermes 4 405B-0.34
Qwen3 Next 80B A3B-0.34
Gemma 4 12B (Non-reasoning)-0.35
Qwen3.5 35B A3B-0.35$0.3170
Qwen3 VL 30B A3B-0.35
Gemini 2.5 Flash-Lite-0.36
Qwen3 VL 235B A22B-0.36
Hermes 4 405B-0.37
Qwen3 Coder Next-0.37$0.4379
GPT-4.1 mini-0.37$0.1159
Nova 2.0 Lite (low)-0.38
GPT-5.4 mini-0.38
Grok 4.1 Fast-0.38
Magistral Medium 1-0.39
Llama 3.1 405B-0.40
Devstral Medium-0.40
Nova Premier-0.41
Qwen3.5 4B-0.42
gpt-oss-20b (low)-0.43
EXAONE 4.0 32B-0.43
K2-V2 (medium)-0.43
G9v3-3B-0.43
Mistral Medium 3.1-0.43$0.2419
Solar Pro 3-0.44$0.2029
NVIDIA Nemotron Nano 9B V2-0.44
GLM-4.7-Flash-0.44
Qwen3 Coder 30B A3B-0.44
Qwen3 235B-0.44
Llama Nemotron Ultra-0.45
ERNIE 4.5 300B A47B-0.45
HyperCLOVA X SEED Think (32B)-0.45
Gemini 2.0 Flash-0.45
Qwen3 VL 32B-0.46
Qwen3 4B 2507-0.46
Mistral Small 4-0.47
Grok 4 Fast-0.47
DeepSeek V3 (Dec)-0.48$0.0355
DiffusionGemma 26B A4B-0.49
Solar Open 100B-0.49
Qwen3.5 9B-0.50
Magistral Small 1.2-0.51$0.5170
Qwen3.5 Omni Flash-0.51
GLM-4.6V-0.51
Ling-flash-2.0-0.51
NVIDIA Nemotron Nano 12B v2 VL-0.52
Hermes 4 70B-0.52
Qwen3 VL 30B A3B-0.52
Devstral Small 2-0.53
K-EXAONE-0.53
Ring-flash-2.0-0.53
Motif-2-12.7B-0.54
GPT-4o (Nov)-0.54
Claude 3.5 Haiku-0.54
NVIDIA Nemotron Nano 9B V2-0.55
Qwen3 30B A3B 2507-0.55
Falcon-H1R-7B-0.56
Qwen3 32B-0.56
Nemotron 3 Nano Omni 30B A3B-0.57
Olmo 3.1 32B Think-0.57
K2-V2 (low)-0.57
Mi:dm K 2.5 Pro-0.57
Nanbeige4.1-3B-0.57
Llama Nemotron Super 49B v1.5-0.58
Mi:dm K 2.5 Pro Preview-0.58
Magistral Small 1-0.59
Qwen3 VL 8B-0.59
Ling 2.6 Flash-0.59
Gemma 4 E2B-0.60
JT-MINI-0.60
Mistral Small 3.2-0.60$0.2532
Qwen3 14B-0.60
Command A-0.61
Mistral Large 2 (Nov)-0.62
GLM-4.5V-0.62
Nova 2.0 Lite-0.63
Qwen3 Omni 30B A3B-0.63
LongCat Flash Lite-0.64
Qwen3 30B-0.64
Llama 4 Scout-0.65$0.0184
Nova 2.0 Omni-0.65
Step3 VL 10B-0.65
GPT-4.1 nano-0.65$0.0707
Mistral Small 3.1-0.65$0.0578
EXAONE 4.0 32B-0.66
Nova Pro-0.66
Qwen2.5 72B-0.66
DeepSeek R1 0528 Qwen3 8B-0.67
Solar Pro 2-0.67
Llama 3.3 Nemotron Super 49B-0.67
Olmo 3 32B Think-0.67
Ministral 3 14B-0.68$0.2314
Celeris-1-0.68$0.3054
GPT-5 nano (minimal)-0.68
Devstral Small (May)-0.69
Qwen3 235B-0.69
Devstral Small-0.70
Llama 3.3 Nemotron Super 49B-0.70
Qwen3.5 4B-0.71
Gemini 2.5 Flash-Lite-0.71
Llama 3.3 70B-0.72
Tri-21B-think Preview-0.73
DeepSeek R1 Distill Llama 70B-0.74
Llama 3.1 Nemotron 70B-0.75
Qwen3 4B 2507-0.75
Sarvam 105B (high)-0.76
Llama 3.1 70B-0.76
Tri-21B-Think-0.77
Phi-4-0.78
Hermes 4 70B-0.78
Ministral 3 8B-0.80$0.2217
Qwen3.6 35B A3B-0.80
Solar Pro 2-0.80
GLM-4.5V-0.80
Gemma 4 E2B-0.81
Nova Lite-0.81
Qwen3 30B-0.81
Gemma 4 E4B-0.82
Granite 4.1 30B-0.82
Qwen3 VL 8B-0.82
Qwen3 8B-0.82
Qwen3 VL 4B-0.82
Nemotron 3 Nano-0.83
Nemotron 3 Nano 4B-0.83
Reka Flash 3-0.83
GLM-4.7-Flash-0.83
Olmo 3.1 32B Instruct-0.85
Claude 3 Haiku-0.85
Granite 4.0 H Small-0.86
LFM2.5-8B-A1B-0.87
Olmo 3 7B Think-0.87
Granite 4.1 8B-0.88
Qwen3 14B-0.88
Qwen3 Omni 30B A3B-0.90
NVIDIA Nemotron Nano 12B v2 VL-0.91
Jamba 1.7 Large-0.91
Sarvam 30B (high)-0.93
Ministral 3 3B-0.94$0.1780
Sarvam M-0.96
Gemma 3 27B-0.97
Ling-mini-2.0-0.97
LFM2 24B A2B-0.98
Llama 3 70B-0.98
Llama 3.1 8B-0.99
Nova Micro-1.00
Qwen3.5 2B-1.01
Phi-4 Mini-1.05
Exaone 4.0 1.2B-1.06
Qwen3 8B-1.07
Qwen3 VL 4B-1.07
Molmo2-8B-1.08
Jamba Reasoning 3B-1.09
MiniCPM5-1B-1.11
Olmo 3 7B-1.13
Gemma 3 12B-1.13
Granite 4.0 Micro-1.17
Exaone 4.0 1.2B-1.17
Qwen3.5 2B-1.17
Llama 3.2 11B (Vision)-1.18
MiniCPM5-1B-1.19
Granite 4.1 3B-1.21
Jamba 1.7 Mini-1.22
Llama 3 8B-1.22
Granite 3.3 8B-1.23
Granite 4.0 H 1B-1.24
LFM2 8B A1B-1.25
LFM2 2.6B-1.25
Qwen3 1.7B-1.25
Apertus 70B Instruct-1.26
LFM2.5-1.2B-Instruct-1.30
Gemma 3n E4B-1.31
Gemma 3 4B-1.32
LFM2.5-1.2B-Thinking-1.33
Granite 4.0 1B-1.33
Qwen3 1.7B-1.37
Gemma 3 270M-1.39
MiniCPM-V 4.6 1.3B-1.39
Apertus 8B Instruct-1.40
Tiny Aya Global-1.40
Llama 3.2 1B-1.40
Granite 4.0 350M-1.41
Gemma 3n E2B-1.42
Qwen3 0.6B-1.42
LFM2.5-VL-1.6B-1.43
LFM2 1.2B-1.43
Granite 4.0 H 350M-1.46
Qwen3 0.6B-1.46
Gemma 3 1B-1.47
Qwen3.5 0.8B-1.48
Mistral 7B-1.50
Qwen3.5 0.8B-1.53
Model / variantSoftMinZCost per task
GPT-5.6 Luna (low)0.26$0.0135
MiMo-V2.50.89$0.0137
GPT-5.6 Luna (medium)0.38$0.0173
GPT-5.6 Luna (Non-reasoning)-0.09$0.0174
Llama 4 Scout-0.65$0.0184
GPT-5.6 Luna (high)0.57$0.0329
DeepSeek V3 (Dec)-0.48$0.0355
Nemotron 3 Nano-0.15$0.0359
gpt-oss-120b (low)-0.34$0.0412
DeepSeek V4 Flash 0731 (max)0.61$0.0435
Hy30.72$0.0438
Ling 3.0 Flash0.62$0.0439
gpt-oss-20b (high)-0.34$0.0442
MiMo-V2.5-Pro1.15$0.0448
GPT-5 mini (high)0.57$0.0453
Gemini 3.1 Flash-Lite0.28$0.0461
GPT-5.6 Luna (xhigh)0.60$0.0470
DeepSeek V4 Pro (max)0.59$0.0554
DeepSeek V4 Pro (high)0.65$0.0557
Mistral Small 3.1-0.65$0.0578
Gemma 4 31B-0.06$0.0609
DeepSeek V3 0324-0.26$0.0628
Llama 4 Maverick-0.34$0.0648
DeepSeek V4 Flash (high)0.41$0.0662
Gemma 4 26B A4B-0.07$0.0668
GPT-5.6 Luna (max)0.64$0.0675
GPT-4.1 nano-0.65$0.0707
Nemotron 3.5 Lightning0.17$0.0841
DeepSeek V4 Flash (max)0.37$0.0845
Step 3.7 Flash0.27$0.0864
Inkling Small0.99$0.0894
Muse Glimmer (high)0.33$0.0959
Mercury 2-0.11$0.1013
MiniMax-M2.70.84$0.1031
gpt-oss-120b (high)0.05$0.1079
GPT-4.1 mini-0.37$0.1159
Mistral Small 40.06$0.1163
Kimi K2.50.80$0.1175
Qwen3 30B A3B 2507-0.24$0.1183
Mistral Large 3-0.25$0.1284
Gemini 3.5 Flash-Lite0.58$0.1292
GPT-5.6 Terra (low)0.61$0.1484
GPT-5.6 Terra (Non-reasoning)-0.04$0.1680
Ministral 3 3B-0.94$0.1780
Grok 4.3 (high)1.27$0.1823
GPT-5.6 Terra (medium)0.65$0.1936
Solar Pro 3-0.44$0.2029
GPT-5.4 nano (xhigh)0.73$0.2076
MiniMax-M31.15$0.2209
Ministral 3 8B-0.80$0.2217
Ministral 3 14B-0.68$0.2314
Mistral Medium 3.1-0.43$0.2419
Mistral Small 3.2-0.60$0.2532
Qwen3.5 122B A10B-0.08$0.2630
Claude 4.5 Haiku0.38$0.2658
Gemini 2.5 Pro0.35$0.2663
Qwen3 Next 80B A3B-0.09$0.2676
Grok Build 0.1 06160.55$0.2791
Gemini 3.7 Flash (low)1.05$0.2807
Qwen3.7 Plus1.29$0.2961
Qwen3 235B A22B 25070.01$0.2984
Celeris-1-0.68$0.3054
Solar Pro 41.03$0.3075
Kimi K2.7 Code0.82$0.3134
Qwen3.5 35B A3B-0.35$0.3170
Qwen3.5 122B A10B0.19$0.3213
GPT-5.5 (Non-reasoning)0.20$0.3376
GPT-5 (high)0.69$0.3388
Nemotron 3 Super0.15$0.3404
GPT-5.6 Terra (high)0.82$0.3574
DeepSeek V4 Pro 0813 (max)0.59$0.3662
GLM-5.11.02$0.3716
Kimi K3 (low)0.64$0.3735
Muse Spark 1.1 (xhigh)1.66$0.3738
Grok 4.3 (Non-reasoning)-0.18$0.3780
GPT-5.1 (high)1.06$0.3826
GPT-5.6 Sol (low)0.88$0.3943
GLM-4.60.02$0.3950
Qwen3.5 9B-0.10$0.4006
GPT-5.5 (low)0.79$0.4015
Qwen3.6 27B0.60$0.4021
GPT-5.6 Sol (Non-reasoning)0.23$0.4037
Qwen3.6 35B A3B0.48$0.4079
Gemini 3.1 Pro Preview1.77$0.4118
Qwen3 Coder Next-0.37$0.4379
Inkling0.87$0.4421
GLM-5.2 (max)1.65$0.4537
Kimi K2.61.35$0.4619
Gemini 3.7 Flash (medium)1.28$0.4684
Qwen3.5 397B A17B0.34$0.4697
GPT-5.6 Terra (xhigh)0.89$0.4965
Qwen3.6 Plus0.91$0.5053
GLM-4.70.29$0.5090
Magistral Small 1.2-0.51$0.5170
Nemotron 3 Ultra0.88$0.5244
Grok 4.5 (high)1.53$0.5272
DeepSeek R1 (Jan)-0.17$0.5301
Mistral Medium 3.50.12$0.5708
Qwen3.6 27B-0.02$0.5896
Claude 4.5 Sonnet0.68$0.6070
Muse Spark 1.2 (xhigh)1.83$0.6126
GPT-5.4 mini (xhigh)0.61$0.6193
GPT-5.6 Sol (medium)0.91$0.6346
Claude Opus 5 (low)1.42$0.6647
GPT-5.5 Instant (June 2026)0.44$0.6674
Qwen3.7 Max1.59$0.6966
Claude Sonnet 5 (Non-reasoning)0.60$0.7065
Gemini 3.7 Flash (high)1.47$0.7208
GPT-5.6 Terra (max)0.96$0.7633
GPT-5.5 (medium)0.92$0.7785
Gemini 3.6 Flash1.41$0.8421
GPT-5.6 Sol (high)0.92$0.9466
Claude Opus 5 (medium)1.56$1.1109
Gemini 3.5 Flash1.39$1.1397
GPT-5.5 (high)0.99$1.1888
Kimi K3 (max)1.73$1.3209
GPT-5.6 Sol (xhigh)0.90$1.3706
Magistral Medium 1.20.16$1.4169
Grok 4.6 (high)1.76$1.4213
Claude Sonnet 4.6 (max)1.05$1.4340
GPT-5.4 (xhigh)0.86$1.6368
GPT-5.5 (xhigh)1.00$1.6981
Qwen3.8 2.4T A95B1.60$1.9605
Claude Opus 5 (high)1.61$1.9678
Qwen3.8 Max1.59$1.9863
GPT-5.6 Sol (max)0.92$2.0597
Claude Sonnet 5 (max)1.64$2.6557
Claude Opus 5 (xhigh)1.65$2.9689
Claude Opus 4.8 (max)1.80$3.0712
Claude Opus 4.7 (max)1.61$3.4450
Claude Opus 5 (max)1.63$3.8996
Claude Fable 5 (with fallback)1.71$4.9043
Gemini 3.5 Flash (medium)1.33
Claude Opus 4.6 (max)1.32
Grok 4.20 0309 v21.17
Claude Opus 4.7 (Non-reasoning, high)1.15
Grok 4.20 03091.11
Motif 31.10
Claude Opus 4.51.08
GPT-5.2 (medium)1.05
Grok 4.3 (medium)1.03
GPT-5.2 (xhigh)1.03
Qwen3.6 Max Preview1.03
Solar Open2 250B1.02
GPT-5.2 Codex (xhigh)1.01
A.X-K20.97
Motif 3 (Beta)0.95
Muse Spark0.91
GPT-5.3 Codex (xhigh)0.90
GLM-50.88
GPT-5.4 (low)0.80
Gemini 3 Pro Preview (high)0.80
Grok 40.79
MiMo-V2-Pro0.78
GPT-5 Codex (high)0.75
GPT-5.5 Instant (May 2026)0.75
GPT-5.1 Codex (high)0.72
Gemini 3 Flash0.70
Grok 4.3 (low)0.62
MiMo-V2-Flash (Feb 2026)0.62
MiMo-V2-Omni-03270.62
Kimi K2.60.60
Gemini 3.5 Flash (minimal)0.59
GPT-5.4 nano0.57
GPT-5 mini (medium)0.57
DeepSeek V3.2 Speciale0.57
Kimi K2 Thinking0.57
GPT-5.1 Codex mini (high)0.56
GLM-5-Turbo0.56
MiMo-V2-Omni0.55
K-EXAONE 2.00.54
Claude Opus 4.6 (high)0.54
Agnes 2.5 Pro Alpha0.51
DeepSeek V3.20.50
Grok 4.1 Fast0.49
JT-4.1 Flash 236B A21B0.48
Claude 4 Sonnet0.48
Grok 4 Fast0.48
KAT-Coder-Pro V20.47
Claude Sonnet 4.6 (Non-reasoning)0.47
Claude Sonnet 4.6 (Non-reasoning, Low Effort)0.46
Gemini 3 Pro Preview (low)0.45
Claude Opus 4.50.45
GPT-5 (medium)0.44
MiniMax-M2.10.44
DeepSeek V3.1 Terminus0.41
GLM-5.10.41
GLM 5V Turbo0.41
G9v3-39A5B0.40
GPT-5 (low)0.40
LongCat 2.00.39
Qwen3.5 Omni Plus0.39
o30.38
Grok 3 mini Reasoning (high)0.35
Hy3-preview0.34
Nex-N2-Pro0.33
Command A+0.32
Kimi K2.50.32
Claude 3.7 Sonnet0.31
DeepSeek V3.2 Exp0.31
Ring-2.6-1T0.29
GPT-5.4 mini (medium)0.28
GLM-5.20.27
Gemini 3 Flash0.26
Claude 4.5 Sonnet0.25
Qwen3.5 397B A17B0.24
DeepSeek V3.10.24
Qwen3.5 27B0.24
Doubao Seed Code0.23
GPT-5.20.22
Gemma 4 31B0.21
MiniMax-M2.50.21
MiMo-V2-Flash0.21
o4-mini (high)0.20
GLM-4.50.20
DeepSeek R1 05280.19
GPT-5.4 (Non-reasoning)0.18
Gemini 2.5 Flash0.18
GLM-50.17
Step 3.5 Flash0.14
Qwen3 Max Thinking0.13
Step 3.5 Flash 26030.12
o10.11
Claude 4 Sonnet0.11
Qwen3.5 27B0.11
KAT-Coder-Pro V10.10
Qwen3.5 35B A3B0.10
Claude 4.5 Haiku0.09
JT-35B-Flash0.09
GPT-5 nano (medium)0.08
Gemini 2.5 Flash (Sep)0.06
Kimi K2 09050.06
Ling 3.0 Tiny0.05
GPT-5 nano (high)0.05
MiniMax-M20.04
Nova 2.0 Pro Preview (low)0.02
DeepSeek V4 Pro0.01
Qwen3 Coder 480B0.01
MiMo-V2.5-Pro0.00
Claude 3.7 Sonnet0.00
Kimi K2-0.00
Grok Code Fast 1-0.01
Hy3-preview-0.02
Qwen3 Max-0.03
Nova 2.0 Pro Preview (medium)-0.04
K-EXAONE-0.04
Gemma 4 12B-0.05
Cogito v2.1-0.05
Qwen3 VL 235B A22B-0.05
Qwen3 Max Thinking (Preview)-0.05
Qwen3 Max (Preview)-0.06
Trinity Large Thinking-0.07
DeepSeek V3.2-0.08
o3-mini (high)-0.09
DeepSeek V3.1-0.09
North Mini Code-0.10
Gemini 2.5 Flash (Sep)-0.10
DeepSeek V3.1 Terminus-0.10
GLM-4.6-0.10
Nova 2.0 Lite (high)-0.11
DeepSeek V3.2 Exp-0.11
Qwen3 235B 2507-0.11
EXAONE 4.5 33B-0.12
HyperNova 60B 2605-0.14
K2 Think V2-0.15
INTELLECT-3-0.16
Nova 2.0 Lite (medium)-0.17
Nemotron Cascade 2 30B A3B-0.17
Ring-1T-0.18
Seed-OSS-36B-Instruct-0.18
GPT-5.1-0.19
Gemini 2.5 Flash-Lite (Sep)-0.19
Qwen3 VL 32B-0.20
Llama Nemotron Super 49B v1.5-0.22
Grok 3-0.22
GLM-4.7-0.22
GLM-4.6V-0.22
DeepSeek V4 Flash-0.23
Nova 2.0 Omni (low)-0.23
ERNIE 5.0 Thinking Preview-0.23
GPT-5 (minimal)-0.23
Gemini 2.5 Flash-Lite (Sep)-0.23
Grok 4.20 0309-0.23
Ling-2.6-1T-0.23
MiMo-V2-Flash-0.24
GPT-5.4 nano-0.24
Apriel-v1.6-15B-Thinker-0.25
GPT-4.1-0.25
Apriel-v1.5-15B-Thinker-0.26
MiniMax M1 80k-0.27
Gemma 4 26B A4B-0.30
Gemma 4 E4B-0.30
GPT-5 mini (minimal)-0.30
Gemini 2.5 Flash-0.30
GLM-4.5-Air-0.30
Grok 4.20 0309 v2-0.31
K2-V2 (high)-0.33
GPT-4o (Aug)-0.33
Mistral Medium 3-0.33
Ling-1T-0.33
Devstral 2-0.34
Nova 2.0 Omni (medium)-0.34
Nova 2.0 Pro Preview-0.34
Hermes 4 405B-0.34
Qwen3 Next 80B A3B-0.34
Gemma 4 12B (Non-reasoning)-0.35
Qwen3 VL 30B A3B-0.35
Gemini 2.5 Flash-Lite-0.36
Qwen3 VL 235B A22B-0.36
Hermes 4 405B-0.37
Nova 2.0 Lite (low)-0.38
GPT-5.4 mini-0.38
Grok 4.1 Fast-0.38
Magistral Medium 1-0.39
Llama 3.1 405B-0.40
Devstral Medium-0.40
Nova Premier-0.41
Qwen3.5 4B-0.42
gpt-oss-20b (low)-0.43
EXAONE 4.0 32B-0.43
K2-V2 (medium)-0.43
G9v3-3B-0.43
NVIDIA Nemotron Nano 9B V2-0.44
GLM-4.7-Flash-0.44
Qwen3 Coder 30B A3B-0.44
Qwen3 235B-0.44
Llama Nemotron Ultra-0.45
ERNIE 4.5 300B A47B-0.45
HyperCLOVA X SEED Think (32B)-0.45
Gemini 2.0 Flash-0.45
Qwen3 VL 32B-0.46
Qwen3 4B 2507-0.46
Mistral Small 4-0.47
Grok 4 Fast-0.47
DiffusionGemma 26B A4B-0.49
Solar Open 100B-0.49
Qwen3.5 9B-0.50
Qwen3.5 Omni Flash-0.51
GLM-4.6V-0.51
Ling-flash-2.0-0.51
NVIDIA Nemotron Nano 12B v2 VL-0.52
Hermes 4 70B-0.52
Qwen3 VL 30B A3B-0.52
Devstral Small 2-0.53
K-EXAONE-0.53
Ring-flash-2.0-0.53
Motif-2-12.7B-0.54
GPT-4o (Nov)-0.54
Claude 3.5 Haiku-0.54
NVIDIA Nemotron Nano 9B V2-0.55
Qwen3 30B A3B 2507-0.55
Falcon-H1R-7B-0.56
Qwen3 32B-0.56
Nemotron 3 Nano Omni 30B A3B-0.57
Olmo 3.1 32B Think-0.57
K2-V2 (low)-0.57
Mi:dm K 2.5 Pro-0.57
Nanbeige4.1-3B-0.57
Llama Nemotron Super 49B v1.5-0.58
Mi:dm K 2.5 Pro Preview-0.58
Magistral Small 1-0.59
Qwen3 VL 8B-0.59
Ling 2.6 Flash-0.59
Gemma 4 E2B-0.60
JT-MINI-0.60
Qwen3 14B-0.60
Command A-0.61
Mistral Large 2 (Nov)-0.62
GLM-4.5V-0.62
Nova 2.0 Lite-0.63
Qwen3 Omni 30B A3B-0.63
LongCat Flash Lite-0.64
Qwen3 30B-0.64
Nova 2.0 Omni-0.65
Step3 VL 10B-0.65
EXAONE 4.0 32B-0.66
Nova Pro-0.66
Qwen2.5 72B-0.66
DeepSeek R1 0528 Qwen3 8B-0.67
Solar Pro 2-0.67
Llama 3.3 Nemotron Super 49B-0.67
Olmo 3 32B Think-0.67
GPT-5 nano (minimal)-0.68
Devstral Small (May)-0.69
Qwen3 235B-0.69
Devstral Small-0.70
Llama 3.3 Nemotron Super 49B-0.70
Qwen3.5 4B-0.71
Gemini 2.5 Flash-Lite-0.71
Llama 3.3 70B-0.72
Tri-21B-think Preview-0.73
DeepSeek R1 Distill Llama 70B-0.74
Llama 3.1 Nemotron 70B-0.75
Qwen3 4B 2507-0.75
Sarvam 105B (high)-0.76
Llama 3.1 70B-0.76
Tri-21B-Think-0.77
Phi-4-0.78
Hermes 4 70B-0.78
Qwen3.6 35B A3B-0.80
Solar Pro 2-0.80
GLM-4.5V-0.80
Gemma 4 E2B-0.81
Nova Lite-0.81
Qwen3 30B-0.81
Gemma 4 E4B-0.82
Granite 4.1 30B-0.82
Qwen3 VL 8B-0.82
Qwen3 8B-0.82
Qwen3 VL 4B-0.82
Nemotron 3 Nano-0.83
Nemotron 3 Nano 4B-0.83
Reka Flash 3-0.83
GLM-4.7-Flash-0.83
Olmo 3.1 32B Instruct-0.85
Claude 3 Haiku-0.85
Granite 4.0 H Small-0.86
LFM2.5-8B-A1B-0.87
Olmo 3 7B Think-0.87
Granite 4.1 8B-0.88
Qwen3 14B-0.88
Qwen3 Omni 30B A3B-0.90
NVIDIA Nemotron Nano 12B v2 VL-0.91
Jamba 1.7 Large-0.91
Sarvam 30B (high)-0.93
Sarvam M-0.96
Gemma 3 27B-0.97
Ling-mini-2.0-0.97
LFM2 24B A2B-0.98
Llama 3 70B-0.98
Llama 3.1 8B-0.99
Nova Micro-1.00
Qwen3.5 2B-1.01
Phi-4 Mini-1.05
Exaone 4.0 1.2B-1.06
Qwen3 8B-1.07
Qwen3 VL 4B-1.07
Molmo2-8B-1.08
Jamba Reasoning 3B-1.09
MiniCPM5-1B-1.11
Olmo 3 7B-1.13
Gemma 3 12B-1.13
Granite 4.0 Micro-1.17
Exaone 4.0 1.2B-1.17
Qwen3.5 2B-1.17
Llama 3.2 11B (Vision)-1.18
MiniCPM5-1B-1.19
Granite 4.1 3B-1.21
Jamba 1.7 Mini-1.22
Llama 3 8B-1.22
Granite 3.3 8B-1.23
Granite 4.0 H 1B-1.24
LFM2 8B A1B-1.25
LFM2 2.6B-1.25
Qwen3 1.7B-1.25
Apertus 70B Instruct-1.26
LFM2.5-1.2B-Instruct-1.30
Gemma 3n E4B-1.31
Gemma 3 4B-1.32
LFM2.5-1.2B-Thinking-1.33
Granite 4.0 1B-1.33
Qwen3 1.7B-1.37
Gemma 3 270M-1.39
MiniCPM-V 4.6 1.3B-1.39
Apertus 8B Instruct-1.40
Tiny Aya Global-1.40
Llama 3.2 1B-1.40
Granite 4.0 350M-1.41
Gemma 3n E2B-1.42
Qwen3 0.6B-1.42
LFM2.5-VL-1.6B-1.43
LFM2 1.2B-1.43
Granite 4.0 H 350M-1.46
Qwen3 0.6B-1.46
Gemma 3 1B-1.47
Qwen3.5 0.8B-1.48
Mistral 7B-1.50
Qwen3.5 0.8B-1.53

What SoftMinZ measures

Nine public benchmarks are scored per model. Each benchmark score is first converted to a z-score across every model measured on that benchmark, so all benchmarks contribute in comparable units regardless of their raw scale:

zb = (xb − μb) / σb   (μb, σb are the mean and sample standard deviation across all measured models)

The index is then the negative natural logarithm of the mean of e−z over the model's benchmarks:

SoftMinZ = −ln [ (1/n) · Σb exp(−zb) ]

This is a smooth soft-minimum of the z-scores: it always sits between the model's worst z-score and its mean z-score, and it slides toward the worst end as the profile becomes more uneven. The practical consequences:

A model is excluded from scoring entirely (no index, rather than a low one) if it is measured on fewer than 7 of the 9 battery benchmarks, or if either of two mandatory measurements is missing: CritPt (physics) and the non-hallucination rate. A model with no physics evaluation or no hallucination measurement is simply not evaluated here.

Why these nine benchmarks

The performance battery is chosen for relevance to scientific computing: science-adjacent reasoning and knowledge (GPQA Diamond, CritPt, Humanity's Last Exam, AA-Omniscience), the code-execution skills real computational work depends on (SciCode, LiveCodeBench, Terminal-Bench Hard), and the trust dimensions that decide whether a model's output can be believed without full re-verification (non-hallucination rate, AA-LCR long-context reasoning).

τ³-Banking is excluded from the performance calculation because it simulates fintech customer support — a domain-specific agent task whose skill profile says nothing about scientific capability. (It does remain in the cost calculation below, where it represents real agentic workflow spend.) GDPval-AA v2 is likewise excluded: its tasks are real-world professional deliverables (analyst memos, marketing plans and the like), not scientific computing, so it has no place in a performance battery for science either — though it too remains in the cost model as agentic workload.

How the cost is computed

The cost benchmarks are chosen for relevance to agentic workflow costs. Artificial Analysis publishes, for each benchmark in its Intelligence Index, a weighted cost per task. Dividing that by the benchmark's Intelligence Index weight recovers the unweighted per-benchmark cost Ci; the four benchmarks that represent agentic workload — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1 and AA-LCR — are then averaged weighted by task count:

Ctask = Σi Ci Ti / Σi Ti

Cheap high-volume knowledge benchmarks (AA-Omniscience, HLE, GPQA Diamond, CritPt, SciCode) are excluded from the cost average because they would otherwise dominate the task count while representing almost no real agentic spend. The result is a raw USD figure per agentic task — list prices only, no discounts. A handful of open-weight models are served by free hosting endpoints and report a $0 cost to AA; they are excluded from the cost axis (and the frontier) rather than treated as genuinely free products, but remain scored and ranked in the table.

The Artificial Analysis Intelligence Index, and why we don't use it

Artificial Analysis' own headline metric, the Intelligence Index v4.1.1, is a plain weighted average of normalised benchmark scores:

II = Σb wb · sb

with category weights Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, and per-benchmark weights GDPval-AA v2 20%, τ³-Banking 14%, Terminal-Bench v2.1 16%, SciCode 8%, HLE 12%, GPQA Diamond 6%, CritPt 6%, AA-Omniscience 12% (itself split: 8% accuracy + 4% non-hallucination) and AA-LCR 6%. GDPval-AA v2 Elo scores are folded in as clamp((Elo − 500) / 2000). (Formula per the AA methodology page.)

It is a useful headline, but for scientific work it has real flaws:

Why the GPT-5.6 lineup scores below its reputation

Across the 18 GPT-5.6 variants scored today, the median non-hallucination rate is 9.2% against a field mean of 24.5% (population σ = 20.8 points) — a z-score of roughly -0.74 on the trust benchmark alone. Frontier models sit far higher: the top of the table carries non-hallucination rates in the 45–70% range. Because SoftMinZ is a soft-minimum, that weak trust score receives the largest Boltzmann weight in the average and drags every GPT-5.6 variant down even where GPQA, HLE and SciCode are strong — the best-placed variant ranks #51 of 478.

Reproducibility

Every scraped score is cross-validated against the JSON-LD metadata on its source page; a benchmark whose parsed values disagree is rejected rather than published. Intelligence Index weights and task counts used in the cost model are parsed live from the AA methodology page on every run (today: live (AA methodology page, parsed this run)); if that page is unreachable, a frozen snapshot of the weights is used and noted here. Today's scrape: agentic cost per task: 132 models; artificial-analysis-intelligence-index: 153 models (validated 20); artificial-analysis-long-context-reasoning: 500 models (validated 20); critpt: 482 models (validated 20); gpqa-diamond: 575 models (validated 20); humanitys-last-exam: 567 models (validated 20); livecodebench: 343 models (validated 7); non-hallucination rate: 479 models; omniscience: 479 models (validated 20); scicode: 567 models (validated 20); terminalbench-hard: 432 models (validated 14). One script generates this page — see the repository.