Generated 2026-08-14 · 478 models scored from 579 measured by Artificial Analysis · raw cost data only
Each point is a model: vertical position is the SoftMinZ performance index, horizontal position the raw cost of one agentic task (log scale). The dashed line is the Pareto frontier — models for which no cheaper model is at least as good. Frontier members are colour- and shape-coded; everything else is grey.
Every scored model. Bold rows are Pareto frontier members. Use the two buttons to switch the ranking — they need no JavaScript. In cost view, models with no published cost data sort last.
| Model / variant | SoftMinZ | Cost per task |
|---|---|---|
| Muse Spark 1.2 (xhigh) | 1.83 | $0.6126 |
| Claude Opus 4.8 (max) | 1.80 | $3.0712 |
| Gemini 3.1 Pro Preview | 1.77 | $0.4118 |
| Grok 4.6 (high) | 1.76 | $1.4213 |
| Kimi K3 (max) | 1.73 | $1.3209 |
| Claude Fable 5 (with fallback) | 1.71 | $4.9043 |
| Muse Spark 1.1 (xhigh) | 1.66 | $0.3738 |
| Claude Opus 5 (xhigh) | 1.65 | $2.9689 |
| GLM-5.2 (max) | 1.65 | $0.4537 |
| Claude Sonnet 5 (max) | 1.64 | $2.6557 |
| Claude Opus 5 (max) | 1.63 | $3.8996 |
| Claude Opus 5 (high) | 1.61 | $1.9678 |
| Claude Opus 4.7 (max) | 1.61 | $3.4450 |
| Qwen3.8 2.4T A95B | 1.60 | $1.9605 |
| Qwen3.8 Max | 1.59 | $1.9863 |
| Qwen3.7 Max | 1.59 | $0.6966 |
| Claude Opus 5 (medium) | 1.56 | $1.1109 |
| Grok 4.5 (high) | 1.53 | $0.5272 |
| Gemini 3.7 Flash (high) | 1.47 | $0.7208 |
| Claude Opus 5 (low) | 1.42 | $0.6647 |
| Gemini 3.6 Flash | 1.41 | $0.8421 |
| Gemini 3.5 Flash | 1.39 | $1.1397 |
| Kimi K2.6 | 1.35 | $0.4619 |
| Gemini 3.5 Flash (medium) | 1.33 | — |
| Claude Opus 4.6 (max) | 1.32 | — |
| Qwen3.7 Plus | 1.29 | $0.2961 |
| Gemini 3.7 Flash (medium) | 1.28 | $0.4684 |
| Grok 4.3 (high) | 1.27 | $0.1823 |
| Grok 4.20 0309 v2 | 1.17 | — |
| MiniMax-M3 | 1.15 | $0.2209 |
| MiMo-V2.5-Pro | 1.15 | $0.0448 |
| Claude Opus 4.7 (Non-reasoning, high) | 1.15 | — |
| Grok 4.20 0309 | 1.11 | — |
| Motif 3 | 1.10 | — |
| Claude Opus 4.5 | 1.08 | — |
| GPT-5.1 (high) | 1.06 | $0.3826 |
| Gemini 3.7 Flash (low) | 1.05 | $0.2807 |
| GPT-5.2 (medium) | 1.05 | — |
| Claude Sonnet 4.6 (max) | 1.05 | $1.4340 |
| Grok 4.3 (medium) | 1.03 | — |
| Solar Pro 4 | 1.03 | $0.3075 |
| GPT-5.2 (xhigh) | 1.03 | — |
| Qwen3.6 Max Preview | 1.03 | — |
| Solar Open2 250B | 1.02 | — |
| GLM-5.1 | 1.02 | $0.3716 |
| GPT-5.2 Codex (xhigh) | 1.01 | — |
| GPT-5.5 (xhigh) | 1.00 | $1.6981 |
| Inkling Small | 0.99 | $0.0894 |
| GPT-5.5 (high) | 0.99 | $1.1888 |
| A.X-K2 | 0.97 | — |
| GPT-5.6 Terra (max) | 0.96 | $0.7633 |
| Motif 3 (Beta) | 0.95 | — |
| GPT-5.6 Sol (max) | 0.92 | $2.0597 |
| GPT-5.6 Sol (high) | 0.92 | $0.9466 |
| GPT-5.5 (medium) | 0.92 | $0.7785 |
| GPT-5.6 Sol (medium) | 0.91 | $0.6346 |
| Muse Spark | 0.91 | — |
| Qwen3.6 Plus | 0.91 | $0.5053 |
| GPT-5.3 Codex (xhigh) | 0.90 | — |
| GPT-5.6 Sol (xhigh) | 0.90 | $1.3706 |
| GPT-5.6 Terra (xhigh) | 0.89 | $0.4965 |
| MiMo-V2.5 | 0.89 | $0.0137 |
| GPT-5.6 Sol (low) | 0.88 | $0.3943 |
| Nemotron 3 Ultra | 0.88 | $0.5244 |
| GLM-5 | 0.88 | — |
| Inkling | 0.87 | $0.4421 |
| GPT-5.4 (xhigh) | 0.86 | $1.6368 |
| MiniMax-M2.7 | 0.84 | $0.1031 |
| GPT-5.6 Terra (high) | 0.82 | $0.3574 |
| Kimi K2.7 Code | 0.82 | $0.3134 |
| GPT-5.4 (low) | 0.80 | — |
| Gemini 3 Pro Preview (high) | 0.80 | — |
| Kimi K2.5 | 0.80 | $0.1175 |
| GPT-5.5 (low) | 0.79 | $0.4015 |
| Grok 4 | 0.79 | — |
| MiMo-V2-Pro | 0.78 | — |
| GPT-5 Codex (high) | 0.75 | — |
| GPT-5.5 Instant (May 2026) | 0.75 | — |
| GPT-5.4 nano (xhigh) | 0.73 | $0.2076 |
| Hy3 | 0.72 | $0.0438 |
| GPT-5.1 Codex (high) | 0.72 | — |
| Gemini 3 Flash | 0.70 | — |
| GPT-5 (high) | 0.69 | $0.3388 |
| Claude 4.5 Sonnet | 0.68 | $0.6070 |
| GPT-5.6 Terra (medium) | 0.65 | $0.1936 |
| DeepSeek V4 Pro (high) | 0.65 | $0.0557 |
| GPT-5.6 Luna (max) | 0.64 | $0.0675 |
| Kimi K3 (low) | 0.64 | $0.3735 |
| Grok 4.3 (low) | 0.62 | — |
| MiMo-V2-Flash (Feb 2026) | 0.62 | — |
| MiMo-V2-Omni-0327 | 0.62 | — |
| Ling 3.0 Flash | 0.62 | $0.0439 |
| GPT-5.6 Terra (low) | 0.61 | $0.1484 |
| DeepSeek V4 Flash 0731 (max) | 0.61 | $0.0435 |
| GPT-5.4 mini (xhigh) | 0.61 | $0.6193 |
| GPT-5.6 Luna (xhigh) | 0.60 | $0.0470 |
| Claude Sonnet 5 (Non-reasoning) | 0.60 | $0.7065 |
| Qwen3.6 27B | 0.60 | $0.4021 |
| Kimi K2.6 | 0.60 | — |
| Gemini 3.5 Flash (minimal) | 0.59 | — |
| DeepSeek V4 Pro 0813 (max) | 0.59 | $0.3662 |
| DeepSeek V4 Pro (max) | 0.59 | $0.0554 |
| Gemini 3.5 Flash-Lite | 0.58 | $0.1292 |
| GPT-5.6 Luna (high) | 0.57 | $0.0329 |
| GPT-5 mini (high) | 0.57 | $0.0453 |
| GPT-5.4 nano | 0.57 | — |
| GPT-5 mini (medium) | 0.57 | — |
| DeepSeek V3.2 Speciale | 0.57 | — |
| Kimi K2 Thinking | 0.57 | — |
| GPT-5.1 Codex mini (high) | 0.56 | — |
| GLM-5-Turbo | 0.56 | — |
| MiMo-V2-Omni | 0.55 | — |
| Grok Build 0.1 0616 | 0.55 | $0.2791 |
| K-EXAONE 2.0 | 0.54 | — |
| Claude Opus 4.6 (high) | 0.54 | — |
| Agnes 2.5 Pro Alpha | 0.51 | — |
| DeepSeek V3.2 | 0.50 | — |
| Grok 4.1 Fast | 0.49 | — |
| JT-4.1 Flash 236B A21B | 0.48 | — |
| Qwen3.6 35B A3B | 0.48 | $0.4079 |
| Claude 4 Sonnet | 0.48 | — |
| Grok 4 Fast | 0.48 | — |
| KAT-Coder-Pro V2 | 0.47 | — |
| Claude Sonnet 4.6 (Non-reasoning) | 0.47 | — |
| Claude Sonnet 4.6 (Non-reasoning, Low Effort) | 0.46 | — |
| Gemini 3 Pro Preview (low) | 0.45 | — |
| Claude Opus 4.5 | 0.45 | — |
| GPT-5.5 Instant (June 2026) | 0.44 | $0.6674 |
| GPT-5 (medium) | 0.44 | — |
| MiniMax-M2.1 | 0.44 | — |
| DeepSeek V4 Flash (high) | 0.41 | $0.0662 |
| DeepSeek V3.1 Terminus | 0.41 | — |
| GLM-5.1 | 0.41 | — |
| GLM 5V Turbo | 0.41 | — |
| G9v3-39A5B | 0.40 | — |
| GPT-5 (low) | 0.40 | — |
| LongCat 2.0 | 0.39 | — |
| Qwen3.5 Omni Plus | 0.39 | — |
| GPT-5.6 Luna (medium) | 0.38 | $0.0173 |
| o3 | 0.38 | — |
| Claude 4.5 Haiku | 0.38 | $0.2658 |
| DeepSeek V4 Flash (max) | 0.37 | $0.0845 |
| Gemini 2.5 Pro | 0.35 | $0.2663 |
| Grok 3 mini Reasoning (high) | 0.35 | — |
| Qwen3.5 397B A17B | 0.34 | $0.4697 |
| Hy3-preview | 0.34 | — |
| Muse Glimmer (high) | 0.33 | $0.0959 |
| Nex-N2-Pro | 0.33 | — |
| Command A+ | 0.32 | — |
| Kimi K2.5 | 0.32 | — |
| Claude 3.7 Sonnet | 0.31 | — |
| DeepSeek V3.2 Exp | 0.31 | — |
| Ring-2.6-1T | 0.29 | — |
| GLM-4.7 | 0.29 | $0.5090 |
| GPT-5.4 mini (medium) | 0.28 | — |
| Gemini 3.1 Flash-Lite | 0.28 | $0.0461 |
| Step 3.7 Flash | 0.27 | $0.0864 |
| GLM-5.2 | 0.27 | — |
| GPT-5.6 Luna (low) | 0.26 | $0.0135 |
| Gemini 3 Flash | 0.26 | — |
| Claude 4.5 Sonnet | 0.25 | — |
| Qwen3.5 397B A17B | 0.24 | — |
| DeepSeek V3.1 | 0.24 | — |
| Qwen3.5 27B | 0.24 | — |
| GPT-5.6 Sol (Non-reasoning) | 0.23 | $0.4037 |
| Doubao Seed Code | 0.23 | — |
| GPT-5.2 | 0.22 | — |
| Gemma 4 31B | 0.21 | — |
| MiniMax-M2.5 | 0.21 | — |
| MiMo-V2-Flash | 0.21 | — |
| GPT-5.5 (Non-reasoning) | 0.20 | $0.3376 |
| o4-mini (high) | 0.20 | — |
| GLM-4.5 | 0.20 | — |
| Qwen3.5 122B A10B | 0.19 | $0.3213 |
| DeepSeek R1 0528 | 0.19 | — |
| GPT-5.4 (Non-reasoning) | 0.18 | — |
| Gemini 2.5 Flash | 0.18 | — |
| Nemotron 3.5 Lightning | 0.17 | $0.0841 |
| GLM-5 | 0.17 | — |
| Magistral Medium 1.2 | 0.16 | $1.4169 |
| Nemotron 3 Super | 0.15 | $0.3404 |
| Step 3.5 Flash | 0.14 | — |
| Qwen3 Max Thinking | 0.13 | — |
| Mistral Medium 3.5 | 0.12 | $0.5708 |
| Step 3.5 Flash 2603 | 0.12 | — |
| o1 | 0.11 | — |
| Claude 4 Sonnet | 0.11 | — |
| Qwen3.5 27B | 0.11 | — |
| KAT-Coder-Pro V1 | 0.10 | — |
| Qwen3.5 35B A3B | 0.10 | — |
| Claude 4.5 Haiku | 0.09 | — |
| JT-35B-Flash | 0.09 | — |
| GPT-5 nano (medium) | 0.08 | — |
| Mistral Small 4 | 0.06 | $0.1163 |
| Gemini 2.5 Flash (Sep) | 0.06 | — |
| Kimi K2 0905 | 0.06 | — |
| gpt-oss-120b (high) | 0.05 | $0.1079 |
| Ling 3.0 Tiny | 0.05 | — |
| GPT-5 nano (high) | 0.05 | — |
| MiniMax-M2 | 0.04 | — |
| Nova 2.0 Pro Preview (low) | 0.02 | — |
| GLM-4.6 | 0.02 | $0.3950 |
| DeepSeek V4 Pro | 0.01 | — |
| Qwen3 Coder 480B | 0.01 | — |
| Qwen3 235B A22B 2507 | 0.01 | $0.2984 |
| MiMo-V2.5-Pro | 0.00 | — |
| Claude 3.7 Sonnet | 0.00 | — |
| Kimi K2 | -0.00 | — |
| Grok Code Fast 1 | -0.01 | — |
| Qwen3.6 27B | -0.02 | $0.5896 |
| Hy3-preview | -0.02 | — |
| Qwen3 Max | -0.03 | — |
| GPT-5.6 Terra (Non-reasoning) | -0.04 | $0.1680 |
| Nova 2.0 Pro Preview (medium) | -0.04 | — |
| K-EXAONE | -0.04 | — |
| Gemma 4 12B | -0.05 | — |
| Cogito v2.1 | -0.05 | — |
| Qwen3 VL 235B A22B | -0.05 | — |
| Qwen3 Max Thinking (Preview) | -0.05 | — |
| Gemma 4 31B | -0.06 | $0.0609 |
| Qwen3 Max (Preview) | -0.06 | — |
| Gemma 4 26B A4B | -0.07 | $0.0668 |
| Trinity Large Thinking | -0.07 | — |
| Qwen3.5 122B A10B | -0.08 | $0.2630 |
| DeepSeek V3.2 | -0.08 | — |
| GPT-5.6 Luna (Non-reasoning) | -0.09 | $0.0174 |
| Qwen3 Next 80B A3B | -0.09 | $0.2676 |
| o3-mini (high) | -0.09 | — |
| DeepSeek V3.1 | -0.09 | — |
| North Mini Code | -0.10 | — |
| Qwen3.5 9B | -0.10 | $0.4006 |
| Gemini 2.5 Flash (Sep) | -0.10 | — |
| DeepSeek V3.1 Terminus | -0.10 | — |
| GLM-4.6 | -0.10 | — |
| Nova 2.0 Lite (high) | -0.11 | — |
| Mercury 2 | -0.11 | $0.1013 |
| DeepSeek V3.2 Exp | -0.11 | — |
| Qwen3 235B 2507 | -0.11 | — |
| EXAONE 4.5 33B | -0.12 | — |
| HyperNova 60B 2605 | -0.14 | — |
| Nemotron 3 Nano | -0.15 | $0.0359 |
| K2 Think V2 | -0.15 | — |
| INTELLECT-3 | -0.16 | — |
| Nova 2.0 Lite (medium) | -0.17 | — |
| Nemotron Cascade 2 30B A3B | -0.17 | — |
| DeepSeek R1 (Jan) | -0.17 | $0.5301 |
| Grok 4.3 (Non-reasoning) | -0.18 | $0.3780 |
| Ring-1T | -0.18 | — |
| Seed-OSS-36B-Instruct | -0.18 | — |
| GPT-5.1 | -0.19 | — |
| Gemini 2.5 Flash-Lite (Sep) | -0.19 | — |
| Qwen3 VL 32B | -0.20 | — |
| Llama Nemotron Super 49B v1.5 | -0.22 | — |
| Grok 3 | -0.22 | — |
| GLM-4.7 | -0.22 | — |
| GLM-4.6V | -0.22 | — |
| DeepSeek V4 Flash | -0.23 | — |
| Nova 2.0 Omni (low) | -0.23 | — |
| ERNIE 5.0 Thinking Preview | -0.23 | — |
| GPT-5 (minimal) | -0.23 | — |
| Gemini 2.5 Flash-Lite (Sep) | -0.23 | — |
| Grok 4.20 0309 | -0.23 | — |
| Ling-2.6-1T | -0.23 | — |
| MiMo-V2-Flash | -0.24 | — |
| GPT-5.4 nano | -0.24 | — |
| Qwen3 30B A3B 2507 | -0.24 | $0.1183 |
| Mistral Large 3 | -0.25 | $0.1284 |
| Apriel-v1.6-15B-Thinker | -0.25 | — |
| GPT-4.1 | -0.25 | — |
| DeepSeek V3 0324 | -0.26 | $0.0628 |
| Apriel-v1.5-15B-Thinker | -0.26 | — |
| MiniMax M1 80k | -0.27 | — |
| Gemma 4 26B A4B | -0.30 | — |
| Gemma 4 E4B | -0.30 | — |
| GPT-5 mini (minimal) | -0.30 | — |
| Gemini 2.5 Flash | -0.30 | — |
| GLM-4.5-Air | -0.30 | — |
| Grok 4.20 0309 v2 | -0.31 | — |
| K2-V2 (high) | -0.33 | — |
| GPT-4o (Aug) | -0.33 | — |
| Mistral Medium 3 | -0.33 | — |
| Ling-1T | -0.33 | — |
| gpt-oss-20b (high) | -0.34 | $0.0442 |
| gpt-oss-120b (low) | -0.34 | $0.0412 |
| Llama 4 Maverick | -0.34 | $0.0648 |
| Devstral 2 | -0.34 | — |
| Nova 2.0 Omni (medium) | -0.34 | — |
| Nova 2.0 Pro Preview | -0.34 | — |
| Hermes 4 405B | -0.34 | — |
| Qwen3 Next 80B A3B | -0.34 | — |
| Gemma 4 12B (Non-reasoning) | -0.35 | — |
| Qwen3.5 35B A3B | -0.35 | $0.3170 |
| Qwen3 VL 30B A3B | -0.35 | — |
| Gemini 2.5 Flash-Lite | -0.36 | — |
| Qwen3 VL 235B A22B | -0.36 | — |
| Hermes 4 405B | -0.37 | — |
| Qwen3 Coder Next | -0.37 | $0.4379 |
| GPT-4.1 mini | -0.37 | $0.1159 |
| Nova 2.0 Lite (low) | -0.38 | — |
| GPT-5.4 mini | -0.38 | — |
| Grok 4.1 Fast | -0.38 | — |
| Magistral Medium 1 | -0.39 | — |
| Llama 3.1 405B | -0.40 | — |
| Devstral Medium | -0.40 | — |
| Nova Premier | -0.41 | — |
| Qwen3.5 4B | -0.42 | — |
| gpt-oss-20b (low) | -0.43 | — |
| EXAONE 4.0 32B | -0.43 | — |
| K2-V2 (medium) | -0.43 | — |
| G9v3-3B | -0.43 | — |
| Mistral Medium 3.1 | -0.43 | $0.2419 |
| Solar Pro 3 | -0.44 | $0.2029 |
| NVIDIA Nemotron Nano 9B V2 | -0.44 | — |
| GLM-4.7-Flash | -0.44 | — |
| Qwen3 Coder 30B A3B | -0.44 | — |
| Qwen3 235B | -0.44 | — |
| Llama Nemotron Ultra | -0.45 | — |
| ERNIE 4.5 300B A47B | -0.45 | — |
| HyperCLOVA X SEED Think (32B) | -0.45 | — |
| Gemini 2.0 Flash | -0.45 | — |
| Qwen3 VL 32B | -0.46 | — |
| Qwen3 4B 2507 | -0.46 | — |
| Mistral Small 4 | -0.47 | — |
| Grok 4 Fast | -0.47 | — |
| DeepSeek V3 (Dec) | -0.48 | $0.0355 |
| DiffusionGemma 26B A4B | -0.49 | — |
| Solar Open 100B | -0.49 | — |
| Qwen3.5 9B | -0.50 | — |
| Magistral Small 1.2 | -0.51 | $0.5170 |
| Qwen3.5 Omni Flash | -0.51 | — |
| GLM-4.6V | -0.51 | — |
| Ling-flash-2.0 | -0.51 | — |
| NVIDIA Nemotron Nano 12B v2 VL | -0.52 | — |
| Hermes 4 70B | -0.52 | — |
| Qwen3 VL 30B A3B | -0.52 | — |
| Devstral Small 2 | -0.53 | — |
| K-EXAONE | -0.53 | — |
| Ring-flash-2.0 | -0.53 | — |
| Motif-2-12.7B | -0.54 | — |
| GPT-4o (Nov) | -0.54 | — |
| Claude 3.5 Haiku | -0.54 | — |
| NVIDIA Nemotron Nano 9B V2 | -0.55 | — |
| Qwen3 30B A3B 2507 | -0.55 | — |
| Falcon-H1R-7B | -0.56 | — |
| Qwen3 32B | -0.56 | — |
| Nemotron 3 Nano Omni 30B A3B | -0.57 | — |
| Olmo 3.1 32B Think | -0.57 | — |
| K2-V2 (low) | -0.57 | — |
| Mi:dm K 2.5 Pro | -0.57 | — |
| Nanbeige4.1-3B | -0.57 | — |
| Llama Nemotron Super 49B v1.5 | -0.58 | — |
| Mi:dm K 2.5 Pro Preview | -0.58 | — |
| Magistral Small 1 | -0.59 | — |
| Qwen3 VL 8B | -0.59 | — |
| Ling 2.6 Flash | -0.59 | — |
| Gemma 4 E2B | -0.60 | — |
| JT-MINI | -0.60 | — |
| Mistral Small 3.2 | -0.60 | $0.2532 |
| Qwen3 14B | -0.60 | — |
| Command A | -0.61 | — |
| Mistral Large 2 (Nov) | -0.62 | — |
| GLM-4.5V | -0.62 | — |
| Nova 2.0 Lite | -0.63 | — |
| Qwen3 Omni 30B A3B | -0.63 | — |
| LongCat Flash Lite | -0.64 | — |
| Qwen3 30B | -0.64 | — |
| Llama 4 Scout | -0.65 | $0.0184 |
| Nova 2.0 Omni | -0.65 | — |
| Step3 VL 10B | -0.65 | — |
| GPT-4.1 nano | -0.65 | $0.0707 |
| Mistral Small 3.1 | -0.65 | $0.0578 |
| EXAONE 4.0 32B | -0.66 | — |
| Nova Pro | -0.66 | — |
| Qwen2.5 72B | -0.66 | — |
| DeepSeek R1 0528 Qwen3 8B | -0.67 | — |
| Solar Pro 2 | -0.67 | — |
| Llama 3.3 Nemotron Super 49B | -0.67 | — |
| Olmo 3 32B Think | -0.67 | — |
| Ministral 3 14B | -0.68 | $0.2314 |
| Celeris-1 | -0.68 | $0.3054 |
| GPT-5 nano (minimal) | -0.68 | — |
| Devstral Small (May) | -0.69 | — |
| Qwen3 235B | -0.69 | — |
| Devstral Small | -0.70 | — |
| Llama 3.3 Nemotron Super 49B | -0.70 | — |
| Qwen3.5 4B | -0.71 | — |
| Gemini 2.5 Flash-Lite | -0.71 | — |
| Llama 3.3 70B | -0.72 | — |
| Tri-21B-think Preview | -0.73 | — |
| DeepSeek R1 Distill Llama 70B | -0.74 | — |
| Llama 3.1 Nemotron 70B | -0.75 | — |
| Qwen3 4B 2507 | -0.75 | — |
| Sarvam 105B (high) | -0.76 | — |
| Llama 3.1 70B | -0.76 | — |
| Tri-21B-Think | -0.77 | — |
| Phi-4 | -0.78 | — |
| Hermes 4 70B | -0.78 | — |
| Ministral 3 8B | -0.80 | $0.2217 |
| Qwen3.6 35B A3B | -0.80 | — |
| Solar Pro 2 | -0.80 | — |
| GLM-4.5V | -0.80 | — |
| Gemma 4 E2B | -0.81 | — |
| Nova Lite | -0.81 | — |
| Qwen3 30B | -0.81 | — |
| Gemma 4 E4B | -0.82 | — |
| Granite 4.1 30B | -0.82 | — |
| Qwen3 VL 8B | -0.82 | — |
| Qwen3 8B | -0.82 | — |
| Qwen3 VL 4B | -0.82 | — |
| Nemotron 3 Nano | -0.83 | — |
| Nemotron 3 Nano 4B | -0.83 | — |
| Reka Flash 3 | -0.83 | — |
| GLM-4.7-Flash | -0.83 | — |
| Olmo 3.1 32B Instruct | -0.85 | — |
| Claude 3 Haiku | -0.85 | — |
| Granite 4.0 H Small | -0.86 | — |
| LFM2.5-8B-A1B | -0.87 | — |
| Olmo 3 7B Think | -0.87 | — |
| Granite 4.1 8B | -0.88 | — |
| Qwen3 14B | -0.88 | — |
| Qwen3 Omni 30B A3B | -0.90 | — |
| NVIDIA Nemotron Nano 12B v2 VL | -0.91 | — |
| Jamba 1.7 Large | -0.91 | — |
| Sarvam 30B (high) | -0.93 | — |
| Ministral 3 3B | -0.94 | $0.1780 |
| Sarvam M | -0.96 | — |
| Gemma 3 27B | -0.97 | — |
| Ling-mini-2.0 | -0.97 | — |
| LFM2 24B A2B | -0.98 | — |
| Llama 3 70B | -0.98 | — |
| Llama 3.1 8B | -0.99 | — |
| Nova Micro | -1.00 | — |
| Qwen3.5 2B | -1.01 | — |
| Phi-4 Mini | -1.05 | — |
| Exaone 4.0 1.2B | -1.06 | — |
| Qwen3 8B | -1.07 | — |
| Qwen3 VL 4B | -1.07 | — |
| Molmo2-8B | -1.08 | — |
| Jamba Reasoning 3B | -1.09 | — |
| MiniCPM5-1B | -1.11 | — |
| Olmo 3 7B | -1.13 | — |
| Gemma 3 12B | -1.13 | — |
| Granite 4.0 Micro | -1.17 | — |
| Exaone 4.0 1.2B | -1.17 | — |
| Qwen3.5 2B | -1.17 | — |
| Llama 3.2 11B (Vision) | -1.18 | — |
| MiniCPM5-1B | -1.19 | — |
| Granite 4.1 3B | -1.21 | — |
| Jamba 1.7 Mini | -1.22 | — |
| Llama 3 8B | -1.22 | — |
| Granite 3.3 8B | -1.23 | — |
| Granite 4.0 H 1B | -1.24 | — |
| LFM2 8B A1B | -1.25 | — |
| LFM2 2.6B | -1.25 | — |
| Qwen3 1.7B | -1.25 | — |
| Apertus 70B Instruct | -1.26 | — |
| LFM2.5-1.2B-Instruct | -1.30 | — |
| Gemma 3n E4B | -1.31 | — |
| Gemma 3 4B | -1.32 | — |
| LFM2.5-1.2B-Thinking | -1.33 | — |
| Granite 4.0 1B | -1.33 | — |
| Qwen3 1.7B | -1.37 | — |
| Gemma 3 270M | -1.39 | — |
| MiniCPM-V 4.6 1.3B | -1.39 | — |
| Apertus 8B Instruct | -1.40 | — |
| Tiny Aya Global | -1.40 | — |
| Llama 3.2 1B | -1.40 | — |
| Granite 4.0 350M | -1.41 | — |
| Gemma 3n E2B | -1.42 | — |
| Qwen3 0.6B | -1.42 | — |
| LFM2.5-VL-1.6B | -1.43 | — |
| LFM2 1.2B | -1.43 | — |
| Granite 4.0 H 350M | -1.46 | — |
| Qwen3 0.6B | -1.46 | — |
| Gemma 3 1B | -1.47 | — |
| Qwen3.5 0.8B | -1.48 | — |
| Mistral 7B | -1.50 | — |
| Qwen3.5 0.8B | -1.53 | — |
| Model / variant | SoftMinZ | Cost per task |
|---|---|---|
| GPT-5.6 Luna (low) | 0.26 | $0.0135 |
| MiMo-V2.5 | 0.89 | $0.0137 |
| GPT-5.6 Luna (medium) | 0.38 | $0.0173 |
| GPT-5.6 Luna (Non-reasoning) | -0.09 | $0.0174 |
| Llama 4 Scout | -0.65 | $0.0184 |
| GPT-5.6 Luna (high) | 0.57 | $0.0329 |
| DeepSeek V3 (Dec) | -0.48 | $0.0355 |
| Nemotron 3 Nano | -0.15 | $0.0359 |
| gpt-oss-120b (low) | -0.34 | $0.0412 |
| DeepSeek V4 Flash 0731 (max) | 0.61 | $0.0435 |
| Hy3 | 0.72 | $0.0438 |
| Ling 3.0 Flash | 0.62 | $0.0439 |
| gpt-oss-20b (high) | -0.34 | $0.0442 |
| MiMo-V2.5-Pro | 1.15 | $0.0448 |
| GPT-5 mini (high) | 0.57 | $0.0453 |
| Gemini 3.1 Flash-Lite | 0.28 | $0.0461 |
| GPT-5.6 Luna (xhigh) | 0.60 | $0.0470 |
| DeepSeek V4 Pro (max) | 0.59 | $0.0554 |
| DeepSeek V4 Pro (high) | 0.65 | $0.0557 |
| Mistral Small 3.1 | -0.65 | $0.0578 |
| Gemma 4 31B | -0.06 | $0.0609 |
| DeepSeek V3 0324 | -0.26 | $0.0628 |
| Llama 4 Maverick | -0.34 | $0.0648 |
| DeepSeek V4 Flash (high) | 0.41 | $0.0662 |
| Gemma 4 26B A4B | -0.07 | $0.0668 |
| GPT-5.6 Luna (max) | 0.64 | $0.0675 |
| GPT-4.1 nano | -0.65 | $0.0707 |
| Nemotron 3.5 Lightning | 0.17 | $0.0841 |
| DeepSeek V4 Flash (max) | 0.37 | $0.0845 |
| Step 3.7 Flash | 0.27 | $0.0864 |
| Inkling Small | 0.99 | $0.0894 |
| Muse Glimmer (high) | 0.33 | $0.0959 |
| Mercury 2 | -0.11 | $0.1013 |
| MiniMax-M2.7 | 0.84 | $0.1031 |
| gpt-oss-120b (high) | 0.05 | $0.1079 |
| GPT-4.1 mini | -0.37 | $0.1159 |
| Mistral Small 4 | 0.06 | $0.1163 |
| Kimi K2.5 | 0.80 | $0.1175 |
| Qwen3 30B A3B 2507 | -0.24 | $0.1183 |
| Mistral Large 3 | -0.25 | $0.1284 |
| Gemini 3.5 Flash-Lite | 0.58 | $0.1292 |
| GPT-5.6 Terra (low) | 0.61 | $0.1484 |
| GPT-5.6 Terra (Non-reasoning) | -0.04 | $0.1680 |
| Ministral 3 3B | -0.94 | $0.1780 |
| Grok 4.3 (high) | 1.27 | $0.1823 |
| GPT-5.6 Terra (medium) | 0.65 | $0.1936 |
| Solar Pro 3 | -0.44 | $0.2029 |
| GPT-5.4 nano (xhigh) | 0.73 | $0.2076 |
| MiniMax-M3 | 1.15 | $0.2209 |
| Ministral 3 8B | -0.80 | $0.2217 |
| Ministral 3 14B | -0.68 | $0.2314 |
| Mistral Medium 3.1 | -0.43 | $0.2419 |
| Mistral Small 3.2 | -0.60 | $0.2532 |
| Qwen3.5 122B A10B | -0.08 | $0.2630 |
| Claude 4.5 Haiku | 0.38 | $0.2658 |
| Gemini 2.5 Pro | 0.35 | $0.2663 |
| Qwen3 Next 80B A3B | -0.09 | $0.2676 |
| Grok Build 0.1 0616 | 0.55 | $0.2791 |
| Gemini 3.7 Flash (low) | 1.05 | $0.2807 |
| Qwen3.7 Plus | 1.29 | $0.2961 |
| Qwen3 235B A22B 2507 | 0.01 | $0.2984 |
| Celeris-1 | -0.68 | $0.3054 |
| Solar Pro 4 | 1.03 | $0.3075 |
| Kimi K2.7 Code | 0.82 | $0.3134 |
| Qwen3.5 35B A3B | -0.35 | $0.3170 |
| Qwen3.5 122B A10B | 0.19 | $0.3213 |
| GPT-5.5 (Non-reasoning) | 0.20 | $0.3376 |
| GPT-5 (high) | 0.69 | $0.3388 |
| Nemotron 3 Super | 0.15 | $0.3404 |
| GPT-5.6 Terra (high) | 0.82 | $0.3574 |
| DeepSeek V4 Pro 0813 (max) | 0.59 | $0.3662 |
| GLM-5.1 | 1.02 | $0.3716 |
| Kimi K3 (low) | 0.64 | $0.3735 |
| Muse Spark 1.1 (xhigh) | 1.66 | $0.3738 |
| Grok 4.3 (Non-reasoning) | -0.18 | $0.3780 |
| GPT-5.1 (high) | 1.06 | $0.3826 |
| GPT-5.6 Sol (low) | 0.88 | $0.3943 |
| GLM-4.6 | 0.02 | $0.3950 |
| Qwen3.5 9B | -0.10 | $0.4006 |
| GPT-5.5 (low) | 0.79 | $0.4015 |
| Qwen3.6 27B | 0.60 | $0.4021 |
| GPT-5.6 Sol (Non-reasoning) | 0.23 | $0.4037 |
| Qwen3.6 35B A3B | 0.48 | $0.4079 |
| Gemini 3.1 Pro Preview | 1.77 | $0.4118 |
| Qwen3 Coder Next | -0.37 | $0.4379 |
| Inkling | 0.87 | $0.4421 |
| GLM-5.2 (max) | 1.65 | $0.4537 |
| Kimi K2.6 | 1.35 | $0.4619 |
| Gemini 3.7 Flash (medium) | 1.28 | $0.4684 |
| Qwen3.5 397B A17B | 0.34 | $0.4697 |
| GPT-5.6 Terra (xhigh) | 0.89 | $0.4965 |
| Qwen3.6 Plus | 0.91 | $0.5053 |
| GLM-4.7 | 0.29 | $0.5090 |
| Magistral Small 1.2 | -0.51 | $0.5170 |
| Nemotron 3 Ultra | 0.88 | $0.5244 |
| Grok 4.5 (high) | 1.53 | $0.5272 |
| DeepSeek R1 (Jan) | -0.17 | $0.5301 |
| Mistral Medium 3.5 | 0.12 | $0.5708 |
| Qwen3.6 27B | -0.02 | $0.5896 |
| Claude 4.5 Sonnet | 0.68 | $0.6070 |
| Muse Spark 1.2 (xhigh) | 1.83 | $0.6126 |
| GPT-5.4 mini (xhigh) | 0.61 | $0.6193 |
| GPT-5.6 Sol (medium) | 0.91 | $0.6346 |
| Claude Opus 5 (low) | 1.42 | $0.6647 |
| GPT-5.5 Instant (June 2026) | 0.44 | $0.6674 |
| Qwen3.7 Max | 1.59 | $0.6966 |
| Claude Sonnet 5 (Non-reasoning) | 0.60 | $0.7065 |
| Gemini 3.7 Flash (high) | 1.47 | $0.7208 |
| GPT-5.6 Terra (max) | 0.96 | $0.7633 |
| GPT-5.5 (medium) | 0.92 | $0.7785 |
| Gemini 3.6 Flash | 1.41 | $0.8421 |
| GPT-5.6 Sol (high) | 0.92 | $0.9466 |
| Claude Opus 5 (medium) | 1.56 | $1.1109 |
| Gemini 3.5 Flash | 1.39 | $1.1397 |
| GPT-5.5 (high) | 0.99 | $1.1888 |
| Kimi K3 (max) | 1.73 | $1.3209 |
| GPT-5.6 Sol (xhigh) | 0.90 | $1.3706 |
| Magistral Medium 1.2 | 0.16 | $1.4169 |
| Grok 4.6 (high) | 1.76 | $1.4213 |
| Claude Sonnet 4.6 (max) | 1.05 | $1.4340 |
| GPT-5.4 (xhigh) | 0.86 | $1.6368 |
| GPT-5.5 (xhigh) | 1.00 | $1.6981 |
| Qwen3.8 2.4T A95B | 1.60 | $1.9605 |
| Claude Opus 5 (high) | 1.61 | $1.9678 |
| Qwen3.8 Max | 1.59 | $1.9863 |
| GPT-5.6 Sol (max) | 0.92 | $2.0597 |
| Claude Sonnet 5 (max) | 1.64 | $2.6557 |
| Claude Opus 5 (xhigh) | 1.65 | $2.9689 |
| Claude Opus 4.8 (max) | 1.80 | $3.0712 |
| Claude Opus 4.7 (max) | 1.61 | $3.4450 |
| Claude Opus 5 (max) | 1.63 | $3.8996 |
| Claude Fable 5 (with fallback) | 1.71 | $4.9043 |
| Gemini 3.5 Flash (medium) | 1.33 | — |
| Claude Opus 4.6 (max) | 1.32 | — |
| Grok 4.20 0309 v2 | 1.17 | — |
| Claude Opus 4.7 (Non-reasoning, high) | 1.15 | — |
| Grok 4.20 0309 | 1.11 | — |
| Motif 3 | 1.10 | — |
| Claude Opus 4.5 | 1.08 | — |
| GPT-5.2 (medium) | 1.05 | — |
| Grok 4.3 (medium) | 1.03 | — |
| GPT-5.2 (xhigh) | 1.03 | — |
| Qwen3.6 Max Preview | 1.03 | — |
| Solar Open2 250B | 1.02 | — |
| GPT-5.2 Codex (xhigh) | 1.01 | — |
| A.X-K2 | 0.97 | — |
| Motif 3 (Beta) | 0.95 | — |
| Muse Spark | 0.91 | — |
| GPT-5.3 Codex (xhigh) | 0.90 | — |
| GLM-5 | 0.88 | — |
| GPT-5.4 (low) | 0.80 | — |
| Gemini 3 Pro Preview (high) | 0.80 | — |
| Grok 4 | 0.79 | — |
| MiMo-V2-Pro | 0.78 | — |
| GPT-5 Codex (high) | 0.75 | — |
| GPT-5.5 Instant (May 2026) | 0.75 | — |
| GPT-5.1 Codex (high) | 0.72 | — |
| Gemini 3 Flash | 0.70 | — |
| Grok 4.3 (low) | 0.62 | — |
| MiMo-V2-Flash (Feb 2026) | 0.62 | — |
| MiMo-V2-Omni-0327 | 0.62 | — |
| Kimi K2.6 | 0.60 | — |
| Gemini 3.5 Flash (minimal) | 0.59 | — |
| GPT-5.4 nano | 0.57 | — |
| GPT-5 mini (medium) | 0.57 | — |
| DeepSeek V3.2 Speciale | 0.57 | — |
| Kimi K2 Thinking | 0.57 | — |
| GPT-5.1 Codex mini (high) | 0.56 | — |
| GLM-5-Turbo | 0.56 | — |
| MiMo-V2-Omni | 0.55 | — |
| K-EXAONE 2.0 | 0.54 | — |
| Claude Opus 4.6 (high) | 0.54 | — |
| Agnes 2.5 Pro Alpha | 0.51 | — |
| DeepSeek V3.2 | 0.50 | — |
| Grok 4.1 Fast | 0.49 | — |
| JT-4.1 Flash 236B A21B | 0.48 | — |
| Claude 4 Sonnet | 0.48 | — |
| Grok 4 Fast | 0.48 | — |
| KAT-Coder-Pro V2 | 0.47 | — |
| Claude Sonnet 4.6 (Non-reasoning) | 0.47 | — |
| Claude Sonnet 4.6 (Non-reasoning, Low Effort) | 0.46 | — |
| Gemini 3 Pro Preview (low) | 0.45 | — |
| Claude Opus 4.5 | 0.45 | — |
| GPT-5 (medium) | 0.44 | — |
| MiniMax-M2.1 | 0.44 | — |
| DeepSeek V3.1 Terminus | 0.41 | — |
| GLM-5.1 | 0.41 | — |
| GLM 5V Turbo | 0.41 | — |
| G9v3-39A5B | 0.40 | — |
| GPT-5 (low) | 0.40 | — |
| LongCat 2.0 | 0.39 | — |
| Qwen3.5 Omni Plus | 0.39 | — |
| o3 | 0.38 | — |
| Grok 3 mini Reasoning (high) | 0.35 | — |
| Hy3-preview | 0.34 | — |
| Nex-N2-Pro | 0.33 | — |
| Command A+ | 0.32 | — |
| Kimi K2.5 | 0.32 | — |
| Claude 3.7 Sonnet | 0.31 | — |
| DeepSeek V3.2 Exp | 0.31 | — |
| Ring-2.6-1T | 0.29 | — |
| GPT-5.4 mini (medium) | 0.28 | — |
| GLM-5.2 | 0.27 | — |
| Gemini 3 Flash | 0.26 | — |
| Claude 4.5 Sonnet | 0.25 | — |
| Qwen3.5 397B A17B | 0.24 | — |
| DeepSeek V3.1 | 0.24 | — |
| Qwen3.5 27B | 0.24 | — |
| Doubao Seed Code | 0.23 | — |
| GPT-5.2 | 0.22 | — |
| Gemma 4 31B | 0.21 | — |
| MiniMax-M2.5 | 0.21 | — |
| MiMo-V2-Flash | 0.21 | — |
| o4-mini (high) | 0.20 | — |
| GLM-4.5 | 0.20 | — |
| DeepSeek R1 0528 | 0.19 | — |
| GPT-5.4 (Non-reasoning) | 0.18 | — |
| Gemini 2.5 Flash | 0.18 | — |
| GLM-5 | 0.17 | — |
| Step 3.5 Flash | 0.14 | — |
| Qwen3 Max Thinking | 0.13 | — |
| Step 3.5 Flash 2603 | 0.12 | — |
| o1 | 0.11 | — |
| Claude 4 Sonnet | 0.11 | — |
| Qwen3.5 27B | 0.11 | — |
| KAT-Coder-Pro V1 | 0.10 | — |
| Qwen3.5 35B A3B | 0.10 | — |
| Claude 4.5 Haiku | 0.09 | — |
| JT-35B-Flash | 0.09 | — |
| GPT-5 nano (medium) | 0.08 | — |
| Gemini 2.5 Flash (Sep) | 0.06 | — |
| Kimi K2 0905 | 0.06 | — |
| Ling 3.0 Tiny | 0.05 | — |
| GPT-5 nano (high) | 0.05 | — |
| MiniMax-M2 | 0.04 | — |
| Nova 2.0 Pro Preview (low) | 0.02 | — |
| DeepSeek V4 Pro | 0.01 | — |
| Qwen3 Coder 480B | 0.01 | — |
| MiMo-V2.5-Pro | 0.00 | — |
| Claude 3.7 Sonnet | 0.00 | — |
| Kimi K2 | -0.00 | — |
| Grok Code Fast 1 | -0.01 | — |
| Hy3-preview | -0.02 | — |
| Qwen3 Max | -0.03 | — |
| Nova 2.0 Pro Preview (medium) | -0.04 | — |
| K-EXAONE | -0.04 | — |
| Gemma 4 12B | -0.05 | — |
| Cogito v2.1 | -0.05 | — |
| Qwen3 VL 235B A22B | -0.05 | — |
| Qwen3 Max Thinking (Preview) | -0.05 | — |
| Qwen3 Max (Preview) | -0.06 | — |
| Trinity Large Thinking | -0.07 | — |
| DeepSeek V3.2 | -0.08 | — |
| o3-mini (high) | -0.09 | — |
| DeepSeek V3.1 | -0.09 | — |
| North Mini Code | -0.10 | — |
| Gemini 2.5 Flash (Sep) | -0.10 | — |
| DeepSeek V3.1 Terminus | -0.10 | — |
| GLM-4.6 | -0.10 | — |
| Nova 2.0 Lite (high) | -0.11 | — |
| DeepSeek V3.2 Exp | -0.11 | — |
| Qwen3 235B 2507 | -0.11 | — |
| EXAONE 4.5 33B | -0.12 | — |
| HyperNova 60B 2605 | -0.14 | — |
| K2 Think V2 | -0.15 | — |
| INTELLECT-3 | -0.16 | — |
| Nova 2.0 Lite (medium) | -0.17 | — |
| Nemotron Cascade 2 30B A3B | -0.17 | — |
| Ring-1T | -0.18 | — |
| Seed-OSS-36B-Instruct | -0.18 | — |
| GPT-5.1 | -0.19 | — |
| Gemini 2.5 Flash-Lite (Sep) | -0.19 | — |
| Qwen3 VL 32B | -0.20 | — |
| Llama Nemotron Super 49B v1.5 | -0.22 | — |
| Grok 3 | -0.22 | — |
| GLM-4.7 | -0.22 | — |
| GLM-4.6V | -0.22 | — |
| DeepSeek V4 Flash | -0.23 | — |
| Nova 2.0 Omni (low) | -0.23 | — |
| ERNIE 5.0 Thinking Preview | -0.23 | — |
| GPT-5 (minimal) | -0.23 | — |
| Gemini 2.5 Flash-Lite (Sep) | -0.23 | — |
| Grok 4.20 0309 | -0.23 | — |
| Ling-2.6-1T | -0.23 | — |
| MiMo-V2-Flash | -0.24 | — |
| GPT-5.4 nano | -0.24 | — |
| Apriel-v1.6-15B-Thinker | -0.25 | — |
| GPT-4.1 | -0.25 | — |
| Apriel-v1.5-15B-Thinker | -0.26 | — |
| MiniMax M1 80k | -0.27 | — |
| Gemma 4 26B A4B | -0.30 | — |
| Gemma 4 E4B | -0.30 | — |
| GPT-5 mini (minimal) | -0.30 | — |
| Gemini 2.5 Flash | -0.30 | — |
| GLM-4.5-Air | -0.30 | — |
| Grok 4.20 0309 v2 | -0.31 | — |
| K2-V2 (high) | -0.33 | — |
| GPT-4o (Aug) | -0.33 | — |
| Mistral Medium 3 | -0.33 | — |
| Ling-1T | -0.33 | — |
| Devstral 2 | -0.34 | — |
| Nova 2.0 Omni (medium) | -0.34 | — |
| Nova 2.0 Pro Preview | -0.34 | — |
| Hermes 4 405B | -0.34 | — |
| Qwen3 Next 80B A3B | -0.34 | — |
| Gemma 4 12B (Non-reasoning) | -0.35 | — |
| Qwen3 VL 30B A3B | -0.35 | — |
| Gemini 2.5 Flash-Lite | -0.36 | — |
| Qwen3 VL 235B A22B | -0.36 | — |
| Hermes 4 405B | -0.37 | — |
| Nova 2.0 Lite (low) | -0.38 | — |
| GPT-5.4 mini | -0.38 | — |
| Grok 4.1 Fast | -0.38 | — |
| Magistral Medium 1 | -0.39 | — |
| Llama 3.1 405B | -0.40 | — |
| Devstral Medium | -0.40 | — |
| Nova Premier | -0.41 | — |
| Qwen3.5 4B | -0.42 | — |
| gpt-oss-20b (low) | -0.43 | — |
| EXAONE 4.0 32B | -0.43 | — |
| K2-V2 (medium) | -0.43 | — |
| G9v3-3B | -0.43 | — |
| NVIDIA Nemotron Nano 9B V2 | -0.44 | — |
| GLM-4.7-Flash | -0.44 | — |
| Qwen3 Coder 30B A3B | -0.44 | — |
| Qwen3 235B | -0.44 | — |
| Llama Nemotron Ultra | -0.45 | — |
| ERNIE 4.5 300B A47B | -0.45 | — |
| HyperCLOVA X SEED Think (32B) | -0.45 | — |
| Gemini 2.0 Flash | -0.45 | — |
| Qwen3 VL 32B | -0.46 | — |
| Qwen3 4B 2507 | -0.46 | — |
| Mistral Small 4 | -0.47 | — |
| Grok 4 Fast | -0.47 | — |
| DiffusionGemma 26B A4B | -0.49 | — |
| Solar Open 100B | -0.49 | — |
| Qwen3.5 9B | -0.50 | — |
| Qwen3.5 Omni Flash | -0.51 | — |
| GLM-4.6V | -0.51 | — |
| Ling-flash-2.0 | -0.51 | — |
| NVIDIA Nemotron Nano 12B v2 VL | -0.52 | — |
| Hermes 4 70B | -0.52 | — |
| Qwen3 VL 30B A3B | -0.52 | — |
| Devstral Small 2 | -0.53 | — |
| K-EXAONE | -0.53 | — |
| Ring-flash-2.0 | -0.53 | — |
| Motif-2-12.7B | -0.54 | — |
| GPT-4o (Nov) | -0.54 | — |
| Claude 3.5 Haiku | -0.54 | — |
| NVIDIA Nemotron Nano 9B V2 | -0.55 | — |
| Qwen3 30B A3B 2507 | -0.55 | — |
| Falcon-H1R-7B | -0.56 | — |
| Qwen3 32B | -0.56 | — |
| Nemotron 3 Nano Omni 30B A3B | -0.57 | — |
| Olmo 3.1 32B Think | -0.57 | — |
| K2-V2 (low) | -0.57 | — |
| Mi:dm K 2.5 Pro | -0.57 | — |
| Nanbeige4.1-3B | -0.57 | — |
| Llama Nemotron Super 49B v1.5 | -0.58 | — |
| Mi:dm K 2.5 Pro Preview | -0.58 | — |
| Magistral Small 1 | -0.59 | — |
| Qwen3 VL 8B | -0.59 | — |
| Ling 2.6 Flash | -0.59 | — |
| Gemma 4 E2B | -0.60 | — |
| JT-MINI | -0.60 | — |
| Qwen3 14B | -0.60 | — |
| Command A | -0.61 | — |
| Mistral Large 2 (Nov) | -0.62 | — |
| GLM-4.5V | -0.62 | — |
| Nova 2.0 Lite | -0.63 | — |
| Qwen3 Omni 30B A3B | -0.63 | — |
| LongCat Flash Lite | -0.64 | — |
| Qwen3 30B | -0.64 | — |
| Nova 2.0 Omni | -0.65 | — |
| Step3 VL 10B | -0.65 | — |
| EXAONE 4.0 32B | -0.66 | — |
| Nova Pro | -0.66 | — |
| Qwen2.5 72B | -0.66 | — |
| DeepSeek R1 0528 Qwen3 8B | -0.67 | — |
| Solar Pro 2 | -0.67 | — |
| Llama 3.3 Nemotron Super 49B | -0.67 | — |
| Olmo 3 32B Think | -0.67 | — |
| GPT-5 nano (minimal) | -0.68 | — |
| Devstral Small (May) | -0.69 | — |
| Qwen3 235B | -0.69 | — |
| Devstral Small | -0.70 | — |
| Llama 3.3 Nemotron Super 49B | -0.70 | — |
| Qwen3.5 4B | -0.71 | — |
| Gemini 2.5 Flash-Lite | -0.71 | — |
| Llama 3.3 70B | -0.72 | — |
| Tri-21B-think Preview | -0.73 | — |
| DeepSeek R1 Distill Llama 70B | -0.74 | — |
| Llama 3.1 Nemotron 70B | -0.75 | — |
| Qwen3 4B 2507 | -0.75 | — |
| Sarvam 105B (high) | -0.76 | — |
| Llama 3.1 70B | -0.76 | — |
| Tri-21B-Think | -0.77 | — |
| Phi-4 | -0.78 | — |
| Hermes 4 70B | -0.78 | — |
| Qwen3.6 35B A3B | -0.80 | — |
| Solar Pro 2 | -0.80 | — |
| GLM-4.5V | -0.80 | — |
| Gemma 4 E2B | -0.81 | — |
| Nova Lite | -0.81 | — |
| Qwen3 30B | -0.81 | — |
| Gemma 4 E4B | -0.82 | — |
| Granite 4.1 30B | -0.82 | — |
| Qwen3 VL 8B | -0.82 | — |
| Qwen3 8B | -0.82 | — |
| Qwen3 VL 4B | -0.82 | — |
| Nemotron 3 Nano | -0.83 | — |
| Nemotron 3 Nano 4B | -0.83 | — |
| Reka Flash 3 | -0.83 | — |
| GLM-4.7-Flash | -0.83 | — |
| Olmo 3.1 32B Instruct | -0.85 | — |
| Claude 3 Haiku | -0.85 | — |
| Granite 4.0 H Small | -0.86 | — |
| LFM2.5-8B-A1B | -0.87 | — |
| Olmo 3 7B Think | -0.87 | — |
| Granite 4.1 8B | -0.88 | — |
| Qwen3 14B | -0.88 | — |
| Qwen3 Omni 30B A3B | -0.90 | — |
| NVIDIA Nemotron Nano 12B v2 VL | -0.91 | — |
| Jamba 1.7 Large | -0.91 | — |
| Sarvam 30B (high) | -0.93 | — |
| Sarvam M | -0.96 | — |
| Gemma 3 27B | -0.97 | — |
| Ling-mini-2.0 | -0.97 | — |
| LFM2 24B A2B | -0.98 | — |
| Llama 3 70B | -0.98 | — |
| Llama 3.1 8B | -0.99 | — |
| Nova Micro | -1.00 | — |
| Qwen3.5 2B | -1.01 | — |
| Phi-4 Mini | -1.05 | — |
| Exaone 4.0 1.2B | -1.06 | — |
| Qwen3 8B | -1.07 | — |
| Qwen3 VL 4B | -1.07 | — |
| Molmo2-8B | -1.08 | — |
| Jamba Reasoning 3B | -1.09 | — |
| MiniCPM5-1B | -1.11 | — |
| Olmo 3 7B | -1.13 | — |
| Gemma 3 12B | -1.13 | — |
| Granite 4.0 Micro | -1.17 | — |
| Exaone 4.0 1.2B | -1.17 | — |
| Qwen3.5 2B | -1.17 | — |
| Llama 3.2 11B (Vision) | -1.18 | — |
| MiniCPM5-1B | -1.19 | — |
| Granite 4.1 3B | -1.21 | — |
| Jamba 1.7 Mini | -1.22 | — |
| Llama 3 8B | -1.22 | — |
| Granite 3.3 8B | -1.23 | — |
| Granite 4.0 H 1B | -1.24 | — |
| LFM2 8B A1B | -1.25 | — |
| LFM2 2.6B | -1.25 | — |
| Qwen3 1.7B | -1.25 | — |
| Apertus 70B Instruct | -1.26 | — |
| LFM2.5-1.2B-Instruct | -1.30 | — |
| Gemma 3n E4B | -1.31 | — |
| Gemma 3 4B | -1.32 | — |
| LFM2.5-1.2B-Thinking | -1.33 | — |
| Granite 4.0 1B | -1.33 | — |
| Qwen3 1.7B | -1.37 | — |
| Gemma 3 270M | -1.39 | — |
| MiniCPM-V 4.6 1.3B | -1.39 | — |
| Apertus 8B Instruct | -1.40 | — |
| Tiny Aya Global | -1.40 | — |
| Llama 3.2 1B | -1.40 | — |
| Granite 4.0 350M | -1.41 | — |
| Gemma 3n E2B | -1.42 | — |
| Qwen3 0.6B | -1.42 | — |
| LFM2.5-VL-1.6B | -1.43 | — |
| LFM2 1.2B | -1.43 | — |
| Granite 4.0 H 350M | -1.46 | — |
| Qwen3 0.6B | -1.46 | — |
| Gemma 3 1B | -1.47 | — |
| Qwen3.5 0.8B | -1.48 | — |
| Mistral 7B | -1.50 | — |
| Qwen3.5 0.8B | -1.53 | — |
Nine public benchmarks are scored per model. Each benchmark score is first converted to a z-score across every model measured on that benchmark, so all benchmarks contribute in comparable units regardless of their raw scale:
The index is then the negative natural logarithm of the mean of e−z over the model's benchmarks:
This is a smooth soft-minimum of the z-scores: it always sits between the model's worst z-score and its mean z-score, and it slides toward the worst end as the profile becomes more uneven. The practical consequences:
A model is excluded from scoring entirely (no index, rather than a low one) if it is measured on fewer than 7 of the 9 battery benchmarks, or if either of two mandatory measurements is missing: CritPt (physics) and the non-hallucination rate. A model with no physics evaluation or no hallucination measurement is simply not evaluated here.
The performance battery is chosen for relevance to scientific computing: science-adjacent reasoning and knowledge (GPQA Diamond, CritPt, Humanity's Last Exam, AA-Omniscience), the code-execution skills real computational work depends on (SciCode, LiveCodeBench, Terminal-Bench Hard), and the trust dimensions that decide whether a model's output can be believed without full re-verification (non-hallucination rate, AA-LCR long-context reasoning).
τ³-Banking is excluded from the performance calculation because it simulates fintech customer support — a domain-specific agent task whose skill profile says nothing about scientific capability. (It does remain in the cost calculation below, where it represents real agentic workflow spend.) GDPval-AA v2 is likewise excluded: its tasks are real-world professional deliverables (analyst memos, marketing plans and the like), not scientific computing, so it has no place in a performance battery for science either — though it too remains in the cost model as agentic workload.
The cost benchmarks are chosen for relevance to agentic workflow costs. Artificial Analysis publishes, for each benchmark in its Intelligence Index, a weighted cost per task. Dividing that by the benchmark's Intelligence Index weight recovers the unweighted per-benchmark cost Ci; the four benchmarks that represent agentic workload — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1 and AA-LCR — are then averaged weighted by task count:
Cheap high-volume knowledge benchmarks (AA-Omniscience, HLE, GPQA Diamond, CritPt, SciCode) are excluded from the cost average because they would otherwise dominate the task count while representing almost no real agentic spend. The result is a raw USD figure per agentic task — list prices only, no discounts. A handful of open-weight models are served by free hosting endpoints and report a $0 cost to AA; they are excluded from the cost axis (and the frontier) rather than treated as genuinely free products, but remain scored and ranked in the table.
Artificial Analysis' own headline metric, the Intelligence Index v4.1.1, is a plain weighted average of normalised benchmark scores:
with category weights Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%, and per-benchmark weights GDPval-AA v2 20%, τ³-Banking 14%, Terminal-Bench v2.1 16%, SciCode 8%, HLE 12%, GPQA Diamond 6%, CritPt 6%, AA-Omniscience 12% (itself split: 8% accuracy + 4% non-hallucination) and AA-LCR 6%. GDPval-AA v2 Elo scores are folded in as clamp((Elo − 500) / 2000). (Formula per the AA methodology page.)
It is a useful headline, but for scientific work it has real flaws:
Across the 18 GPT-5.6 variants scored today, the median non-hallucination rate is 9.2% against a field mean of 24.5% (population σ = 20.8 points) — a z-score of roughly -0.74 on the trust benchmark alone. Frontier models sit far higher: the top of the table carries non-hallucination rates in the 45–70% range. Because SoftMinZ is a soft-minimum, that weak trust score receives the largest Boltzmann weight in the average and drags every GPT-5.6 variant down even where GPQA, HLE and SciCode are strong — the best-placed variant ranks #51 of 478.
Every scraped score is cross-validated against the JSON-LD metadata on its source page; a benchmark whose parsed values disagree is rejected rather than published. Intelligence Index weights and task counts used in the cost model are parsed live from the AA methodology page on every run (today: live (AA methodology page, parsed this run)); if that page is unreachable, a frozen snapshot of the weights is used and noted here. Today's scrape: agentic cost per task: 132 models; artificial-analysis-intelligence-index: 153 models (validated 20); artificial-analysis-long-context-reasoning: 500 models (validated 20); critpt: 482 models (validated 20); gpqa-diamond: 575 models (validated 20); humanitys-last-exam: 567 models (validated 20); livecodebench: 343 models (validated 7); non-hallucination rate: 479 models; omniscience: 479 models (validated 20); scicode: 567 models (validated 20); terminalbench-hard: 432 models (validated 14). One script generates this page — see the repository.