| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | claude-opus-5-5 | max | rsc | 57.6 | $5.98 | – | 61.4% | 59.6% | $4.00 | $20.00 | 95.3 |
| Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Anthropic | claude-opus-5-5 | xhigh | rsc | 56.0 | $3.46 | – | 57.5% | 59.6% | $4.00 | $20.00 | 81.5 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | claude-sonnet-5-5 | max | rsc | 56.0 | $7.60 | – | 55.0% | 63.6% | $2.00 | $10.00 | 145.1 |
| Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Anthropic | claude-opus-5-5 | high | rsc | 53.6 | $1.82 | – | 55.6% | 56.6% | $4.00 | $20.00 | 76.1 |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | claude-fable-5-1 | max | rsc | 53.4 | $7.63 | 93.7% | 59.1% | 52.0% | $10.00 | $50.00 | 68.9 |
| Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Anthropic | claude-fable-5-1 | xhigh | rsc | 53.2 | $5.98 | 93.4% | 58.7% | 55.1% | $10.00 | $50.00 | 60.8 |
| GPT-6 Astra (max) | OpenAI | gpt-6-astra | max | rsc | 52.7 | $3.26 | 96.1% | 54.7% | 59.1% | $10.00 | $50.00 | 61.7 |
| GPT-6 Astra (xhigh) | OpenAI | gpt-6-astra | xhigh | rsc | 52.4 | $2.31 | 96.3% | 54.6% | 59.6% | $10.00 | $50.00 | 55.6 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Anthropic | claude-sonnet-5-5 | xhigh | rsc | 51.9 | $2.74 | – | 50.0% | 57.1% | $2.00 | $10.00 | 99.2 |
| Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Anthropic | claude-opus-5-5 | medium | rsc | 51.2 | $1.34 | – | 54.7% | 52.5% | $4.00 | $20.00 | 74.1 |
| Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Anthropic | claude-fable-5-1 | high | rsc | 51.2 | $3.91 | 90.6% | 55.9% | 52.0% | $10.00 | $50.00 | 50.4 |
| GPT-6 Astra (high) | OpenAI | gpt-6-astra | high | rsc | 50.9 | $1.73 | 94.9% | 53.1% | 54.0% | $10.00 | $50.00 | 54.9 |
| Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | claude-opus-5 | max | rsc | 50.8 | $5.86 | 93.2% | 54.9% | 49.0% | $5.00 | $25.00 | – |
| Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | claude-opus-5 | xhigh | rsc | 49.7 | $4.88 | 93.7% | 54.4% | 46.5% | $5.00 | $25.00 | – |
| Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | claude-fable-5 | max | rsc | 49.6 | $8.75 | 92.6% | 55.5% | 42.4% | $10.00 | $50.00 | – |
| GPT-6 Astra (medium) | OpenAI | gpt-6-astra | medium | rsc | 49.6 | $1.54 | 93.9% | 52.7% | 49.5% | $10.00 | $50.00 | 53.3 |
| Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Anthropic | claude-fable-5-1 | medium | rsc | 48.9 | $2.98 | 88.6% | 53.8% | 44.9% | $10.00 | $50.00 | 50.8 |
| Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | claude-opus-5 | high | rsc | 48.1 | $3.61 | 93.7% | 52.8% | 46.0% | $5.00 | $25.00 | – |
| Muse Spark 1.3 (max) | Meta | muse-spark-1-3 | max | rsc | 48.1 | $1.60 | 93.5% | 48.7% | 33.3% | $1.25 | $4.25 | 189.5 |
| GPT-6 Sol (max) | OpenAI | gpt-6-sol | max | rsc | 47.5 | $1.05 | – | 47.9% | 43.9% | $2.00 | $10.00 | 85.4 |
| GPT-5.6 Sol (max) | OpenAI | gpt-5-6-sol | max | rsc | 47.0 | $1.99 | 94.1% | 49.5% | 39.9% | $4.00 | $20.00 | – |
| Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Anthropic | claude-fable-5-1 | low | rsc | 46.8 | $2.37 | 88.1% | 48.9% | 40.4% | $10.00 | $50.00 | 50.2 |
| Claude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Anthropic | claude-sonnet-5-5 | high | rsc | 46.7 | $1.08 | – | 45.8% | 43.9% | $2.00 | $10.00 | 95.4 |
| Grok 4.7 (xhigh) | SpaceXAI | grok-4-7 | xhigh | rsc | 46.4 | $3.74 | – | 43.1% | 25.8% | $2.00 | $6.00 | 72.3 |
| Grok 4.7 (high) | SpaceXAI | grok-4-7 | high | rsc | 46.3 | $2.73 | – | 42.3% | 24.7% | $2.00 | $6.00 | 79.2 |
| MiMo-V2.6-Pro | Xiaomi | mimo-v2-6-pro | | rsc | 46.3 | $0.13 | – | 49.4% | 34.8% | $0.43 | $0.87 | 44.3 |
| GPT-6 Astra (low) | OpenAI | gpt-6-astra | low | rsc | 45.8 | $0.82 | 93.1% | 49.2% | 41.9% | $10.00 | $50.00 | 51.2 |
| Qwen3.8 Max (0902) | Alibaba | qwen3-8-max-0902 | 0902 | rsc | 45.4 | $5.41 | 92.8% | 43.1% | 38.9% | $2.00 | $6.00 | 38.3 |
| Muse Spark 1.3 (xhigh) | Meta | muse-spark-1-3 | xhigh | rsc | 45.1 | $1.37 | 94.1% | 47.5% | 16.7% | $1.25 | $4.25 | 148.6 |
| Claude Opus 5 (Adaptive Reasoning, Medium Effort) | Anthropic | claude-opus-5 | medium | rsc | 44.8 | $2.19 | 91.9% | 51.3% | 34.3% | $5.00 | $25.00 | – |
| GLM-5.3 (max) | Z AI | glm-5-3 | max | rsc | 44.8 | $2.01 | 91.7% | 42.3% | 41.9% | $1.40 | $4.40 | 74.8 |
| Grok 4.6 (high) | SpaceXAI | grok-4-6 | high | rsc | 44.3 | $1.86 | 94.9% | 42.9% | 21.2% | $2.00 | $6.00 | 72.3 |
| Grok 4.6 (xhigh) | SpaceXAI | grok-4-6 | xhigh | rsc | 44.2 | $2.32 | 93.5% | 44.1% | 17.2% | $2.00 | $6.00 | 70.3 |
| GPT-6 Sol (xhigh) | OpenAI | gpt-6-sol | xhigh | rsc | 44.1 | $0.52 | – | 46.3% | 30.3% | $2.00 | $10.00 | 77.3 |
| GPT-5.6 Sol (xhigh) | OpenAI | gpt-5-6-sol | xhigh | rsc | 44.0 | $1.18 | 93.1% | 47.3% | 24.7% | $4.00 | $20.00 | – |
| Step 5 Preview | StepFun | step-5-preview | | rsc | 43.7 | $0.72 | – | 46.5% | 33.3% | $1.00 | $2.70 | 84.1 |
| Kimi K3 (max) | Kimi | kimi-k3 | max | rsc | 43.6 | $2.00 | 93.5% | 46.9% | 12.6% | $3.00 | $15.00 | – |
| Grok 4.6 (medium) | SpaceXAI | grok-4-6 | medium | rsc | 42.8 | $1.50 | 93.5% | 42.1% | 13.1% | $2.00 | $6.00 | 60.5 |
| GPT-6 Sol (high) | OpenAI | gpt-6-sol | high | rsc | 42.8 | $0.38 | – | 44.1% | 26.3% | $2.00 | $10.00 | 74.4 |
| GPT-5.6 Sol (high) | OpenAI | gpt-5-6-sol | high | rsc | 42.3 | $0.81 | 92.8% | 46.0% | 20.7% | $4.00 | $20.00 | – |
| Claude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Anthropic | claude-opus-5-5 | low | rsc | 42.3 | $0.55 | – | 48.3% | 31.3% | $4.00 | $20.00 | 75.3 |
| GPT-5.6 Terra (max) | OpenAI | gpt-5-6-terra | max | rsc | 42.1 | $1.40 | 92.5% | 42.9% | 35.4% | $2.00 | $12.00 | 109.0 |
| GLM 5.3 Flash | Z AI | glm-5-3-flash | max | rsc | 41.8 | $0.25 | 91.2% | 39.9% | 32.8% | $0.15 | $0.50 | 49.2 |
| Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | claude-opus-4-8 | max | rsc | 41.8 | $4.08 | 92.0% | 48.7% | 21.7% | $5.00 | $25.00 | – |
| Gemini 3.8 Flash (high) | Google | gemini-3-8-flash | high | rsc | 40.9 | $1.24 | 95.3% | 47.8% | 19.7% | $0.75 | $3.75 | 238.1 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Anthropic | claude-sonnet-5-5 | medium | rsc | 40.7 | $0.59 | – | 39.8% | 29.8% | $2.00 | $10.00 | 87.7 |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | claude-opus-4-7 | max | rsc | 40.7 | – | 91.4% | 42.3% | – | $5.00 | $25.00 | – |
| Qwen3.8 Max | Alibaba | qwen3-8-max | 0803 | rsc | 40.2 | $2.67 | 92.7% | 43.0% | 18.7% | $2.00 | $6.00 | – |
| Qwen3.8 2.4T A95B | Alibaba | qwen3-8-2-4t-a95b | | rsc | 39.9 | $2.16 | 93.5% | 42.4% | 11.1% | $2.00 | $6.00 | 38.5 |
| Qwen3.8-Flash-Next | Alibaba | qwen3-8-flash-next | | rsc | 39.8 | $0.37 | 92.3% | 38.0% | 25.3% | $0.15 | $0.47 | 53.0 |
| GPT-6 Sol (medium) | OpenAI | gpt-6-sol | medium | rsc | 39.8 | $0.25 | – | 41.0% | 18.7% | $2.00 | $10.00 | – |
| Gemini 3.8 Flash (medium) | Google | gemini-3-8-flash | medium | rsc | 39.8 | $0.93 | 93.5% | 42.1% | 19.7% | $0.75 | $3.75 | – |
| Gemini 3.7 Flash (medium) | Google | gemini-3-7-flash | medium | rsc | 39.6 | – | 92.1% | 39.0% | – | $0.75 | $3.75 | – |
| Muse Spark 1.2 (xhigh) | Meta | muse-spark-1-2 | xhigh | rsc | 39.6 | $0.97 | 90.4% | 45.5% | 7.1% | $1.25 | $4.25 | – |
| DeepSeek V4.1 Flash (Reasoning, Max Effort) | DeepSeek | deepseek-v4-1-flash | max | rsc | 39.5 | $0.27 | – | 39.2% | 26.8% | $0.30 | $1.20 | 217.6 |
| Claude Opus 5 (Adaptive Reasoning, Low Effort) | Anthropic | claude-opus-5 | low | rsc | 39.4 | $1.10 | 88.9% | 43.4% | 26.3% | $5.00 | $25.00 | – |
| GPT-5.6 Sol (medium) | OpenAI | gpt-5-6-sol | medium | rsc | 39.2 | $0.50 | 92.6% | 42.2% | 14.6% | $4.00 | $20.00 | – |
| Gemini 3.7 Flash (high) | Google | gemini-3-7-flash | high | rsc | 39.1 | $0.93 | 94.5% | 47.9% | 13.6% | $0.75 | $3.75 | – |
| GPT-5.4 (xhigh) | OpenAI | gpt-5-4 | xhigh | rsc | 39.0 | – | 92.0% | 43.7% | – | $2.50 | $15.00 | – |
| Grok 4.5 (high) | SpaceXAI | grok-4-5 | high | rsc | 38.8 | $1.04 | 93.1% | 42.7% | 10.6% | $2.00 | $6.00 | – |
| GPT-5.5 (xhigh) | OpenAI | gpt-5-5 | xhigh | rsc | 38.4 | $2.63 | 93.5% | 45.8% | 14.6% | $5.00 | $30.00 | – |
| Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | Anthropic | claude-sonnet-5 | max | rsc | 38.2 | $5.09 | 91.1% | 41.3% | 14.1% | $2.00 | $10.00 | 81.7 |
| GPT-5.6 Terra (xhigh) | OpenAI | gpt-5-6-terra | xhigh | rsc | 38.0 | $0.63 | 90.8% | 41.9% | 10.1% | $2.00 | $12.00 | 100.4 |
| MiMo-V2.6-Flash | Xiaomi | mimo-v2-6-flash | | rsc | 37.9 | $0.062 | – | 35.1% | 22.7% | $0.14 | $0.28 | 42.3 |
| GPT-5.6 Luna (max) | OpenAI | gpt-5-6-luna | max | rsc | 37.3 | $0.18 | 91.1% | 39.5% | 11.6% | $0.20 | $1.20 | – |
| GPT-6 Luna (max) | OpenAI | gpt-6-luna | max | rsc | 37.3 | $0.068 | – | 38.5% | 12.6% | $0.10 | $0.50 | 143.3 |
| GPT-5.5 (high) | OpenAI | gpt-5-5 | high | rsc | 37.0 | $1.54 | 93.2% | 45.0% | 9.1% | $5.00 | $30.00 | – |
| Gemini 3.7 Flash (low) | Google | gemini-3-7-flash | low | rsc | 36.9 | – | 90.1% | 35.1% | – | $0.75 | $3.75 | – |
| DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | DeepSeek | deepseek-v4-pro | max | rsc | 36.0 | $0.67 | 92.8% | 41.0% | 14.1% | $1.32 | $3.96 | 90.7 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Anthropic | claude-sonnet-5-5 | low | rsc | 35.8 | $0.41 | – | 36.2% | 20.7% | $2.00 | $10.00 | 89.5 |
| Grok 4.6 (low) | SpaceXAI | grok-4-6 | low | rsc | 35.1 | $0.48 | 87.9% | 27.6% | 3.0% | $2.00 | $6.00 | 62.5 |
| DeepSeek V4 Flash Vision (Reasoning, Max Effort) | DeepSeek | deepseek-v4-flash-vision | max | rsc | 34.8 | $0.31 | 91.3% | 34.5% | 12.1% | $0.44 | $1.32 | 227.5 |
| GPT-5.6 Luna (xhigh) | OpenAI | gpt-5-6-luna | xhigh | rsc | 34.6 | $0.085 | 89.5% | 37.0% | 3.5% | $0.20 | $1.20 | – |
| Kimi K3 (low) | Kimi | kimi-k3 | low | rsc | 34.5 | – | 84.2% | 25.0% | 12.6% | $3.00 | $15.00 | – |
| Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | claude-sonnet-5 | xhigh | rsc | 34.4 | $2.87 | – | 39.0% | 7.1% | $2.00 | $10.00 | 73.2 |
| DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | DeepSeek | deepseek-v4-flash | max | rsc | 34.3 | $0.22 | 90.8% | 38.6% | 12.1% | $0.44 | $1.32 | – |
| GLM-5.3 (low) | Z AI | glm-5-3 | low | rsc | 34.3 | $0.85 | – | 36.6% | 34.8% | $1.40 | $4.40 | 78.9 |
| GPT-5.6 Terra (high) | OpenAI | gpt-5-6-terra | high | rsc | 34.2 | $0.34 | 89.6% | 38.5% | 1.5% | $2.00 | $12.00 | 94.6 |
| Gemini 3.6 Flash (high) | Google | gemini-3-6-flash | high | rsc | 34.0 | $0.93 | 92.8% | 40.8% | 7.1% | $0.75 | $3.75 | – |
| GPT-6 Sol (low) | OpenAI | gpt-6-sol | low | rsc | 33.9 | $0.13 | – | 34.9% | 9.1% | $2.00 | $10.00 | 79.7 |
| GPT-6 Luna (xhigh) | OpenAI | gpt-6-luna | xhigh | rsc | 33.9 | $0.042 | – | 34.3% | 8.1% | $0.10 | $0.50 | 145.6 |
| GPT-5.5 (medium) | OpenAI | gpt-5-5 | medium | rsc | 33.8 | $0.90 | 92.6% | 42.4% | 5.1% | $5.00 | $30.00 | – |
| Muse Spark 1.1 (xhigh) | Meta | muse-spark-1-1 | xhigh | rsc | 33.7 | $1.38 | 89.8% | 46.2% | 6.1% | $1.25 | $4.25 | – |
| GLM-5.2 (max) | Z AI | glm-5-2 | max | rsc | 33.7 | $1.47 | 89.5% | 41.1% | 1.0% | $1.40 | $4.40 | – |
| Qwen3.8 27B (xhigh) | Alibaba | qwen3-8-27b | xhigh | rsc | 33.7 | $1.01 | 90.5% | 33.9% | 5.6% | $0.50 | $3.00 | 45.5 |
| Gemini 3.5 Flash (medium) | Google | gemini-3-5-flash | medium | rsc | 33.6 | – | 92.1% | 41.3% | – | $1.50 | $9.00 | – |
| Motif 3 | Motif Technologies | motif-3 | | rsc | 33.6 | – | 83.4% | 37.0% | – | – | – | – |
| GPT-5.6 Sol (low) | OpenAI | gpt-5-6-sol | low | rsc | 33.5 | $0.26 | 89.8% | 39.4% | 1.0% | $4.00 | $20.00 | – |
| Gemini 3.8 Flash (low) | Google | gemini-3-8-flash | low | rsc | 33.5 | – | 92.0% | 37.1% | 10.1% | $0.75 | $3.75 | – |
| Gemini 3.5 Flash (high) | Google | gemini-3-5-flash | high | rsc | 32.6 | $1.56 | 92.2% | 42.7% | 6.6% | $1.50 | $9.00 | – |
| GPT-5.3 Codex (xhigh) | OpenAI | gpt-5-3-codex | xhigh | rsc | 32.5 | – | 91.5% | 42.5% | – | $1.75 | $14.00 | 113.8 |
| Motif 3 (Beta) | Motif Technologies | motif-0714 | beta | rsc | 32.3 | – | 86.9% | 40.4% | – | – | – | – |
| GPT-6 Luna (high) | OpenAI | gpt-6-luna | high | rsc | 32.1 | $0.029 | – | 32.9% | 4.5% | $0.10 | $0.50 | 144.8 |
| GPT-5.6 Luna (high) | OpenAI | gpt-5-6-luna | high | rsc | 32.1 | $0.044 | 89.2% | 33.4% | 2.5% | $0.20 | $1.20 | – |
| Claude Opus 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | claude-opus-4-6 | max | rsc | 31.9 | – | 89.6% | 39.9% | – | $5.00 | $25.00 | – |
| Claude Sonnet 5 (Adaptive Reasoning, High Effort) | Anthropic | claude-sonnet-5 | high | rsc | 31.7 | $1.79 | – | 35.7% | 5.1% | $2.00 | $10.00 | 62.8 |
| Muse Spark | Meta | muse-spark | | rsc | 31.3 | – | 88.4% | 40.7% | – | – | – | – |
| Claude Opus 4.7 (Non-reasoning, High Effort) | Anthropic | claude-opus-4-7 | high | rsc | 30.9 | – | 88.5% | 33.3% | – | $5.00 | $25.00 | – |
| GPT-5.5 (low) | OpenAI | gpt-5-5 | low | rsc | 30.7 | – | 91.0% | 32.7% | – | $5.00 | $30.00 | – |
| K2 Horizon 375B A23B | Institute of Foundation Models | k2-horizon-375b-a23b | | rsc | 30.5 | – | 87.3% | 32.0% | 1.5% | – | – | – |
| DeepSeek V4 Pro 0424 (Reasoning, Max Effort) | DeepSeek | deepseek-v4-pro-0424 | max | rsc | 30.4 | $0.12 | 88.8% | 37.5% | 14.6% | $0.43 | $0.87 | – |
| GPT-5.2 (xhigh) | OpenAI | gpt-5-2 | xhigh | rsc | 30.4 | – | 90.3% | 37.7% | – | $1.75 | $14.00 | – |
| DeepSeek V4 Pro 0424 (Reasoning, High Effort) | DeepSeek | deepseek-v4-pro-0424 | high | rsc | 30.1 | – | 90.5% | 35.2% | – | $0.43 | $0.87 | – |
| GPT-5.6 Terra (medium) | OpenAI | gpt-5-6-terra | medium | rsc | 30.1 | $0.18 | 87.2% | 33.3% | 1.0% | $2.00 | $12.00 | 97.3 |
| Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | claude-sonnet-4-6 | max | rsc | 30.1 | $2.49 | 87.5% | 33.6% | 3.0% | $3.00 | $15.00 | – |
| Gemini 3.1 Pro Preview | Google | gemini-3-1-pro-preview | | rsc | 29.7 | $0.67 | 94.1% | 47.0% | 4.0% | $2.00 | $12.00 | 131.4 |
| GPT-6 Luna (medium) | OpenAI | gpt-6-luna | medium | rsc | 29.5 | $0.018 | – | 28.3% | 2.5% | $0.10 | $0.50 | – |
| Qwen3.7 Max | Alibaba | qwen3-7-max | | rsc | 29.5 | $1.15 | 92.3% | 40.5% | 1.5% | $2.50 | $7.50 | – |
| MiniMax-M3 | MiniMax | minimax-m3 | | rsc | 29.2 | $0.51 | 92.9% | 39.0% | 2.0% | $0.30 | $1.20 | 137.6 |
| Claude Opus 4.5 (Reasoning) | Anthropic | claude-opus-4-5 | thinking | rsc | 29.1 | – | 86.6% | 30.1% | – | $5.00 | $25.00 | – |
| MiMo-V2-Pro | Xiaomi | mimo-v2-pro | | rsc | 28.6 | – | 87.0% | 30.4% | – | – | – | – |
| GPT-5.2 Codex (xhigh) | OpenAI | gpt-5-2-codex | xhigh | rsc | 28.5 | – | 89.9% | 35.7% | – | $1.75 | $14.00 | – |
| Qwen3.6 Max Preview | Alibaba | qwen3-6-max | | rsc | 28.4 | – | 88.8% | 30.8% | – | $1.30 | $7.80 | – |
| GPT-5.6 Sol (Non-reasoning) | OpenAI | gpt-5-6-sol | non-reasoning | rsc | 28.3 | – | 79.0% | 16.7% | – | $4.00 | $20.00 | – |
| Nex-N2-Pro (based on Qwen3.5-397B-A17B) | Nex AGI | nex-n2-pro | based on qwen3.5-397b-a17b | rsc | 28.2 | – | 89.2% | 33.7% | – | – | – | – |
| Solar Pro 4 | Upstage | solar-pro4 | | rsc | 28.2 | – | 89.1% | 29.2% | 0.5% | $0.30 | $1.20 | 84.0 |
| GPT-6 Sol (Non-reasoning) | OpenAI | gpt-6-sol | non-reasoning | rsc | 28.1 | $0.33 | – | 18.4% | 13.1% | $2.00 | $10.00 | 77.7 |
| Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) | Anthropic | claude-sonnet-5 | medium | rsc | 28.1 | $1.0 | – | 30.0% | 2.0% | $2.00 | $10.00 | 58.4 |
| Gemini 3 Pro Preview (high) | Google | gemini-3-pro-preview | high | rsc | 28.0 | – | 90.8% | 39.7% | – | $2.00 | $12.00 | – |
| GLM-5 (Reasoning) | Z AI | glm-5 | | rsc | 27.9 | – | 82.0% | 29.3% | – | $1.00 | $3.20 | – |
| Inkling Small | Thinking Machines | inkling-small | | rsc | 27.8 | – | 89.5% | 33.3% | 1.0% | $0.30 | $1.20 | 226.6 |
| GPT-5.4 (low) | OpenAI | gpt-5-4 | low | rsc | 27.6 | – | 87.1% | 30.8% | – | $2.50 | $15.00 | – |
| Qwen3.8 27B (medium) | Alibaba | qwen3-8-27b | medium | rsc | 27.6 | $1.13 | 84.5% | 14.1% | 5.1% | $0.50 | $3.00 | 47.6 |
| GPT-5.6 Terra (low) | OpenAI | gpt-5-6-terra | low | rsc | 27.5 | $0.14 | 84.3% | 29.2% | 1.5% | $2.00 | $12.00 | 93.9 |
| JT-4.1 Flash 236B A21B | China Mobile | jt-4-1-flash-236b-a21b | non-reasoning | rsc | 27.3 | – | 84.5% | 17.8% | – | – | – | – |
| Grok Build 0.1 0616 | SpaceXAI | grok-build-0-1-0616 | | rsc | 27.2 | – | 89.5% | 38.3% | – | $1.00 | $2.00 | – |
| Qwen3.6 Plus | Alibaba | qwen3-6-plus | | rsc | 27.0 | – | 88.2% | 27.8% | – | $0.50 | $3.00 | – |
| Kimi K2.6 | Kimi | kimi-k2-6 | | rsc | 27.0 | $0.80 | 91.1% | 37.5% | 0.5% | $0.95 | $4.00 | – |
| Quasar 438B (max, based on GLM-5.2) | Multiverse Computing | quasar-438b | max | rsc | 26.7 | $2.02 | 73.2% | 18.7% | 1.0% | $0.60 | $1.80 | 160.9 |
| GLM-5-Turbo | Z AI | glm-5-turbo | | rsc | 26.6 | – | 84.7% | 27.8% | – | – | – | – |
| GPT-5.2 (medium) | OpenAI | gpt-5-2 | medium | rsc | 26.5 | – | 86.4% | 26.7% | – | $1.75 | $14.00 | – |
| Apodex 1.1 | Apodex | apodex-1-1 | | rsc | 26.4 | $0.46 | 86.4% | 34.1% | 0.0% | $0.30 | $3.00 | – |
| Claude Opus 4.6 (Non-reasoning, High Effort) | Anthropic | claude-opus-4-6 | high | rsc | 26.4 | – | 84.0% | 19.1% | – | $5.00 | $25.00 | – |
| Gemini 3 Flash Preview (Reasoning) | Google | gemini-3-flash-preview | | rsc | 26.3 | – | 89.8% | 36.6% | – | $0.50 | $3.00 | – |
| Qwen3.8 27B (low) | Alibaba | qwen3-8-27b | low | rsc | 26.2 | $1.05 | 84.5% | 14.0% | 2.5% | $0.50 | $3.00 | 51.4 |
| GLM-5.1 (Reasoning) | Z AI | glm-5-1 | | rsc | 26.1 | $1.02 | 86.8% | 30.1% | 2.0% | $1.28 | $4.07 | – |
| GPT-5.5 Instant (June 2026) | OpenAI | gpt-5-5-instant-06-26 | june 2026 | rsc | 26.0 | $0.69 | 82.3% | 19.9% | 12.6% | $5.00 | $30.00 | 128.7 |
| DeepSeek V4 Flash 0420 (Reasoning, High Effort) | DeepSeek | deepseek-v4-flash-0420 | high | rsc | 26.0 | – | 86.7% | 30.3% | 3.0% | $0.11 | $0.24 | – |
| MiMo-V2.5-Pro | Xiaomi | mimo-v2-5-pro | | rsc | 26.0 | $0.054 | 86.6% | 35.7% | 0.0% | $0.43 | $0.87 | 29.0 |
| Kimi K2.7 Code | Kimi | kimi-k2-7-code | | rsc | 25.8 | $0.54 | 89.6% | 35.0% | 1.0% | $0.95 | $4.00 | – |
| Grok 4.20 0309 v2 (Reasoning) | SpaceXAI | grok-4-20 | | rsc | 25.7 | – | 91.1% | 34.5% | – | $1.25 | $2.50 | – |
| K2 Horizon MoVA 36B A4B | Institute of Foundation Models | k2-horizon-mova-36b-a4b | | rsc | 25.3 | – | 82.2% | 23.4% | 0.0% | – | – | – |
| Hy3 | Tencent | hy3 | | rsc | 25.3 | $0.072 | 89.7% | 33.5% | 0.5% | $0.14 | $0.55 | 95.6 |
| Grok 4.20 0309 (Reasoning) | SpaceXAI | grok-4-20-0309 | | rsc | 25.2 | – | 88.5% | 32.4% | – | $2.00 | $6.00 | – |
| MiMo-V2.5 | Xiaomi | mimo-v2-5-0424 | | rsc | 25.2 | – | 84.9% | 27.2% | 0.0% | $0.14 | $0.28 | 45.4 |
| Qwen3.7 Plus | Alibaba | qwen3-7-plus | | rsc | 25.2 | $0.22 | 90.0% | 35.6% | 1.0% | $0.40 | $1.60 | 57.5 |
| MiMo-V2-Omni-0327 | Xiaomi | mimo-v2-omni-0327 | | rsc | 25.1 | – | 85.5% | 22.6% | – | – | – | – |
| GPT-5.6 Luna (medium) | OpenAI | gpt-5-6-luna | medium | rsc | 25.0 | $0.016 | 85.9% | 25.8% | 0.5% | $0.20 | $1.20 | – |
| Inkling (xhigh) | Thinking Machines | inkling | xhigh | rsc | 25.0 | – | 87.2% | 31.9% | 1.0% | $1.00 | $4.05 | 186.2 |
| Ling 3.0 Flash | InclusionAI | ling-3-0-flash | | rsc | 24.9 | – | 85.5% | 23.7% | 0.0% | $0.075 | $0.22 | 378.8 |
| GPT-5 Codex (high) | OpenAI | gpt-5-codex | high | rsc | 24.9 | – | 83.7% | 27.8% | – | $1.25 | $10.00 | – |
| Grok 4.3 (high) | SpaceXAI | grok-4-3 | high | rsc | 24.9 | $0.21 | 90.1% | 37.2% | 0.0% | $1.25 | $2.50 | – |
| Grok 4.3 (medium) | SpaceXAI | grok-4-3 | medium | rsc | 24.8 | – | 89.0% | 30.0% | – | $1.25 | $2.50 | – |
| Solar Open2 250B | Upstage | solar-open2-250b | | rsc | 24.7 | – | 85.7% | 28.5% | – | – | – | – |
| GPT-5.1 (high) | OpenAI | gpt-5-1 | high | rsc | 24.7 | – | 87.3% | 28.5% | – | $1.25 | $10.00 | – |
| Claude Sonnet 4.6 (Non-reasoning, High Effort) | Anthropic | claude-sonnet-4-6 | non-reasoning | rsc | 24.7 | – | 79.9% | 13.3% | – | $3.00 | $15.00 | – |
| DeepSeek V4.1 Flash (Non-Reasoning) | DeepSeek | deepseek-v4-1-flash | non-reasoning | rsc | 24.7 | $0.15 | – | 10.8% | 5.6% | $0.30 | $1.20 | 217.5 |
| Ling-3.0-flash-VL | InclusionAI | ling-3-0-flash-vl | | rsc | 24.6 | – | 86.2% | 22.0% | 0.0% | $0.075 | $0.22 | 149.7 |
| Grok 4.3 (low) | SpaceXAI | grok-4-3 | low | rsc | 24.3 | – | 84.3% | 18.4% | – | $1.25 | $2.50 | – |
| Claude Sonnet 5 (Adaptive Reasoning, Low Effort) | Anthropic | claude-sonnet-5 | low | rsc | 24.3 | $0.51 | – | 21.9% | 2.5% | $2.00 | $10.00 | 60.1 |
| GLM-5.1 (Non-reasoning) | Z AI | glm-5-1 | non-reasoning | rsc | 24.2 | – | 83.9% | 27.9% | – | $1.38 | $4.40 | – |
| DeepSeek V4 Flash 0420 (Reasoning, Max Effort) | DeepSeek | deepseek-v4-flash-0420 | max | rsc | 24.2 | $0.11 | 89.4% | 34.8% | 2.5% | $0.13 | $0.28 | – |
| GPT-5.4 mini (xhigh) | OpenAI | gpt-5-4-mini | xhigh | rsc | 24.1 | $0.45 | 87.5% | 28.1% | 2.0% | $0.75 | $4.50 | – |
| MiMo-V2-Omni | Xiaomi | mimo-v2-omni | | rsc | 23.9 | – | 82.8% | 22.1% | – | – | – | – |
| Gemini 3.5 Flash (minimal) | Google | gemini-3-5-flash | minimal | rsc | 23.8 | – | 82.8% | 24.1% | – | $1.50 | $9.00 | – |
| GPT-5.1 Codex (high) | OpenAI | gpt-5-1-codex | high | rsc | 23.7 | – | 86.0% | 25.7% | – | $1.25 | $10.00 | – |
| Claude Opus 4.5 (Non-reasoning) | Anthropic | claude-opus-4-5 | non-reasoning | rsc | 23.7 | – | 81.0% | 13.2% | – | $5.00 | $25.00 | – |
| Kimi K2.6 (Non-reasoning) | Kimi | kimi-k2-6 | non-reasoning | rsc | 23.6 | – | 78.8% | 19.6% | – | $0.95 | $4.00 | – |
| GLM 5V Turbo (Reasoning) | Z AI | glm-5v-turbo | | rsc | 23.5 | – | 80.9% | 17.1% | – | – | – | – |
| Kimi K2.5 (Reasoning) | Kimi | kimi-k2-5 | | rsc | 23.5 | – | 87.9% | 30.7% | – | $0.60 | $2.75 | – |
| Claude Sonnet 4.6 (Non-reasoning, Low Effort) | Anthropic | claude-sonnet-4-6 | non-reasoning | rsc | 23.3 | – | 79.7% | 11.2% | – | $3.00 | $15.00 | – |
| Claude Sonnet 5 (Non-reasoning, High Effort) | Anthropic | claude-sonnet-5 | non-reasoning | rsc | 23.2 | – | 80.0% | 19.0% | – | $2.00 | $10.00 | 58.3 |
| GPT-5.5 (Non-reasoning) | OpenAI | gpt-5-5 | non-reasoning | rsc | 23.2 | – | 76.8% | 13.7% | – | $5.00 | $30.00 | – |
| GPT-5 (high) | OpenAI | gpt-5 | high | rsc | 23.0 | – | 85.4% | 28.5% | – | $1.25 | $10.00 | – |
| Nemotron 3 Ultra 550B A55B (Reasoning) | NVIDIA | nemotron-3-ultra-550b-a55b | | rsc | 22.9 | $0.60 | 86.7% | 28.4% | 0.5% | $0.60 | $2.50 | 214.9 |
| Qwen3.5 27B (Reasoning) | Alibaba | qwen3-5-27b | | rsc | 22.9 | – | 85.8% | 23.9% | – | $0.30 | $2.40 | – |
| GPT-5 (medium) | OpenAI | gpt-5 | medium | rsc | 22.9 | – | 84.2% | 25.4% | – | $1.25 | $10.00 | – |
| Claude 4.1 Opus (Reasoning) | Anthropic | claude-4-1-opus | thinking | rsc | 22.8 | – | 80.9% | 12.5% | – | $15.00 | $75.00 | – |
| MiniMax-M2.5 | MiniMax | minimax-m2-5 | | rsc | 22.8 | – | 84.8% | 20.5% | – | $0.30 | $1.20 | – |
| MiniMax-M2.7 | MiniMax | minimax-m2-7 | | rsc | 22.8 | – | 87.4% | 29.6% | 0.0% | $0.30 | $1.20 | – |
| Hy3-preview (Reasoning) | Tencent | hy3-preview | | rsc | 22.7 | – | 86.7% | 27.8% | – | $0.063 | $0.21 | – |
| A.X-K2 | SK Telecom | a-x-k2 | | rsc | 22.7 | – | 85.7% | 29.6% | – | – | – | – |
| GPT-5.5 Instant (May 2026) | OpenAI | gpt-5-5-instant-05-26 | may 2026 | rsc | 22.7 | – | 84.6% | 21.6% | – | $5.00 | $30.00 | – |
| Ling-3.0-flash-Fin | InclusionAI | ling-3-0-flash-fin | | rsc | 22.6 | – | – | 22.6% | 0.0% | $0.075 | $0.22 | 165.2 |
| Grok 4 | SpaceXAI | grok-4 | | rsc | 22.5 | – | 87.7% | 26.7% | – | $3.00 | $15.00 | – |
| MiMo-V2-Flash (Feb 2026) | Xiaomi | mimo-v2-0206 | feb 2026 | rsc | 22.4 | – | 83.5% | 22.1% | – | – | – | – |
| GLM-5.2 (Non-reasoning) | Z AI | glm-5-2 | non-reasoning | rsc | 22.4 | – | 68.6% | 9.8% | – | $1.40 | $4.40 | – |
| Gemini 3 Pro Preview (low) | Google | gemini-3-pro-preview | low | rsc | 22.3 | – | 88.7% | 29.5% | – | $2.00 | $12.00 | – |
| GLM-4.7 (Reasoning) | Z AI | glm-4-7 | | rsc | 22.2 | – | 85.9% | 27.4% | – | $0.60 | $2.20 | – |
| Gemini 3.5 Flash-Lite | Google | gemini-3-5-flash-lite | high | rsc | 22.2 | $0.12 | 83.8% | 18.8% | 1.0% | $0.30 | $2.50 | 334.8 |
| Kimi K2 Thinking | Kimi | kimi-k2-thinking | | rsc | 22.0 | – | 83.8% | 23.8% | – | $0.60 | $2.50 | – |
| o3-pro | OpenAI | o3-pro | | rsc | 21.9 | – | 84.5% | – | – | $20.00 | $80.00 | – |
| G9v3-39A5B | AI9Stars | g9v3-39a5b | | rsc | 21.8 | – | 80.5% | 17.5% | – | $0.0 | $0.0 | – |
| GLM-5 (Non-reasoning) | Z AI | glm-5 | non-reasoning | rsc | 21.8 | – | 66.6% | 7.6% | – | $1.00 | $3.20 | – |
| KAT Coder Pro V2 | KwaiKAT | kat-coder-pro-v2 | non-reasoning | rsc | 21.7 | – | 85.5% | 16.1% | – | $0.30 | $1.20 | – |
| DeepSeek V3.2 (Reasoning) | DeepSeek | deepseek-v3-2 | reasoning | rsc | 21.5 | – | 84.0% | 24.6% | – | $0.28 | $0.42 | – |
| Qwen3.5 397B A17B (Non-reasoning) | Alibaba | qwen3-5-397b-a17b | non-reasoning | rsc | 21.4 | – | 86.1% | 19.8% | – | $0.60 | $3.60 | 86.3 |
| Qwen3.6 27B (Reasoning) | Alibaba | qwen3-6-27b | | rsc | 21.4 | $0.62 | 84.2% | 23.1% | 0.0% | $0.60 | $3.60 | – |
| Qwen3 Max Thinking | Alibaba | qwen3-max-thinking | | rsc | 21.3 | – | 86.1% | 28.0% | – | – | – | – |
| GPT-5.6 Luna (low) | OpenAI | gpt-5-6-luna | low | rsc | 21.0 | $0.0098 | 83.5% | 19.8% | 0.0% | $0.20 | $1.20 | – |
| MiniMax-M2.1 | MiniMax | minimax-m2-1 | | rsc | 20.9 | – | 83.0% | 23.2% | – | $0.30 | $1.20 | – |
| GPT-6 Luna (low) | OpenAI | gpt-6-luna | low | rsc | 20.9 | $0.0045 | – | 20.3% | 0.0% | $0.10 | $0.50 | 141.5 |
| DeepSeek V4 Pro 0424 (Non-reasoning) | DeepSeek | deepseek-v4-pro-0424 | non-reasoning | rsc | 20.8 | – | 71.7% | 8.2% | – | $0.43 | $0.87 | – |
| MiMo-V2-Flash (Reasoning) | Xiaomi | mimo-v2-flash | reasoning | rsc | 20.8 | – | 84.6% | 22.8% | – | $0.10 | $0.30 | – |
| GPT-5 (low) | OpenAI | gpt-5 | low | rsc | 20.8 | – | 80.8% | 19.6% | – | $1.25 | $10.00 | – |
| GPT-5.6 Terra (Non-reasoning) | OpenAI | gpt-5-6-terra | non-reasoning | rsc | 20.8 | $0.14 | 74.6% | 11.4% | 0.5% | $2.00 | $12.00 | 90.7 |
| GPT-5.4 nano (xhigh) | OpenAI | gpt-5-4-nano | xhigh | rsc | 20.7 | $0.18 | 81.7% | 28.3% | 0.5% | $0.20 | $1.25 | – |
| Claude 4.5 Sonnet (Reasoning) | Anthropic | claude-4-5-sonnet | thinking | rsc | 20.7 | $0.56 | 83.4% | 17.8% | 0.0% | $3.00 | $15.00 | – |
| Claude 4 Opus (Reasoning) | Anthropic | claude-4-opus | thinking | rsc | 20.6 | – | 79.6% | 12.3% | – | $15.00 | $75.00 | – |
| GPT-5 mini (medium) | OpenAI | gpt-5-mini | medium | rsc | 20.6 | – | 80.3% | 15.9% | – | $0.25 | $2.00 | – |
| K2 Horizon 7B | Institute of Foundation Models | k2-horizon-7b | | rsc | 20.6 | – | 75.3% | 18.2% | 1.0% | – | – | – |
| Qwen3.5 Omni Plus | Alibaba | qwen3-5-omni-plus | non-reasoning | rsc | 20.4 | – | 82.6% | 14.9% | – | $0.40 | $4.80 | 86.9 |
| GPT-5.1 Codex mini (high) | OpenAI | gpt-5-1-codex-mini | high | rsc | 20.4 | – | 81.3% | 18.5% | – | $0.25 | $2.00 | – |
| Grok 4.1 Fast (Reasoning) | SpaceXAI | grok-4-1-fast | reasoning | rsc | 20.4 | – | 85.3% | 19.3% | – | – | – | – |
| o3 | OpenAI | o3 | | rsc | 20.2 | – | 82.7% | 20.1% | – | $2.00 | $8.00 | 170.5 |
| Qwen3.8 27B (Non-reasoning) | Alibaba | qwen3-8-27b | non-reasoning | rsc | 20.2 | $2.49 | 81.8% | 12.1% | 0.0% | $0.50 | $3.00 | 50.9 |
| GPT-5.4 nano (medium) | OpenAI | gpt-5-4-nano | medium | rsc | 20.0 | – | 76.1% | 15.9% | – | $0.20 | $1.25 | – |
| Qwen3.6 27B (Non-reasoning) | Alibaba | qwen3-6-27b | non-reasoning | rsc | 19.8 | – | 82.9% | 15.1% | – | $0.60 | $3.60 | – |
| GPT-5.4 mini (medium) | OpenAI | gpt-5-4-mini | medium | rsc | 19.7 | – | 82.3% | 18.6% | – | $0.75 | $4.50 | – |
| K-EXAONE 2.0 0803 | LG AI Research | k-exaone-2-0-0803 | | rsc | 19.7 | – | 82.9% | 18.6% | – | – | – | – |
| Step 3.7 Flash | StepFun | step-3-7-flash | high | rsc | 19.5 | – | 80.9% | 21.4% | – | $0.20 | $1.15 | 197.3 |
| Kimi K2.5 (Non-reasoning) | Kimi | kimi-k2-5 | non-reasoning | rsc | 19.4 | – | 78.9% | 13.2% | – | $0.60 | $3.00 | – |
| Qwen3.5 27B (Non-reasoning) | Alibaba | qwen3-5-27b | non-reasoning | rsc | 19.4 | – | 84.2% | 13.9% | – | $0.30 | $2.40 | – |
| Claude 4.5 Sonnet (Non-reasoning) | Anthropic | claude-4-5-sonnet | non-reasoning | rsc | 19.3 | – | 72.7% | 7.2% | – | $3.00 | $15.00 | – |
| Qwen3.5 35B A3B (Reasoning) | Alibaba | qwen3-5-35b-a3b | | rsc | 19.3 | – | 84.5% | 21.0% | – | $0.25 | $2.00 | – |
| LongCat 2.0 | LongCat | longcat-2-0 | | rsc | 19.1 | $0.059 | 78.0% | 33.7% | 0.0% | $0.30 | $1.20 | – |
| Gemma 4 31B (Reasoning) | Google | gemma-4-31b | | rsc | 19.0 | – | 85.7% | 23.6% | 0.0% | $0.0 | $0.0 | 35.0 |
| Claude 4 Sonnet (Reasoning) | Anthropic | claude-4-sonnet | thinking | rsc | 18.9 | – | 77.7% | 10.7% | – | – | – | – |
| DeepSeek V4 Flash 0420 (Non-reasoning) | DeepSeek | deepseek-v4-flash-0420 | non-reasoning | rsc | 18.9 | – | 71.6% | 7.8% | – | $0.092 | $0.19 | – |
| JT-35B-Flash | China Mobile | jt-35b-flash | non-reasoning | rsc | 18.7 | – | 82.9% | 6.4% | – | – | – | – |
| MiniMax-M2 | MiniMax | minimax-m2 | | rsc | 18.6 | – | 77.7% | 13.7% | – | $0.30 | $1.20 | – |
| KAT-Coder-Pro V1 | KwaiKAT | kat-coder-pro-v1 | non-reasoning | rsc | 18.6 | – | 76.4% | 33.6% | – | – | – | – |
| Claude 4.1 Opus (Non-reasoning) | Anthropic | claude-4-1-opus | non-reasoning | rsc | 18.6 | – | – | – | – | $15.00 | $75.00 | – |
| GLM-4.6 (Reasoning) | Z AI | glm-4-6 | reasoning | rsc | 18.5 | – | 78.0% | 14.5% | – | $0.55 | $2.20 | – |
| Qwen3.5 397B A17B (Reasoning) | Alibaba | qwen3-5-397b-a17b | | rsc | 18.4 | $0.47 | 89.3% | 29.0% | 0.0% | $0.60 | $3.60 | 88.8 |
| MiMo-V2.5-Pro (Non-reasoning) | Xiaomi | mimo-v2-5-pro | non-reasoning | rsc | 18.3 | – | 76.2% | 14.8% | – | $0.43 | $0.87 | 27.7 |
| GPT-6 Luna (Non-reasoning) | OpenAI | gpt-6-luna | non-reasoning | rsc | 18.3 | $0.011 | – | 8.6% | 1.5% | $0.10 | $0.50 | 139.3 |
| Qwen3.6 35B A3B (Reasoning) | Alibaba | qwen3-6-35b-a3b | | rsc | 18.2 | $0.48 | 84.1% | 22.2% | 0.0% | $0.38 | $2.25 | 138.7 |
| GPT-5.4 (Non-reasoning) | OpenAI | gpt-5-4 | non-reasoning | rsc | 18.2 | – | 74.8% | 11.3% | – | $2.50 | $15.00 | – |
| Grok 4 Fast (Reasoning) | SpaceXAI | grok-4-fast | reasoning | rsc | 17.9 | – | 84.7% | 19.1% | – | $0.20 | $0.50 | – |
| Gemini 3 Flash Preview (Non-reasoning) | Google | gemini-3-flash-preview | non-reasoning | rsc | 17.9 | – | 81.2% | 15.0% | – | $0.50 | $3.00 | – |
| Qwen3.5 122B A10B (Non-reasoning) | Alibaba | qwen3-5-122b-a10b | non-reasoning | rsc | 17.7 | – | 82.7% | 15.9% | – | $0.40 | $3.20 | 148.6 |
| Claude 3.7 Sonnet (Reasoning) | Anthropic | claude-3-7-sonnet | thinking | rsc | 17.7 | – | 77.2% | 9.7% | – | – | – | – |
| Muse Glimmer (high) | Meta | muse-glimmer | high | rsc | 17.5 | $0.057 | 83.5% | 22.0% | 0.5% | $0.32 | $1.35 | 161.9 |
| GLM-4.7 (Non-reasoning) | Z AI | glm-4-7 | non-reasoning | rsc | 17.4 | – | 66.4% | 6.4% | – | $0.60 | $2.20 | – |
| Hy3-preview (Non-reasoning) | Tencent | hy3-preview | non-reasoning | rsc | 17.0 | – | 73.2% | 7.0% | – | $0.063 | $0.21 | – |
| Ling-2.6-1T | InclusionAI | ling-2-6-1t | non-reasoning | rsc | 17.0 | – | 75.2% | 8.7% | – | $0.30 | $2.50 | – |
| GPT-5.2 (Non-reasoning) | OpenAI | gpt-5-2 | non-reasoning | rsc | 17.0 | – | 71.2% | 8.0% | – | $1.75 | $14.00 | – |
| Step 3.5 Flash 2603 | StepFun | step-3-5-flash | | rsc | 17.0 | – | 82.6% | 24.5% | – | $0.10 | $0.30 | – |
| Doubao Seed Code | ByteDance Seed | doubao-seed-code | | rsc | 16.9 | – | 76.4% | 14.1% | – | – | – | – |
| Claude 4.5 Haiku (Reasoning) | Anthropic | claude-4-5-haiku | reasoning | rsc | 16.9 | $0.28 | 67.2% | 10.4% | 0.0% | $1.00 | $5.00 | 104.4 |
| GPT-5 mini (high) | OpenAI | gpt-5-mini | high | rsc | 16.8 | $0.053 | 82.8% | 21.5% | 0.0% | $0.25 | $2.00 | – |
| Gemma 4 26B A4B (Reasoning) | Google | gemma-4-26b-a4b | | rsc | 16.7 | – | 79.2% | 19.3% | – | $0.10 | $0.37 | – |
| o4-mini (high) | OpenAI | o4-mini | high | rsc | 16.7 | – | 78.4% | 16.5% | – | $1.10 | $4.40 | – |
| Ring-2.6-1T | InclusionAI | ring-2-6-1t | xhigh | rsc | 16.6 | $0.29 | 85.7% | 21.6% | 0.5% | $0.30 | $2.50 | 118.4 |
| Step 3.5 Flash | StepFun | step-3-5-flash-0202 | | rsc | 16.6 | – | 83.1% | 21.1% | – | $0.10 | $0.30 | – |
| Claude 4 Opus (Non-reasoning) | Anthropic | claude-4-opus | non-reasoning | rsc | 16.6 | – | 70.1% | 6.2% | – | $15.00 | $75.00 | – |
| Claude 4 Sonnet (Non-reasoning) | Anthropic | claude-4-sonnet | non-reasoning | rsc | 16.6 | – | 68.3% | 4.3% | – | – | – | – |
| DeepSeek V3.2 Exp (Reasoning) | DeepSeek | deepseek-v3-2-0925 | | rsc | 16.6 | – | 79.7% | 14.9% | – | $0.28 | $0.42 | – |
| Qwen3 Max Thinking (Preview) | Alibaba | qwen3-max-thinking-preview | preview | rsc | 16.3 | – | 77.6% | 12.7% | – | $1.20 | $6.00 | – |
| Gemini 2.5 Pro | Google | gemini-2-5-pro | | rsc | 16.1 | $0.23 | 84.4% | 22.5% | 0.0% | $1.25 | $10.00 | – |
| MiMo-V2-Flash (Non-reasoning) | Xiaomi | mimo-v2-flash | non-reasoning | rsc | 16.0 | – | 65.6% | 8.6% | – | – | – | – |
| DeepSeek V3.2 (Non-reasoning) | DeepSeek | deepseek-v3-2 | non-reasoning | rsc | 16.0 | – | 75.1% | 11.2% | – | $0.28 | $0.42 | – |
| K2 Horizon 3.7B | Institute of Foundation Models | k2-horizon-3-7b | | rsc | 15.6 | – | 69.2% | 13.9% | 0.0% | – | – | – |
| Qwen3 Max | Alibaba | qwen3-max | non-reasoning | rsc | 15.6 | – | 76.4% | 11.9% | – | $1.20 | $6.00 | – |
| Qwen3.5 122B A10B (Reasoning) | Alibaba | qwen3-5-122b-a10b | | rsc | 15.6 | $0.32 | 85.7% | 25.2% | 0.0% | $0.40 | $3.20 | 133.1 |
| Gemini 3.1 Flash-Lite | Google | gemini-3-1-flash-lite | high | rsc | 15.6 | $0.039 | 82.2% | 17.2% | 0.5% | $0.25 | $1.50 | – |
| GPT-5.6 Luna (Non-reasoning) | OpenAI | gpt-5-6-luna | non-reasoning | rsc | 15.5 | $0.010 | 64.5% | 7.2% | 1.0% | $0.20 | $1.20 | – |
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) | Google | gemini-2-5-flash-preview-09-2025 | reasoning | rsc | 15.5 | – | 79.3% | 13.8% | – | – | – | – |
| Claude 4.5 Haiku (Non-reasoning) | Anthropic | claude-4-5-haiku | non-reasoning | rsc | 15.4 | – | 64.6% | 4.2% | – | $1.00 | $5.00 | 91.3 |
| Kimi K2 0905 | Kimi | kimi-k2-0905 | non-reasoning | rsc | 15.3 | – | 76.7% | 6.4% | – | $0.60 | $2.50 | – |
| Ling 3.0 Tiny | InclusionAI | ling-3-0-tiny | | rsc | 15.3 | – | 73.4% | 9.3% | 0.0% | $0.0 | $0.0 | 30.7 |
| Claude 3.7 Sonnet (Non-reasoning) | Anthropic | claude-3-7-sonnet | non-reasoning | rsc | 15.3 | – | 65.6% | 4.2% | – | $3.00 | $15.00 | – |
| Qwen3.6 35B A3B (Non-reasoning) | Alibaba | qwen3-6-35b-a3b | non-reasoning | rsc | 15.2 | – | 81.7% | 13.9% | – | $0.38 | $2.25 | 141.5 |
| o1 | OpenAI | o1 | | rsc | 15.2 | – | 74.7% | 7.0% | – | $15.00 | $60.00 | – |
| Qwen3.5 35B A3B (Non-reasoning) | Alibaba | qwen3-5-35b-a3b | non-reasoning | rsc | 15.1 | – | 81.9% | 13.4% | – | $0.25 | $2.00 | – |
| Gemini 2.5 Pro Preview (Mar' 25) | Google | gemini-2-5-pro-03-25 | mar' 25 | rsc | 15.0 | – | 83.6% | 18.0% | – | – | – | – |
| GLM-4.6 (Non-reasoning) | Z AI | glm-4-6 | non-reasoning | rsc | 14.9 | – | 63.2% | 5.5% | – | $0.57 | $2.20 | – |
| GLM-4.7-Flash (Reasoning) | Z AI | glm-4-7-flash | | rsc | 14.9 | – | 58.1% | 7.6% | – | $0.070 | $0.40 | – |
| Granite 4.2 30B | IBM | granite-4-2-30b | | rsc | 14.8 | – | 64.4% | 11.2% | – | $0.16 | $0.65 | 78.2 |
| DeepSeek V3.1 Terminus (Reasoning) | DeepSeek | deepseek-v3-1-terminus | reasoning | rsc | 14.8 | – | 79.2% | 16.4% | 0.0% | $1.64 | $2.75 | – |
| Grok 3 mini Reasoning (high) | SpaceXAI | grok-3-mini-reasoning | high | rsc | 14.6 | – | 79.1% | 11.0% | – | $0.30 | $0.50 | – |
| Grok 4.20 0309 (Non-reasoning) | SpaceXAI | grok-4-20-0309 | non-reasoning | rsc | 14.6 | – | 78.5% | 24.5% | – | $2.00 | $6.00 | – |
| Gemini 2.5 Pro Preview (May' 25) | Google | gemini-2-5-pro-05-06 | may' 25 | rsc | 14.5 | – | 82.2% | 19.9% | – | $1.25 | $10.00 | – |
| DeepSeek V3.2 Speciale | DeepSeek | deepseek-v3-2-speciale | | rsc | 14.5 | – | 87.1% | 28.7% | – | – | – | – |
| K-EXAONE (Reasoning) | LG AI Research | k-exaone | | rsc | 14.4 | – | 78.3% | 13.9% | – | – | – | – |
| ERNIE 5.0 Thinking Preview | Baidu | ernie-5-0-thinking-preview | | rsc | 14.3 | – | 77.7% | 13.3% | – | – | – | – |
| Grok 4.20 0309 v2 (Non-reasoning) | SpaceXAI | grok-4-20 | non-reasoning | rsc | 14.2 | – | 77.6% | 27.9% | – | $1.25 | $2.50 | – |
| Mistral Medium 3.5 | Mistral | mistral-medium-3-5 | high | rsc | 14.2 | $0.50 | 74.8% | 13.8% | 0.0% | $1.50 | $7.50 | 168.0 |
| Gemma 4 12B (Reasoning) | Google | gemma-4-12b | | rsc | 14.2 | – | 75.3% | 15.7% | – | $0.10 | $0.30 | 151.6 |
| Nova 2.0 Pro Preview (medium) | Amazon | nova-2-0-pro-preview | medium | rsc | 14.2 | – | 78.5% | 9.4% | – | $1.25 | $10.00 | 124.5 |
| Grok Code Fast 1 | SpaceXAI | grok-code-fast-1 | | rsc | 14.1 | – | 72.7% | 8.0% | – | – | – | – |
| Grok 4.3 (Non-reasoning) | SpaceXAI | grok-4-3 | non-reasoning | rsc | 14.0 | $0.17 | 65.8% | 6.8% | 0.0% | $1.25 | $2.50 | – |
| DeepSeek V3.1 Terminus (Non-reasoning) | DeepSeek | deepseek-v3-1-terminus | non-reasoning | rsc | 13.9 | – | 75.1% | 8.7% | – | $0.27 | $1.00 | – |
| Gemma 4 31B (Non-reasoning) | Google | gemma-4-31b | non-reasoning | rsc | 13.9 | – | 76.3% | 11.8% | – | $0.14 | $0.40 | 40.8 |
| DeepSeek V3.2 Exp (Non-reasoning) | DeepSeek | deepseek-v3-2-0925 | non-reasoning | rsc | 13.9 | – | 73.8% | 9.0% | – | $0.28 | $0.42 | – |
| Apriel-v1.5-15B-Thinker | ServiceNow | apriel-v1-5-15b-thinker | | rsc | 13.8 | – | 71.3% | 12.1% | – | $0.0 | $0.0 | – |
| Mercury 2 | Inception | mercury-2 | high | rsc | 13.8 | – | 77.0% | 17.1% | 0.0% | $0.25 | $0.75 | – |
| DeepSeek V3.1 (Non-reasoning) | DeepSeek | deepseek-v3-1 | non-reasoning | rsc | 13.7 | – | 73.5% | 6.7% | – | $0.57 | $1.68 | – |
| Nova 2.0 Omni (medium) | Amazon | nova-2-0-omni | medium | rsc | 13.6 | – | 76.0% | 7.0% | – | $0.30 | $2.50 | – |
| DeepSeek V3.1 (Reasoning) | DeepSeek | deepseek-v3-1 | reasoning | rsc | 13.5 | – | 77.9% | 14.3% | – | $0.59 | $1.69 | – |
| Qwen3 VL 235B A22B (Reasoning) | Alibaba | qwen3-vl-235b-a22b | reasoning | rsc | 13.4 | – | 77.2% | 11.9% | – | $0.40 | $4.00 | – |
| Apriel-v1.6-15B-Thinker | ServiceNow | apriel-v1-6-15b-thinker | | rsc | 13.4 | – | 73.3% | 10.8% | – | $0.0 | $0.0 | – |
| Nova 2.0 Lite (high) | Amazon | nova-2-0-lite | high | rsc | 13.4 | – | 81.1% | 11.6% | – | $0.30 | $2.50 | 208.5 |
| GPT-5.1 (Non-reasoning) | OpenAI | gpt-5-1 | non-reasoning | rsc | 13.3 | – | 64.3% | 5.3% | – | $1.25 | $10.00 | – |
| Qwen3.5 9B (Non-reasoning) | Alibaba | qwen3-5-9b | non-reasoning | rsc | 13.3 | – | 78.6% | 9.4% | – | $0.17 | $0.25 | 76.7 |
| EXAONE 4.5 33B | LG AI Research | exaone-4-5-33b | | rsc | 13.2 | – | 79.4% | 12.9% | – | – | – | – |
| Command A+ | Cohere | command-a-plus | | rsc | 13.1 | $0.0 | 76.1% | 12.0% | 0.5% | $0.0 | $0.0 | 191.0 |
| Gemma 4 26B A4B (Non-reasoning) | Google | gemma-4-26b-a4b | non-reasoning | rsc | 13.1 | – | 71.4% | 11.5% | – | $0.13 | $0.40 | 99.4 |
| Qwen3.5 4B (Reasoning) | Alibaba | qwen3-5-4b | | rsc | 13.1 | – | 77.1% | 9.9% | – | $0.030 | $0.15 | 26.7 |
| DeepSeek R1 0528 (May '25) | DeepSeek | deepseek-r1 | may '25 | rsc | 13.1 | – | 81.3% | 15.8% | – | $1.35 | $3.00 | – |
| Gemini 2.5 Flash (Reasoning) | Google | gemini-2-5-flash | reasoning | rsc | 13.1 | – | 79.0% | 12.1% | – | $0.30 | $2.50 | – |
| GPT-5 nano (high) | OpenAI | gpt-5-nano | high | rsc | 13.0 | – | 67.6% | 9.5% | – | $0.050 | $0.40 | – |
| Nemotron 3.5 Lightning | NVIDIA | nemotron-3-5-lightning | | rsc | 12.9 | $0.10 | 74.3% | 10.6% | 0.5% | $0.070 | $0.22 | 263.1 |
| Nemotron 3 Super 120B A12B (Reasoning) | NVIDIA | nvidia-nemotron-3-super-120b-a12b | | rsc | 12.8 | $1.64 | 80.0% | 20.8% | 0.0% | $0.30 | $0.90 | 165.3 |
| Nova 2.0 Pro Preview (low) | Amazon | nova-2-0-pro-preview | low | rsc | 12.8 | – | 75.1% | 5.2% | – | $1.25 | $10.00 | 127.9 |
| GLM-4.5 (Reasoning) | Z AI | glm-4-5 | | rsc | 12.8 | – | 78.2% | 13.0% | – | – | – | – |
| Kimi K2 | Kimi | kimi-k2 | non-reasoning | rsc | 12.7 | – | 76.6% | 7.4% | – | $0.57 | $2.30 | – |
| Qwen3 235B A22B 2507 (Reasoning) | Alibaba | qwen3-235b-a22b-instruct-2507-reasoning | | rsc | 12.7 | $0.081 | 79.0% | 15.9% | 0.0% | $0.23 | $2.30 | – |
| GPT-4.1 | OpenAI | gpt-4-1 | non-reasoning | rsc | 12.7 | – | 66.6% | 4.2% | – | $2.00 | $8.00 | – |
| Qwen3 Max (Preview) | Alibaba | qwen3-max-preview | non-reasoning | rsc | 12.6 | – | 76.4% | 10.1% | – | $1.20 | $6.00 | – |
| Nova 2.0 Lite (medium) | Amazon | nova-2-0-lite | medium | rsc | 12.5 | – | 76.8% | 9.0% | – | $0.30 | $2.50 | 212.5 |
| GPT-5 nano (medium) | OpenAI | gpt-5-nano | medium | rsc | 12.5 | – | 67.0% | 8.7% | – | $0.050 | $0.40 | – |
| Qwen3.5 Omni Flash | Alibaba | qwen3-5-omni-flash | non-reasoning | rsc | 12.5 | – | 74.2% | 7.6% | – | $0.10 | $0.80 | 226.5 |
| o3-mini | OpenAI | o3-mini | | rsc | 12.5 | – | 74.8% | 7.9% | – | $1.10 | $4.40 | – |
| MiniCPM5-2B | OpenBMB | minicpm5-2b | | rsc | 12.5 | – | 70.2% | 8.9% | 0.0% | – | – | – |
| o1-pro | OpenAI | o1-pro | | rsc | 12.4 | – | – | – | – | $150 | $600 | – |
| Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) | Google | gemini-2-5-flash-preview-09-2025 | non-reasoning | rsc | 12.4 | – | 76.6% | 8.7% | – | – | – | – |
| Mercury 2.5 | Inception | mercury-2-5 | | rsc | 12.3 | $0.12 | – | 11.8% | 0.0% | $0.25 | $0.75 | 746.4 |
| JT-MINI | China Mobile | jt-mini | non-reasoning | rsc | 12.2 | – | 67.6% | 6.4% | – | – | – | – |
| Grok 3 | SpaceXAI | grok-3 | non-reasoning | rsc | 12.1 | – | 69.3% | 4.1% | – | $4.00 | $20.00 | – |
| Seed-OSS-36B-Instruct | ByteDance Seed | seed-oss-36b-instruct | | rsc | 12.1 | – | 72.6% | 9.9% | – | $0.21 | $0.57 | – |
| Qwen3 235B A22B 2507 Instruct | Alibaba | qwen3-235b-a22b-instruct-2507 | non-reasoning | rsc | 12.0 | – | 75.3% | 11.1% | – | $0.23 | $0.92 | – |
| Qwen3 Coder 480B A35B Instruct | Alibaba | qwen3-coder-480b-a35b-instruct | non-reasoning | rsc | 11.9 | – | 61.8% | 4.5% | – | $1.50 | $7.50 | – |
| Qwen3 VL 32B (Reasoning) | Alibaba | qwen3-vl-32b | reasoning | rsc | 11.9 | – | 73.3% | 10.1% | – | $0.16 | $0.64 | – |
| Magistral Medium 1.2 | Mistral | magistral-medium-2509 | | rsc | 11.8 | – | 73.9% | 10.3% | – | – | – | – |
| Sonar Reasoning Pro | Perplexity | sonar-reasoning-pro | | rsc | 11.8 | – | – | – | – | – | – | – |
| Nova 2.0 Lite (low) | Amazon | nova-2-0-lite | low | rsc | 11.8 | – | 69.8% | 4.0% | – | $0.30 | $2.50 | 213.0 |
| HyperNova 60B 2605 (high, based on gpt-oss-120b) | Multiverse Computing | hypernova-60b | high | rsc | 11.7 | – | 73.3% | 16.6% | – | – | – | – |
| MiniMax M1 80k | MiniMax | minimax-m1-80k | | rsc | 11.7 | – | 69.7% | 8.9% | – | $0.55 | $2.20 | – |
| GPT-5.4 nano (Non-Reasoning) | OpenAI | gpt-5-4-nano | non-reasoning | rsc | 11.7 | – | 55.8% | 4.1% | – | $0.20 | $1.25 | – |
| Nemotron Cascade 2 30B A3B | NVIDIA | nemotron-cascade-2-30b-a3b | | rsc | 11.7 | – | 75.8% | 12.0% | – | – | – | – |
| Gemini 2.5 Flash Preview (Reasoning) | Google | gemini-2-5-flash-04-2025 | | rsc | 11.7 | – | 69.8% | 12.1% | – | – | – | – |
| gpt-oss-120b (high) | OpenAI | gpt-oss-120b | high | rsc | 11.6 | $0.11 | 78.2% | 19.6% | 0.0% | $0.15 | $0.59 | 216.0 |
| K2 Think V2 | Institute of Foundation Models | k2-think-v2 | high | rsc | 11.5 | – | 71.3% | 10.1% | – | – | – | – |
| LongCat Flash Lite | LongCat | longcat-flash-lite | non-reasoning | rsc | 11.5 | – | 63.6% | 5.8% | – | – | – | – |
| GPT-5 (minimal) | OpenAI | gpt-5 | minimal | rsc | 11.4 | – | 67.3% | 6.0% | – | $1.25 | $10.00 | – |
| DeepSeek R1 (Jan '25) | DeepSeek | deepseek-r1-0120 | jan '25 | rsc | 11.4 | $0.26 | 70.8% | 8.5% | 0.0% | $2.00 | $4.00 | – |
| o1-preview | OpenAI | o1-preview | | rsc | 11.4 | – | 76.5% | – | – | $16.50 | $66.00 | – |
| HyperCLOVA X SEED Think (32B) | Naver | hyperclova-x-seed-think-32b | 32b | rsc | 11.4 | – | 61.5% | 5.5% | – | – | – | – |
| Grok 4.1 Fast (Non-reasoning) | SpaceXAI | grok-4-1-fast | non-reasoning | rsc | 11.3 | – | 63.7% | 5.1% | – | – | – | – |
| Mistral Small 4 (Reasoning) | Mistral | mistral-small-4 | | rsc | 11.3 | $0.015 | 76.9% | 9.9% | 0.0% | $0.15 | $0.60 | 172.2 |
| GLM-4.6V (Reasoning) | Z AI | glm-4-6v | reasoning | rsc | 11.2 | – | 71.9% | 9.6% | – | $0.30 | $0.90 | – |
| K-EXAONE (Non-reasoning) | LG AI Research | k-exaone | non-reasoning | rsc | 11.2 | – | 69.5% | 5.7% | – | – | – | – |
| Qwen3 Next 80B A3B (Reasoning) | Alibaba | qwen3-next-80b-a3b | reasoning | rsc | 11.2 | – | 75.9% | 12.6% | – | $0.15 | $1.20 | 195.1 |
| Qwen3.5 9B (Reasoning) | Alibaba | qwen3-5-9b | | rsc | 11.2 | $0.21 | 80.6% | 14.9% | 0.5% | $0.14 | $0.20 | 33.2 |
| GPT-5.4 mini (Non-Reasoning) | OpenAI | gpt-5-4-mini | non-reasoning | rsc | 11.1 | – | 60.6% | 5.9% | – | $0.75 | $4.50 | – |
| Nova 2.0 Omni (low) | Amazon | nova-2-0-omni | low | rsc | 11.1 | – | 69.9% | – | – | $0.30 | $2.50 | – |
| GLM-4.5-Air | Z AI | glm-4-5-air | | rsc | 11.1 | – | 73.3% | 7.0% | – | $0.17 | $0.98 | – |
| Granite 4.2 8B | IBM | granite-4-2-8b | | rsc | 11.1 | $0.024 | 63.1% | 9.7% | 0.0% | $0.060 | $0.25 | 67.2 |
| Grok 4 Fast (Non-reasoning) | SpaceXAI | grok-4-fast | non-reasoning | rsc | 11.1 | – | 60.6% | 4.5% | – | $0.20 | $0.50 | – |
| Mi:dm K 2.5 Pro | Korea Telecom | mi-dm-k-2-5-pro-dec28 | | rsc | 11.0 | – | 70.1% | 8.1% | – | – | – | – |
| o3-mini (high) | OpenAI | o3-mini | high | rsc | 11.0 | – | 77.3% | 12.0% | 0.0% | $1.10 | $4.40 | – |
| Ring-1T | InclusionAI | ring-1t | | rsc | 10.9 | – | 77.4% | 11.1% | – | – | – | – |
| G9v3-3B | AI9Stars | g9v3-3b | | rsc | 10.8 | – | 43.8% | 4.5% | – | $0.0 | $0.0 | – |
| Trinity Large Thinking | Arcee AI | trinity-large-thinking | | rsc | 10.8 | $0.12 | 75.2% | 15.8% | 0.5% | $0.25 | $0.80 | 341.8 |
| Qwen3.5 4B (Non-reasoning) | Alibaba | qwen3-5-4b | non-reasoning | rsc | 10.8 | – | 71.2% | 8.0% | – | $0.030 | $0.15 | 25.9 |
| INTELLECT-3 (based on GLM-4.5-Air) | Prime Intellect | intellect-3 | based on glm-4.5-air | rsc | 10.6 | – | 76.1% | 13.1% | – | – | – | – |
| GLM-4.7-Flash (Non-reasoning) | Z AI | glm-4-7-flash | non-reasoning | rsc | 10.6 | – | 45.2% | 5.0% | – | $0.070 | $0.40 | – |
| GPT-5 (ChatGPT) | OpenAI | gpt-5-chatgpt | non-reasoning | rsc | 10.4 | – | 68.6% | 6.6% | – | – | – | – |
| Solar Open 100B (Reasoning) | Upstage | solar-open-100b | high | rsc | 10.4 | – | 65.7% | 10.4% | – | – | – | – |
| Grok 3 Reasoning Beta | SpaceXAI | grok-3-reasoning | | rsc | 10.4 | – | – | – | – | – | – | – |
| Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) | Google | gemini-2-5-flash-lite-preview-09-2025 | reasoning | rsc | 10.4 | – | 70.9% | 7.0% | – | $0.10 | $0.40 | – |
| Nemotron 3 Nano Omni 30B A3B Reasoning | NVIDIA | nemotron-3-nano-omni-30b-a3b | | rsc | 10.3 | – | 46.9% | 4.8% | – | $0.20 | $1.09 | 267.2 |
| gpt-oss-120b (low) | OpenAI | gpt-oss-120b | low | rsc | 10.2 | – | 67.2% | 5.9% | – | $0.15 | $0.54 | 217.5 |
| GPT-4.1 mini | OpenAI | gpt-4-1-mini | non-reasoning | rsc | 10.2 | – | 66.4% | 5.0% | – | $0.40 | $1.60 | – |
| Llama 4 Maverick | Meta | llama-4-maverick | non-reasoning | rsc | 10.0 | – | 67.1% | 4.9% | 0.0% | $0.26 | $0.91 | 105.5 |
| MiniMax M1 40k | MiniMax | minimax-m1-40k | | rsc | 10.0 | – | 68.2% | 7.8% | – | – | – | – |
| Nova 2.0 Pro Preview (Non-reasoning) | Amazon | nova-2-0-pro-preview | non-reasoning | rsc | 10.0 | – | 63.6% | 3.9% | – | $1.25 | $10.00 | 132.0 |
| gpt-oss-20b (low) | OpenAI | gpt-oss-20b | low | rsc | 10.0 | – | 61.1% | 5.3% | – | $0.070 | $0.21 | 258.7 |
| Qwen3 VL 235B A22B Instruct | Alibaba | qwen3-vl-235b-a22b-instruct | non-reasoning | rsc | 9.9 | – | 71.2% | 6.6% | – | $0.40 | $1.60 | – |
| North Mini Code | Cohere | north-mini-code | | rsc | 9.9 | $0.0 | 75.7% | 11.1% | 0.5% | $0.0 | $0.0 | 82.1 |
| GPT-5 mini (minimal) | OpenAI | gpt-5-mini | minimal | rsc | 9.9 | – | 68.7% | 5.1% | – | $0.25 | $2.00 | – |
| K2-V2 (high) | Institute of Foundation Models | k2-v2 | high | rsc | 9.9 | – | 68.1% | 10.5% | – | – | – | – |
| Gemini 2.5 Flash (Non-reasoning) | Google | gemini-2-5-flash | non-reasoning | rsc | 9.9 | – | 68.3% | 4.7% | – | $0.30 | $2.50 | – |
| Qwen3 30B A3B 2507 (Reasoning) | Alibaba | qwen3-30b-a3b-2507-reasoning | | rsc | 9.8 | $0.077 | 70.7% | 10.3% | 0.0% | $0.20 | $2.40 | – |
| o1-mini | OpenAI | o1-mini | | rsc | 9.8 | – | 60.3% | 3.6% | – | – | – | – |
| DeepSeek V3 0324 | DeepSeek | deepseek-v3-0324 | non-reasoning | rsc | 9.7 | $0.11 | 65.5% | 4.7% | 0.0% | $0.84 | $1.18 | – |
| Ling 2.6 Flash | InclusionAI | ling-2-6-flash | non-reasoning | rsc | 9.7 | – | 59.3% | 6.3% | – | – | – | – |
| Qwen3 Next 80B A3B Instruct | Alibaba | qwen3-next-80b-a3b-instruct | non-reasoning | rsc | 9.6 | – | 73.8% | 7.6% | – | $0.15 | $1.20 | 182.3 |
| Tri-21B-think Preview | Trillion Labs | tri-21b-think-preview | | rsc | 9.6 | – | 53.8% | 5.8% | – | – | – | – |
| Qwen3 Coder 30B A3B Instruct | Alibaba | qwen3-coder-30b-a3b-instruct | non-reasoning | rsc | 9.6 | – | 51.6% | 3.8% | – | $0.45 | $2.25 | – |
| GPT-4.5 (Preview) | OpenAI | gpt-4-5-preview | non-reasoning | rsc | 9.6 | – | – | – | – | – | – | – |
| DiffusionGemma 26B A4B | Google | diffusiongemma-26b-a4b | | rsc | 9.5 | – | 66.9% | 10.8% | – | – | – | – |
| Qwen3 235B A22B (Reasoning) | Alibaba | qwen3-235b-a22b-instruct | reasoning | rsc | 9.5 | – | 70.0% | 11.0% | – | $0.70 | $8.40 | – |
| QwQ 32B | Alibaba | qwq-32b | | rsc | 9.5 | – | 59.3% | 7.3% | – | $0.66 | $1.00 | – |
| Qwen3 VL 30B A3B (Reasoning) | Alibaba | qwen3-vl-30b-a3b | reasoning | rsc | 9.5 | – | 72.0% | 8.9% | – | $0.20 | $2.40 | – |
| Gemini 2.0 Flash Thinking Experimental (Jan '25) | Google | gemini-2-0-flash-thinking-exp-0121 | jan '25 | rsc | 9.4 | – | 70.1% | 6.3% | – | – | – | – |
| Gemma 4 12B (Non-reasoning) | Google | gemma-4-12b | non-reasoning | rsc | 9.4 | – | 66.1% | 6.3% | – | $0.10 | $0.30 | 144.7 |
| Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) | Google | gemini-2-5-flash-lite-preview-09-2025 | non-reasoning | rsc | 9.3 | – | 65.1% | 5.1% | – | $0.10 | $0.40 | – |
| Mistral Large 3 | Mistral | mistral-large-3 | non-reasoning | rsc | 9.3 | $0.031 | 68.0% | 4.2% | 0.0% | $0.50 | $1.50 | 76.9 |
| Qwen3 Coder Next | Alibaba | qwen3-coder-next | non-reasoning | rsc | 9.2 | $0.55 | 73.7% | 10.1% | 0.0% | $0.35 | $1.20 | 129.6 |
| Motif-2-12.7B-Reasoning | Motif Technologies | motif-2-12-7b-reasoning | | rsc | 9.2 | – | 69.5% | 8.8% | – | – | – | – |
| Mistral Medium 3.1 | Mistral | mistral-medium-3-1 | non-reasoning | rsc | 9.2 | – | 58.8% | 4.7% | 0.0% | – | – | – |
| Ling-1T | InclusionAI | ling-1t | non-reasoning | rsc | 9.2 | – | 71.9% | 7.3% | – | – | – | – |
| Nova Premier | Amazon | nova-premier | non-reasoning | rsc | 9.2 | – | 56.9% | 4.2% | – | – | – | – |
| Solar Pro 2 (Preview) (Reasoning) | Upstage | solar-pro-2-preview | high | rsc | 9.1 | – | 57.8% | 6.1% | – | – | – | – |
| Granite 4.2 3B | IBM | granite-4-2-3b | | rsc | 9.1 | $0.0060 | 55.9% | 6.6% | 0.0% | $0.030 | $0.12 | 231.9 |
| Magistral Medium 1 | Mistral | magistral-medium | | rsc | 9.1 | – | 67.9% | 9.8% | – | – | – | – |
| Mistral Medium 3 | Mistral | mistral-medium-3 | non-reasoning | rsc | 9.0 | – | 57.8% | 4.0% | – | $0.40 | $2.00 | – |
| K2-V2 (medium) | Institute of Foundation Models | k2-v2 | medium | rsc | 9.0 | – | 59.8% | 4.4% | – | – | – | – |
| Llama Nemotron Super 49B v1.5 (Reasoning) | NVIDIA | llama-nemotron-super-49b-v1-5 | reasoning | rsc | 9.0 | – | 74.8% | 7.3% | – | – | – | – |
| Devstral Medium | Mistral | devstral-medium | non-reasoning | rsc | 9.0 | – | 49.2% | 3.8% | – | – | – | – |
| Mistral Small 4 (Non-reasoning) | Mistral | mistral-small-4 | non-reasoning | rsc | 9.0 | – | 57.1% | 3.8% | – | $0.15 | $0.60 | 154.5 |
| Tri-21B-Think | Trillion Labs | tri-21b-think-v0-5 | | rsc | 9.0 | – | 60.1% | 5.9% | – | – | – | – |
| gpt-oss-20b (high) | OpenAI | gpt-oss-20b | high | rsc | 9.0 | $0.012 | 68.8% | 11.0% | 0.0% | $0.070 | $0.18 | 187.8 |
| GPT-4o (March 2025, chatgpt-4o-latest) | OpenAI | gpt-4o-chatgpt-03-25 | non-reasoning | rsc | 9.0 | – | 65.5% | 4.0% | – | – | – | – |
| Gemini 2.0 Flash (Feb '25) | Google | gemini-2-0-flash | non-reasoning | rsc | 8.9 | – | 62.3% | 4.3% | – | – | – | – |
| Claude 3.5 Haiku | Anthropic | claude-3-5-haiku | non-reasoning | rsc | 8.9 | – | 40.8% | 3.6% | – | – | – | – |
| Llama 3.3 Nemotron Super 49B v1 (Reasoning) | NVIDIA | llama-3-3-nemotron-super-49b | reasoning | rsc | 8.9 | – | 64.3% | 5.6% | – | – | – | – |
| Gemma 4 E4B (Reasoning) | Google | gemma-4-e4b | | rsc | 8.9 | – | 57.6% | 3.8% | – | $0.020 | $0.10 | 72.8 |
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) | NVIDIA | nvidia-nemotron-3-nano-30b-a3b | reasoning | rsc | 8.9 | $0.017 | 75.7% | 11.4% | 0.0% | $0.050 | $0.20 | 193.2 |
| Qwen3 4B 2507 (Reasoning) | Alibaba | qwen3-4b-2507-instruct-reasoning | | rsc | 8.8 | – | 66.7% | 6.2% | – | – | – | – |
| MiniCPM5-1B (Reasoning) | OpenBMB | minicpm5-1b | | rsc | 8.8 | – | 27.8% | 6.5% | – | – | – | – |
| Sarvam 105B (high) | Sarvam | sarvam-105b | high | rsc | 8.8 | – | 73.8% | 11.0% | – | $0.042 | $0.17 | – |
| Gemini 2.0 Pro Experimental (Feb '25) | Google | gemini-2-0-pro-experimental-02-05 | non-reasoning | rsc | 8.7 | – | 62.2% | 6.2% | – | – | – | – |
| Nova 2.0 Lite (Non-reasoning) | Amazon | nova-2-0-lite | non-reasoning | rsc | 8.7 | – | 60.3% | 2.9% | – | $0.30 | $2.50 | 225.1 |
| Devstral Small (May '25) | Mistral | devstral-small-2505 | non-reasoning | rsc | 8.7 | – | 43.4% | 4.0% | – | – | – | – |
| Claude 3 Opus | Anthropic | claude-3-opus | non-reasoning | rsc | 8.7 | – | 48.9% | 2.8% | – | $15.00 | $75.00 | – |
| MiniCPM5-1B (Non-reasoning) | OpenBMB | minicpm5-1b | non-reasoning | rsc | 8.7 | – | 26.9% | 4.5% | – | – | – | – |
| Sonar Reasoning | Perplexity | sonar-reasoning | | rsc | 8.7 | – | 62.3% | – | – | – | – | – |
| Gemini 2.5 Flash Preview (Non-reasoning) | Google | gemini-2-5-flash-04-2025 | non-reasoning | rsc | 8.7 | – | 59.4% | 3.8% | – | – | – | – |
| Devstral 2 | Mistral | devstral-2 | non-reasoning | rsc | 8.6 | – | 59.4% | 3.6% | 0.0% | – | – | – |
| Magistral Small 1.2 | Mistral | magistral-small-2509 | | rsc | 8.6 | – | 66.3% | 6.4% | – | $0.50 | $1.50 | – |
| Qwen3 32B (Reasoning) | Alibaba | qwen3-32b-instruct | reasoning | rsc | 8.6 | – | 66.8% | 7.4% | 0.0% | $0.16 | $0.64 | – |
| Gemini 2.5 Flash-Lite (Reasoning) | Google | gemini-2-5-flash-lite | reasoning | rsc | 8.5 | – | 62.5% | 6.8% | – | $0.10 | $0.40 | – |
| DeepSeek V3 (Dec '24) | DeepSeek | deepseek-v3 | non-reasoning | rsc | 8.5 | $0.020 | 55.7% | 2.9% | 0.0% | $0.32 | $0.89 | – |
| GPT-4o (Nov '24) | OpenAI | gpt-4o | non-reasoning | rsc | 8.4 | – | 54.3% | 2.4% | – | $2.50 | $10.00 | – |
| Nanbeige4.1-3B | Nanbeige | nanbeige4-1-3b | | rsc | 8.4 | – | 84.9% | 10.9% | – | – | – | – |
| LFM2.5-2.6B | Liquid AI | lfm2-5-2-6b | | rsc | 8.4 | – | 55.8% | 6.2% | – | $0.0 | $0.0 | – |
| Qwen3 VL 32B Instruct | Alibaba | qwen3-vl-32b-instruct | non-reasoning | rsc | 8.4 | – | 67.1% | 6.8% | – | $0.16 | $0.64 | – |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | deepseek-r1-distill-qwen-32b | | rsc | 8.4 | – | 61.5% | 4.6% | – | – | – | – |
| GLM-4.6V (Non-reasoning) | Z AI | glm-4-6v | non-reasoning | rsc | 8.4 | – | 56.6% | 3.7% | – | $0.30 | $0.90 | – |
| Qwen3 235B A22B (Non-reasoning) | Alibaba | qwen3-235b-a22b-instruct | non-reasoning | rsc | 8.3 | – | 61.3% | 4.2% | – | $0.70 | $2.80 | – |
| Mistral Small 3.2 | Mistral | mistral-small-3-2 | non-reasoning | rsc | 8.2 | – | 50.5% | 4.3% | 0.0% | $0.075 | $0.20 | – |
| Magistral Small 1 | Mistral | magistral-small | | rsc | 8.2 | – | 64.1% | 7.6% | – | – | – | – |
| Gemini 2.0 Flash (experimental) | Google | gemini-2-0-flash-experimental | non-reasoning | rsc | 8.2 | – | 63.6% | 4.1% | – | $0.0 | $0.0 | – |
| EXAONE 4.0 32B (Reasoning) | LG AI Research | exaone-4-0-32b | reasoning | rsc | 8.2 | – | 73.9% | 11.4% | – | – | – | – |
| Qwen3 VL 8B (Reasoning) | Alibaba | qwen3-vl-8b | reasoning | rsc | 8.2 | – | 57.9% | 3.8% | – | $0.18 | $2.10 | – |
| Qwen3 14B (Reasoning) | Alibaba | qwen3-14b-instruct | reasoning | rsc | 8.2 | – | 60.4% | 4.5% | 0.0% | $0.35 | $4.20 | – |
| Nova 2.0 Omni (Non-reasoning) | Amazon | nova-2-0-omni | non-reasoning | rsc | 8.2 | – | 55.5% | 3.8% | – | $0.30 | $2.50 | – |
| DeepSeek R1 0528 Qwen3 8B | DeepSeek | deepseek-r1-qwen3-8b | | rsc | 8.1 | – | 61.2% | 5.9% | – | – | – | – |
| Llama 4 Scout | Meta | llama-4-scout | non-reasoning | rsc | 8.1 | – | 58.7% | 3.8% | 0.0% | $0.19 | $0.68 | 131.4 |
| Qwen2.5 Max | Alibaba | qwen-2-5-max | non-reasoning | rsc | 8.0 | – | 58.7% | 3.8% | – | – | – | – |
| Qwen3 VL 30B A3B Instruct | Alibaba | qwen3-vl-30b-a3b-instruct | non-reasoning | rsc | 7.9 | – | 69.5% | 6.3% | – | $0.20 | $0.80 | – |
| Hermes 4 - Llama-3.1 70B (Reasoning) | Nous Research | hermes-4-llama-3-1-70b | reasoning | rsc | 7.9 | – | 69.9% | 8.8% | – | – | – | – |
| Gemini 1.5 Pro (Sep '24) | Google | gemini-1-5-pro | non-reasoning | rsc | 7.9 | – | 58.9% | 4.6% | – | – | – | – |
| Solar Pro 2 (Preview) (Non-reasoning) | Upstage | solar-pro-2-preview | non-reasoning | rsc | 7.9 | – | 54.4% | 3.7% | – | – | – | – |
| DeepSeek R1 Distill Llama 70B | DeepSeek | deepseek-r1-distill-llama-70b | | rsc | 7.9 | – | 40.2% | 5.1% | – | $0.70 | $1.10 | – |
| Claude 3.5 Sonnet (Oct '24) | Anthropic | claude-35-sonnet | non-reasoning | rsc | 7.9 | – | 59.9% | 3.7% | – | $3.00 | $15.00 | – |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | deepseek-r1-distill-qwen-14b | | rsc | 7.8 | – | 48.4% | 4.1% | – | – | – | – |
| Falcon-H1R-7B | TII UAE | falcon-h1r-7b | | rsc | 7.8 | – | 66.1% | 11.0% | – | – | – | – |
| GPT-4.1 nano | OpenAI | gpt-4-1-nano | non-reasoning | rsc | 7.8 | – | 51.2% | 3.8% | – | $0.10 | $0.40 | – |
| Solar Pro 3 | Upstage | solar-pro-3 | high | rsc | 7.8 | $0.079 | 72.4% | 10.3% | 0.0% | $0.15 | $0.60 | – |
| Ling-flash-2.0 | InclusionAI | ling-flash-2-0 | non-reasoning | rsc | 7.8 | – | 65.7% | 6.2% | – | $0.14 | $0.57 | – |
| Gemma 4 E2B (Reasoning) | Google | gemma-4-e2b | | rsc | 7.8 | – | 43.3% | 4.8% | – | – | – | – |
| Qwen3 Omni 30B A3B (Reasoning) | Alibaba | qwen3-omni-30b-a3b | reasoning | rsc | 7.8 | – | 72.6% | 7.5% | – | $0.25 | $0.97 | 111.0 |
| GPT-4o (Aug '24) | OpenAI | gpt-4o-2024-08-06 | non-reasoning | rsc | 7.7 | – | 52.1% | 2.3% | – | $2.50 | $10.00 | – |
| Qwen2.5 Instruct 72B | Alibaba | qwen2-5-72b-instruct | non-reasoning | rsc | 7.7 | – | 49.1% | 3.6% | – | $0.47 | $0.49 | – |
| Sonar | Perplexity | sonar | non-reasoning | rsc | 7.7 | – | 47.1% | 4.9% | – | – | – | – |
| Step3 VL 10B | StepFun | step3-vl-10b | | rsc | 7.7 | – | 69.0% | 10.8% | – | – | – | – |
| Llama 3.3 Instruct 70B | Meta | llama-3-3-instruct-70b | non-reasoning | rsc | 7.7 | – | 49.8% | 3.6% | – | $0.71 | $0.72 | 94.8 |
| Qwen3 30B A3B (Reasoning) | Alibaba | qwen3-30b-a3b-instruct | reasoning | rsc | 7.6 | – | 61.6% | 6.2% | – | $0.20 | $2.40 | – |
| Sonar Pro | Perplexity | sonar-pro | non-reasoning | rsc | 7.6 | – | 57.8% | 6.6% | – | – | – | – |
| Devstral Small (Jul '25) | Mistral | devstral-small | non-reasoning | rsc | 7.6 | – | 41.4% | 3.8% | – | – | – | – |
| QwQ 32B-Preview | Alibaba | QwQ-32B-Preview | | rsc | 7.6 | – | 55.7% | 3.9% | – | – | – | – |
| GLM-4.5V (Reasoning) | Z AI | glm-4-5v | reasoning | rsc | 7.6 | – | 68.4% | 6.3% | – | $0.60 | $1.80 | – |
| Mistral Large 2 (Nov '24) | Mistral | mistral-large-2 | non-reasoning | rsc | 7.6 | – | 48.6% | 3.3% | – | – | – | – |
| Devstral Small 2 | Mistral | devstral-small-2 | non-reasoning | rsc | 7.5 | – | 53.2% | 3.5% | 0.0% | – | – | – |
| Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) | NVIDIA | llama-3-1-nemotron-ultra-253b-v1 | reasoning | rsc | 7.5 | – | 72.8% | 7.4% | – | – | – | – |
| Qwen3 30B A3B 2507 Instruct | Alibaba | qwen3-30b-a3b-2507 | non-reasoning | rsc | 7.5 | – | 65.9% | 6.9% | – | $0.20 | $0.80 | – |
| ERNIE 4.5 300B A47B | Baidu | ernie-4-5-300b-a47b | non-reasoning | rsc | 7.5 | – | 81.1% | 3.3% | – | $0.28 | $1.10 | – |
| Hermes 4 - Llama-3.1 405B (Reasoning) | Nous Research | hermes-4-llama-3-1-405b | reasoning | rsc | 7.5 | – | 72.7% | 10.9% | – | $1.00 | $3.00 | 44.1 |
| Solar Pro 2 (Reasoning) | Upstage | solar-pro-2 | high | rsc | 7.5 | – | 68.7% | 7.4% | – | – | – | – |
| NVIDIA Nemotron Nano 12B v2 VL (Reasoning) | NVIDIA | nvidia-nemotron-nano-12b-v2-vl | reasoning | rsc | 7.5 | – | 57.2% | 5.5% | – | – | – | – |
| Gemma 4 E4B (Non-reasoning) | Google | gemma-4-e4b | non-reasoning | rsc | 7.5 | – | 54.9% | 4.8% | – | $0.020 | $0.10 | 73.5 |
| Granite 4.1 30B | IBM | granite-4-1-30b | non-reasoning | rsc | 7.4 | – | 48.1% | 4.1% | – | – | – | – |
| NVIDIA Nemotron Nano 9B V2 (Reasoning) | NVIDIA | nvidia-nemotron-nano-9b-v2 | reasoning | rsc | 7.4 | – | 57.0% | 4.9% | – | $0.040 | $0.16 | 90.8 |
| Hermes 4 - Llama-3.1 405B (Non-reasoning) | Nous Research | hermes-4-llama-3-1-405b | non-reasoning | rsc | 7.4 | – | 53.6% | 4.2% | – | $1.00 | $3.00 | 43.5 |
| Gemini 2.0 Flash-Lite (Feb '25) | Google | gemini-2-0-flash-lite-001 | non-reasoning | rsc | 7.4 | – | 53.5% | 3.4% | – | – | – | – |
| NVIDIA Nemotron 3 Nano 4B | NVIDIA | nvidia-nemotron-3-nano-4b | | rsc | 7.4 | – | 51.3% | 4.9% | – | – | – | – |
| Llama Nemotron Super 49B v1.5 (Non-reasoning) | NVIDIA | llama-nemotron-super-49b-v1-5 | non-reasoning | rsc | 7.4 | – | 48.1% | 4.3% | – | – | – | – |
| Qwen3 32B (Non-reasoning) | Alibaba | qwen3-32b-instruct | non-reasoning | rsc | 7.3 | – | 53.5% | 4.1% | – | $0.16 | $0.64 | – |
| GPT-4o (May '24) | OpenAI | gpt-4o-2024-05-13 | non-reasoning | rsc | 7.3 | – | 52.6% | 1.8% | – | $5.00 | $15.00 | – |
| Gemini 2.0 Flash-Lite (Preview) | Google | gemini-2-0-flash-lite-preview | non-reasoning | rsc | 7.3 | – | 54.2% | 4.1% | – | – | – | – |
| K2-V2 (low) | Institute of Foundation Models | k2-v2 | low | rsc | 7.3 | – | 54.1% | 3.6% | – | – | – | – |
| Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) | NVIDIA | llama-3-1-nemotron-nano-4b | reasoning | rsc | 7.3 | – | 40.8% | 5.2% | – | – | – | – |
| Kimi Linear 48B A3B Instruct | Kimi | kimi-linear-48b-a3b-instruct | non-reasoning | rsc | 7.3 | – | 41.2% | 2.5% | – | – | – | – |
| Llama 3.1 Instruct 405B | Meta | llama-3-1-instruct-405b | non-reasoning | rsc | 7.3 | – | 51.5% | 4.0% | – | – | – | – |
| Qwen3 8B (Reasoning) | Alibaba | qwen3-8b-instruct | reasoning | rsc | 7.3 | – | 58.9% | 3.9% | 0.0% | $0.18 | $2.10 | – |
| Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) | NVIDIA | llama-3-3-nemotron-super-49b | non-reasoning | rsc | 7.3 | – | 51.7% | 3.8% | – | – | – | – |
| Qwen3 VL 8B Instruct | Alibaba | qwen3-vl-8b-instruct | non-reasoning | rsc | 7.3 | – | 42.7% | 2.7% | – | $0.18 | $0.70 | – |
| Qwen3 4B (Reasoning) | Alibaba | qwen3-4b-instruct | reasoning | rsc | 7.2 | – | 52.2% | 4.4% | – | – | – | – |
| LFM2.5-8B-A1B | Liquid AI | lfm2-5-8b-a1b | | rsc | 7.2 | – | 51.3% | 6.9% | – | – | – | – |
| Claude 3.5 Sonnet (June '24) | Anthropic | claude-35-sonnet-june-24 | non-reasoning | rsc | 7.2 | – | 56.0% | 3.2% | – | $3.00 | $15.00 | – |
| Llama 3.1 Tulu3 405B | Allen Institute for AI | tulu3-405b | non-reasoning | rsc | 7.2 | – | 51.6% | 3.3% | – | – | – | – |
| GPT-4o (ChatGPT) | OpenAI | gpt-4o-chatgpt | non-reasoning | rsc | 7.2 | – | 51.1% | 2.7% | – | – | – | – |
| Ring-flash-2.0 | InclusionAI | ring-flash-2-0 | | rsc | 7.2 | – | 72.5% | 9.6% | – | $0.14 | $0.57 | – |
| Pixtral Large | Mistral | pixtral-large-2411 | non-reasoning | rsc | 7.1 | – | 50.5% | 2.8% | – | – | – | – |
| Olmo 3.1 32B Think | Allen Institute for AI | olmo-3-1-32b-think | | rsc | 7.1 | – | 59.1% | 6.3% | – | $0.0 | $0.0 | – |
| Mistral Small 3.1 | Mistral | mistral-small-3-1 | non-reasoning | rsc | 7.1 | $0.035 | 45.4% | 4.3% | 0.0% | $0.11 | $0.17 | – |
| Grok 2 (Dec '24) | SpaceXAI | grok-2-1212 | non-reasoning | rsc | 7.1 | – | 51.0% | 3.1% | – | – | – | – |
| GPT-5 nano (minimal) | OpenAI | gpt-5-nano | minimal | rsc | 7.1 | – | 42.8% | 4.0% | – | $0.050 | $0.40 | – |
| Gemini 1.5 Flash (Sep '24) | Google | gemini-1-5-flash-sep-24 | non-reasoning | rsc | 7.1 | – | 46.3% | 3.2% | – | – | – | – |
| Qwen3 VL 4B (Reasoning) | Alibaba | qwen3-vl-4b | reasoning | rsc | 7.0 | – | 49.4% | 4.6% | – | – | – | – |
| GPT-4 Turbo | OpenAI | gpt-4-turbo | non-reasoning | rsc | 7.0 | – | – | 3.1% | – | $10.00 | $30.00 | – |
| Solar Pro 2 (Non-reasoning) | Upstage | solar-pro-2 | non-reasoning | rsc | 7.0 | – | 56.1% | 3.7% | – | – | – | – |
| Nova Pro | Amazon | nova-pro | non-reasoning | rsc | 7.0 | – | 49.9% | 3.2% | – | $0.80 | $3.20 | – |
| Command A | Cohere | command-a | non-reasoning | rsc | 7.0 | – | 52.7% | 4.0% | – | $2.50 | $10.00 | 59.9 |
| Qwen3.5 2B (Reasoning) | Alibaba | qwen3-5-2b | | rsc | 6.9 | – | 45.6% | 2.6% | – | – | – | – |
| Llama 3.1 Nemotron Instruct 70B | NVIDIA | llama-3-1-nemotron-instruct-70b | non-reasoning | rsc | 6.9 | – | 46.5% | 4.2% | – | – | – | – |
| Llama 3.1 Instruct 8B | Meta | llama-3-1-instruct-8b | non-reasoning | rsc | 6.9 | – | 25.9% | 5.3% | – | $0.020 | $0.050 | – |
| Grok Beta | SpaceXAI | grok-beta | non-reasoning | rsc | 6.9 | – | 47.1% | 4.5% | – | – | – | – |
| Qwen2.5 Instruct 32B | Alibaba | qwen2.5-32b-instruct | non-reasoning | rsc | 6.9 | – | 46.6% | 4.0% | – | – | – | – |
| NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) | NVIDIA | nvidia-nemotron-3-nano-30b-a3b | non-reasoning | rsc | 6.8 | – | 39.9% | 4.6% | – | $0.050 | $0.20 | 233.7 |
| NVIDIA Nemotron Nano 9B V2 (Non-reasoning) | NVIDIA | nvidia-nemotron-nano-9b-v2 | non-reasoning | rsc | 6.8 | – | 55.7% | 4.6% | – | $0.050 | $0.20 | 159.7 |
| Mistral Large 2 (Jul '24) | Mistral | mistral-large-2407 | non-reasoning | rsc | 6.8 | – | 47.2% | 2.9% | – | $2.00 | $6.00 | – |
| Qwen3 4B 2507 Instruct | Alibaba | qwen3-4b-2507-instruct | non-reasoning | rsc | 6.7 | – | 51.7% | 4.5% | – | – | – | – |
| Qwen2.5 Coder Instruct 32B | Alibaba | qwen2-5-coder-32b-instruct | non-reasoning | rsc | 6.7 | – | 41.7% | 3.5% | – | – | – | – |
| Qwen3 14B (Non-reasoning) | Alibaba | qwen3-14b-instruct | non-reasoning | rsc | 6.7 | – | 47.0% | 4.1% | – | $0.35 | $1.40 | – |
| GPT-4 | OpenAI | gpt-4 | non-reasoning | rsc | 6.7 | – | 34.9% | – | – | $30.00 | $60.00 | – |
| GLM-4.5V (Non-reasoning) | Z AI | glm-4-5v | non-reasoning | rsc | 6.7 | – | 57.3% | 3.5% | – | $0.60 | $1.80 | – |
| Mistral Small 3 | Mistral | mistral-small-3 | non-reasoning | rsc | 6.7 | – | 46.2% | 3.8% | – | $0.050 | $0.080 | – |
| Gemini 2.5 Flash-Lite (Non-reasoning) | Google | gemini-2-5-flash-lite | non-reasoning | rsc | 6.7 | – | 47.4% | 3.7% | – | $0.10 | $0.40 | – |
| Nova Lite | Amazon | nova-lite | non-reasoning | rsc | 6.7 | – | 43.3% | 4.3% | – | $0.060 | $0.24 | – |
| GPT-4o mini | OpenAI | gpt-4o-mini | non-reasoning | rsc | 6.7 | – | 42.6% | 4.2% | – | $0.15 | $0.60 | – |
| Hermes 4 - Llama-3.1 70B (Non-reasoning) | Nous Research | hermes-4-llama-3-1-70b | non-reasoning | rsc | 6.7 | – | 49.1% | 3.6% | – | – | – | – |
| Qwen3 30B A3B (Non-reasoning) | Alibaba | qwen3-30b-a3b-instruct | non-reasoning | rsc | 6.6 | – | 51.5% | 4.6% | – | $0.20 | $0.80 | – |
| DeepSeek-V2.5 (Dec '24) | DeepSeek | deepseek-v2-5 | non-reasoning | rsc | 6.6 | – | 42.3% | – | – | – | – | – |
| Qwen3 4B (Non-reasoning) | Alibaba | qwen3-4b-instruct | non-reasoning | rsc | 6.6 | – | 39.8% | 3.3% | – | – | – | – |
| Llama 3.1 Instruct 70B | Meta | llama-3-1-instruct-70b | non-reasoning | rsc | 6.6 | – | 40.9% | 4.5% | – | $0.56 | $0.56 | – |
| Granite 4.1 8B | IBM | granite-4-1-8b | non-reasoning | rsc | 6.6 | – | 43.3% | 3.8% | – | – | – | – |
| Sarvam 30B (high) | Sarvam | sarvam-30b | high | rsc | 6.6 | – | 63.3% | 7.5% | – | $0.026 | $0.11 | – |
| Gemini 2.0 Flash Thinking Experimental (Dec '24) | Google | gemini-2-0-flash-thinking-exp-1219 | dec '24 | rsc | 6.6 | – | – | – | – | – | – | – |
| DeepSeek-V2.5 | DeepSeek | deepseek-v2-5-sep-2024 | non-reasoning | rsc | 6.6 | – | – | – | – | – | – | – |
| Olmo 3.1 32B Instruct | Allen Institute for AI | olmo-3-1-32b-instruct | non-reasoning | rsc | 6.5 | – | 53.9% | 5.0% | – | – | – | – |
| Mistral Saba | Mistral | mistral-saba | non-reasoning | rsc | 6.5 | – | 42.4% | 4.3% | – | – | – | – |
| DeepSeek R1 Distill Llama 8B | DeepSeek | deepseek-r1-distill-llama-8b | | rsc | 6.5 | – | 30.2% | – | – | – | – | – |
| Gemma 4 E2B (Non-reasoning) | Google | gemma-4-e2b | non-reasoning | rsc | 6.5 | – | 40.5% | 4.7% | – | – | – | – |
| Olmo 3 32B Think | Allen Institute for AI | olmo-3-32b-think | | rsc | 6.5 | – | 61.0% | 6.4% | – | – | – | – |
| Gemini 1.5 Pro (May '24) | Google | gemini-1-5-pro-may-24 | non-reasoning | rsc | 6.4 | – | 37.1% | 3.5% | – | – | – | – |
| R1 1776 | Perplexity | r1-1776 | | rsc | 6.4 | – | – | – | – | – | – | – |
| Qwen2.5 Turbo | Alibaba | qwen-turbo | non-reasoning | rsc | 6.4 | – | 41.0% | 4.1% | – | $0.050 | $0.20 | – |
| Reka Flash (Sep '24) | Reka AI | reka-flash | non-reasoning | rsc | 6.4 | – | – | – | – | $0.20 | $0.80 | – |
| Llama 3.2 Instruct 90B (Vision) | Meta | llama-3-2-instruct-90b-vision | non-reasoning | rsc | 6.4 | – | 43.2% | 4.5% | – | – | – | – |
| Solar Mini | Upstage | solar-mini | non-reasoning | rsc | 6.4 | – | – | – | – | $0.15 | $0.15 | – |
| Celeris-1 | Celeris | celeris-1 | non-reasoning | rsc | 6.3 | $0.050 | 63.1% | 6.8% | 0.0% | $0.20 | $0.70 | 2186.1 |
| Grok-1 | SpaceXAI | grok-1 | non-reasoning | rsc | 6.3 | – | – | – | – | – | – | – |
| Qwen2 Instruct 72B | Alibaba | qwen2-72b-instruct | non-reasoning | rsc | 6.3 | – | 37.1% | 3.7% | – | – | – | – |
| Phi-4 Mini Instruct | Microsoft | phi-4-mini | non-reasoning | rsc | 6.3 | – | 33.1% | 4.5% | – | $0.0 | $0.0 | 47.6 |
| EXAONE 4.0 32B (Non-reasoning) | LG AI Research | exaone-4-0-32b | non-reasoning | rsc | 6.3 | – | 62.8% | 5.0% | – | – | – | – |
| Qwen3.5 2B (Non-reasoning) | Alibaba | qwen3-5-2b | non-reasoning | rsc | 6.2 | – | 43.8% | 5.0% | – | – | – | – |
| Gemini 1.5 Flash-8B | Google | gemini-1-5-flash-8b | non-reasoning | rsc | 6.2 | – | 35.9% | 4.7% | – | – | – | – |
| Qwen3.5 0.8B (Reasoning) | Alibaba | qwen3-5-0-8b | | rsc | 6.1 | – | 11.1% | 1.1% | – | – | – | – |
| DeepHermes 3 - Mistral 24B Preview (Non-reasoning) | Nous Research | deephermes-3-mistral-24b-preview | non-reasoning | rsc | 6.1 | – | 38.2% | 3.8% | – | – | – | – |
| Jamba 1.7 Large | AI21 Labs | jamba-1-7-large | non-reasoning | rsc | 6.1 | – | 39.0% | 3.7% | – | – | – | – |
| Granite 4.0 H Small | IBM | granite-4-0-h-small | non-reasoning | rsc | 6.0 | – | 41.6% | 3.8% | – | $0.060 | $0.25 | – |
| Ministral 3 14B | Mistral | ministral-3-14b | non-reasoning | rsc | 6.0 | $0.019 | 57.2% | 4.6% | 0.0% | $0.20 | $0.20 | 91.3 |
| Jamba 1.5 Large | AI21 Labs | jamba-1-5-large | non-reasoning | rsc | 6.0 | – | 42.7% | 4.1% | – | $2.00 | $8.00 | – |
| Qwen3 Omni 30B A3B Instruct | Alibaba | qwen3-omni-30b-a3b-instruct | non-reasoning | rsc | 6.0 | – | 62.0% | 4.6% | – | $0.25 | $0.97 | 109.8 |
| Hermes 3 - Llama-3.1 70B | Nous Research | hermes-3-llama-3-1-70b | non-reasoning | rsc | 6.0 | – | 40.1% | 4.0% | – | $0.70 | $0.70 | – |
| Qwen3 8B (Non-reasoning) | Alibaba | qwen3-8b-instruct | non-reasoning | rsc | 6.0 | – | 45.2% | 1.9% | – | $0.18 | $0.70 | – |
| DeepSeek-Coder-V2 | DeepSeek | deepseek-coder-v2 | non-reasoning | rsc | 6.0 | – | – | – | – | – | – | – |
| OLMo 2 32B | Allen Institute for AI | olmo-2-32b | non-reasoning | rsc | 6.0 | – | 32.8% | 3.6% | – | – | – | – |
| Jamba 1.6 Large | AI21 Labs | jamba-1-6-large | non-reasoning | rsc | 6.0 | – | 38.7% | 3.6% | – | – | – | – |
| LFM2 24B A2B | Liquid AI | lfm2-24b-a2b | non-reasoning | rsc | 5.9 | – | 47.4% | 4.2% | – | – | – | – |
| Gemini 1.5 Flash (May '24) | Google | gemini-1-5-flash-may-24 | non-reasoning | rsc | 5.9 | – | 32.4% | 4.4% | – | – | – | – |
| Phi-4 | Microsoft | phi-4 | non-reasoning | rsc | 5.9 | – | 57.5% | 3.8% | – | $0.13 | $0.50 | 44.1 |
| Claude 3 Sonnet | Anthropic | claude-3-sonnet | non-reasoning | rsc | 5.9 | – | 40.0% | 3.6% | – | – | – | – |
| Nova Micro | Amazon | nova-micro | non-reasoning | rsc | 5.9 | – | 35.8% | 4.6% | – | $0.035 | $0.14 | 260.3 |
| Granite 4.1 3B | IBM | granite-4-1-3b | non-reasoning | rsc | 5.9 | – | 31.4% | 3.4% | – | – | – | – |
| Mistral Small (Sep '24) | Mistral | mistral-small | non-reasoning | rsc | 5.8 | – | 38.1% | 4.3% | – | – | – | – |
| Gemini 1.0 Ultra | Google | gemini-1-0-ultra | non-reasoning | rsc | 5.8 | – | – | – | – | – | – | – |
| Phi-3 Mini Instruct 3.8B | Microsoft | phi-3-mini | non-reasoning | rsc | 5.8 | – | 31.9% | 5.0% | – | – | – | – |
| NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) | NVIDIA | nvidia-nemotron-nano-12b-v2-vl | non-reasoning | rsc | 5.8 | – | 43.9% | 4.3% | – | $0.20 | $0.60 | 207.2 |
| Gemma 3n E4B Instruct Preview (May '25) | Google | gemma-3n-e4b-preview-0520 | non-reasoning | rsc | 5.8 | – | 27.8% | 4.8% | – | – | – | – |
| Phi-4 Multimodal Instruct | Microsoft | phi-4-multimodal | non-reasoning | rsc | 5.8 | – | 31.5% | 5.0% | – | $0.0 | $0.0 | 17.2 |
| Qwen2.5 Coder Instruct 7B | Alibaba | qwen2-5-coder-7b-instruct | non-reasoning | rsc | 5.8 | – | 33.9% | 4.9% | – | – | – | – |
| Mistral Large (Feb '24) | Mistral | mistral-large | non-reasoning | rsc | 5.8 | – | 35.1% | 3.5% | – | $4.00 | $12.00 | – |
| Mixtral 8x22B Instruct | Mistral | mistral-8x22b-instruct | non-reasoning | rsc | 5.7 | – | 33.2% | 4.0% | – | – | – | – |
| Llama 2 Chat 7B | Meta | llama-2-chat-7b | non-reasoning | rsc | 5.7 | – | 22.7% | 5.3% | – | $0.050 | $0.25 | – |
| Llama 3.2 Instruct 3B | Meta | llama-3-2-instruct-3b | non-reasoning | rsc | 5.7 | – | 25.5% | 5.3% | – | – | – | – |
| MiniCPM-V 4.6 1.3B | OpenBMB | minicpm-v-4-6-1-3b | non-reasoning | rsc | 5.7 | – | 30.5% | 5.1% | – | – | – | – |
| Jamba Reasoning 3B | AI21 Labs | jamba-reasoning-3b | | rsc | 5.7 | – | 33.3% | 3.8% | – | – | – | – |
| Qwen3 VL 4B Instruct | Alibaba | qwen3-vl-4b-instruct | non-reasoning | rsc | 5.7 | – | 37.1% | 3.6% | – | – | – | – |
| Qwen1.5 Chat 110B | Alibaba | qwen1.5-110b-chat | non-reasoning | rsc | 5.7 | – | 28.9% | – | – | – | – | – |
| Reka Flash 3 | Reka AI | reka-flash-3 | | rsc | 5.6 | – | 52.9% | 4.4% | – | $0.20 | $0.80 | 95.9 |
| Olmo 3 7B Think | Allen Institute for AI | olmo-3-7b-think | | rsc | 5.6 | – | 51.6% | 6.0% | – | – | – | – |
| Claude 2.1 | Anthropic | claude-21 | non-reasoning | rsc | 5.6 | – | 31.9% | 3.9% | – | – | – | – |
| Claude 3 Haiku | Anthropic | claude-3-haiku | non-reasoning | rsc | 5.6 | – | 37.4% | 4.1% | – | $0.25 | $1.25 | – |
| OLMo 2 7B | Allen Institute for AI | olmo-2-7b | non-reasoning | rsc | 5.6 | – | 28.8% | 5.4% | – | – | – | – |
| Molmo 7B-D | Allen Institute for AI | molmo-7b-d | non-reasoning | rsc | 5.6 | – | 24.0% | 5.1% | – | – | – | – |
| Ling-mini-2.0 | InclusionAI | ling-mini-2-0 | non-reasoning | rsc | 5.5 | – | 56.2% | 5.1% | – | – | – | – |
| DeepSeek R1 Distill Qwen 1.5B | DeepSeek | deepseek-r1-distill-qwen-1-5b | | rsc | 5.5 | – | 9.8% | 3.1% | – | – | – | – |
| Claude 2.0 | Anthropic | claude-2 | non-reasoning | rsc | 5.5 | – | 34.4% | – | – | – | – | – |
| DeepSeek-V2-Chat | DeepSeek | deepseek-v2 | non-reasoning | rsc | 5.5 | – | – | – | – | – | – | – |
| Mistral Small (Feb '24) | Mistral | mistral-small-2402 | non-reasoning | rsc | 5.5 | – | 30.2% | 4.1% | – | – | – | – |
| Mistral Medium | Mistral | mistral-medium | non-reasoning | rsc | 5.5 | – | 34.9% | 3.5% | – | – | – | – |
| GPT-3.5 Turbo | OpenAI | gpt-3-5-turbo | non-reasoning | rsc | 5.5 | – | 29.7% | – | – | $0.50 | $1.50 | – |
| Ministral 3 8B | Mistral | ministral-3-8b | non-reasoning | rsc | 5.5 | $0.011 | 47.1% | 4.3% | 0.0% | $0.15 | $0.15 | 104.6 |
| Llama 3 Instruct 70B | Meta | llama-3-instruct-70b | non-reasoning | rsc | 5.5 | – | 37.9% | 4.5% | – | $0.65 | $2.75 | – |
| Qwen Chat 72B | Alibaba | qwen-chat-72b | non-reasoning | rsc | 5.4 | – | – | – | – | – | – | – |
| Arctic Instruct | Snowflake | arctic-instruct | non-reasoning | rsc | 5.4 | – | – | – | – | – | – | – |
| LFM 40B | Liquid AI | lfm-40b | non-reasoning | rsc | 5.4 | – | 32.7% | 4.9% | – | – | – | – |
| Llama 3.2 Instruct 11B (Vision) | Meta | llama-3-2-instruct-11b-vision | non-reasoning | rsc | 5.4 | – | 22.1% | 5.5% | – | $0.34 | $0.34 | 14.1 |
| Qwen3.5 0.8B (Non-reasoning) | Alibaba | qwen3-5-0-8b | non-reasoning | rsc | 5.4 | – | 23.6% | 5.1% | – | – | – | – |
| PALM-2 | Google | palm-2 | non-reasoning | rsc | 5.4 | – | – | – | – | – | – | – |
| Gemini 1.0 Pro | Google | gemini-1-0-pro | non-reasoning | rsc | 5.3 | – | 27.7% | 4.2% | – | – | – | – |
| DeepSeek Coder V2 Lite Instruct | DeepSeek | deepseek-coder-v2-lite | non-reasoning | rsc | 5.3 | – | 31.9% | 5.4% | – | – | – | – |
| Sarvam M (Reasoning, based on Mistral Small 3.1) | Sarvam | sarvam-m | high | rsc | 5.3 | – | 41.6% | 3.1% | – | $0.0 | $0.0 | – |
| DeepSeek LLM 67B Chat (V1) | DeepSeek | deepseek-llm-67b-chat | non-reasoning | rsc | 5.3 | – | – | – | – | – | – | – |
| Llama 2 Chat 70B | Meta | llama-2-chat-70b | non-reasoning | rsc | 5.3 | – | 32.7% | 5.2% | – | – | – | – |
| Llama 2 Chat 13B | Meta | llama-2-chat-13b | non-reasoning | rsc | 5.3 | – | 32.1% | 4.8% | – | – | – | – |
| Command-R+ (Apr '24) | Cohere | command-r-plus-04-2024 | non-reasoning | rsc | 5.3 | – | 32.3% | 4.6% | – | – | – | – |
| OpenChat 3.5 (1210) | OpenChat | openchat-35 | non-reasoning | rsc | 5.3 | – | 23.0% | 4.8% | – | – | – | – |
| DBRX Instruct | Databricks | dbrx | non-reasoning | rsc | 5.3 | – | 33.1% | 2.9% | – | – | – | – |
| Exaone 4.0 1.2B (Reasoning) | LG AI Research | exaone-4-0-1-2b | reasoning | rsc | 5.3 | – | 51.5% | 6.0% | – | – | – | – |
| Olmo 3 7B Instruct | Allen Institute for AI | olmo-3-7b-instruct | non-reasoning | rsc | 5.2 | – | 40.0% | 5.8% | – | $0.10 | $0.20 | – |
| Exaone 4.0 1.2B (Non-reasoning) | LG AI Research | exaone-4-0-1-2b | non-reasoning | rsc | 5.2 | – | 42.4% | 5.7% | – | – | – | – |
| LFM2.5-1.2B-Thinking | Liquid AI | lfm2-5-1-2b-thinking | | rsc | 5.2 | – | 33.9% | 6.2% | – | – | – | – |
| Jamba 1.7 Mini | AI21 Labs | jamba-1-7-mini | non-reasoning | rsc | 5.2 | – | 32.2% | 4.4% | – | – | – | – |
| LFM2 2.6B | Liquid AI | lfm2-2-6b | non-reasoning | rsc | 5.2 | – | 30.6% | 5.5% | – | – | – | – |
| LFM2.5-1.2B-Instruct | Liquid AI | lfm2-5-1-2b-instruct | non-reasoning | rsc | 5.2 | – | 32.6% | 6.7% | – | – | – | – |
| Jamba 1.5 Mini | AI21 Labs | jamba-1-5-mini | non-reasoning | rsc | 5.2 | – | 30.2% | 5.1% | – | $0.20 | $0.40 | – |
| Granite 4.0 H 1B | IBM | granite-4-0-h-nano-1b | non-reasoning | rsc | 5.2 | – | 26.3% | 5.0% | – | – | – | – |
| Qwen3 1.7B (Reasoning) | Alibaba | qwen3-1.7b-instruct | reasoning | rsc | 5.2 | – | 35.6% | 4.6% | – | – | – | – |
| Jamba 1.6 Mini | AI21 Labs | jamba-1-6-mini | non-reasoning | rsc | 5.2 | – | 30.0% | 4.3% | – | – | – | – |
| Mixtral 8x7B Instruct | Mistral | mixtral-8x7b-instruct | non-reasoning | rsc | 5.1 | – | 29.2% | 4.7% | – | $0.45 | $0.70 | – |
| Gemma 3 270M | Google | gemma-3-270m | non-reasoning | rsc | 5.1 | – | 22.4% | 3.6% | – | – | – | – |
| Apertus 70B Instruct | Swiss AI Initiative | apertus-70b-instruct | non-reasoning | rsc | 5.1 | – | 27.2% | 5.5% | – | $0.82 | $2.92 | – |
| Granite 4.0 Micro | IBM | granite-4-0-micro | non-reasoning | rsc | 5.1 | – | 33.6% | 5.0% | – | – | – | – |
| DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) | Nous Research | deephermes-3-llama-3-1-8b-preview | non-reasoning | rsc | 5.1 | – | 27.0% | 4.3% | – | – | – | – |
| Qwen Chat 14B | Alibaba | qwen-chat-14b | non-reasoning | rsc | 5.0 | – | – | – | – | – | – | – |
| Claude Instant | Anthropic | claude-instant | non-reasoning | rsc | 5.0 | – | 33.0% | 3.5% | – | – | – | – |
| Command-R (Mar '24) | Cohere | command-r-03-2024 | non-reasoning | rsc | 5.0 | – | 28.4% | 4.8% | – | – | – | – |
| Llama 65B | Meta | llama-65b | non-reasoning | rsc | 5.0 | – | – | – | – | – | – | – |
| Mistral 7B Instruct | Mistral | mistral-7b-instruct | non-reasoning | rsc | 5.0 | – | 17.7% | 4.5% | – | $0.15 | $0.20 | – |
| Granite 4.0 1B | IBM | granite-4-0-nano-1b | non-reasoning | rsc | 5.0 | – | 28.1% | 4.8% | – | – | – | – |
| Molmo2-8B | Allen Institute for AI | molmo2-8b | non-reasoning | rsc | 5.0 | – | 42.5% | 4.3% | – | – | – | – |
| LFM2 8B A1B | Liquid AI | lfm2-8b-a1b | non-reasoning | rsc | 4.9 | – | 34.4% | 4.9% | – | – | – | – |
| Granite 3.3 8B (Non-reasoning) | IBM | granite-3-3-8b-instruct | non-reasoning | rsc | 4.9 | – | 33.8% | 4.2% | – | $0.030 | $0.25 | – |
| Qwen3 1.7B (Non-reasoning) | Alibaba | qwen3-1.7b-instruct | non-reasoning | rsc | 4.9 | – | 28.3% | 5.3% | – | – | – | – |
| Gemma 3 27B Instruct | Google | gemma-3-27b-instruct | non-reasoning | rsc | 4.9 | $0.0 | 42.8% | 4.4% | 0.0% | $0.0 | $0.0 | – |
| Ministral 3 3B | Mistral | ministral-3-3b | non-reasoning | rsc | 4.8 | $0.0078 | 35.8% | 5.4% | 0.0% | $0.10 | $0.10 | 238.9 |
| Qwen3 0.6B (Non-reasoning) | Alibaba | qwen3-0.6b-instruct | non-reasoning | rsc | 4.8 | – | 23.1% | 4.9% | – | – | – | – |
| Qwen3 0.6B (Reasoning) | Alibaba | qwen3-0.6b-instruct | reasoning | rsc | 4.8 | – | 23.9% | 5.6% | – | – | – | – |
| Tiny Aya Global | Cohere | tiny-aya-global | non-reasoning | rsc | 4.8 | – | 30.5% | 5.2% | – | $0.0 | $0.0 | 131.5 |
| Gemma 3 1B Instruct | Google | gemma-3-1b-instruct | non-reasoning | rsc | 4.8 | – | 23.7% | 5.3% | – | $0.0 | $0.0 | – |
| Gemma 3 4B Instruct | Google | gemma-3-4b-instruct | non-reasoning | rsc | 4.8 | – | 29.1% | 5.3% | – | $0.0 | $0.0 | – |
| Gemma 3n E2B Instruct | Google | gemma-3n-e2b-instruct | non-reasoning | rsc | 4.8 | – | 22.9% | 4.2% | – | $0.0 | $0.0 | – |
| Gemma 3n E4B Instruct | Google | gemma-3n-e4b-instruct | non-reasoning | rsc | 4.8 | – | 29.6% | 4.5% | – | – | – | – |
| Granite 4.0 350M | IBM | granite-4-0-350m | non-reasoning | rsc | 4.8 | – | 26.1% | 5.5% | – | – | – | – |
| Granite 4.0 H 350M | IBM | granite-4-0-h-350m | non-reasoning | rsc | 4.8 | – | 25.7% | 6.4% | – | – | – | – |
| LFM2 1.2B | Liquid AI | lfm2-1-2b | non-reasoning | rsc | 4.8 | – | 22.8% | 5.6% | – | – | – | – |
| LFM2.5-VL-1.6B | Liquid AI | lfm2-5-vl-1-6b | non-reasoning | rsc | 4.8 | – | 28.9% | 5.1% | – | – | – | – |
| Llama 3 Instruct 8B | Meta | llama-3-instruct-8b | non-reasoning | rsc | 4.8 | – | 29.6% | 5.1% | – | $0.045 | $0.14 | – |
| Llama 3.2 Instruct 1B | Meta | llama-3-2-instruct-1b | non-reasoning | rsc | 4.8 | – | 19.6% | 5.5% | – | – | – | – |
| Apertus 8B Instruct | Swiss AI Initiative | apertus-8b-instruct | non-reasoning | rsc | 4.8 | – | 25.6% | 5.0% | – | $0.10 | $0.20 | – |
| Gemma 3 12B Instruct | Google | gemma-3-12b-instruct | non-reasoning | rsc | 3.8 | $0.0 | 34.9% | 4.2% | 0.0% | $0.0 | $0.0 | – |
| K2 Horizon 0.9B | Institute of Foundation Models | k2-horizon-0-9b | | rsc | 3.0 | – | 29.3% | 5.4% | 0.0% | – | – | – |
| Cogito v2.1 (Reasoning) | Deep Cogito | cogito-v2-1 | reasoning | rsc | – | – | 76.8% | 12.0% | – | $1.25 | $1.25 | – |
| Gemini 3 Deep Think | Google | gemini-3-deep-think | | rsc | – | – | – | – | – | – | – | – |
| Mi:dm K 2.5 Pro Preview | Korea Telecom | midm-250-pro-rsnsft | | rsc | – | – | 72.2% | 9.1% | – | – | – | – |
| EXAONE 4.5 33B (Non-reasoning) | LG AI Research | exaone-4-5-33b | non-reasoning | rsc | – | – | – | – | – | – | – | – |
| GPT-3.5 Turbo (0613) | OpenAI | gpt-3-5-turbo-0613 | non-reasoning | rsc | – | – | – | – | – | – | – | – |
| GPT-4o Realtime (Dec '24) | OpenAI | gpt-4o-realtime-dec-24 | non-reasoning | rsc | – | – | – | – | – | – | – | – |
| GPT-4o mini Realtime (Dec '24) | OpenAI | gpt-4o-mini-realtime-dec-24 | non-reasoning | rsc | – | – | – | – | – | – | – | – |
| GPT-5.4 Pro (xhigh) | OpenAI | gpt-5-4-pro | xhigh | rsc | – | – | – | – | – | $30.00 | $180 | – |
| GPT-5.5 Pro (xhigh) | OpenAI | gpt-5-5-pro | xhigh | rsc | – | – | – | – | – | – | – | – |