AI Model Leaderboard
Every model ranked by one number — the Artificial Analysis Intelligence Index, which averages ~10 independent benchmarks into a single capability score.
Capability only — cost is not part of the score.
| # | Model | Intelligence Index |
|---|---|---|
| 1 | Claude Fable 5 Anthropic | 59.9 |
| 2 | GPT-5.6 Sol OpenAI | 58.9 |
| 3 | Kimi K3 Moonshot AI | 57.1 |
| 4 | Claude Opus 4.8 Anthropic | 55.7 |
| 5 | GPT-5.6 Terra OpenAI | 55.0 |
| 6 | GPT-5.5 OpenAI | 54.8 |
| 7 | Grok 4.5 xAI | 53.8 |
| 8 | Claude Opus 4.7 Anthropic | 53.5 |
| 9 | Claude Sonnet 5 Anthropic | 53.4 |
| 10 | GPT-5.4 OpenAI | 51.4 |
| 11 | GPT-5.6 Luna OpenAI | 51.2 |
| 12 | GLM-5.2 Z.AI | 51.1 |
| 13 | Muse Spark 1.1 Meta | 50.6 |
| 14 | Gemini 3.5 Flash Google | 50.2 |
| 15 | Gemini 3.6 Flash Google | 50.1 |
| 16 | Gemini 3.1 Pro Google | 46.5 |
| 17 | Qwen3.7 Max Alibaba | 46.0 |
| 18 | MiniMax M3 MiniMax | 44.4 |
| 19 | DeepSeek V4 Pro DeepSeek | 44.3 |
| 19 | GPT-5.3 Codex OpenAI | 44.3 |
Scores are the Artificial Analysis Intelligence Index (0–100, higher is better), an aggregate of ~10 independent benchmarks including GPQA Diamond, SciCode, Terminal-Bench and Humanity's Last Exam. Data captured July 23, 2026; one representative configuration per model. Want cost factored in too? Use the interactive compare tool.