Source: Artificial Analysis β Composite average pass@1 across SWE-Bench-Pro-Hard-AA, Terminal-Bench v2, and SWE-Atlas-QnA.
Auto-updated weekly Β· Last update: 2026-08-10 12:06 UTC
| # | Agent | Model | Provider | Score |
|---|---|---|---|---|
| 1 | Claude Code | Opus 5 | xhigh | 67 |
| 2 | Codex | GPT-5.6 Sol | max | 67 |
| 3 | Claude Code | Fable 5 | max | 66 |
| 4 | Grok Build | Grok 4.5 | high | 64 |
| 5 | Kimi Code CLI | Kimi K3 | 61 | |
| 6 | Opencode | Muse Spark 1.1 | xhigh | 54 |
| 7 | Opencode | Gemini 3.6 Flash | high | 46 |
| 8 | Claude Code | GLM-5.2 | 43 | |
| 9 | Cursor CLI | Composer 2.5 Fast | 38 | |
| 10 | Claude Code | DeepSeek V4 Pro | high | 31 |
β± Time per Task
Mean wall clock time per task (lower is better)
| # | Agent | Wall Time |
|---|---|---|
| 1 | Cursor CLI - Composer 2.5 Fast (Cursor) | 6.8m |
| 2 | Codex - GPT-5.6 Sol (max) (OpenAI) | 10.2m |
| 3 | Opencode - Gemini 3.6 Flash (high) (Google) | 10.4m |
| 4 | Opencode - Muse Spark 1.1 (xhigh) (Meta) | 12.6m |
| 5 | Grok Build - Grok 4.5 (high) (SpaceXAI) | 16.5m |
| 6 | Claude Code - DeepSeek V4 Pro (high) (DeepSeek) | 17.9m |
| 7 | Claude Code - Fable 5 (max) (with fallback) (Anthropic) | 23.4m |
| 8 | Claude Code - Opus 5 (xhigh) (Anthropic) | 23.6m |
| 9 | Kimi Code CLI - Kimi K3 (Moonshot AI) | 23.8m |
| 10 | Claude Code - GLM-5.2 (Novita) | 25.1m |
π° Cost per Task
Mean API cost per task in USD (lower is better)
| # | Agent | Cost (USD) |
|---|---|---|
| 1 | Claude Code - DeepSeek V4 Pro (high) (DeepSeek) | $0.27 |
| 2 | Cursor CLI - Composer 2.5 Fast (Cursor) | $0.55 |
| 3 | Opencode - Muse Spark 1.1 (xhigh) (Meta) | $1.43 |
| 4 | Opencode - Gemini 3.6 Flash (high) (Google) | $2.08 |
| 5 | Grok Build - Grok 4.5 (high) (SpaceXAI) | $2.59 |
| 6 | Kimi Code CLI - Kimi K3 (Moonshot AI) | $3.18 |
| 7 | Claude Code - GLM-5.2 (Novita) | $6.51 |
| 8 | Codex - GPT-5.6 Sol (max) (OpenAI) | $7.08 |
| 9 | Claude Code - Opus 5 (xhigh) (Anthropic) | $8.23 |
| 10 | Claude Code - Fable 5 (max) (with fallback) (Anthropic) | $11.7 |
About the Benchmarks
- SWE-Bench-Pro-Hard-AA β Code generation, 150 questions (Scale AI)
- Terminal-Bench v2 β Agentic terminal use, 84 questions (Laude Institute)
- SWE-Atlas-QnA β Technical Q&A, 124 questions (Scale AI)
The index represents the average pass@1 across 3 runs of each benchmark.
Data scraped weekly by an AI Agent. For the latest results, visit the original page.