Source: Artificial Analysis β€” Composite average pass@1 across SWE-Bench-Pro-Hard-AA, Terminal-Bench v2, and SWE-Atlas-QnA.
Auto-updated weekly Β· Last update: 2026-08-10 12:06 UTC

#AgentModelProviderScore
1Claude CodeOpus 5xhigh67
2CodexGPT-5.6 Solmax67
3Claude CodeFable 5max66
4Grok BuildGrok 4.5high64
5Kimi Code CLIKimi K361
6OpencodeMuse Spark 1.1xhigh54
7OpencodeGemini 3.6 Flashhigh46
8Claude CodeGLM-5.243
9Cursor CLIComposer 2.5 Fast38
10Claude CodeDeepSeek V4 Prohigh31

⏱ Time per Task

Mean wall clock time per task (lower is better)

#AgentWall Time
1Cursor CLI - Composer 2.5 Fast (Cursor)6.8m
2Codex - GPT-5.6 Sol (max) (OpenAI)10.2m
3Opencode - Gemini 3.6 Flash (high) (Google)10.4m
4Opencode - Muse Spark 1.1 (xhigh) (Meta)12.6m
5Grok Build - Grok 4.5 (high) (SpaceXAI)16.5m
6Claude Code - DeepSeek V4 Pro (high) (DeepSeek)17.9m
7Claude Code - Fable 5 (max) (with fallback) (Anthropic)23.4m
8Claude Code - Opus 5 (xhigh) (Anthropic)23.6m
9Kimi Code CLI - Kimi K3 (Moonshot AI)23.8m
10Claude Code - GLM-5.2 (Novita)25.1m

πŸ’° Cost per Task

Mean API cost per task in USD (lower is better)

#AgentCost (USD)
1Claude Code - DeepSeek V4 Pro (high) (DeepSeek)$0.27
2Cursor CLI - Composer 2.5 Fast (Cursor)$0.55
3Opencode - Muse Spark 1.1 (xhigh) (Meta)$1.43
4Opencode - Gemini 3.6 Flash (high) (Google)$2.08
5Grok Build - Grok 4.5 (high) (SpaceXAI)$2.59
6Kimi Code CLI - Kimi K3 (Moonshot AI)$3.18
7Claude Code - GLM-5.2 (Novita)$6.51
8Codex - GPT-5.6 Sol (max) (OpenAI)$7.08
9Claude Code - Opus 5 (xhigh) (Anthropic)$8.23
10Claude Code - Fable 5 (max) (with fallback) (Anthropic)$11.7

About the Benchmarks

  • SWE-Bench-Pro-Hard-AA β€” Code generation, 150 questions (Scale AI)
  • Terminal-Bench v2 β€” Agentic terminal use, 84 questions (Laude Institute)
  • SWE-Atlas-QnA β€” Technical Q&A, 124 questions (Scale AI)

The index represents the average pass@1 across 3 runs of each benchmark.


Data scraped weekly by an AI Agent. For the latest results, visit the original page.