# The best AI coding agents in 2026, by independent rankings

Most leaderboards rank models. These rank the agents you actually install, plus what developers say they use.

Published 2026-10-06 by Richard Kaminsky and Mitchell Lipyansky, the co-founders of Poly. Canonical: https://usepoly.co/best-ai-coding-agents
Poly is a multiplayer AI coding workspace: a shared room where your team works with one AI agent, together. Free to start: https://usepoly.co/

**As of October 2026, Claude Code and OpenAI's Codex are the strongest AI coding agents on the independent benchmarks that test agents rather than bare models. On Terminal-Bench 4.0, Claude Code with Opus 5.5 leads at 64.8%, followed by Claude Code with Sonnet 5.5 at 61.8% and Codex with GPT-6 Astra or GPT-6.1 Sol at 58.2%. Artificial Analysis's coding-agent index ranks Claude Code with Sonnet 5.5 first at 68, ahead of Claude Code with Opus 5.5 at 66, Google's Antigravity CLI at 64 and Codex at 63. Developers agree: in Stack Overflow's 2026 survey, 65.5% of agent users had used Claude Code, ahead of GitHub Copilot at 58.7%.**

## Key takeaways

- Only a few leaderboards test the agent product, meaning the harness and the model together: Terminal-Bench, Artificial Analysis and SWE-rebench. Most others rank models in one generic harness.
- The harness matters: Grok 4.7 scores 37.6% in xAI's own Grok Build but 28.8% in a generic harness on Vals AI.
- Best value at the top: Codex with GPT-6.1 Sol tied GPT-6 Astra's 58.2% on Terminal-Bench at about a fifth of the cost.
- Free agents exist, but the free models trail: Codex with GPT-6 Luna, the model on ChatGPT Free, scores 16.4% on Terminal-Bench.
- At work, JetBrains counts 39% of developers using Claude Code, 21% GitHub Copilot, 16% Codex and 12% Cursor.

## Terminal-Bench 4.0

Terminal-Bench gives each agent 66 hard tasks in a terminal, with up to eight hours each, and reports the share solved ([tbench.ai](https://www.tbench.ai/)).

*Terminal-Bench 4.0, top agent and model pairs, tbench.ai, October 6, 2026.*

| Rank | Agent | Model | Tasks solved |
| --- | --- | --- | --- |
| 1 | Claude Code | Opus 5.5 (max) | 64.8% |
| 2 | Claude Code | Sonnet 5.5 (max) | 61.8% |
| 3 | Codex | GPT-6 Astra (max) | 58.2% |
| 3 | Codex | GPT-6.1 Sol (max) | 58.2% |
| 5 | Claude Code | Fable 5.1 (max) | 57.9% |
| 6 | Claude Code | Opus 5 (xhigh) | 53.9% |
| 7 | Codex | GPT-6 Sol (max) | 49.4% |
| 10 | Grok Build | Grok 4.7 (xhigh) | 37.6% |

## Artificial Analysis coding-agent index

Artificial Analysis averages three benchmarks, DeepSWE, Terminal-Bench 4.0 and Scale AI's SWE-Atlas Q&A, for each agent and model pair ([Artificial Analysis](https://artificialanalysis.ai/agents/coding-agents)):

1. Claude Code with Sonnet 5.5: 68
2. Claude Code with Opus 5.5: 66
3. Antigravity CLI with Gemini 4 Argon: 64 (a model not yet publicly available)
4. Codex with GPT-6.1 Sol: 63
5. Claude Code with Fable 5.1: 62
6. Codex with GPT-6 Astra: 62
7. Grok Build with Grok 4.7: 56
8. Muse Code with Meta's Muse Spark 1.3: 54

## What do the other leaderboards say?

- **SWE-rebench** runs agent products on fresh GitHub problems. Its May to July 2026 window put JetBrains Junie, Claude Code and Codex near the top, with Cursor running its own Composer model further back ([SWE-rebench](https://swe-rebench.com/)).
- **Vals AI** runs every model in one harness on Terminal-Bench 4.0: Opus 5.5 65.15%, Sonnet 5.5 64.14%, GPT-6 Astra 59.60% ([Vals AI](https://www.vals.ai/benchmarks/terminal-bench-4)).
- **Arena's WebDev board**, from human votes, ranks Opus 5.5 first at 1815, ahead of GPT-6 Astra and Sonnet 5.5.
- **SWE-bench Verified** has had no new entries since February 2026, and **SWE-Bench Pro** doesn't yet list the current Claude or GPT models, so both are out of date.

## The agents worth knowing

- **Claude Code** (Anthropic): terminal, VS Code, JetBrains, desktop and web. From Claude Pro at $20 a month ($17 billed annually); no free plan.
- **Codex** (OpenAI): app, CLI, IDE and cloud. On every ChatGPT plan, Free included with GPT-6 Luna in the desktop app; Plus is $20. The CLI is open source (Apache 2.0).
- **GitHub Copilot:** editor, CLI and a cloud agent. Free with a limited allowance; Pro $10, Pro+ $39, Max $100.
- **Cursor:** an AI editor with agents and a CLI. Free Hobby plan with limited agent requests; Pro $20.
- **Devin** (Cognition): Devin Desktop, the former Windsurf, plus cloud and CLI agents. A free light quota; Pro $20.
- **Google Antigravity:** an agent IDE and CLI with a free weekly quota; Google AI Pro at $19.99 adds Claude Sonnet 5.5.
- **Kiro** (AWS): an IDE and CLI with 50 free credits a month; Pro $20.
- **Grok Build** (xAI): a terminal agent on every Grok plan, free included, and open source under Apache 2.0.
- **JetBrains Junie:** JetBrains' agent, generally available since June 2026.
- **Open source:** OpenCode, Cline and Kilo Code run with your own keys or local models.

Detailed comparisons: [Claude Code vs Codex](/claude-code-vs-codex), [Claude Code vs GitHub Copilot](/claude-code-vs-copilot), [Devin vs Claude Code](/devin-vs-claude-code) and [Claude Code alternatives](/claude-code-alternatives).

## What is the best free AI coding agent?

Codex on ChatGPT Free, Google Antigravity's free quota, Grok Build on Grok's free plan, Kiro's 50 monthly credits and Copilot Free all run a real agent at no cost, and open-source agents run free with local models. The trade-off is the model: the free tiers run smaller models that score far below the leaders. More in [The best free AI for coding](/best-free-ai-for-coding).

## What do developers actually use?

- **Stack Overflow's 2026 survey** (30,903 responses; 12,255 answered the agent question) asked which agents people had used in the past year: Claude Code 65.5%, GitHub Copilot 58.7%, Codex 29.5%, Cursor 25.5%, Google Antigravity 16.0% ([Stack Overflow](https://survey.stackoverflow.co/2026/ai/data)).
- **JetBrains' 2026 survey** of more than 15,000 developers, May to July: 39% use Claude Code at work, up from 18% in January, and 31% call it their main tool; GitHub Copilot fell to 21%, and Codex rose from 3% to 16% ([JetBrains](https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/)).

## How to choose

- **The strongest results:** Claude Code with Opus 5.5 or Sonnet 5.5, or Codex with GPT-6.1 Sol.
- **The most for the money:** Codex with GPT-6.1 Sol, or Claude Code with Sonnet 5.5.
- **You live in an editor:** Cursor or GitHub Copilot.
- **Free:** Codex on ChatGPT Free, Antigravity or Grok Build.
- **Open source:** OpenCode or Cline.
- **A team working with one agent:** Poly (usepoly.co) runs Claude Code and Codex, the two leaders on these boards, in a shared room where everyone sees each step and any member can approve a change before it happens. Free to start. [What is Poly?](/what-is-poly)

## Common questions

**What is the best AI coding agent in 2026?**

On independent agent benchmarks, Claude Code: it leads Terminal-Bench 4.0 with Opus 5.5 at 64.8% and Artificial Analysis's index with Sonnet 5.5 at 68. Codex with GPT-6.1 Sol or GPT-6 Astra is close behind.

**What is the best free AI coding agent?**

Codex on ChatGPT Free, Google Antigravity's free quota and Grok Build on Grok's free plan all run a real agent at no cost, but their free models score well below the paid leaders.

**Is Claude Code better than Codex?**

On Terminal-Bench 4.0 and Artificial Analysis's index, Claude Code's best configurations score a few points higher. Codex with GPT-6.1 Sol reaches 58.2% at a much lower cost.

**Which AI coding agent do most developers use?**

Claude Code. In Stack Overflow's 2026 survey 65.5% of agent users had used it, and JetBrains found 39% of developers use it at work, ahead of GitHub Copilot at 21%.

**Why do AI coding leaderboards disagree?**

Most rank models in one generic harness, not the agents people install, and they use different tasks. The same model can score very differently in its own agent than in a generic one.

More guides: https://usepoly.co/guides · Security: https://usepoly.co/security · Pricing: https://usepoly.co/pricing
