Comparison

Codex vs Cursor in 2026: OpenAI's agent or the AI editor?

One is an agent that comes with ChatGPT; the other is an editor that runs everyone's models. And OpenAI is pulling its models out of Cursor.

Codex and Cursor tackle the same work from opposite ends. Codex is OpenAI's coding agent, included in every ChatGPT plan from Free up, running OpenAI's models in the ChatGPT app, a command-line tool, IDE extensions and the cloud. Cursor is an AI editor, owned by SpaceX since August 2026, that runs many companies' models, including its own Composer and xAI's Grok, from $20 a month. Only Codex appears on the independent agent benchmarks: with GPT-6.1 Sol it scores 58.2% on Terminal-Bench 4.0 and 63 on Artificial Analysis's index. And the two are separating: OpenAI plans to stop supplying its models to Cursor, with a proposed transition period to November 12, 2026, though the Codex extension still runs inside Cursor.

Key takeaways

  • Price: Codex comes with ChatGPT, from Free to Plus at $20 and Pro at $100 to $500. Cursor has a free Hobby plan, Pro at $20, Pro+ at $60 and Ultra at $200.
  • Usage: Codex meters a five-hour window plus weekly limits, then credits. Cursor gives two monthly pools, then bills on-demand at API prices.
  • Models: Codex runs OpenAI's models only. Cursor runs Composer, Grok, Claude, Gemini and others.
  • Benchmarks: Codex is on Terminal-Bench 4.0 and Artificial Analysis's index. Cursor is on neither and publishes its own numbers.
  • Adoption: Ramp's data shows 40% of businesses paying for coding agents use Codex and 35% use Cursor, behind Claude Code at 87%.
Codex and Cursor, from OpenAI's and Cursor's pricing pages and docs and the Terminal-Bench leaderboard, October 7, 2026.
CodexCursor
Made byOpenAICursor, owned by SpaceX
What it isA coding agentAn AI editor with agents
ModelsOpenAI'sComposer, Grok, Claude, Gemini and others
Free optionChatGPT Free (GPT-6 Luna)Hobby, limited agent requests
Paid fromChatGPT Plus, $20 (Go, $8)Pro, $20
Heavy useChatGPT Pro, $100 to $500Pro+ $60, Ultra $200
TeamsBusiness $20 to $125 a seatTeams $40, Premium $120 a seat
Where it runsChatGPT app, CLI, IDE extensions, cloudCursor editor, CLI, cloud agents, iPhone
Open sourceThe CLI (Apache 2.0)No
Terminal-Bench 4.058.2%Not listed

How do Codex and Cursor work?

Codex is an agent: you describe a task, and it reads the code, runs commands and makes changes, in the ChatGPT desktop app's Codex mode, the Codex CLI, extensions for VS Code and its forks, or in OpenAI's cloud. One harness, the code that wraps the model, powers all of them through OpenAI's open-source App Server (OpenAI).

Cursor is an editor built on VS Code, with agents inside it. Cursor 3, released in April 2026, added an Agents Window that runs several local and cloud agents in parallel; there is also a Cursor CLI, cloud agents in Cursor's virtual machines, and remote control from an iPhone.

How much do Codex and Cursor cost?

Codex (OpenAI):

  • ChatGPT Free and Go ($8): GPT-6 Luna in the desktop app.
  • Plus, $20: Codex everywhere, including the cloud.
  • Pro, $100, $200 or $500: much higher allowances and no five-hour limit.
  • Business: $20 to $25 a standard seat, $100 to $125 premium (Codex for teams).

Cursor (Cursor):

How do their usage limits work?

Codex gives Plus and standard Business users an allowance per five-hour window, measured in messages that vary by model, and "weekly limits may also apply." Local and cloud work share it. Past the allowance, Plus and Pro users buy credits or switch to an API key. Fast modes spend the allowance faster.

Cursor gives each plan two monthly pools: one for its own models (Composer 2.5 and Grok) and one for other companies' models at their API prices. When a pool runs out you pay on demand at the same rates, and Cursor says requests are "never downgraded." Its own guide says daily agent users typically spend $60 to $100 a month in total, and power users $200 or more. Since July 2026, Auto mode bills at the price of whichever model it routes to.

What is the harness, and why does it matter?

The harness is everything around the model: its tools, how it manages context and how it edits files. OpenAI tunes one harness for its own models. Cursor tunes one harness for many models, giving each the edit format it was trained on, patches for OpenAI's and string replacement for Anthropic's, and measures changes by "Keep Rate," how much of the agent's code survives in your codebase (Cursor).

The harness changes results: the same model can score very differently in different agents (the rankings).

Which is better on benchmarks?

  • Codex: with GPT-6.1 Sol or GPT-6 Astra, 58.2% on Terminal-Bench 4.0; with GPT-6.1 Sol, 63 on Artificial Analysis's coding-agent index.
  • Cursor: not listed on either. It reports its own results for Composer 2.5, which is built on Moonshot's Kimi K2.5, on older and in-house benchmarks that don't compare directly (Cursor).
  • SWE-rebench, which ran agent products on fresh problems from May to July 2026, had Codex with GPT-5.6 Sol at 58.0% and Cursor with Composer 2.5 at 51.7%.

How do their features compare?

  • Cloud agents: both. Codex Cloud uses reusable environments; Cursor's cloud agents can work across several repositories and run automations on a schedule.
  • Code review: Codex reviews GitHub pull requests with @codex review or automatically. Cursor's Bugbot reviews on GitHub, GitLab, Bitbucket and Azure DevOps (AI code review tools).
  • Integrations: Codex works from Slack through @ChatGPT and from Linear. Cursor connects to Slack, Linear, Jira, GitHub and Microsoft Teams.
  • Instructions: Codex reads AGENTS.md files; Cursor reads its own rules and AGENTS.md.
  • Both have subagents, MCP servers, skills and plugins.

Are OpenAI's models leaving Cursor?

Yes, if the plan holds. OpenAI says it is winding down its contract supplying models to Cursor, with a proposed transition period to November 12, 2026 and no confirmed end date yet. After that, your own OpenAI API key works in Cursor's local Chat and Agent but not in Tab, Auto, cloud agents or the CLI, and the Codex extension runs inside Cursor with a ChatGPT plan (OpenAI).

What about Claude Code and Antigravity?

Claude Code leads both independent agent leaderboards and is the most widely adopted coding agent in Ramp's spending data; see Claude Code vs Codex and Claude Code vs Cursor. Google Antigravity is an agent-first editor and CLI with a free weekly quota, and Google AI Pro at $19.99 raises the limits.

Which should you choose?

  • You already pay for ChatGPT: Codex, included in your plan.
  • You want one editor that runs many companies' models: Cursor.
  • You want OpenAI's models inside Cursor after November: the Codex extension in Cursor, or your own API key for local chat.
  • You want an open-source agent: the Codex CLI is Apache 2.0; Cursor isn't open source.
  • A team working with one agent: Poly (usepoly.co) runs Codex and Claude Code in a shared browser room where everyone sees each step and any member can approve a change before it happens. Free to start. What is Poly?

Common questions

Is Codex better than Cursor?

On independent agent benchmarks only Codex is listed, at 58.2% on Terminal-Bench 4.0. Cursor's strength is one editor that runs many companies' models; Codex is included with ChatGPT.

Can I use Codex in Cursor?

Yes. OpenAI's Codex IDE extension runs in Cursor with a ChatGPT plan, and it's one of OpenAI's suggested ways to keep using its models in Cursor after its contract with Cursor winds down.

Which is cheaper, Codex or Cursor?

Codex comes with ChatGPT, including the free plan, and Plus at $20 covers most individual use. Cursor Pro is also $20, and Cursor says daily agent users typically spend $60 to $100 a month in total.

Will Cursor still have OpenAI models?

Not for long, if OpenAI's plan holds: it proposed a transition period to November 12, 2026. Afterwards your own OpenAI key works in Cursor's local Chat and Agent, but not in Tab, Auto, cloud agents or the CLI.

What is a coding agent harness?

Everything around the model: its tools, context handling and how it edits files. OpenAI tunes one harness for its own models; Cursor tunes one for many models, which is why the same model can perform differently in each.