Claude vs ChatGPT for Coding: Which Is Better on Real Projects?
Claude Code vs OpenAI Codex on real project work: benchmarks, pricing, context windows, and which one to pick for debugging, refactors, tests, and reviews.
Disclosure: We may earn a commission when you buy through links on this page, at no extra cost to you. This never affects our ratings. Learn more.

Table of contents
Short answer: as of September 2026, Claude is the better pick for serious work on real codebases, mainly because of Claude Code and the current Opus 5.5 model. ChatGPT, through OpenAI's Codex agent, is close behind and wins on price in a couple of important places: its API coding model is cheaper, and you can try Codex on the free plan.
That's the headline. The details matter more, because "better at coding" means different things when you're untangling a 400-file monorepo versus shipping a landing page. Below, we break it down by the tasks you actually do on a project, using published benchmarks, vendor docs, and current pricing. We haven't run our own head-to-head project tests for this piece; everything here comes from documented sources, linked inline. (See how we test for our general approach.)
If you want the everyday chatbot comparison instead (writing, research, planning), that's in our ChatGPT vs Claude vs Gemini roundup. This one is strictly about code.
Our pick
Claude
4.7Best for: Long agentic sessions, large-codebase refactors, and terminal-heavy work · Free; Pro $20/mo; Max $100 or $200/mo (Claude Code on paid plans)
Claude wins for coding on real projects. Opus 5.5 posts the highest vendor-reported scores on the agentic benchmarks that most resemble project work (Terminal-Bench 4.0 and CursorBench), it's the default in Claude Code, and Claude Code runs everywhere you'd want it: terminal, VS Code, JetBrains, desktop, web, and CI. Pick ChatGPT and Codex instead if API cost is your main constraint, if you want to try an agent without paying, or if your team already lives in ChatGPT.
What you're actually comparing in 2026
Nobody serious codes by pasting snippets into a chat window anymore. Both companies now sell an agent that reads your repo, edits files, runs commands, and opens pull requests. The chat apps still matter for quick questions, but the agents are where the real comparison happens.
- Claude side: Claude Code, available as a CLI, a VS Code (and Cursor) extension, a JetBrains plugin, a desktop app, and on the web at claude.ai/code. It also plugs into GitHub Actions, GitLab CI/CD, Slack, and Chrome for debugging live web apps.
- ChatGPT side: Codex, available as a CLI, an IDE extension, a desktop app, and Codex cloud, plus a GitHub code review integration.
The models underneath
Anthropic released Claude Opus 5.5 on September 22, 2026. Anthropic says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. According to the Claude Code model docs, Opus 5.5 is now the default model in Claude Code on most providers, with Fable 5.1 available as an explicit opt-in for the longest, hardest sessions, and Sonnet 5 for day-to-day work.
On OpenAI's side, GPT-6 Astra arrived in early September as the flagship. Then on September 22, OpenAI shipped GPT-6 Sol and GPT-6 Luna. OpenAI's model docs describe Sol as "built to power complex coding and agentic workflows," and it's rolling out in Codex, replacing the GPT-5.6 generation.
| Feature | WinnerClaude | ChatGPT |
|---|---|---|
| Our rating | 4.7 | 4.5 |
| Pricing | Free; Pro $20/mo; Max $100 or $200/mo (Claude Code on paid plans) | Free; Go $8/mo; Plus $20/mo; Pro from $100/mo (Codex on all plans) |
| Best for | Long agentic sessions, large-codebase refactors, and terminal-heavy work | Cheaper API coding, trying an agent for free, and parallel cloud tasks |
| Agentic coding tool | Claude Code | Codex |
| Default or recommended coding model | Opus 5.5 | GPT-6 Sol |
| Top-tier model | Fable 5.1 | GPT-6 Astra |
| Context window (flagship coding models) | 1M tokens | 1.05M tokens |
| API price, main coding model (input / output per 1M tokens) | $4 / $20 (Opus 5.5) | $2 / $10 (GPT-6 Sol) |
| API price, top model | $10 / $50 (Fable 5.1) | $10 / $50 (GPT-6 Astra) |
| Agent on free plan | No | Yes |
| Cheapest plan with the agent | Pro, $20/mo | Free (limited) |
The benchmarks, and why you shouldn't over-read them
First, a caveat you'll see everywhere else ignored: SWE-bench Verified is no longer a useful scoreboard. OpenAI stopped reporting it in early 2026, citing flawed tests and training-data contamination, and moved to SWE-bench Pro. Top models now cluster in the 90s on Verified, which tells you very little. If a comparison article leans on Verified numbers in late 2026, be skeptical.
The more interesting numbers are agentic benchmarks, where the model has to plan, run commands, and recover from mistakes. Here's what Anthropic published with the Opus 5.5 launch. The GPT-6 Astra figures are OpenAI's own reported results, as cited by Anthropic:
| Benchmark (vendor-reported) | Claude Opus 5.5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | 54.4% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | not listed | 41.7% |
Source: Anthropic, Claude Opus 5.5 announcement, September 22, 2026. Anthropic notes Opus 5.5 was run at xhigh effort and Astra at high effort for Terminal-Bench, each model's highest reported score. DataCamp's writeup lists Astra at 57.7% on Terminal-Bench 4.0, a small discrepancy worth noting.
Three things to keep in mind:
- These are vendor numbers. Anthropic chose which benchmarks to show, and every lab does the same.
- The Terminal-Bench gap is real, but FrontierCode is nearly a tie. Roughly nine points on Terminal-Bench is meaningful. A one-point gap on FrontierCode isn't.
- GPT-6 Sol isn't in the table. It shipped the same day as Opus 5.5, and at the time of writing we couldn't find a published Terminal-Bench 4.0 score for it. Given that Sol is the model most Codex users will actually run, this is the biggest open question in the comparison.
Five real project tasks, side by side
Here's how the two stack up on the jobs that fill an actual sprint. For each, we've included a prompt you can reuse in either tool.
1. Understanding a large, unfamiliar codebase
Both flagships now offer roughly a million tokens of context: 1M for Opus 5.5 and Sonnet 5 on the Anthropic API, and 1.05M for GPT-6 Astra, Sol, and Luna. Raw window size is basically a wash.
What differs is how each agent handles long sessions. Claude Code auto-compacts near the limit (about 967K tokens by default, per the docs) and supports a CLAUDE.md file with project conventions, plus auto memory that saves learnings between sessions. It'll also read an existing AGENTS.md. Codex with GPT-6 Astra introduced searchable notes that persist through a session instead of repeatedly summarizing old context, which should help it keep track of constraints you set early in a long refactor.
One pricing gotcha on the OpenAI side: Astra prompts with more than 272K input tokens are billed at 2x input and 1.5x output for the full request. If you're pushing whole repos through the API, that adds up.
I'm new to this repo. Before changing anything:
1. Map the top-level architecture: entry points, main modules, and how
requests flow from the HTTP layer to the database.
2. List the 5 files I'd most need to understand to add a new API endpoint.
3. Point out anything surprising: dead code, circular imports, config that
overrides defaults in non-obvious places.
Don't edit files. Cite file paths and line numbers for every claim.Edge: slight Claude, mostly for the mature memory and project-instruction tooling. If you mostly work through the API on huge inputs, it's closer, since Sol is cheaper per token.
2. Debugging a failure you can reproduce
This is where Terminal-Bench is most relevant, because debugging means running things, reading output, and trying again. Opus 5.5's lead there lines up with what the task demands. Claude Code also lets you pipe logs straight in, which is handy for production issues:
tail -200 app.log | claude -p "Find the root cause of the 502s after 14:00 and propose a fix"Codex handles the same loop through its CLI or IDE extension. The difference is less about capability and more about which one keeps going sensibly on the fourth failed attempt, which is exactly what agentic benchmarks try to measure.
A good debugging prompt forces the agent to prove the bug before it patches it:
// Failing test the agent should write first, then make pass
import { describe, it, expect } from "vitest";
import { applyDiscount } from "../src/pricing";
describe("applyDiscount", () => {
it("never returns a negative total when the coupon exceeds the cart", () => {
expect(applyDiscount(20, { type: "fixed", amount: 25 })).toBe(0);
});
});Edge: Claude, based on the published agentic numbers.
3. Multi-file refactoring
Refactors punish agents that lose track of constraints halfway through. Claude Code can spawn subagents that split a big change across parts of the codebase, with a lead agent merging the results, and Anthropic pitches Fable 5.1 specifically for "tasks larger than a single session." Codex has subagents too, and its cloud mode is well suited to kicking off several isolated tasks in parallel and reviewing the diffs later.
Migrate every call site of the legacy `fetchJson` helper to the typed
`apiClient`. Constraints:
- Keep public function signatures unchanged.
- Don't touch files under /generated.
- Run `pnpm typecheck && pnpm test` after each module and stop if anything fails.
Show me the plan before editing.Edge: Claude for one long, coherent refactor. Codex is appealing if you'd rather fire off ten small, independent cleanups in the cloud.
4. Writing tests
Both agents write decent tests when you tell them what "good" means. The trick is to demand behavior coverage, not line coverage, and to make the agent run the suite. The Claude Code docs literally use this as their example: "write tests for the auth module, run them, and fix any failures."
# Ask for property-style edge cases, not just the happy path
import pytest
from invoices import split_amount
@pytest.mark.parametrize("total,parts", [(100, 3), (1, 2), (0, 5)])
def test_split_sums_to_total(total, parts):
shares = split_amount(total, parts)
assert sum(shares) == total
assert max(shares) - min(shares) <= 1Edge: roughly even. Test writing is well within both tools' abilities; your prompt quality matters more here. Our prompt-writing guide covers the basics.
5. Frontend work and code review
For frontend, Claude Code's Chrome integration lets the agent look at and debug a running web app rather than guessing from JSX alone. On OpenAI's side, several trackers rank GPT-6 Astra highly on frontend-focused arenas, but we couldn't confirm those rankings from a primary source, so we're not leaning on them.
For code review, both have first-party GitHub integrations: Claude Code's GitHub Code Review and Actions support, and Codex's GitHub code review. Either will catch the obvious stuff. Neither replaces a human who knows why the code is shaped the way it is.
Review this PR as a senior engineer. Prioritize, in order:
1. Correctness bugs and unhandled error paths
2. Security issues (injection, authz gaps, secrets in code)
3. Anything that makes the next change harder
Skip style nits the linter already catches. Max 10 comments.Edge: even, with a nod to Claude for live browser debugging.
Pricing for developers
Subscriptions first. Both companies now use very similar tiers, which makes this easier than it used to be.
| Feature | WinnerClaude | ChatGPT |
|---|---|---|
| Our rating | 4.7 | 4.5 |
| Pricing | Free; Pro $20/mo; Max $100 or $200/mo (Claude Code on paid plans) | Free; Go $8/mo; Plus $20/mo; Pro from $100/mo (Codex on all plans) |
| Best for | Long agentic sessions, large-codebase refactors, and terminal-heavy work | Cheaper API coding, trying an agent for free, and parallel cloud tasks |
| Free plan | Yes, no Claude Code | Yes, limited Codex |
| Low-cost tier | None | Go, $8/mo |
| Standard paid tier | Pro, $20/mo ($17/mo annual) | Plus, $20/mo |
| Heavy-use tier | Max 5x $100/mo, Max 20x $200/mo | Pro 5x or 20x, from $100/mo |
| Team seats | $25/mo standard, $125/mo premium (monthly) | Business from $20/user/mo (annual) |
Sources: Claude pricing, Claude Max details, and Codex pricing.
Both run on rolling five-hour usage windows. OpenAI publishes rough message ranges per model; for example, Plus users get roughly 15 to 150 GPT-6 Sol messages per five-hour window in Codex depending on task size. Anthropic describes Pro as "at least 5x more usage per 5-hour session than Free" and Max as 5x or 20x Pro. In practice, if you run an agent for hours a day, you'll end up on one of the $100 to $200 tiers either way. (If you're weighing ChatGPT's paid tiers for general use, see ChatGPT Free vs Plus.)
On the API, OpenAI is cheaper for the everyday coding model: GPT-6 Sol costs $2 in and $10 out per million tokens, versus $4 and $20 for Opus 5.5. Anthropic's Sonnet 5 matches Sol at $2 and $10, though. At the top end, Fable 5.1 and GPT-6 Astra are identically priced at $10 and $50.
| Feature | Claude | ChatGPT |
|---|---|---|
| Agentic CLI tool | Yes | Yes |
| VS Code extension | Yes | Yes |
| JetBrains support | Yes | Yes |
| Desktop app for agent sessions | Yes | Yes |
| Cloud or web agent sessions | Yes | Yes |
| GitHub code review integration | Yes | Yes |
| Agent available on free plan | No | Yes |
| Scheduled or recurring agent tasks | Yes | No |
Where each one fits
Claude Our pick
4.7- Pricing
- Free; Pro $20/mo; Max $100 or $200/mo (Claude Code on paid plans)
- Best for
- Long agentic sessions, large-codebase refactors, and terminal-heavy work
Choose Claude if you're doing sustained agent work on a real codebase: long debugging sessions, cross-cutting refactors, CI automation. Opus 5.5 leads the agentic benchmarks Anthropic published, it's the default in Claude Code, and the tooling around it (CLAUDE.md, hooks, skills, subagents, routines, MCP servers) is deep.
Watch out for: no Claude Code on the free plan, and Opus 5.5 API calls cost twice what GPT-6 Sol does. Fable 5.1 may draw on extra usage credits on some plans.
ChatGPT
4.5- Pricing
- Free; Go $8/mo; Plus $20/mo; Pro from $100/mo (Codex on all plans)
- Best for
- Cheaper API coding, trying an agent for free, and parallel cloud tasks
Choose ChatGPT if cost is the constraint or you want to try an agent before paying. Codex is included on every plan, including Free and the $8 Go tier, and GPT-6 Sol is priced at $2/$10 on the API. Codex cloud is great for queuing several independent tasks and reviewing diffs later.
Watch out for: GPT-6 Sol is brand new and lacks public agentic benchmark numbers we could verify. Long Astra API prompts past 272K tokens are billed at a premium.
Already using GitHub Copilot? It now offers GPT-6 Sol on its higher-tier plans, so you may not need a separate subscription for inline completions. Our GitHub Copilot vs Cursor comparison covers its current plans.
The verdict
Claude takes it. The published agentic evidence favors Opus 5.5, especially the roughly nine-point Terminal-Bench 4.0 lead over GPT-6 Astra, and Claude Code is the more complete package for working inside a real repository across terminal, IDE, and CI. For most professional developers paying $20 to $200 a month, that's where we'd put the money today.
ChatGPT is a very close second, not a distant one. If you're price-sensitive on the API, just starting out, or already paying for ChatGPT for everything else, Codex with GPT-6 Sol is a strong, cheaper choice. And since Sol launched the same day as Opus 5.5, this verdict is the one most likely to shift once independent benchmark results arrive. We'll update this page when they do.
Frequently asked questions
Is Claude Code free?
No. As of September 2026, Claude Code requires a paid plan, starting with Pro at $20 a month ($17 billed annually), or pay-as-you-go API usage. Codex, by contrast, has a limited free tier.
Which model should I pick inside Claude Code?
The default, Opus 5.5, is the right choice for most work. Anthropic's docs suggest switching to Fable 5.1 (/model fable) for tasks bigger than a single session and Sonnet 5 for quicker day-to-day edits.
Are SWE-bench Verified scores still worth checking?
Not really. OpenAI dropped it over flawed tests and contamination, and top models are clustered near the ceiling. Look at SWE-bench Pro and agentic benchmarks like Terminal-Bench instead, and treat vendor-reported figures with some caution.
Can I use both?
Yes, and plenty of developers do. Both agents read an AGENTS.md file for project instructions, so you can keep one set of conventions and switch tools depending on the task or which usage window you've burned through.
Which is cheaper for API-based coding agents?
For the main coding models, OpenAI: GPT-6 Sol is $2/$10 per million tokens versus $4/$20 for Opus 5.5. Anthropic's Sonnet 5 matches Sol's price, and the top models (Fable 5.1 and GPT-6 Astra) cost the same at $10/$50.

Written by
Panoptix Editorial Team
Our editors test AI tools hands-on for weeks before we publish a word. We pay for our own subscriptions and never accept payment for rankings.
How we test AI tools →

