ChatGPT vs Claude for coding: test them on your own code

By Leonardo Cantoni, founder of Rileva · Published 9 October 2026 · Tool details checked 9 October 2026

Key takeaways

  • There's no stable answer to "ChatGPT or Claude for coding". Models change every few months, and the result depends on your language, your codebase and the kind of work.
  • Separate two decisions: the model (who writes better code for you) and the tool (Codex or Claude Code, IDE plugins, agents). People often mix them up.
  • Run 3 short tests on code like yours: a bug hunt, a refactor under constraints, and a design decision. All three prompts are below, ready to paste.
  • Score with things you can check, like tests passing and bugs found, not with how confident the answer sounds.
A code editor split down the middle, the same Python function on both sides with different lines highlighted, and the title "ChatGPT vs Claude for coding: test it yourself".

Search "ChatGPT vs Claude for coding" and you'll find strong opinions, usually based on one project and one model version that has since been replaced. Developers who use both tend to say the same thing: it depends on the task. That's true, but not helpful unless you know which parts of the task make the difference.

This post gives you a way to find out in under an hour, on your own code.

First, separate the model from the tool

When people say "Claude is better for coding" or "ChatGPT is better for coding", they often mean the whole product, not the model.

  • ChatGPT Plus ($20/month) includes OpenAI's coding agent, Codex. OpenAI includes Codex in every ChatGPT plan, Free and Go too, with more usage on Plus (OpenAI Codex pricing).
  • Claude Pro ($20/month billed monthly) includes Claude Code, which shares your plan's usage limits (claude.com/pricing).

A coding agent that reads your repo, runs your tests and opens a pull request is a different thing from a chat window where you paste a function. If you mostly work through an agent in your terminal, choose the agent workflow you like and test the model inside it. If you mostly paste code into a chat to debug, review or design, the tests below fit directly.

What about benchmarks? Public leaderboards such as SWE-bench and LMArena are worth a look. Note the date, what each one measures, and which model versions were tested. They measure someone else's tasks. Your codebase, language and conventions are the test that matters for you.

The 3-test method

Pick a language you work in every day. For each test, open a fresh chat in each model (no history, no custom instructions you wouldn't use in real work) and paste exactly the same prompt.

Why fresh chats: a model that has spent an hour in your conversation is anchored to your earlier choices. You want each one's first independent take.

Why three different tests: models often differ less on "write a function" and more on judgment: what counts as a bug, how much to change, which trade-off to pick.

Test 1: the bug hunt

Use a function with real bugs. Here's one in Python with several. Swap in your own buggy code if you have some.

Find every bug in this function. For each bug give: the line, what goes wrong,
an input that triggers it, and the fix. Then give the corrected function.
Don't change the style or rename things unless a fix needs it.

def paginate(items, page, per_page=10, seen=[]):
    start = page * per_page
    end = start + per_page
    for item in items[start:end]:
        seen.append(item["id"])
    total_pages = len(items) // per_page
    return items[start:end], total_pages

What to check:

  • Did it find the mutable default argument (seen=[] keeps growing across calls)?
  • Did it catch that total_pages drops the last partial page (25 items at 10 per page gives 2, not 3)?
  • Did it ask, or state an assumption, about whether page starts at 0 or 1?
  • Did it handle per_page=0, or at least mention it?
  • Did it stay in scope, or rewrite everything?

Count the real bugs found, and the made-up ones. A model that invents a bug costs you time too.

Test 2: the refactor under constraints

This tests whether it respects your constraints, not just whether it can write async code.

Refactor this so it's fast for 200 ids without hammering the API:
at most 5 requests in flight at once. Keep results in the same order as ids.
If one request fails, keep the others and report which ids failed.
No new dependencies. Explain the trade-offs in 3 bullets.

async function getUserData(ids) {
  const results = [];
  for (const id of ids) {
    const res = await fetch(`/api/users/${id}`);
    const data = await res.json();
    if (data) results.push(data);
  }
  return results;
}

What to check:

  • Is concurrency really capped at 5? Promise.all over everything at once is a common miss.
  • Is order preserved, including when some requests fail?
  • Are failures reported per id, not swallowed?
  • Did it notice that res.ok is never checked, so a 404 or 500 gets parsed as data?
  • Did it respect "no new dependencies"?

Then run both versions against a mock API. Five minutes of testing beats any amount of reading.

Test 3: the design decision

This is where you'll often see the most differences, and where a second opinion is most useful.

We're a 3-person team building a B2B SaaS on Postgres. We expect 50 to 500
customer companies in the next two years, a few of them large. Should we keep
all tenants in one schema with a tenant_id column, or use one schema per tenant?
Pick one and say why. Then list what would change your mind, and the first
two things we'd regret if we chose wrong.

What to check:

  • Did it commit to an answer, or give a balanced overview?
  • Are the reasons specific to your numbers (team size, tenant count, a few large tenants)?
  • Did it raise the things that actually bite: migrations across many schemas, row-level security, noisy neighbours, per-tenant backups or data-residency requests?
  • Is "what would change my mind" concrete?

There isn't one correct answer here, so don't score which side it picked. Score how well it reasoned about your situation.

How to score: a 10-minute sheet

Model AModel B
Test 1: real bugs found / made-up bugs
Test 1: stayed in scope?
Test 2: constraints respected (cap, order, failures, no deps)/4/4
Test 2: code ran first time against the mock?
Test 3: committed + specific + named real risks/3/3
Which answer would you merge or act on with fewer edits?

If one model clearly comes out ahead on your tests, you have your answer, for now. If they split, for example one better at bug hunts and the other at design calls, you've learned something more useful than a ranking: which model to ask for which job, and that a second opinion is worth having on design decisions.

Rather not juggle tabs? Paste one of the tests into Council and get 3 labs' answers side by side. One tap to Council.

What 3 labs did with these exact prompts

We ran all three prompts through Rileva's Council, which asks 3 AI models from 3 different labs at once and writes a summary of where they agree and differ. For coding questions, the lineup includes models from both OpenAI and Anthropic, so you can see ChatGPT's and Claude's model families answer side by side, plus a third lab. These are single runs, shown as they came out. They show how the answers differ. They don't prove that one model is better at coding, and your own run may get a different lineup.

Lineup in these runs (8 Oct 2026): A: GPT-6 Astra · B: Claude Sonnet 5.5 · C: Grok 4.7, left to right as shown, the same in all three runs. Your run may differ.

Test 1 results: all three found all four bugs

Screenshot of a Rileva Council summary from 8 October 2026 on a Python pagination bug hunt: the 3 models consulted (GPT-6 Astra, Claude Sonnet 5.5, Grok 4.7) with response times, the short answer, and a "Where they agree" table of five issues, each with the line, what goes wrong, an input that triggers it and the fix.

All three answers found all four issues we built into the function: the mutable default, the dropped partial page, the unstated 0- or 1-based page, and per_page=0. We ran each corrected function against our own checks (25 items, three pages, no leaking seen, per_page=0 rejected), and all three passed.

So the differences weren't about finding bugs. They were about what counts as one:

  • Page numbering. A chose 1-based pages and said "that distinction cannot be determined from the code alone". B and C kept 0-based pages; B because "start = page * per_page implies it". Same code, two readings of the contract. That's exactly the question to settle before you merge a fix.
  • Extras. All three flagged unchecked negative values, and we confirmed these are real: a negative page silently slices from the end of the list. A also wanted strict integer checks, which is a stricter contract more than a bug. C counted a fifth issue: a missing "id" updates seen partway, then crashes. A and B treated that as depending on what callers are allowed to pass.
  • Examples. One example input was slightly off: it used plain numbers instead of dicts, so it fails differently from what the answer described. You only catch that kind of thing by running it.

The summary (written by GPT-6 Astra, also one of the three) kept 0-based pages because "the function alone does not establish a 1-based contract", even though its own seat had chosen 1-based. For the missing-"id" case it leaned toward C's fix: "collect all IDs before extending seen, so a missing ID leaves it unchanged."

Screenshot of the same Council summary, 8 October 2026: "Where they differ" (1-based versus 0-based pages, integer validation, and a missing id partly updating the seen list), then "What it depends on", listing the corrected function's assumptions.

Test 2 results: all three passed every constraint

Screenshot of the "Where they differ" section of a Rileva Council summary from 8 October 2026 on refactoring a JavaScript fetch loop: no substantive disagreement, then how each of the three models set the concurrency cap and ordered failures.

We ran each answer's code, unedited, against a mock API with random delays and three failing ids (two 500s and a 404). All three:

  • never had more than 5 requests in flight;
  • kept results in the same order as ids;
  • reported each failed id;
  • checked res.ok;
  • added no dependencies and ran first time.

The differences were small, and the summary's "Where they differ" section said so: "There is no substantive disagreement. [A] fixes the cap at five and stores failures by index; [B] exposes configurable concurrency and sorts failures afterward; [C] validates array input but leaves failures in completion order."

Two points from its "What it depends on" section are worth checking in your own code. All three answers changed the return value from an array to an object, so "Callers must now read { results, failed } rather than an array." And a cap of five "limits concurrency—not requests per second; a rate-limited API may also require pacing." It also noted that "All three perspectives flag hanging requests as a risk: add a timeout and bounded backoff if your API requires them."

Test 3 results: one choice, three different worries

Screenshot of the short answer in a Rileva Council summary from 8 October 2026: choose one shared Postgres schema, with tenant_id on every tenant-owned table and row-level security, because one migration path is worth more to a three-person team than separate schemas.

All three chose the same design. The summary's short answer: "Choose one shared schema, with tenant_id on every tenant-owned table and Postgres row-level security (RLS). For a three-person team serving 50–500 companies, one migration path is worth more than the limited isolation that separate schemas provide."

Where they agreed: shared tables are less operational work than rolling every schema change out across 50 to 500 copies. And "Separate schemas do not isolate performance": tenants still share CPU, I/O and connections.

Where they differed: "There is no disagreement on the choice; the meaningful differences concern safeguards and what failure would hurt most." [A] "prioritizes recovery and extraction", [B] "prioritizes isolation and infrastructure scaling", and [C] "prioritizes engineering capacity". This is the useful part on a design question: ask one model and you'd likely get one of these worries; asking three put all of them on the table.

Screenshot of the "What it depends on" section of the same summary, 8 October 2026: what would change its mind, the first two regrets if shared tables turn out to be wrong, and a closing note that a cross-tenant data leak is the more urgent hazard.

What it depends on: the two regrets it named are worth putting in your design doc:

  1. "We cannot restore just this customer cleanly." "You may need to restore the entire database into a temporary environment, extract the affected tenant, and reconcile subsequent writes."
  2. "Moving this customer out touches everything." "Extraction requires consistent movement of related rows, jobs, and ongoing writes, followed by validation and cutover."

It added that "A cross-tenant data leak is the more urgent implementation hazard, not evidence that shared tables were inherently the wrong choice."

The summaries are written by a model, not by us. In these runs that was GPT-6 Astra, which was also one of the three that answered and isn't told which answer is its own. The code checks are ours: we ran each answer's code and checked every claimed bug by hand. You can read all three full answers in the app.

So which should you use?

It depends on four things more than on the brand:

  1. How you work. Agent in the terminal or chat window? If agent, test the agent you'd actually use (Codex or Claude Code), not the chat.
  2. Your stack. Results on Python scripts don't transfer automatically to a large TypeScript monorepo or embedded C. Test in your language.
  3. The kind of task. Generating boilerplate, hunting bugs, reviewing a pull request and making architecture calls are different skills. You may get different results on each.
  4. How much you need a second opinion. For design decisions, comparing two or three independent answers often matters more than which model you start with.
A 2-by-2 grid: agent or chat on one axis, generating code or making decisions on the other, with a suggested way to test in each quadrant.

A practical setup many developers end up with: one main tool for daily work, and a quick second opinion for bugs that resist and for decisions that are expensive to undo.

FAQ

Is ChatGPT or Claude better for coding? It depends on your language, codebase and the kind of task, and the answer changes as new models ship. Test both on three tasks from your own work, such as a bug hunt, a constrained refactor and a design decision, and score them on things you can check, like tests passing.

Should I use Codex or Claude Code? They're coding agents from OpenAI and Anthropic. Codex comes with every ChatGPT plan, with more usage on Plus; Claude Code comes with Claude's paid plans, including Pro. If you work mainly through an agent, try both on the same real ticket in your own repo. That tells you more than any comparison article.

Are coding benchmarks useful? Yes, as a starting point. Public leaderboards like SWE-bench and LMArena show how models do on their test sets, but your codebase isn't their test set. Check the date and the model versions, then test on your own code.

Can AI find all the bugs in my code? No. Models miss bugs and sometimes report bugs that aren't there. Treat AI review as a fast first pass, run your tests, and verify every claimed bug before you change anything.

How do I compare AI models for coding side by side? Paste the same prompt into each model in a fresh chat and compare the answers against a checklist. Or use a tool that asks several models at once, such as Rileva's Council, which asks 3 models from 3 different labs and summarises where they differ.

Run Test 1 on your own code

Code review from 3 labs at once Paste a function you're not sure about. Council asks 3 AI models from 3 different labs at once and shows where they agree, where they differ, and what it depends on. 2 free answers, no signup. Review my code → It opens with the prompt in the composer and waits for you to press send. Signed in, Council is already on; as a guest, it's one tap to Council. Don't paste secrets or code you're not allowed to share.

Leonardo Cantoni, founder of Rileva

Keep readingRileva Blog