GPT vs Claude vs Gemini vs Grok: the same prompt, four models

· 8 min read

"Which AI is best?" gets a new answer every few weeks, and most of those answers come from benchmarks that look nothing like your work. There's a more useful approach: give the same prompt to GPT, Claude, Gemini and Grok, put the answers side by side, and judge them on the tasks you actually do. This guide gives you five prompts, a simple scorecard, and a way to run the comparison without juggling four browser tabs.

The short version

No model wins everything. Which one is "best" depends on the task, your tone and your budget. The best way to find out is to run your own 4–5 real prompts through all four and keep score. It takes about half an hour, and what you learn holds up far longer than any single leaderboard.

Why the same prompt matters

Most comparisons you'll read online break at least one of these rules, which makes them less useful than they look:

The four model families

Each provider offers a range of models, from large flagship models to small fast ones. Comparing a flagship with a fast model isn't a fair fight, so compare within a tier. These are the models available in Chat Tree today:

ProviderMost capableBalancedFast & cheap
OpenAI (GPT)GPT-6 AstraGPT-6 SolGPT-6 Luna
Anthropic (Claude)Claude Fable 5.1, Opus 5.5Claude Sonnet 5Claude Haiku 4.5
Google (Gemini)—Gemini 2.5 FlashGemini 2.5 Flash-Lite
xAI (Grok)Grok 4.7Grok 4.3—

A practical rule: start with the balanced tier of each family. Only move up to a flagship model if the balanced one clearly falls short on a particular task. Flagship models are slower and use up premium allowances faster.

Five prompts to run in all four models

Each prompt tests a different skill. Replace the bracketed parts with your own material. That's the whole point: a model that does well on your email and your code is the one you should use.

1. Explain a concept you already know well

Explain the CAP theorem to a senior frontend developer who has never run a database in production. Use one concrete example and keep it under 250 words.

What to look for:

  • Is it correct? You chose a topic you know, so you can check.
  • Did it respect the audience and the 250-word limit?
  • Is the example specific, or just a generic metaphor?

2. Rewrite something of yours

Rewrite the email below so it is firm but friendly, keeps every fact, and is at most half the length. [paste a real email you wrote]

What to look for:

  • Did it keep every fact, or quietly drop or invent some?
  • Does it still sound like you, or like a press release?
  • Did it actually halve the length?

3. Fix real code

This function returns the wrong result for an empty list. Find the bug, explain it in two sentences, and give the corrected function. [paste a small function from your project]

What to look for:

  • Did it find the actual bug, or rewrite everything?
  • Does the fix run? Paste it in and try it.
  • Did it follow the format you asked for (two sentences, then code)?

4. A decision with trade-offs

We are a 4-person team with one Postgres database. Should we add Redis for caching now or wait? List the three questions that should decide it, then give a recommendation.

What to look for:

  • Does it take a position, or just say "it depends"?
  • Are the three questions the ones an experienced engineer would ask?
  • Did it use the constraint (a 4-person team) or ignore it?

5. A question it can’t fully answer

What were the main changes in version 3.2 of [a niche library or internal tool you know]? If you are not sure, say so.

What to look for:

  • Did it admit uncertainty, or make up plausible-sounding details?
  • If it answered, is it actually right?

A simple scorecard

Score each answer from 0 to 2 on four criteria. That gives a maximum of 8 per prompt, or 40 over all five.

Criterion012
CorrectWrong or made upMinor errorsFully right
Followed instructionsIgnored the format or limitsPartlyExactly
UsefulGenericSome real insightReady to use
ToneOff-puttingAcceptableWhat you wanted

Don't just add up the scores. Look at where each model lost points. A model that loses points on tone is easy to fix with a better prompt. A model that loses points on correctness for your kind of work isn't.

Patterns to watch for

Models get updated all the time, so we won't tell you which one "wins". By the time you read this, it may have changed. These are the differences that tend to matter most when you put answers side by side:

How to run the comparison without four tabs

The manual way is four browser tabs, four subscriptions and a lot of copy-pasting. You end up with four separate chats and nothing to connect them. It works, but you probably won't do it more than once.

In Chat Tree the comparison is part of the conversation:

  1. Ask your prompt once, with any model.
  2. Re-ask the same message with a different model. Each answer becomes a sibling branch of the same question, so the input is exactly the same.
  3. Open tree view to see GPT, Claude, Gemini and Grok next to each other, and jump between them with one click.
  4. Keep going from the best answer. The other branches stay where they are if you want to come back to them.

Nothing is overwritten, so every answer stays in the tree for as long as you need it. If you've ever lost a good answer to a regenerate somewhere else, see how to get back a ChatGPT answer after Regenerate.

So which one should you use?

After running this with your own prompts, you'll usually find that you don't need to pick just one. Most people end up with a default model for everyday questions and one or two others they go to for specific tasks, like one for code and another for writing. The real question is how easily you can switch. With every model in one place, switching is just another branch.

FAQ

Is GPT or Claude better?

It depends on the task and changes with each release. Run the same prompts from your real work through both and compare them on correctness, instruction-following, usefulness and tone. The scorecard above takes about 30 minutes.

Is it fair to compare a free model with a paid one?

Not really. Compare within a tier (flagship vs flagship, fast vs fast), otherwise you're mostly measuring price.

Can I compare all four without paying for four subscriptions?

Yes. Chat Tree gives you GPT, Claude, Gemini and Grok in one workspace. The free plan includes the fast models and a monthly allowance of premium-model messages. See pricing.