GPT vs Claude vs Gemini vs Grok: the same prompt, four models
"Which AI is best?" gets a new answer every few weeks, and most of those answers come from benchmarks that look nothing like your work. There's a more useful approach: give the same prompt to GPT, Claude, Gemini and Grok, put the answers side by side, and judge them on the tasks you actually do. This guide gives you five prompts, a simple scorecard, and a way to run the comparison without juggling four browser tabs.
The short version
No model wins everything. Which one is "best" depends on the task, your tone and your budget. The best way to find out is to run your own 4–5 real prompts through all four and keep score. It takes about half an hour, and what you learn holds up far longer than any single leaderboard.
Why the same prompt matters
Most comparisons you'll read online break at least one of these rules, which makes them less useful than they look:
- Identical input. The exact same wording, with no extra context for one model. Even one added sentence changes the answer.
- Fresh context. Each model starts from the same point. A model deep in a long chat is working with a different prompt than one starting from scratch.
- Your tasks, not trivia. Riddles and trick questions are fun but tell you little. Use prompts from your real work.
- More than one prompt. Answers vary between runs, so one lucky or unlucky reply proves little. Judge each model on several prompts, not a single one.
The four model families
Each provider offers a range of models, from large flagship models to small fast ones. Comparing a flagship with a fast model isn't a fair fight, so compare within a tier. These are the models available in Chat Tree today:
| Provider | Most capable | Balanced | Fast & cheap |
|---|---|---|---|
| OpenAI (GPT) | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna |
| Anthropic (Claude) | Claude Fable 5.1, Opus 5.5 | Claude Sonnet 5 | Claude Haiku 4.5 |
| Google (Gemini) | — | Gemini 2.5 Flash | Gemini 2.5 Flash-Lite |
| xAI (Grok) | Grok 4.7 | Grok 4.3 | — |
A practical rule: start with the balanced tier of each family. Only move up to a flagship model if the balanced one clearly falls short on a particular task. Flagship models are slower and use up premium allowances faster.
Five prompts to run in all four models
Each prompt tests a different skill. Replace the bracketed parts with your own material. That's the whole point: a model that does well on your email and your code is the one you should use.
1. Explain a concept you already know well
Explain the CAP theorem to a senior frontend developer who has never run a database in production. Use one concrete example and keep it under 250 words.What to look for:
- Is it correct? You chose a topic you know, so you can check.
- Did it respect the audience and the 250-word limit?
- Is the example specific, or just a generic metaphor?
2. Rewrite something of yours
Rewrite the email below so it is firm but friendly, keeps every fact, and is at most half the length. [paste a real email you wrote]What to look for:
- Did it keep every fact, or quietly drop or invent some?
- Does it still sound like you, or like a press release?
- Did it actually halve the length?
3. Fix real code
This function returns the wrong result for an empty list. Find the bug, explain it in two sentences, and give the corrected function. [paste a small function from your project]What to look for:
- Did it find the actual bug, or rewrite everything?
- Does the fix run? Paste it in and try it.
- Did it follow the format you asked for (two sentences, then code)?
4. A decision with trade-offs
We are a 4-person team with one Postgres database. Should we add Redis for caching now or wait? List the three questions that should decide it, then give a recommendation.What to look for:
- Does it take a position, or just say "it depends"?
- Are the three questions the ones an experienced engineer would ask?
- Did it use the constraint (a 4-person team) or ignore it?
5. A question it can’t fully answer
What were the main changes in version 3.2 of [a niche library or internal tool you know]? If you are not sure, say so.What to look for:
- Did it admit uncertainty, or make up plausible-sounding details?
- If it answered, is it actually right?
A simple scorecard
Score each answer from 0 to 2 on four criteria. That gives a maximum of 8 per prompt, or 40 over all five.
| Criterion | 0 | 1 | 2 |
|---|---|---|---|
| Correct | Wrong or made up | Minor errors | Fully right |
| Followed instructions | Ignored the format or limits | Partly | Exactly |
| Useful | Generic | Some real insight | Ready to use |
| Tone | Off-putting | Acceptable | What you wanted |
Don't just add up the scores. Look at where each model lost points. A model that loses points on tone is easy to fix with a better prompt. A model that loses points on correctness for your kind of work isn't.
Patterns to watch for
Models get updated all the time, so we won't tell you which one "wins". By the time you read this, it may have changed. These are the differences that tend to matter most when you put answers side by side:
- Length. Some models write essays when you wanted a paragraph. If you often need short answers, instruction-following matters more than raw intelligence.
- Hedging and honesty. Prompt 5 is the most telling. A model that confidently makes up details is a risk you'll pay for later.
- Taking a position. On trade-off questions, some models commit to a recommendation and some list considerations forever. Decide which one you want.
- Code style. A minimal fix and a full rewrite can both be "correct". For real projects, the minimal fix is almost always better.
- Agreement. When three models agree and one gives a very different answer, check that answer first. It's either wrong or the only one that noticed something.
How to run the comparison without four tabs
The manual way is four browser tabs, four subscriptions and a lot of copy-pasting. You end up with four separate chats and nothing to connect them. It works, but you probably won't do it more than once.
In Chat Tree the comparison is part of the conversation:
- Ask your prompt once, with any model.
- Re-ask the same message with a different model. Each answer becomes a sibling branch of the same question, so the input is exactly the same.
- Open tree view to see GPT, Claude, Gemini and Grok next to each other, and jump between them with one click.
- Keep going from the best answer. The other branches stay where they are if you want to come back to them.
Nothing is overwritten, so every answer stays in the tree for as long as you need it. If you've ever lost a good answer to a regenerate somewhere else, see how to get back a ChatGPT answer after Regenerate.
So which one should you use?
After running this with your own prompts, you'll usually find that you don't need to pick just one. Most people end up with a default model for everyday questions and one or two others they go to for specific tasks, like one for code and another for writing. The real question is how easily you can switch. With every model in one place, switching is just another branch.
FAQ
Is GPT or Claude better?
It depends on the task and changes with each release. Run the same prompts from your real work through both and compare them on correctness, instruction-following, usefulness and tone. The scorecard above takes about 30 minutes.
Is it fair to compare a free model with a paid one?
Not really. Compare within a tier (flagship vs flagship, fast vs fast), otherwise you're mostly measuring price.
Can I compare all four without paying for four subscriptions?
Yes. Chat Tree gives you GPT, Claude, Gemini and Grok in one workspace. The free plan includes the fast models and a monthly allowance of premium-model messages. See pricing.