The best engine doesn't win the race: how to pick an AI model
Choosing an AI model by benchmark score is like choosing a car by its engine. What a harness is, why the pairing matters, and how to choose.
On this page
Every few weeks a new AI model tops a leaderboard, and whether you work solo or on a team, the same question comes up: should you switch?
On its own, that's the wrong question. The model is only half of what you're choosing. The other half is the harness it runs in, and the harness decides how much of the model's power actually reaches your work.

The model is the engineLink to this section
A model takes text in and gives text back. That's where the raw ability lives: reasoning, writing code, understanding language.
But a model on its own can't open your files, run your tests or remember what it did yesterday. It's an engine on a test bench: lots of power, going nowhere.
The harness is the carLink to this section
The harness is everything built around the model that turns its power into finished work. It has four main parts:
| Car part | Harness part | What it does |
|---|---|---|
| Steering and wheels | Tools | Lets the model act: read files, edit code, run commands, search |
| Dashboard | Instructions | The system prompt and rules that tell it how to behave in your project |
| Fuel system | Context | What it gets to see: files, memory, conversation history |
| Brakes | Guardrails | Permissions and sandboxes that limit what it's allowed to touch |

You already use harnesses. Claude Code, Codex and Gemini CLI are all harnesses: a model plus tools, instructions, context and guardrails, packaged so it can work in your codebase.
Why the pairing mattersLink to this section
Drop a Ferrari engine into a Corolla and it still runs. It just can't use what it has. The gearbox, cooling and chassis weren't built for that much power, so most of it never reaches the road.

AI models work the same way. Each lab tunes its models alongside its own harness: the exact tool definitions, the format for editing files, the way context is fed in. The model gets very good at driving that particular car.
Put the same model in a harness built around a different one and it still works, but power leaks out at every joint:
- It calls tools in a slightly different shape than the harness expects.
- It edits files in a style the harness handles less reliably.
- It gets context laid out differently from what it was tuned on.
None of this shows up on a leaderboard. Benchmarks test the engine on a test bench, not your car on your road.
How to chooseLink to this section

Pick your car firstLink to this section
Start with the harness, because that's where you'll spend your day, whether you work alone or with a team. Ask:
- Where does the work happen? The terminal, your editor, pull requests, CI.
- What does it need to connect to? Your repo host, issue tracker, internal tools.
- What will actually get used? By you, or by everyone on your team. The best tool nobody opens is worth nothing.
Fit the engine built for itLink to this section
Then use the model that harness was built around:
| Harness | Model built for it |
|---|---|
| Claude Code | Claude |
| Codex | GPT |
| Gemini CLI | Gemini |
If your harness supports many models, don't assume the top benchmark model will be the best one inside it. Test the pairing.
Test on your own trackLink to this section
A small, honest test on your own work beats any leaderboard. It takes an afternoon:
- Pick 3 to 5 real tasks from your backlog: a bug fix, a small feature, a refactor, a missing test.
- Run every task in each pair you're considering, with the same prompt and the same starting commit.
- Score what matters to you, not what's easy to measure:
| Question | Why it matters |
|---|---|
| Did it finish without help? | Shows whether it can work on its own |
| How many times did you correct it? | Every correction is your time |
| Would you merge the result? | The only quality bar that counts |
| How long did it take, and what did it cost? | Speed and spend at your real volume |
- Pick the pair that wins on your tasks. Re-run the same test when a major new model or harness version ships.
The takeawayLink to this section
Benchmarks test the engine. Your work happens on the road.
Choose the car you'll actually drive, put in the engine built for it, and judge the result on your own track.