All of Learn

The best engine doesn't win the race: how to pick an AI model

Choosing an AI model by benchmark score is like choosing a car by its engine. What a harness is, why the pairing matters, and how to choose.

By Ansumana Badjie4 min read
On this page

Every few weeks a new AI model tops a leaderboard, and whether you work solo or on a team, the same question comes up: should you switch?

On its own, that's the wrong question. The model is only half of what you're choosing. The other half is the harness it runs in, and the harness decides how much of the model's power actually reaches your work.

Slide titled 'How do you pick an AI model?' reading: The best engine doesn't win the race. The best car does. Below it: Model (the engine) plus Harness (the car) equals Result (how far you get).

The model is the engineLink to this section

A model takes text in and gives text back. That's where the raw ability lives: reasoning, writing code, understanding language.

But a model on its own can't open your files, run your tests or remember what it did yesterday. It's an engine on a test bench: lots of power, going nowhere.

The harness is the carLink to this section

The harness is everything built around the model that turns its power into finished work. It has four main parts:

Car partHarness partWhat it does
Steering and wheelsToolsLets the model act: read files, edit code, run commands, search
DashboardInstructionsThe system prompt and rules that tell it how to behave in your project
Fuel systemContextWhat it gets to see: files, memory, conversation history
BrakesGuardrailsPermissions and sandboxes that limit what it's allowed to touch
Diagram of a car outline with a small engine inside. The engine is the model (reasoning, code, language). Around it: Tools are the steering and wheels (read, edit, run, search), Instructions are the dashboard (system prompt, rules), Context is the fuel system (files, memory, history) and Guardrails are the brakes (permissions, sandbox).
Everything around the model is the harness.

You already use harnesses. Claude Code, Codex and Gemini CLI are all harnesses: a model plus tools, instructions, context and guardrails, packaged so it can work in your codebase.

Why the pairing mattersLink to this section

Drop a Ferrari engine into a Corolla and it still runs. It just can't use what it has. The gearbox, cooling and chassis weren't built for that much power, so most of it never reaches the road.

Two cards compared. Ferrari engine in a Ferrari body, marked matched: gearbox, cooling and chassis built around the engine, full power reaches the road, power bar full. Ferrari engine in a Corolla body, marked mismatched: great engine, wrong car, power is lost in every part that wasn't built for it, power bar less than half full.

AI models work the same way. Each lab tunes its models alongside its own harness: the exact tool definitions, the format for editing files, the way context is fed in. The model gets very good at driving that particular car.

Put the same model in a harness built around a different one and it still works, but power leaks out at every joint:

  • It calls tools in a slightly different shape than the harness expects.
  • It edits files in a style the harness handles less reliably.
  • It gets context laid out differently from what it was tuned on.

None of this shows up on a leaderboard. Benchmarks test the engine on a test bench, not your car on your road.

How to chooseLink to this section

Slide titled 'So how do you choose?' with three steps. 01: Pick your car first. Which harness will your team actually live in? 02: Fit the engine built for it: Claude with Claude Code, GPT with Codex, Gemini with Gemini CLI. 03: Test on your own track. Your repo, your tasks. Not a leaderboard.

Pick your car firstLink to this section

Start with the harness, because that's where you'll spend your day, whether you work alone or with a team. Ask:

  • Where does the work happen? The terminal, your editor, pull requests, CI.
  • What does it need to connect to? Your repo host, issue tracker, internal tools.
  • What will actually get used? By you, or by everyone on your team. The best tool nobody opens is worth nothing.

Fit the engine built for itLink to this section

Then use the model that harness was built around:

HarnessModel built for it
Claude CodeClaude
CodexGPT
Gemini CLIGemini

If your harness supports many models, don't assume the top benchmark model will be the best one inside it. Test the pairing.

Test on your own trackLink to this section

A small, honest test on your own work beats any leaderboard. It takes an afternoon:

  1. Pick 3 to 5 real tasks from your backlog: a bug fix, a small feature, a refactor, a missing test.
  2. Run every task in each pair you're considering, with the same prompt and the same starting commit.
  3. Score what matters to you, not what's easy to measure:
QuestionWhy it matters
Did it finish without help?Shows whether it can work on its own
How many times did you correct it?Every correction is your time
Would you merge the result?The only quality bar that counts
How long did it take, and what did it cost?Speed and spend at your real volume
  1. Pick the pair that wins on your tasks. Re-run the same test when a major new model or harness version ships.

The takeawayLink to this section

Benchmarks test the engine. Your work happens on the road.

Choose the car you'll actually drive, put in the engine built for it, and judge the result on your own track.