The Parity Index · live

Cheaper models are quietly beating the expensive ones.

Every router ranks models by price. We rank them by whether a cheaper model, tuned by our patent-pending proof, actually held or beat the model you are paying for, on your real prompts. Across 16,764 blind tests, it did 85.5% of the time.

This is not the same prompt sent to two models. Our patent-pending proof optimises how the cheaper model is used for your task first, then a blind judge checks the result. That step is the product.

1,230 prompts145 companies11 current models30 to 60% typical saving
JSON extractionbeat it
JSON extractionheld quality
JSON extractionstayed put
Prose & writingheld quality
Prose & writingstayed put
Structured + prosebeat it
Structured + prosestayed put
Long-form HTMLbeat it
Long-form HTMLheld quality
JSON extractionbeat it
JSON extractionheld quality
Prose & writingbeat it
Prose & writingheld quality
Prose & writingstayed put
Structured + proseheld quality
Structured + prosestayed put
Long-form HTMLbeat it
Long-form HTMLstayed put
JSON extractionbeat it
JSON extractionstayed put
Prose & writingbeat it
Prose & writingheld quality
Structured + prosebeat it
Structured + proseheld quality
Structured + prosestayed put
Long-form HTMLheld quality
JSON extractionheld quality
JSON extractionstayed put
JSON extractionbeat it
JSON extractionheld quality
JSON extractionstayed put
Prose & writingheld quality
Prose & writingstayed put
Structured + prosebeat it
Structured + prosestayed put
Long-form HTMLbeat it
Long-form HTMLheld quality
JSON extractionbeat it
JSON extractionheld quality
Prose & writingbeat it
Prose & writingheld quality
Prose & writingstayed put
Structured + proseheld quality
Structured + prosestayed put
Long-form HTMLbeat it
Long-form HTMLstayed put
JSON extractionbeat it
JSON extractionstayed put
Prose & writingbeat it
Prose & writingheld quality
Structured + prosebeat it
Structured + proseheld quality
Structured + prosestayed put
Long-form HTMLheld quality
JSON extractionheld quality
JSON extractionstayed put

The leaderboard

Cheaper models winning on quality right now.

How often each cheaper model, once tuned by Parity's patent-pending proof, held or beat the more expensive one it was tested against, on real production prompts, ranked. A green badge means it produced genuinely better work, not just as good.

1
gemma-3n-e4b-itbeat 1%290 tests · 14 cos
0%
2
gpt-4o-minibeat 9%110 tests · 4 cos
0%
3
mistral-nemobeat 6%515 tests · 14 cos
0%
4
gemini-2.0-flash-lite-001beat 7%3,767 tests · 43 cos
0%
5
gpt-4.1-nanobeat 5%455 tests · 9 cos
0%
6
qwen3-235b-a22b-2507beat 9%138 tests · 3 cos
0%
7
gpt-4.1-minibeat 28%167 tests · 5 cos
0%
8
gemini-2.5-flash-lite-preview-09-2025beat 14%421 tests · 1 cos
0%
9
gemini-2.5-flash-litebeat 15%250 tests · 7 cos
0%

Head to head

Examples where a cheaper model beat a frontier model.

Each of these is a real head to head on production prompts, and what you are looking at is how often the cheaper model came back with better work than the more expensive one it was tested against, once our patent-pending proof had tuned it for the task. Aggregate across many companies.

gemini-2.0-flash-001beatgpt-4o-mini

The cheaper model produced better work 45% of the time, and held or beat it 85% of the time, across 40 blind comparisons.

gpt-5.4-nanobeatgpt-4o

The cheaper model produced better work 36% of the time, and held or beat it 36% of the time, across 50 blind comparisons.

gpt-4.1-minibeatgpt-4o

The cheaper model produced better work 32% of the time, and held or beat it 89% of the time, across 142 blind comparisons.

gemini-2.5-flashbeatgpt-5.1

The cheaper model produced better work 30% of the time, and held or beat it 78% of the time, across 50 blind comparisons.

qwen3-235b-a22b-2507beatgpt-4o-mini

The cheaper model produced better work 30% of the time, and held or beat it 98% of the time, across 40 blind comparisons.

claude-haiku-4.5beatgpt-5.1

The cheaper model produced better work 28% of the time, and held or beat it 80% of the time, across 40 blind comparisons.

gemini-2.5-flash-lite-preview-09-2025beatgpt-4o

The cheaper model produced better work 21% of the time, and held or beat it 82% of the time, across 225 blind comparisons.

gpt-5-minibeatgpt-4o

The cheaper model produced better work 20% of the time, and held or beat it 40% of the time, across 75 blind comparisons.

We also publish where cheaper models failed.

If we only showed you the wins this would not be worth much, so every comparison we run lands in one of three outcomes and you can see all three. The ones that fail never switch, they stay on the model you are already paying for.

77.5% matched. Same quality as the model in use.

8% beat it. Genuinely better than the dearer model.

14.5% did not hold. These prompts stay on the original.

Start from your model

Which model are you paying for now?

Pick the model you run today and you can see which cheaper ones have actually been tested against it, and how often they held. These are the real comparisons we have run, so some have more behind them than others, and where the sample is still small we say so.

gpt-4.1-miniheld or beat 89% · came back better 32%

142 blind comparisons

gemini-2.5-flash-lite-preview-09-2025held or beat 82% · came back better 21%

225 blind comparisons

gpt-5-miniheld or beat 40% · came back better 20%

75 blind comparisons

gpt-5.4-nanoheld or beat 36% · came back better 36%

50 blind comparisons, early data on a small sample

What it is worth

What that looks like on your bill.

A percentage is easy to skim past, so put your own number in and see the range on the work a cheaper model passes on. It is a band rather than a figure because it genuinely varies prompt by prompt.

$$5,000

On the work that passes, per month

$1,500 to $3,000

Over a year

$18,000 to $36,000

This is the 30 to 60% band applied to your number, and it only covers the prompts a cheaper model actually passes on. Prompts that fail stay on the model you are using now, so treat this as the range on proven work rather than a quote. Your own report gives you the real figure for your traffic.

By kind of work

Which work is safest to move.

Not all prompts are equal. This is how often a cheaper model held or beat the original, by the kind of work. Your own mix decides your own number.

1
Long-form HTML75 tests
0%
2
JSON extraction1,697 tests
0%
3
Prose & writing4,620 tests
0%
4
Structured + prose1,175 tests
0%

Why this ranking exists

No other router keeps this number.

A router chooses a cheaper model on price, latency and uptime, before it has seen the answer. A degraded reply and a good one both come back the same way, so the quality question is never measured. It just becomes your problem later.

Parity is the first router that proves quality when you switch to a cheaper model. Cheaper candidates run against your real prompts, a blind judge decides whether the cheaper output matched or beat your current one, only a proven model is ever offered as a switch, and fallback to your original is instant. This ranking is the aggregate of every one of those judgements. It is data only Parity can produce, because only Parity runs the proof.

How it is measured

  • The cheaper model is tuned first, not sent the raw prompt. This is the part that matters. Our patent-pending proof works out how to get the cheaper model to do your task well, then tests that. The same prompt sent blindly to a cheaper model would often fall short. The result is what Parity makes the cheaper model produce, which is exactly why you cannot get this number anywhere else.
  • On real prompts, not a benchmark. Every comparison uses a company's real production traffic, never a public test set that can be gamed.
  • A blind judge in your own model class. The judge does not know which output is the cheaper one, and the bar is set by how much your current model already disagrees with itself run to run.
  • Three separate dimensions. Format, categorical correctness and meaning are each checked, so a cheaper model has to hold on all of them, not just look similar.
  • Aggregate and anonymous. The ranking pools results across companies. We never publish which model was proven for any single company, or any prompt content.

The only number that decides it is yours.

This ranking is the average across many companies. Whether a cheaper model holds on your prompts is specific to you. Connect your traffic or start with a JSONL export of past requests, and see your own proof, before anything switches. No proof, no charge. Not for coding agents.

Aggregate figures as of August 22, 2026. Proof runs March 27, 2026 to August 3, 2026. Refreshed monthly.