LLM benchmark
For nine discrete, bounded tasks we benchmark the model we serve against Claude Opus 4.8. Seven we run ourselves — six objective tasks on the same 100 inputs, plus a translate gist/comprehension eval (n=50 per language direction) — with every input, output, and score published to download. Two are cited (summarize, answer): rather than redo them we reproduce the most reliable public leaderboard for each and add our own cost column, linking out so you can check the source. The model we advertise lands at ≈Opus accuracy at ~1–5% of its cost.
How to read it. 🔵 blue = the highest-accuracy model (the winner, or the cited source's #1); 🟢 green = the Bellalingua pick (the model we serve). Both can sit on one row. The cost column is always % of Opus cost (our cost-plus charge ÷ Opus's). The accuracy column is % of Opus accuracy on the tasks we run and on cited boards that include Opus; where a cited board did not test Opus 4.8 (summarize), we show that source's own native metric instead and say so.
JSON repair · Opus accuracy 1 = 100% reference · metric: valid-JSON match (parse + semantic) · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 🟢 Qwen3 30B-A3B Instruct | 100% | 1% |
| 🔵 Nova Lite | 100% | 1% |
| 🔵 Gemini 2.5 Flash-Lite | 100% | 1.2% |
| 🔵 Llama 4 Maverick | 100% | 2.6% |
| 🔵 Qwen2.5 72B Instruct | 100% | 2.9% |
| 🔵 Gemini 2.5 Flash | 100% | 6% |
| 🔵 GPT-5 nano | 100% | 11% |
| 🔵 Claude Haiku 4.5 | 100% | 19.6% |
| 🔵 Claude Opus 4.8 | 100% | 100% |
| Mistral Small 3.2 24B | 98% | 1% |
| DeepSeek V4 Flash | 97% | 3.4% |
| Llama 4 Scout | 92% | 1.5% |
| Ministral 8B | 59% | 1.5% |
What you pay — qwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Classify · Opus accuracy 0.99 = 100% reference · metric: exact-match · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 Claude Opus 4.8 | 100% | 100% |
| 🟢 Qwen3 30B-A3B Instruct | 99% | 1.1% |
| Qwen2.5 72B Instruct | 98% | 5% |
| Nova Lite | 97% | 0.8% |
| Gemini 2.5 Flash | 96% | 3.8% |
| Claude Haiku 4.5 | 96% | 17.5% |
| Gemini 2.5 Flash-Lite | 93.9% | 1.1% |
| Mistral Small 3.2 24B | 93.9% | 1.1% |
| Llama 4 Maverick | 93.9% | 3.4% |
| Ministral 8B | 77.8% | 2.2% |
What you pay — qwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Keywords · Opus accuracy 0.4912 = 100% reference · metric: F1@k · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 🟢 Gemini 2.5 Flash | 104.1% | 5.2% |
| Claude Opus 4.8 | 100% | 100% |
| Qwen2.5 72B Instruct | 97.6% | 2.4% |
| Nova Lite | 96.9% | 0.6% |
| Llama 4 Scout | 96.2% | 1.2% |
| Gemini 2.5 Flash-Lite | 95.2% | 0.9% |
| Mistral Small 3.2 24B | 94.2% | 0.8% |
| Llama 4 Maverick | 93.7% | 3% |
| Qwen3 30B-A3B Instruct | 89.3% | 0.8% |
| Claude Haiku 4.5 | 84.5% | 13.7% |
| Ministral 8B | 77.1% | 1.2% |
What you pay — google/gemini-2.5-flash: $0.318/M input · $2.65/M output metered cost + 6% — the actual provider rate ($0.3/$2.5 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Sentiment · Opus accuracy 0.96 = 100% reference · metric: exact-match · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 DeepSeek V4 Flash | 101% | 4.8% |
| Claude Opus 4.8 | 100% | 100% |
| 🟢 Gemini 2.5 Flash-Lite | 99% | 1.1% |
| Qwen3 30B-A3B Instruct | 99% | 1.3% |
| Gemini 2.5 Flash | 99% | 3.6% |
| Llama 4 Maverick | 97.9% | 3.7% |
| Claude Haiku 4.5 | 97.9% | 16.7% |
| Qwen2.5 72B Instruct | 91.7% | 4.8% |
| Mistral Small 3.2 24B | 90.6% | 1.1% |
| Ministral 8B | 83.3% | 2.1% |
| Nova Lite | 65.6% | 0.7% |
What you pay — google/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Extract · Opus accuracy 0.97 = 100% reference · metric: field-match · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 Gemini 2.5 Flash | 102.1% | 10.4% |
| 🟢 Qwen3 30B-A3B Instruct | 101% | 0.9% |
| Mistral Small 3.2 24B | 100.3% | 1.2% |
| Claude Opus 4.8 | 100% | 100% |
| Nova Lite | 99.3% | 0.9% |
| Llama 4 Maverick | 98.6% | 3% |
| Gemini 2.5 Flash-Lite | 97.6% | 1.9% |
| DeepSeek V4 Flash | 96.7% | 2% |
| Llama 4 Scout | 94.6% | 1.6% |
| Claude Haiku 4.5 | 94.5% | 23.7% |
| Ministral 8B | 93.5% | 1.5% |
What you pay — qwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Proofread · Opus accuracy 0.6193 = 100% reference · metric: source-aware GLEU · n=100
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 Claude Haiku 4.5 | 103.1% | 16.1% |
| 🟢 Gemini 2.5 Flash-Lite | 101% | 1.1% |
| Claude Opus 4.8 | 100% | 100% |
| Nova Lite | 99.9% | 0.7% |
| GPT-5 nano | 99.6% | 23.6% |
| Mistral Small 3.2 24B | 99.1% | 0.9% |
| Gemini 2.5 Flash | 98.6% | 5.8% |
| Llama 4 Maverick | 98% | 2.8% |
| Qwen3 30B-A3B Instruct | 93.2% | 0.9% |
| Llama 4 Scout | 90.6% | 1.4% |
| Ministral 8B | 75.7% | 1.2% |
What you pay — google/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.
Translate · Opus accuracy 0.9857 = 100% reference · metric: Belebele 4-way MC comprehension (foreign→English; fixed reader) · n=350
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 🟢 Mistral Small 3.2 24B | 101.2% | 0.8% |
| Nova Lite | 100% | 0.7% |
| Llama 4 Scout | 100% | 1.1% |
| Claude Opus 4.8 | 100% | 100% |
| Gemini 2.5 Flash-Lite | 99.4% | 1.3% |
| DeepSeek V4 Flash | 99.4% | 1.4% |
| Llama 4 Maverick | 99.4% | 2.3% |
| Qwen2.5 72B Instruct | 99.1% | 2.1% |
| Claude Haiku 4.5 | 99.1% | 15% |
| Gemini 2.5 Flash | 98.8% | 7% |
| Ministral 8B | 98.3% | 0.8% |
| Qwen3 30B-A3B Instruct | 97.4% | 0.8% |
What you pay — mistralai/mistral-small-3.2-24b-instruct: $0.106/M input · $0.318/M output metered cost + 6% — the actual provider rate ($0.1/$0.3 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
How to read the picks. Blue = highest measured accuracy; green = our served pick. Both land on Mistral Small 3.2 24B — it has the highest accuracy and the lowest cost in the table, and it's the most consistent across all seven directions (≥98% in every language — the highest floor here — and 100% on low-resource Swahili, where several models drop to 80–88%). At n=350 the gaps between the top models are within statistical noise, but Mistral Small 3.2 24B leads on accuracy, cost, and cross-language consistency at once.
Per-direction accuracy (each language → English, n=50)
| Model | Spanish | French | German | Portuguese | Chinese | Swahili | Polish |
|---|---|---|---|---|---|---|---|
| 🔵 🟢 Mistral Small 3.2 24B | 100% | 100% | 100% | 98% | 100% | 100% | 100% |
| Nova Lite | 100% | 98% | 100% | 98% | 96% | 98% | 100% |
| Llama 4 Scout | 100% | 98% | 98% | 100% | 98% | 98% | 98% |
| Claude Opus 4.8 | 98% | 100% | 100% | 98% | 96% | 100% | 98% |
| Gemini 2.5 Flash-Lite | 98% | 100% | 100% | 94% | 96% | 100% | 98% |
| DeepSeek V4 Flash | 98% | 96% | 100% | 96% | 100% | 100% | 96% |
| Llama 4 Maverick | 98% | 98% | 100% | 96% | 96% | 98% | 100% |
| Qwen2.5 72B Instruct | 100% | 100% | 100% | 98% | 98% | 88% | 100% |
| Claude Haiku 4.5 | 100% | 98% | 98% | 98% | 96% | 96% | 98% |
| Gemini 2.5 Flash | 98% | 96% | 98% | 98% | 96% | 100% | 96% |
| Ministral 8B | 100% | 98% | 100% | 98% | 96% | 88% | 98% |
| Qwen3 30B-A3B Instruct | 100% | 100% | 98% | 96% | 98% | 80% | 100% |
Download the full n=50/direction test data (CSV) ↓ · XLSX ↓ every passage, the candidate translation, the comprehension question + options + gold answer, the reader’s pick, and what it cost — re-run and check the math.
What we measured — and what we didn't. We did not grade these translations on fluency or publishable polish — whether the result reads like a native writer wrote it. For an agent translating a support ticket or a foreign blog post, that's the wrong test. What matters is understandability: does the translation carry the meaning accurately enough that you can read it and make the right decision? That's the gist, and it's what we scored.
How we scored the gist. We used a reading-comprehension benchmark (Belebele, built on FLORES-200 passages). For each item the candidate model translates a foreign-language passage into English; then one fixed reader (gemini-2.5-flash, held constant across every model) answers a 4-way multiple-choice question about that passage using only the translation. If the meaning survived, the question is answerable; if it was lost or distorted, it isn't. The score is the share answered correctly — how reliably you can pull accurate information out of the output. Claude Opus 4.8's translation is scored the same way and pinned at 100%.
Incomplete runs are removed above 2% empty responses (relaxed from the strict zero used for the n=100 objective tasks — at n=350, sub-2% provider flake shouldn't eject a model; empties count as incorrect). The fixed reader also appears as a translation candidate and scored mid-pack (not top), confirming no self-favoritism.
Summarize · cited Vectara HHEM hallucination leaderboard · metric: factual-consistency rate (Vectara HHEM, native) · as of 2026-06-25
| Model | factual-consistency rate (source) | % of Opus cost |
|---|---|---|
| 🔵 Finix S1 32B | 98.2% | — |
| GPT-5.4 nano | 96.9% | 5.1% |
| 🟢 Gemini 2.5 Flash-Lite | 96.7% | 1.8% |
| Phi-4 | 96.3% | 0.7% |
| Llama 3.3 70B Instruct | 95.9% | 1.5% |
| Snowflake Arctic Instruct | 95.7% | — |
| Gemma 3 12B | 95.6% | 0.7% |
| Mistral Large 2411 | 95.5% | — |
| Qwen3 8B | 95.2% | 1.6% |
| Nova Pro | 94.9% | 14.1% |
What you pay — google/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
How to read the picks. We cite Vectara’s HHEM leaderboard rather than run our own summarization eval — it’s the most reliable public faithfulness benchmark and it already covers the cheap models we serve. Vectara did not evaluate Claude Opus 4.8; the nearest Anthropic model on their board is Opus 4.7 (88.0%). The accuracy column is Vectara’s own factual-consistency metric (not normalized to Opus). The cost column is OUR cost-plus charge as a % of Opus 4.8’s cost. Opus stays the cost benchmark, not a target — follow the source link and decide.
View the source leaderboard ↗ third-party accuracy (their own metric), our cost column — we cite the source rather than redo it; follow the link and check it yourself.
Cited table — not run by us. Accuracy = Vectara’s factual-consistency rate (higher = fewer hallucinations on their summarization set). Cost = the model’s combined OpenRouter list price (input + output per-million-token) ×1.06, as a % of Opus 4.8’s combined list price (pulled live 2026-06-25); “—” where a model isn’t served on OpenRouter. 🔵 = Vectara’s #1; 🟢 = the model we serve.
Answer · cited FACTS Grounding leaderboard (Google DeepMind / Kaggle) · metric: grounding score (DeepMind FACTS Grounding) · as of 2026-06-25
| Model | % of Opus accuracy | % of Opus cost |
|---|---|---|
| 🔵 🟢 Gemma 4 26B A4B | 116.4% | 1.4% |
| Gemma 4 31B | 116.1% | 1.7% |
| GPT-5.2 | 109.6% | 55.7% |
| GPT-5.5 | 109.2% | 123.7% |
| Gemini 2.5 Pro | 106.9% | 39.8% |
| Llama 3 Grounded LM | 103.3% | — |
| GPT-5.4 | 101.2% | 61.8% |
| Gemini 3.5 Flash | 100.9% | 37.1% |
| Gemini 2.5 Flash | 100.7% | 9.9% |
| GPT-5 | 100.1% | 39.8% |
| Claude Opus 4.8 | 100% | 100% |
What you pay — google/gemma-4-26b-a4b-it: $0.1723/M input · $0.636/M output metered cost + 6% — the actual provider rate ($0.1625/$0.6 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.
How to read the picks. We cite DeepMind’s FACTS Grounding leaderboard — the exact-task authority for grounded question answering — rather than run our own. Claude Opus 4.8 IS on this board (#11), so the accuracy column is the normal % of Opus accuracy. Our served pick, Gemma 4 26B A4B, is the board’s #1 — above GPT-5.5, Gemini 2.5 Pro, and every Claude — at ~1% of Opus’s cost. (Our summarize pick, gemini-2.5-flash-lite, is NOT competitive on this exact grounded-QA task, so we did not green it here — its Vectara #3 is summarization-faithfulness, a different task.) Opus is the cost benchmark, not a target — follow the source link and decide.
View the source leaderboard ↗ third-party accuracy (their own metric), our cost column — we cite the source rather than redo it; follow the link and check it yourself.
Cited table — not run by us. Accuracy = % of Opus accuracy (the model’s FACTS grounding score ÷ Opus 4.8’s, 69.5%); the Opus 4.8 row is shown as the 100% reference even though it ranks #11. Cost = the model’s combined OpenRouter list price (input + output per-million-token) ×1.06, as a % of Opus 4.8’s combined list price (pulled live 2026-06-25); “—” where a model isn’t served on OpenRouter. 🔵 = FACTS #1; 🟢 = the model we serve (both on Gemma 4 26B A4B).
Methodology
What you pay (the "What you pay" line under each task).
You're charged on actual usage at cost + 6% — input tokens × the
input rate + output tokens × the output rate. The displayed rate is a ceiling in two ways, so
you typically pay less than shown: tokens — you authorize a ceiling and are billed only
for the tokens actually delivered (the output length isn't known until the model responds); and
rate — the per-token rate shown is the highest across OpenRouter's providers
for this model, and OpenRouter routes to the lowest-cost available provider, so the rate you're charged
is usually below it. Rates are pulled live from OpenRouter (2026-06-26) and the markup is the
live llm.markup_pct billing setting, so the page can't drift from what you're actually
billed. The "% of Opus cost" column is the same charge expressed relative to Opus — the value
story; the per-token rates are the absolute price.
Reference = Claude Opus 4.8 via OpenRouter, pinned at 100%. The six objective tasks run n=100 inputs each. Per-task metric: JSON repair = valid-JSON match (parse + semantic), Classify = exact-match, Keywords = F1@k, Sentiment = exact-match, Extract = field-match, Proofread = source-aware GLEU, Translate = Belebele 4-way multiple-choice comprehension (foreign→English, n=50 per language direction; a fixed reader answers from the translation only). Cost is the real provider per-call charge × 1.06 (cost-plus), expressed as a % of Opus's cost; "% of Opus accuracy" = model accuracy ÷ Opus accuracy. Every model gets the same prompt per item, temperature 0, each example in isolation.
Rows from an incomplete run are removed — for the objective tasks that means any empty response; for translate (n=350 total) the threshold is relaxed to 2% empty so sub-2% provider flake doesn't eject a strong model (empties still count as incorrect). The published download for each task we run is the complete per-row data, so the table is fully checkable.
Cited tasks (summarize, answer). We don't run these — we reproduce the most reliable public leaderboard for each (summarize → Vectara HHEM factual-consistency; answer → DeepMind FACTS Grounding) and link to it. Their accuracy column is the source's own metric; our cost column is the model's combined OpenRouter list price (input + output per-million-token) × 1.06 as a % of Opus 4.8's, pulled live at build time (“—” where a model isn't on OpenRouter). Where a source did not evaluate Opus 4.8 (Vectara), we show its native metric rather than fabricate an Opus-relative number, and disclose it. We cite the most reliable dataset rather than redo it; Opus stays the cost benchmark, not a target — follow each source link and decide for yourself.