Relaystation logo Relaystation

LLM benchmark

For nine discrete, bounded tasks we benchmark the model we serve against Claude Opus 4.8. Seven we run ourselves — six objective tasks on the same 100 inputs, plus a translate gist/comprehension eval (n=50 per language direction) — with every input, output, and score published to download. Two are cited (summarize, answer): rather than redo them we reproduce the most reliable public leaderboard for each and add our own cost column, linking out so you can check the source. The model we advertise lands at ≈Opus accuracy at ~1–5% of its cost.

How to read it. 🔵 blue = the highest-accuracy model (the winner, or the cited source's #1); 🟢 green = the Bellalingua pick (the model we serve). Both can sit on one row. The cost column is always % of Opus cost (our cost-plus charge ÷ Opus's). The accuracy column is % of Opus accuracy on the tasks we run and on cited boards that include Opus; where a cited board did not test Opus 4.8 (summarize), we show that source's own native metric instead and say so.

JSON repair · Opus accuracy 1 = 100% reference · metric: valid-JSON match (parse + semantic) · n=100

Model % of Opus accuracy % of Opus cost
🔵 🟢 Qwen3 30B-A3B Instruct 100% 1%
🔵 Nova Lite 100% 1%
🔵 Gemini 2.5 Flash-Lite 100% 1.2%
🔵 Llama 4 Maverick 100% 2.6%
🔵 Qwen2.5 72B Instruct 100% 2.9%
🔵 Gemini 2.5 Flash 100% 6%
🔵 GPT-5 nano 100% 11%
🔵 Claude Haiku 4.5 100% 19.6%
🔵 Claude Opus 4.8 100% 100%
Mistral Small 3.2 24B 98% 1%
DeepSeek V4 Flash 97% 3.4%
Llama 4 Scout 92% 1.5%
Ministral 8B 59% 1.5%

What you payqwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Classify · Opus accuracy 0.99 = 100% reference · metric: exact-match · n=100

Model % of Opus accuracy % of Opus cost
🔵 Claude Opus 4.8 100% 100%
🟢 Qwen3 30B-A3B Instruct 99% 1.1%
Qwen2.5 72B Instruct 98% 5%
Nova Lite 97% 0.8%
Gemini 2.5 Flash 96% 3.8%
Claude Haiku 4.5 96% 17.5%
Gemini 2.5 Flash-Lite 93.9% 1.1%
Mistral Small 3.2 24B 93.9% 1.1%
Llama 4 Maverick 93.9% 3.4%
Ministral 8B 77.8% 2.2%

What you payqwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Keywords · Opus accuracy 0.4912 = 100% reference · metric: F1@k · n=100

Model % of Opus accuracy % of Opus cost
🔵 🟢 Gemini 2.5 Flash 104.1% 5.2%
Claude Opus 4.8 100% 100%
Qwen2.5 72B Instruct 97.6% 2.4%
Nova Lite 96.9% 0.6%
Llama 4 Scout 96.2% 1.2%
Gemini 2.5 Flash-Lite 95.2% 0.9%
Mistral Small 3.2 24B 94.2% 0.8%
Llama 4 Maverick 93.7% 3%
Qwen3 30B-A3B Instruct 89.3% 0.8%
Claude Haiku 4.5 84.5% 13.7%
Ministral 8B 77.1% 1.2%

What you paygoogle/gemini-2.5-flash: $0.318/M input · $2.65/M output metered cost + 6% — the actual provider rate ($0.3/$2.5 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Sentiment · Opus accuracy 0.96 = 100% reference · metric: exact-match · n=100

Model % of Opus accuracy % of Opus cost
🔵 DeepSeek V4 Flash 101% 4.8%
Claude Opus 4.8 100% 100%
🟢 Gemini 2.5 Flash-Lite 99% 1.1%
Qwen3 30B-A3B Instruct 99% 1.3%
Gemini 2.5 Flash 99% 3.6%
Llama 4 Maverick 97.9% 3.7%
Claude Haiku 4.5 97.9% 16.7%
Qwen2.5 72B Instruct 91.7% 4.8%
Mistral Small 3.2 24B 90.6% 1.1%
Ministral 8B 83.3% 2.1%
Nova Lite 65.6% 0.7%

What you paygoogle/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Extract · Opus accuracy 0.97 = 100% reference · metric: field-match · n=100

Model % of Opus accuracy % of Opus cost
🔵 Gemini 2.5 Flash 102.1% 10.4%
🟢 Qwen3 30B-A3B Instruct 101% 0.9%
Mistral Small 3.2 24B 100.3% 1.2%
Claude Opus 4.8 100% 100%
Nova Lite 99.3% 0.9%
Llama 4 Maverick 98.6% 3%
Gemini 2.5 Flash-Lite 97.6% 1.9%
DeepSeek V4 Flash 96.7% 2%
Llama 4 Scout 94.6% 1.6%
Claude Haiku 4.5 94.5% 23.7%
Ministral 8B 93.5% 1.5%

What you payqwen/qwen3-30b-a3b-instruct-2507: $0.1378/M input · $0.5512/M output metered cost + 6% — the actual provider rate ($0.13/$0.52 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Proofread · Opus accuracy 0.6193 = 100% reference · metric: source-aware GLEU · n=100

Model % of Opus accuracy % of Opus cost
🔵 Claude Haiku 4.5 103.1% 16.1%
🟢 Gemini 2.5 Flash-Lite 101% 1.1%
Claude Opus 4.8 100% 100%
Nova Lite 99.9% 0.7%
GPT-5 nano 99.6% 23.6%
Mistral Small 3.2 24B 99.1% 0.9%
Gemini 2.5 Flash 98.6% 5.8%
Llama 4 Maverick 98% 2.8%
Qwen3 30B-A3B Instruct 93.2% 0.9%
Llama 4 Scout 90.6% 1.4%
Ministral 8B 75.7% 1.2%

What you paygoogle/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

Download the full n=100 test data (CSV) ↓ · XLSX ↓ every input, the expected answer, each model’s output + score, and what it cost — check the math.

Translate · Opus accuracy 0.9857 = 100% reference · metric: Belebele 4-way MC comprehension (foreign→English; fixed reader) · n=350

Model % of Opus accuracy % of Opus cost
🔵 🟢 Mistral Small 3.2 24B 101.2% 0.8%
Nova Lite 100% 0.7%
Llama 4 Scout 100% 1.1%
Claude Opus 4.8 100% 100%
Gemini 2.5 Flash-Lite 99.4% 1.3%
DeepSeek V4 Flash 99.4% 1.4%
Llama 4 Maverick 99.4% 2.3%
Qwen2.5 72B Instruct 99.1% 2.1%
Claude Haiku 4.5 99.1% 15%
Gemini 2.5 Flash 98.8% 7%
Ministral 8B 98.3% 0.8%
Qwen3 30B-A3B Instruct 97.4% 0.8%

What you paymistralai/mistral-small-3.2-24b-instruct: $0.106/M input · $0.318/M output metered cost + 6% — the actual provider rate ($0.1/$0.3 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

How to read the picks. Blue = highest measured accuracy; green = our served pick. Both land on Mistral Small 3.2 24B — it has the highest accuracy and the lowest cost in the table, and it's the most consistent across all seven directions (≥98% in every language — the highest floor here — and 100% on low-resource Swahili, where several models drop to 80–88%). At n=350 the gaps between the top models are within statistical noise, but Mistral Small 3.2 24B leads on accuracy, cost, and cross-language consistency at once.

Per-direction accuracy (each language → English, n=50)
Model SpanishFrenchGermanPortugueseChineseSwahiliPolish
🔵 🟢 Mistral Small 3.2 24B 100%100%100%98%100%100%100%
Nova Lite 100%98%100%98%96%98%100%
Llama 4 Scout 100%98%98%100%98%98%98%
Claude Opus 4.8 98%100%100%98%96%100%98%
Gemini 2.5 Flash-Lite 98%100%100%94%96%100%98%
DeepSeek V4 Flash 98%96%100%96%100%100%96%
Llama 4 Maverick 98%98%100%96%96%98%100%
Qwen2.5 72B Instruct 100%100%100%98%98%88%100%
Claude Haiku 4.5 100%98%98%98%96%96%98%
Gemini 2.5 Flash 98%96%98%98%96%100%96%
Ministral 8B 100%98%100%98%96%88%98%
Qwen3 30B-A3B Instruct 100%100%98%96%98%80%100%

Download the full n=50/direction test data (CSV) ↓ · XLSX ↓ every passage, the candidate translation, the comprehension question + options + gold answer, the reader’s pick, and what it cost — re-run and check the math.

What we measured — and what we didn't. We did not grade these translations on fluency or publishable polish — whether the result reads like a native writer wrote it. For an agent translating a support ticket or a foreign blog post, that's the wrong test. What matters is understandability: does the translation carry the meaning accurately enough that you can read it and make the right decision? That's the gist, and it's what we scored.

How we scored the gist. We used a reading-comprehension benchmark (Belebele, built on FLORES-200 passages). For each item the candidate model translates a foreign-language passage into English; then one fixed reader (gemini-2.5-flash, held constant across every model) answers a 4-way multiple-choice question about that passage using only the translation. If the meaning survived, the question is answerable; if it was lost or distorted, it isn't. The score is the share answered correctly — how reliably you can pull accurate information out of the output. Claude Opus 4.8's translation is scored the same way and pinned at 100%.

Incomplete runs are removed above 2% empty responses (relaxed from the strict zero used for the n=100 objective tasks — at n=350, sub-2% provider flake shouldn't eject a model; empties count as incorrect). The fixed reader also appears as a translation candidate and scored mid-pack (not top), confirming no self-favoritism.

Summarize · cited Vectara HHEM hallucination leaderboard · metric: factual-consistency rate (Vectara HHEM, native) · as of 2026-06-25

Model factual-consistency rate (source) % of Opus cost
🔵 Finix S1 32B 98.2%
GPT-5.4 nano 96.9% 5.1%
🟢 Gemini 2.5 Flash-Lite 96.7% 1.8%
Phi-4 96.3% 0.7%
Llama 3.3 70B Instruct 95.9% 1.5%
Snowflake Arctic Instruct 95.7%
Gemma 3 12B 95.6% 0.7%
Mistral Large 2411 95.5%
Qwen3 8B 95.2% 1.6%
Nova Pro 94.9% 14.1%

What you paygoogle/gemini-2.5-flash-lite: $0.106/M input · $0.424/M output metered cost + 6% — the actual provider rate ($0.1/$0.4 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

How to read the picks. We cite Vectara’s HHEM leaderboard rather than run our own summarization eval — it’s the most reliable public faithfulness benchmark and it already covers the cheap models we serve. Vectara did not evaluate Claude Opus 4.8; the nearest Anthropic model on their board is Opus 4.7 (88.0%). The accuracy column is Vectara’s own factual-consistency metric (not normalized to Opus). The cost column is OUR cost-plus charge as a % of Opus 4.8’s cost. Opus stays the cost benchmark, not a target — follow the source link and decide.

View the source leaderboard ↗ third-party accuracy (their own metric), our cost column — we cite the source rather than redo it; follow the link and check it yourself.

Cited table — not run by us. Accuracy = Vectara’s factual-consistency rate (higher = fewer hallucinations on their summarization set). Cost = the model’s combined OpenRouter list price (input + output per-million-token) ×1.06, as a % of Opus 4.8’s combined list price (pulled live 2026-06-25); “—” where a model isn’t served on OpenRouter. 🔵 = Vectara’s #1; 🟢 = the model we serve.

Answer · cited FACTS Grounding leaderboard (Google DeepMind / Kaggle) · metric: grounding score (DeepMind FACTS Grounding) · as of 2026-06-25

Model % of Opus accuracy % of Opus cost
🔵 🟢 Gemma 4 26B A4B 116.4% 1.4%
Gemma 4 31B 116.1% 1.7%
GPT-5.2 109.6% 55.7%
GPT-5.5 109.2% 123.7%
Gemini 2.5 Pro 106.9% 39.8%
Llama 3 Grounded LM 103.3%
GPT-5.4 101.2% 61.8%
Gemini 3.5 Flash 100.9% 37.1%
Gemini 2.5 Flash 100.7% 9.9%
GPT-5 100.1% 39.8%
Claude Opus 4.8 100% 100%

What you paygoogle/gemma-4-26b-a4b-it: $0.1723/M input · $0.636/M output metered cost + 6% — the actual provider rate ($0.1625/$0.6 per M) ×1.06, billed on tokens delivered. We advertise the highest provider rate, so you pay this or less. Live OpenRouter pricing, 2026-06-26.

How to read the picks. We cite DeepMind’s FACTS Grounding leaderboard — the exact-task authority for grounded question answering — rather than run our own. Claude Opus 4.8 IS on this board (#11), so the accuracy column is the normal % of Opus accuracy. Our served pick, Gemma 4 26B A4B, is the board’s #1 — above GPT-5.5, Gemini 2.5 Pro, and every Claude — at ~1% of Opus’s cost. (Our summarize pick, gemini-2.5-flash-lite, is NOT competitive on this exact grounded-QA task, so we did not green it here — its Vectara #3 is summarization-faithfulness, a different task.) Opus is the cost benchmark, not a target — follow the source link and decide.

View the source leaderboard ↗ third-party accuracy (their own metric), our cost column — we cite the source rather than redo it; follow the link and check it yourself.

Cited table — not run by us. Accuracy = % of Opus accuracy (the model’s FACTS grounding score ÷ Opus 4.8’s, 69.5%); the Opus 4.8 row is shown as the 100% reference even though it ranks #11. Cost = the model’s combined OpenRouter list price (input + output per-million-token) ×1.06, as a % of Opus 4.8’s combined list price (pulled live 2026-06-25); “—” where a model isn’t served on OpenRouter. 🔵 = FACTS #1; 🟢 = the model we serve (both on Gemma 4 26B A4B).

Methodology

What you pay (the "What you pay" line under each task). You're charged on actual usage at cost + 6% — input tokens × the input rate + output tokens × the output rate. The displayed rate is a ceiling in two ways, so you typically pay less than shown: tokens — you authorize a ceiling and are billed only for the tokens actually delivered (the output length isn't known until the model responds); and rate — the per-token rate shown is the highest across OpenRouter's providers for this model, and OpenRouter routes to the lowest-cost available provider, so the rate you're charged is usually below it. Rates are pulled live from OpenRouter (2026-06-26) and the markup is the live llm.markup_pct billing setting, so the page can't drift from what you're actually billed. The "% of Opus cost" column is the same charge expressed relative to Opus — the value story; the per-token rates are the absolute price.

Reference = Claude Opus 4.8 via OpenRouter, pinned at 100%. The six objective tasks run n=100 inputs each. Per-task metric: JSON repair = valid-JSON match (parse + semantic), Classify = exact-match, Keywords = F1@k, Sentiment = exact-match, Extract = field-match, Proofread = source-aware GLEU, Translate = Belebele 4-way multiple-choice comprehension (foreign→English, n=50 per language direction; a fixed reader answers from the translation only). Cost is the real provider per-call charge × 1.06 (cost-plus), expressed as a % of Opus's cost; "% of Opus accuracy" = model accuracy ÷ Opus accuracy. Every model gets the same prompt per item, temperature 0, each example in isolation.

Rows from an incomplete run are removed — for the objective tasks that means any empty response; for translate (n=350 total) the threshold is relaxed to 2% empty so sub-2% provider flake doesn't eject a strong model (empties still count as incorrect). The published download for each task we run is the complete per-row data, so the table is fully checkable.

Cited tasks (summarize, answer). We don't run these — we reproduce the most reliable public leaderboard for each (summarize → Vectara HHEM factual-consistency; answer → DeepMind FACTS Grounding) and link to it. Their accuracy column is the source's own metric; our cost column is the model's combined OpenRouter list price (input + output per-million-token) × 1.06 as a % of Opus 4.8's, pulled live at build time (“—” where a model isn't on OpenRouter). Where a source did not evaluate Opus 4.8 (Vectara), we show its native metric rather than fabricate an Opus-relative number, and disclose it. We cite the most reliable dataset rather than redo it; Opus stays the cost benchmark, not a target — follow each source link and decide for yourself.

← Back to Bellalingua