Sign in

Benchmark

Translation quality, language by language

Held-out WASH test sets, 17 languages. Our engine — Gemma grounded in the WASH translation memory — against commercial MT and frontier LLMs. Filter languages, sort any column.

These are automated metrics. Scores come from a neural quality model (XCOMET-XL) and blind LLM judges — not human reviewers. They estimate quality; they don't replace professional evaluation.

Human validation, in progress. We're actively working with native-speaking domain/sector professionals who score the engine's output against human validators in our evaluation platform. Those expert ratings — not just automated metrics — will anchor future versions of this benchmark.

XCOMET-XL

Reference-based — compares each translation to a human reference. The primary signal. Higher is better.

Languages
EngineMYHA
Our engine85.191.893.882.593.182.290.688.393.173.477.573.481.0
Gemini 3 Flash84.690.893.482.092.282.790.588.091.373.077.174.279.7
DeepL83.892.493.582.592.881.490.489.392.569.275.369.876.8
Gemma 4 26B83.690.693.281.591.980.788.887.490.771.376.472.578.8
Claude Sonnet 4.683.690.892.780.292.481.689.487.391.271.976.471.877.5
Google Translate83.591.293.382.492.581.088.588.891.574.177.766.374.3
Lara82.991.693.480.592.081.489.786.790.268.377.369.674.6
ModernMT78.389.492.179.090.478.485.881.986.159.270.960.966.0
Tiny Aya + TMX78.689.866.871.867.375.7
Tiny Aya75.488.664.370.064.771.7
Best in columnLowest in columnOur engineClick a column to sort

Tiny Aya (a second base model) was scored on a subset of languages — its rows show only where it ran and carry no corpus mean, so they stay out of the ranking. See the head-to-head below.

Burmese (my) and Hausa (ha) are shown but not yet scored — we're actively benchmarking them and will add their results soon. They don't count toward the means.

Second base model — Tiny Aya

Could a different base LLM beat Gemma? We ran Cohere Tiny Aya through the same translation-memory retrieval as our engine — a clean base-swap — and scored it head-to-head. Gemma + TMX still wins every language.

Engineidswnehiurbnam
Our engine93.182.281.077.573.473.478.1
Gemma 4 26B90.780.778.876.472.571.371.4
Tiny Aya + TMX89.878.675.771.867.366.865.1
Tiny Aya88.675.471.770.064.764.360.3

Gemma + TMX beats Tiny Aya + TMX in every language. The memory's lift is base-independent, but Gemma is the stronger pairing — so it stays our engine.

Tiny Aya + TMX still beats plain Tiny Aya everywhere (+1.1 to +4.8 XCOMET): retrieval helps whatever base it grounds. If a future base model overtakes Gemma, this is where it will show up first.

Cohere Tiny Aya (3.35B) · XCOMET-XL (reference-based), ×100, higher is better · added 2026-07-18.

COMET-blind languages — Akan and Haitian Creole

The COMET metric can't score these (its backbone has no Twi or Creole coverage): every engine clusters near the metric floor and can't be told apart. They're judged separately by large language models ranking blind — lower mean rank is better.

Akan / Twi

Akan is off-distribution for the COMET metric too, so it's scored by the same two blind LLM judges over 255 segments. Both judges self-reported competence and reasoned about real Twi specifics, and they agree strongly (Spearman ρ = 0.83).

EngineClaude Opus 4.8Gemini 3.1 Pro
Mean rankWin %Mean rankWin %
Gemini 3 Flash1.7953%1.4468%
Google Translate2.1633%2.1726%
Lara3.365%3.354%
ModernMT4.604%4.751%
Claude Sonnet 4.64.802%4.882%
Our engine4.942%4.940%
Gemma 4 26B6.240%6.370%

Gemini wins Akan decisively and our engine sits mid-pack — but it clearly beats its own base model (about 1.3 ranks better), so the translation memory still lifts a weak base here. As with Haitian Creole, closing the gap is the base model's job; we show it rather than hide it.

Haitian Creole

Haitian Creole is off-distribution for the COMET metric backbone, so it's scored separately by two large language models judging blind. Each ranks every engine's output per segment; lower mean rank is better, and win % is the share of segments where it placed first.

EngineClaude Opus 4.8Gemini 3.1 Pro
Mean rankWin %Mean rankWin %
Gemini 3 Flash3.3927.5%3.1634%
Google Translate3.8216.3%3.5817.6%
Lara3.9015%3.8316.3%
DeepL4.1812.4%4.376.5%
Gemma 4 26B5.475.9%5.904.6%
Claude Sonnet 4.65.487.2%5.603.9%
Our engine5.617.8%5.477.8%
ModernMT5.993.9%5.925.9%
Kreyòl-MT(specialist)7.163.9%7.183.3%

Both judges agree Gemini leads the field, and our engine sits mid-pack on Haitian Creole. Retrieval over the thin Haitian Creole memory gives little lift here — base Gemma and Gemma + TMX land within a few tenths of a rank of each other — so this gap is the base model's, not something the memory can yet close. We show it rather than hide it.

How it was measured

Dataset
baobabtech/tmx-benchmark-results
Sample
153–300 segments per language
Generated
2026-06-27