Home / Evaluation

Evaluation method

Every score on this site comes from one screening run on 2026-08-04. Here is exactly what it measured and what it did not.

Translation pairs

Why 872 pairs have no score

ReasonModels
FLORES-200 has no test set for the language (most of the long tail)818
CTranslate2 (HPLT) models: measured, but that harness was found broken, so scores are withheld25
Not part of the screening run25
Same-language models (a score would only measure copying)4

Language verification

53 unscored models passed a GlotLID language-identification check on their output: the model writes the language it claims. That is evidence of identity, not of quality, so it is shown as a separate "language-verified" badge and never mixed into the bands.

Caveats

What is next

Re-score every pair on the full FLORES-200 devtest split (1,012 sentences), fix the CTranslate2 harness and publish the HPLT scores, and add word-error-rate results for the speech models. Scores will be published either way, including the weak ones.