LQM is a language-agnostic framework for diagnosing machine translation errors across six linguistic levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics. It captures dialect, cultural, and contextual issues that broad evaluation schemes can miss.
The project evaluates six large language models on a bidirectional corpus of 3,850 sentences spanning seven Arabic dialects, using expert span-level annotations and severity-weighted quality scores. This companion explorer lets you filter the annotations using the LQM or MQM taxonomy and inspect highlighted error spans.
In this explorer, we present LQM framework applied on the following languages/dialects: Egyptian Arabic Dialect (EGY), English (ENG), Jordanian Arabic Dialect (JOR), Mauritanian Arabic Dialect (MAU), Moroccan Arabic Dialect (MOR), Palestinian Arabic Dialect (PAL), Emerati Arabic Dialect (UAE), and Yemeni Arabic Dialect (YEM).