A new peer reviewed benchmark from Abu Dhabi's Mohamed bin Zayed University of AI shows top models still slip on everyday Arabic across 13 national dialects.
For most of the AI era, a model that "supported Arabic" usually meant it could handle Modern Standard Arabic, the formal, newspaper-ready form that almost nobody speaks at home. A new peer-reviewed benchmark, ArabCulture-Dialogue, scores the way 400 million Arabic speakers actually talk.
Researchers at Abu Dhabi's Mohamed bin Zayed University of AI (MBZUAI) built the test with 26 native speakers across 13 Arabic-speaking countries. It covers Modern Standard Arabic plus 13 national dialects, spanning 12 daily-life topics, from weddings and food to parenting, agriculture, arts, and games, and 54 finer subtopics. Models face three tasks: pick the culturally appropriate reply, translate between MSA and a named dialect, and continue a conversation in that dialect on request.
The numbers, presented at ACL 2026 in San Diego, are humbling. The strongest models scored in the mid-90s on choosing the culturally right reply. Asked to actually produce the correct target-country dialect, they succeeded in only about half of cases, with North African and Emirati dialogues the most challenging.
The benchmark is positioned as a measurement tool, not a fix: a way to score the dialect gap so it can no longer hide behind high MSA test scores.