Dialects, cultural references, and image text pairings make Arabic a hard case for hate detection models, a new benchmark argues.
Content warning: this article discusses a dataset of hateful imagery. Examples are described in research terms only.
Most hateful-meme detectors are trained on English. A new Arabic benchmark argues the under-service is now a content-moderation problem, not a research footnote.
AHA-Memes, posted to arXiv this month, ships 5,000 manually annotated Arabic memes and roughly 66,000 silver-labeled ones, labeled not just "hateful" but by the rhetorical move behind the post: name-calling, dehumanization, slur-laden imagery, mockery of a group's identity. The authors call it the attack-strategy taxonomy, and it tracks how a post tries to do harm, not only whether it does.
Arabic hate hides in dialect, religious reference, and visual shorthand that English-trained late-fusion models do not co-read. Those systems score image and text separately and combine the results at the end. The benchmark tests text-only, image-only, and late-fusion systems, plus open- and closed-weight vision-language models under zero-shot, fine-tuned, and few-shot in-context setups. The authors argue Arabic tooling is thin, and the dataset quantifies the gap.
The team released the dataset, guidelines, and evaluation scripts on Hugging Face. The paper opens with a content warning: the examples are disturbing by design, and the work does not promise a fix. Arabic hate is operational now, and the field's tools were not built for it.