A hard, discriminating benchmark for Modern Standard Arabic — models graded by strong LLM judges across generation, reasoning, knowledge, safety, and multi-turn dialogue.
The Average is the item-weighted mean across 13 capability & domain dimensions — every benchmark item counts equally, so a small bucket never outweighs a large one. Knowledge & STEM includes medicine and biology. Category cells are shaded by score, and the gold-outlined cell is the top score in each task column. Click any column header to sort.