---
title: >-
  Mahak (مِحَكّ) — Open community benchmark evaluating LLM Arabic fluency on
  authentic everyday tasks — by Jad Madi (@jadmadi)
description: >-
  Open evaluation benchmark platform that benchmarks leading AI models on native
  Arabic fluency, nuanced cultural comprehension, instruction following, and
  dialectal subtleties across legal contracts, correspondence, and literature.
href: /project/mahak/
author: Jad Madi
category: Arabic Tech
domain: Arabic Tech & Benchmarks
status: Live Platform
license: Open Data / Waqf
year: '2026'
---
# Mahak (مِحَكّ) — Open community benchmark evaluating LLM Arabic fluency on authentic everyday tasks

Open evaluation benchmark platform that benchmarks leading AI models on native Arabic fluency, nuanced cultural comprehension, instruction following, and dialectal subtleties across legal contracts, correspondence, and literature.

## Project Overview

- **Website**: https://jadmadi.net/project/mahak/
- **Live Platform**: https://mahak.waqf.dev
- **Documentation**: https://mahak.waqf.dev/en/matrix/
- **Domain**: Arabic Tech & Benchmarks
- **Category**: Arabic Tech
- **Status**: Live Platform
- **Year**: 2026
- **License**: Open Data / Waqf (waqf)
- **Author**: [@JadMadi](https://x.com/jadmadi)

## The Problem It Solves

General LLM benchmarks (MMLU, Chatbot Arena) rely heavily on translated English benchmarks that fail to evaluate authentic Arabic syntactic flow, legal terminology, and cultural nuance.

## The Solution

A community-annotated benchmark scoring models across authentic Arabic prompts with public Elo ranking matrices and native agent evaluation interfaces.

## Key Features & Capabilities

- Blind side-by-side comparison interface and public Elo ranking leaderboard
- Live public matrix ranking 65 frontier and open models across 25 authentic tasks
- Go CLI (mahak-bench) for headless prompt runs, recovery manifests, and batch synchronization
- Model Context Protocol (MCP) server for automated agent benchmark evaluation
- Evaluation domains: formal correspondence, contracts, customer support, literature, and instruction following

## Technical Specifications & Architecture

- **Languages**: TypeScript, Go, Astro
- **Frameworks / Libraries**: Cloudflare Workers, D1, Tailwind CSS
- **Architecture**: Edge-rendered ranking matrix, blind side-by-side voter engine, Go CLI, and MCP server
- **License Model**: Open Data / Waqf

## Frequently Asked Questions

### How does Mahak differ from standard benchmarks like ArabicMMLU?

ArabicMMLU tests translated multiple-choice knowledge. Mahak evaluates authentic generation, dialectal nuances, formal legal drafting, and tone adaptation.

### Where can I access the live leaderboards?

The public matrix and side-by-side arena are live at https://mahak.waqf.dev.

## Links & References

- Project Page on Crosscurrent Code: https://jadmadi.net/project/mahak/
- Full Projects Directory: https://jadmadi.net/projects/
- Author Profile: https://jadmadi.net/about/
