Arabic is really a family of languages living under one name. Modern Standard Arabic is what you read in the news or a textbook, but it's rarely how people actually talk to each other. In the UAE, day-to-day conversation, humor, negotiation, and storytelling happen in Emirati Arabic, a Gulf dialect with its own vocabulary, its own rhythm, and a culture wrapped tightly around it. Emirati poetry, especially nabati poetry, along with proverbs and short anecdotes, carries meaning that doesn't survive a literal, word-for-word reading. A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means.
That's the gap Falcon-Emirati-7B is built to close. It's a dialect-specialized model on top of Falcon-H1-Arabic, aimed at understanding and generating Emirati Arabic the way a native speaker would: the vocabulary, the tone, and the cultural context behind it.
Falcon-Emirati-7B is built on Falcon-H1-Arabic
Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that set new benchmarks for the language earlier this year. Falcon-H1-Arabic uses the Falcon-H1 hybrid architecture: State Space Models (Mamba) and Transformer attention running in parallel inside every block, with their outputs fused before each block's projection. That combination gives the linear-time efficiency of Mamba on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic. The family spans three scales (3B, 7B, and 34B parameters) with context windows up to 128K and 256K tokens, and it was trained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data.
That gave us a strong starting point: a model that already understood Arabic broadly, handled long context well, and had some dialectal exposure baked in. Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model doesn't pick up on its own.
We built Falcon-Emirati-7B on the 7B variant specifically. It's the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical. The 34B model would likely push quality a bit further, but at a training and serving cost that doesn't make sense for a dialect-specialized chat model, and the 3B model doesn't leave enough headroom for the depth of cultural and linguistic understanding we were after. 7B gave us the best balance of quality against training and inference cost.
Turning a general Arabic model into an Emirati-dialect specialist is genuinely difficult
Turning a general Arabic model into an Emirati-dialect specialist is genuinely difficult for several reasons:
Emirati is mostly a spoken dialect. It shows up far less in writing online than MSA or even other Gulf and Levantine dialects, so there just isn't as much raw text to learn from. Meaning is often non-literal. Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary. There's no established playbook. There isn't a well-documented recipe for how much dialectal data is enough, how to mix it with MSA and general Arabic, or which training stage (continued pre-training, SFT, or preference optimization) matters most for picking up a dialect.
That last point shaped how we worked. A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle.
We built a dedicated Emirati data pipeline
We built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic's pretraining, drawing on three complementary sources.
We crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from MSA. This is where we got our ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.
Alongside the dialectal text, we pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms. This doesn't teach the model to write in dialect, but it teaches the model what it's talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.
Authentic dialectal text alone wasn't enough to cover the range of topics a chat model actually needs to handle day to day. So we generated a large amount of synthetic Emirati-dialect data to fill the gaps. We didn't just let a generator model improvise in “Gulf-ish” Arabic. We constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that's grammatically fine but sounds off to anyone who actually speaks the dialect.
Since there's no standard recipe for MSA-to-dialect adaptation, we treated the training strategy itself as something to figure out experimentally. We ran ablations on how much dialectal data to inject and at which stage of training, how to balance authentic crawled data against synthetic data without the model overfitting to synthetic patterns, and how much MSA cultural context was actually needed to keep it culturally grounded rather than just fluent on the surface. At each step we leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone don't capture naturalness, tone, or cultural fit.
We tracked progress throughout training with two complementary approaches

We tracked progress throughout training with two complementary approaches:
Emirati native speakers reviewed model outputs directly, judging not just whether an answer was correct but whether it sounded right: naturalness, tone, cultural appropriateness. These are the things a benchmark score won't tell you but a native ear catches immediately.
For quantitative tracking, we used Alyah (الياه, “North Star”), a benchmark we and the community released specifically to evaluate Emirati-dialect capability in Arabic LLMs. Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry: the categories where dialect and culture matter most and where generic Arabic models tend to struggle. Full details on Alyah's construction are available in our benchmark blog post, and background on the base model family is available in the Falcon-H1-Arabic announcement.
Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other Arabic and multilingual model we compared it against, including several models many times its size.
Alyah accuracy (%), instruction-tuned models. Falcon-H1-Arabic family models are excluded from this comparison since Falcon-Emirati-7B is built on top of them.
A couple of things jump out from this comparison. Size alone doesn't buy you dialect competence. Some of the largest multilingual models here score well below smaller, more dialect-aware ones, which tells you Emirati proficiency has to be trained for on purpose, not picked up as a side effect of scale. The models that do best also tend to be Arabic-native or Arabic-focused to begin with, which lines up with what we saw during our own ablations: general Arabic and dialect coverage is a necessary starting point, but it still takes targeted, dialect-specific work to close the rest of the gap, particularly on the hardest parts of Alyah, like poetry, heritage knowledge, and the language-and-dialect category itself.
This matches what came out of the Alyah benchmark release more broadly: even strong models show real degradation once you move into genuinely dialectal, culturally embedded content. That gap doesn't close on its own with bigger models. It takes data and evaluation built specifically for the dialect.
Multiple-choice accuracy tells you whether a model can recognize the right answer among four options. It doesn't tell you whether the model will actually produce Emirati Arabic when someone talks to it in that dialect. So alongside Alyah, we ran a second evaluation: open-ended generation on the same 1,173 Alyah questions, scored by an LLM judge (Gemini 3.7 Flash) against five models—Falcon-Emirati-7B, ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat, and Fanar-2-27B-Instruct—chosen as the strongest competing models from the Alyah leaderboard.
The judge scored each answer on two separate dimensions: whether the content was correct, and independently, whether the answer actually came back in Emirati dialect rather than MSA. We report both a partial-credit score (the judge's graded assessment) and a stricter pass/fail version, plus how often each model abstained instead of answering.
LLM-judged correctness on the 1,173 Alyah questions, open-ended generation, Gemini 3.7 as judge.
LLM-judged dialect fidelity on the same questions: does the answer actually come back in Emirati, or does the model default to MSA?
Falcon-Emirati-7B leads on correctness, but the real gap is in the second chart. On dialect fidelity, Falcon-Emirati-7B scores 0.52 (partial credit) against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct. That's close to two orders of magnitude at the low end. In practice, this means the other models often know the right answer but say it in Modern Standard Arabic by default, even when asked directly in Emirati. Falcon-Emirati-7B is the only one of the five that reliably answers back in the dialect it was asked in.
Fanar-2-27B-Instruct stands out for a second reason: it abstains far more than any other model, declining to answer 26.2% of the time, versus under 5% for every other model in the comparison. Combined with its correctness score of 0.27 (partial credit), the lowest of the five, it suggests a model that is both less willing and less able to engage with Emirati-specific content.
Dialect fidelity by Alyah category, partial credit. Falcon-Emirati-7B is the only model that consistently switches into Emirati; the others stay in MSA across nearly every category.
Breaking dialect fidelity down by category makes the pattern even clearer. It holds across every single category in Alyah, from everyday greetings to poetry, which suggests this isn't a narrow trick learned for a handful of question types. It's a general shift in what register the model defaults to when Emirati is the expected register. The one place competing models do relatively better—Greetings & Daily Expressions—is also the category where Emirati and MSA overlap the most, so it's the easiest place for a generic Arabic model to accidentally sound right.
As a third lens on the same question, we ran head-to-head pairwise judging: for every Alyah question, the judge (Gemini 3.7 Flash) was shown Falcon-Emirati-7B's answer next to a competing model's answer, blind to which was which.
Falcon-Emirati-7B Technical Specifications
| Attribute | Falcon-Emirati-7B |
|---|---|
| Developer | Technology Innovation Institute (TII) |
| Release date | Oct. 6, 2026 |
| Parameters | 7 billion |
| Primary specialization | Emirati Arabic |
| Base model | Falcon-H1-Arabic |
| Architecture | Hybrid Mamba + Transformer |
| Model type | Language model |
| Input/output focus | Understanding and generation |
| Training data | Emirati web/forums; MSA culture data; synthetic dialect examples |
| Context length | Not specified for 7B variant |
| Layers | Not published |
| Hidden size | Not published |
| Vocabulary size | Not published |
| Training hardware | Not published |
| License | Not specified |
Most Read
Nobody has commented on this yet.
💬 Be the first to comment