Tuesday, October 6, 2026 Newsletter Advertise
Breaking
AI

TII Releases Falcon-Emirati-7B AI Model for Emirati Arabic Dialect

Researchers at TII released Falcon-Emirati-7B, a 7-billion-parameter language model specialized in Emirati Arabic dialect, vocabulary, and cultural context.

TII Releases Falcon-Emirati-7B AI Model for Emirati Arabic Dialect. Source: Hugging Face

Researchers at the Technology Innovation Institute announced Falcon-Emirati-7B on October 6, 2026, a 7-billion-parameter artificial intelligence model designed to understand and generate Emirati Arabic.

Model architecture and data pipeline

Falcon-Emirati-7B is built on top of the 7B variant of Falcon-H1-Arabic. The underlying architecture uses a combination of State Space Models (Mamba) and Transformer attention running in parallel inside each block.

To adapt the model to the dialect, the development team created a pipeline using three primary data sources: authentic Emirati dialect web content, Modern Standard Arabic materials on Emirati heritage and identity, and synthetic dialect data generated using strict grammatical rules and vocabulary glossaries.

Benchmark performance and dialect fidelity

The researchers tested Falcon-Emirati-7B on Alyah, a native multiple-choice benchmark consisting of 1,173 samples across categories such as poetry, heritage, and daily expressions. Falcon-Emirati-7B achieved an accuracy score of 84.83 percent on Alyah.

In open-ended generation evaluated by Gemini 3.7 Flash as an LLM judge, Falcon-Emirati-7B achieved a dialect fidelity score of 0.52 partial credit. In comparison, ALLaM-7B-Instruct-preview scored 0.05, gemma-3-27b-it scored 0.03, and Jais-2-8B-Chat scored 0.02.

The model also scored 85.57 percent on the UAE multiple-choice task of the ArabCulture-Dialogue benchmark across 283 scenarios, outperforming ALLaM-7B at 83.39 percent, Jais-2-8B at 73.79 percent, and Fanar-2-27B at 71.50 percent.

Availability and limitations

Falcon-Emirati-7B has been made available to users through TII's chat platform.

The researchers noted that Falcon-Emirati-7B can reflect training data biases and may misinterpret rare expressions or localized references, advising evaluation prior to deployment in high-stakes environments.

Key facts and where they come from
  • Falcon-Emirati-7B is built on top of Falcon-H1-Arabic.
    Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that already set new benchmarks for the language earlier this year.
  • Falcon-Emirati-7B scored 84.83% on the Alyah benchmark.
    Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other Arabic and multilingual model we compared it against
  • Falcon-Emirati-7B achieved a 0.52 dialect fidelity score on open-ended Alyah questions.
    On dialect fidelity, Falcon-Emirati-7B scores 0.52 (partial credit) against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat
  • Falcon-Emirati-7B scored 85.57% on the UAE portion of ArabCulture-Dialogue.
    Falcon-Emirati-7B scored 85.57%, the highest among the four models tested, ahead of ALLaM-7B (83.39%), Jais-2-8B (73.79%), and Fanar-2-27B (71.50%).

Read the original from Hugging Face →

The TechUpscale Brief

The day's cyber, AI and tech news in one short email, every weekday morning. Free. Unsubscribe anytime.

I'm most interested in

More AI