AiGpu

Models·

TII Launches Falcon-ASR: 1.6B Parameter Arabic Speech Model Leads Emirati Dialect Benchmarks

Technology Innovation Institute unveils Falcon-ASR, a compact 1.6B parameter automatic speech recognition model that sets new accuracy records for Arabic — especially the Emirati dialect — while supporting multilingual transcription with word-level timestamps.

Falcon-ASR model card showing Arabic speech recognition benchmarks and Emirati dialect performance

The Technology Innovation Institute (TII) in Abu Dhabi has released Falcon-ASR, a 1.6 billion parameter automatic speech recognition model purpose-built for Arabic with a strong emphasis on the Emirati dialect. The model also handles English, French, Spanish, and Portuguese, and provides word-level timestamps for precise alignment between audio and text.

On the Open Universal Arabic ASR Leaderboard maintained by the ELM Research Center, Falcon-ASR achieved an average word error rate (WER) of 20.92% across six standard Arabic test sets, improving on the previous best published result of 23.17%. Its character error rate (CER) stands at 8.79%. The benchmark includes diverse conditions such as broadcast news, telephone speech, and dialectal variations.

Emirati dialect focus delivers measurable gains

Beyond public benchmarks, TII conducted an internal evaluation on held-out Emirati and Gulf recordings with human-validated transcripts. Falcon-ASR recorded a 22.73% WER and 10.19% CER, outperforming significantly larger models including Qwen3-Omni (30B parameters) and Audar-ASR-V1-Turbo (2.35B parameters). The 4.07 percentage-point WER advantage over the next-best system highlights the value of targeted dialectal training data.

Training combined Emirati, Modern Standard Arabic, other Gulf dialects, and English corpora to capture code-switching and real-world speaking styles. The architecture remains efficient at 1.6B parameters, making it practical for deployment on modest GPU infrastructure without sacrificing accuracy.

Why it matters for GPU / AI infrastructure: Falcon-ASR demonstrates that carefully curated, dialect-specific datasets can yield state-of-the-art results at a fraction of the parameter count of general-purpose multilingual models. For teams running inference on GPU clouds, this translates to lower compute costs, reduced latency, and easier scaling — especially for Arabic-language contact centers, media monitoring, and accessibility applications in the GCC.

  • aigpu
  • ai gpu
  • ai gpu cloud
  • aigpu dubai
  • falcon-asr
  • arabic-asr
  • speech-recognition
  • tii-uae

By AiGpu Editorial · Editorial rewrite based on public reporting (Hugging Face Blog)

← All articles