---
## ๐งญ What it is
**Audar-ASR-V1-Turbo** is an **Arabic-first generative speech-recognition model** โ the accuracy tier of
the Audar-ASR family. It recasts transcription as **audio-conditioned next-token prediction** over a
unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is
built on a permissively-licensed open-weight audio-LLM foundation and adapted in-house โ the contribution is the adaptation (the data curriculum and the alignment rubric), not the foundation:
- ๐งฑ **Large-scale bilingual pretraining** โ 300,000+ hours of labeled audio, primarily Arabic and English, spanning MSA, Gulf,
Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
- ๐ฏ **Dialect-targeted fine-tuning** โ hardness sampling and multi-task conditioning focused on proper
nouns, code-switching, and dialect-faithful orthography.
- ๐ง **KTO preference alignment** โ Kahneman-Tversky Optimization on accented dialectal Arabic, with
unpaired binary-desirability labels from trained native annotators across the Gulf, Levantine,
Egyptian, and Maghrebi dialects, along five axes: verbatim accuracy, diacritic correctness,
code-switch handling, named-entity preservation, and output formatting.
The result is **state-of-the-art dialectal Arabic ASR** โ the lowest average WER *and* CER of any
evaluated system on the *Open Universal Arabic ASR Leaderboard*. It transcribes MSA and every major
Arabic dialect, code-switched ArabicโEnglish, and English, across **30 languages** in total.
> Built on a **permissively-licensed open-weight audio-LLM foundation**; the adaptation, data, and
> alignment are Audar's. Full method and results: [Audar-ASR-V1 Technical Report](https://github.com/AudarAI/Audar-ASR-V1/blob/main/report/Audar-ASR-V1-Technical-Report.pdf).
## Model summary
Model
Audar-ASR-V1-Turbo โ Arabic-first generative ASR (accuracy tier)
built on an open-weight audio-LLM foundation; adapted via a 4-stage curriculum โ 300k+ hrs bilingual pretraining โ multi-task fine-tuning โ dialect PEFT โ KTO alignment
Decoder parameters
2,031,739,904 (2.03B)
Audio encoder parameters
317,477,504 (0.32B)
Total parameters
2,349,217,408 (2.35B, bf16)
Audio input
16 kHz mono; 30 s context (longer audio is chunked/streamed)
Languages
Arabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more
Runtime
GGUF / llama.cpp โ CPU ยท GPU ยท edge
License
AudarAI Community License v1.0
## ๐ Benchmarks
Arabic dialectal ASR is **hard** โ heavily dialectal, conversational, code-switched speech is the
frontier for every system. On the *Open Universal Arabic ASR Leaderboard*, **Audar-ASR-V1-Turbo ranks
#1 of 37 systems** with the **lowest average WER (23.2 %) and the lowest average CER (9.2 %)** of any
model evaluated โ and it is the single best system on
**SADA**, **MASC-clean**, **MGB-2** and **Casablanca**.
### Open Universal Arabic ASR Leaderboard โ full standings
*Per-dataset **WER %** across all six leaderboard test sets, plus the two composite averages. Lower is
better; **Avg WER** is the ranking metric. Audar rows show the leaderboard maintainers' independent reproduction (Aug 2026) under the
leaderboard's current normalization; other rows are as previously published by the leaderboard and
may shift slightly when the full board is recomputed under the updated normalization. **Ours in bold.***
| # | Model | **Avg WER** | Avg CER | SADA | CV-18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca |
| --: | --- | --: | --: | --: | --: | --: | --: | --: | --: |
| **1** | **Audar-ASR-V1-Turbo (Ours)** | **23.17** | **9.20** | **28.92** | 8.09 | **16.73** | 27.19 | **11.08** | **47.02** |
| 2 | CohereLabs/cohere-transcribe-arabic-07-2026 | 25.87 | 11.80 | 37.47 | **5.82** | 19.60 | **27.07** | 15.54 | 49.71 |
| 3 | omnilingual-asr/omniASR_LLM_7B | 28.32 | 12.52 | 41.61 | 8.75 | 19.69 | 29.29 | 14.13 | 56.46 |
| 4 | omnilingual-asr/omniASR_LLM_3B | 29.96 | 13.77 | 46.18 | 9.15 | 19.90 | 30.03 | 14.22 | 60.27 |
| 5 | omnilingual-asr/omniASR_LLM_1B | 29.96 | 13.40 | 43.84 | 9.55 | 20.03 | 30.26 | 15.34 | 60.68 |
| 6 | CohereLabs/cohere-transcribe-03-2026 | 30.67 | 16.37 | 60.11 | 8.17 | **8.66** | **19.01** | 25.33 | 62.71 |
| 7 | Qwen/Qwen3-Omni-30B-A3B-Instruct | 30.71 | 13.67 | 44.82 | 11.46 | 21.47 | 30.85 | 13.09 | 62.55 |
| 8 | nvidia-conformer-ctc-large-arabic (lm) | 32.91 | 13.84 | 44.52 | 8.80 | 23.74 | 34.29 | 17.20 | 68.90 |
| 9 | omnilingual-asr/omniASR_LLM_300M | 32.96 | 14.84 | 51.38 | 12.03 | 20.66 | 32.45 | 16.58 | 64.64 |
| 10 | google/gemma-4-E4B-it | 32.98 | 13.71 | 43.40 | 19.65 | 24.86 | 33.59 | 17.72 | 58.63 |
| 11 | Qwen/Qwen3-ASR-1.7B | 33.36 | 12.33 | 45.53 | 16.90 | 24.37 | 34.29 | 16.57 | 64.47 |
| 12 | mistralai/Voxtral-Small-24B-2507 | 34.47 | 15.29 | 50.82 | 15.25 | 23.96 | 34.43 | 16.03 | 66.30 |
| 13 | nvidia-conformer-ctc-large-arabic (greedy) | 34.74 | 13.37 | 47.26 | 10.60 | 24.12 | 35.64 | 19.69 | 71.13 |
| 14 | google/gemma-4-E2B-it | 35.87 | 15.34 | 46.23 | 23.76 | 27.47 | 36.15 | 20.72 | 60.87 |
| 15 | openai/whisper-large-v3 | 36.86 | 17.21 | 55.96 | 17.83 | 24.66 | 34.63 | 16.26 | 71.81 |
| 16 | omnilingual-asr/omniASR_CTC_3B | 37.78 | 19.79 | 69.85 | 14.19 | 21.48 | 34.60 | 18.96 | 67.58 |
| 17 | omnilingual-asr/omniASR_CTC_7B | 38.12 | 20.91 | 72.69 | 12.47 | 21.08 | 35.04 | 20.43 | 67.02 |
| 18 | facebook/seamless-m4t-v2-large | 38.16 | 17.03 | 62.52 | 21.70 | 25.04 | 33.24 | 20.23 | 66.25 |
| 19 | omnilingual-asr/omniASR_CTC_1B | 39.29 | 20.47 | 71.42 | 17.55 | 22.76 | 35.73 | 19.96 | 68.32 |
| 20 | openai/whisper-large-v3-turbo | 40.05 | 18.87 | 60.36 | 25.73 | 25.51 | 37.16 | 17.75 | 73.79 |
| 21 | openai/whisper-large-v2 | 40.20 | 19.55 | 57.46 | 21.77 | 27.25 | 38.55 | 25.17 | 71.01 |
| 22 | Qwen/Qwen3-ASR-0.6B | 42.19 | 16.23 | 53.75 | 28.28 | 31.34 | 42.63 | 25.45 | 71.68 |
| 23 | openai/whisper-large | 42.57 | 20.49 | 63.24 | 26.04 | 28.89 | 40.79 | 24.28 | 72.18 |
| 24 | mistralai/Voxtral-Mini-3B-2507 | 42.58 | 19.90 | 63.65 | 22.12 | 28.37 | 41.27 | 22.56 | 77.52 |
| 25 | asafaya/hubert-large-arabic-transcribe | 45.50 | 17.35 | 67.82 | 8.01 | 32.94 | 50.16 | 37.51 | 76.53 |
| 26 | openai/whisper-medium | 45.57 | 22.27 | 67.71 | 28.07 | 29.99 | 42.91 | 29.32 | 75.44 |
| 27 | nvidia-Parakeet-ctc-1.1b-concat | 46.54 | 23.88 | 70.70 | 26.34 | 30.49 | 45.95 | 24.94 | 80.80 |
| 28 | omnilingual-asr/omniASR_CTC_300M | 46.65 | 21.86 | 78.11 | 27.90 | 28.40 | 43.26 | 26.85 | 75.35 |
| 29 | nvidia-Parakeet-ctc-1.1b-universal | 51.96 | 25.19 | 73.58 | 40.01 | 36.16 | 50.03 | 30.68 | 81.30 |
| 30 | microsoft/VibeVoice-ASR | 52.99 | 28.95 | 69.83 | 44.25 | 32.95 | 52.43 | 25.10 | 93.37 |
| 31 | facebook/mms-1b-all | 54.54 | 21.45 | 77.48 | 26.52 | 38.82 | 57.33 | 39.16 | 87.95 |
| 32 | openai/whisper-small | 55.13 | 21.68 | 78.02 | 24.18 | 35.93 | 56.36 | 48.64 | 87.64 |
| 33 | whitefox123/w2v-bert-2.0-arabic-4 | 58.13 | 27.62 | 87.34 | 41.79 | 37.82 | 53.28 | 40.66 | 87.88 |
| 34 | jonatasgrosman/wav2vec2-large-xlsr-53-arabic | 60.98 | 25.61 | 86.82 | 23.00 | 42.75 | 64.27 | 56.29 | 92.72 |
| 35 | speechbrain/asr-wav2vec2-commonvoice-14-ar | 65.74 | 30.93 | 88.54 | 29.17 | 49.10 | 69.57 | 64.37 | 93.68 |
**Bold** = best in column. The 37th system, our sibling edge model [**Audar-ASR-V1-Flash** (0.78B)](https://huggingface.co/audarai/Audar-ASR-V1-Flash), enters at 32.04 avg WER โ see its card for the full row. Audar-ASR-V1-Turbo owns both composite averages and leads on SADA, MASC-clean, MGB-2 and Casablanca;
the recent Cohere and OmniASR systems are the closest competitors, each strongest on a subset of the
conversational and clean-read sets. Casablanca (Moroccan Darija) is the hardest set for every system.
### Emirati Arabic
| Set | WER % | CER % |
|---|---|---|
| **Emirati** (Mixat, full 1,585-clip test) | **19.4** | **7.3** |
On Emirati, the **real recognition error is โ 7.3 %** โ near-parity with spontaneous English โ while the
residual up to 19.4 % WER is largely **orthographic convention** (near-miss spelling of the *same*
word, e.g. ุงูุชูโุงูุชูุง, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.
## ๐ Benchmark-parity inference (qwen-asr) โ recommended
Our leaderboard numbers were produced with the [`qwen-asr`](https://pypi.org/project/qwen-asr/) package, which implements this model's I/O protocol natively โ and were independently reproduced by the leaderboard maintainers with this exact code:
```python
# pip install qwen-asr torch
import torch
from qwen_asr import Qwen3ASRModel
model = Qwen3ASRModel.from_pretrained(
"audarai/Audar-ASR-V1-Turbo",
dtype=torch.bfloat16, device_map="cuda:0",
max_inference_batch_size=16, max_new_tokens=256,
)
results = model.transcribe(audio=["clip.wav"], language=["Arabic"])
print(results[0].text)
```
Protocol handling is **mandatory**, not optional:
- `language="Arabic"` makes the package prefill `language Arabic` into the prompt, so the model never free-runs language identification.
- The model's no-speech verdict (`language None`) is mapped to an empty transcript; without this, non-speech audio (music, silence) can produce repetition loops.
- `max_new_tokens=256` and bf16 are the exact decode settings behind our published numbers.
If you use raw `transformers` (below), you **must** strip the `language ` output prefix yourself and expect degraded scores on non-speech-heavy data.
## ๐ป GGUF inference (llama.cpp)
Turbo runs on **llama.cpp** via the multimodal (`mtmd`) path โ a quantized **decoder** GGUF plus a
**BF16 audio projector** (`mmproj`). Build a recent llama.cpp (with Qwen3-ASR support), then:
```bash
./llama-mtmd-cli \
-m Audar-ASR-V1-Turbo-Q8_0.gguf \
--mmproj mmproj-Audar-ASR-V1-Turbo.gguf \
--audio clip.wav \
-sys "ูุฑูุบ ุงูููุงู ุงูุนุฑุจู ุงูุชุงูู." \
--temp 0
```
> โ ๏ธ The **audio projector (`mmproj`) must stay BF16** (its `ClippableLinear` is numerically
> sensitive). The **decoder** quantizes normally.
Prefer a managed endpoint? The Audar-ASR family is also available via the
[**Audar API/SDK**](https://www.audarai.com) โ streaming, speaker-attributed transcription, and
diarization, production-hosted.
### GGUF variants
| File | Approx. size | Notes |
|---|---|---|
| `Audar-ASR-V1-Turbo-Q4_K_M.gguf` | ~1.28 GB | Smallest; constrained hardware |
| `Audar-ASR-V1-Turbo-Q8_0.gguf` | ~2.16 GB | Near-lossless (recommended) |
| `Audar-ASR-V1-Turbo.gguf` (BF16) | ~4.07 GB | Full precision decoder |
| `mmproj-Audar-ASR-V1-Turbo.gguf` | ~0.64 GB | **BF16 audio encoder โ required, keep BF16** |
## ๐ค Transformers (full-precision safetensors)
The **full-precision bf16 weights** are published at the **repo root** โ the reference
checkpoint the GGUF and W4A16 builds are derived from (**2,349,217,408** params, `safetensors`). Standard
๐ค Transformers, loaded with `trust_remote_code=True` (the repo ships the self-contained Qwen3-ASR code).
```python
# pip install "transformers==4.57.6" torch librosa
import torch, librosa
from transformers import AutoProcessor, AutoModelForCausalLM
repo = "audarai/Audar-ASR-V1-Turbo"
proc = AutoProcessor.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True,
dtype=torch.bfloat16, device_map="cuda:0",
).eval()
SYSTEM = "ูุฑูุบ ุงูููุงู ุงูุนุฑุจู ุงูุชุงูู." # "Transcribe the following Arabic speech."
audio, _ = librosa.load("clip.wav", sr=16000, mono=True)
conv = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": [{"type": "audio"}]}] # audio placeholder (a list, not "