Darija is spoken by millions and still treated as a rounding error in most language resources. Darija Open Dataset (DODa) is an attempt to change that: an open, collaborative Darija ⇆ English resource built specifically for NLP.
What it contains
- On the order of 500,000 semantic rows (human + controlled synthetic)
- Semantic and syntactic categorisation
- Multiple spellings for the same word (Darija is not a single orthography)
- Verb–noun and masculine–feminine correspondences
- Conjugation of hundreds of verbs across tenses
- Dual script: Arabic and Latin, matching how people actually write
- Aligned utterances across Arabic Darija, Arabizi, and English
Version 1 started around 10,000 entries. Version 2 crossed 100,000. The latest expansion reaches 500,350 rows through controlled synthetic generation grounded in DODa — written up in How We Expanded DODa to 500K Darija Examples.
Why translation, not only raw text
Training a model directly on raw Darija has a place. Translation is the other lever: it lets Darija borrow the tooling, evaluation sets, and applications already built for English instead of starting from zero every time.
Papers
- Moroccan Dialect -Darija- Open Dataset (2021)
- The Evolution of Darija Open Dataset: Introducing Version 2 (2024)
The dataset is collaborative on purpose. The Moroccan IT community is the point, not a footnote.