Case Study: How We Collected 2,000 Hours of Moroccan Arabic Conversational Speech Data
The Challenge: A Rare Language, a Tight Deadline
When a leading AI research lab approached us in late 2024, they had a problem that most speech data providers couldn't solve: they needed 2,000 hours of conversational Moroccan Arabic (Darija) speech data for training a new multilingual voice recognition model. Not Modern Standard Arabic. Not Gulf Arabic. Darija — a language with no standardized written form, heavy code-switching with French and Spanish, and significant regional variation across Morocco.
Three other providers had already declined the project. The challenges were clear:
- No existing datasets to bootstrap from — Moroccan Arabic is classified as a "low-resource" language in NLP research.
- Conversational, not read speech — they needed natural dialogue, not scripted monologues.
- Demographic diversity — balanced across age (18–65), gender, region (Casablanca, Rabat, Fès, Marrakech, Tangier), and socioeconomic background.
- 90-day delivery window — their model training cycle was locked to a quarterly release.
Our Approach: A Three-Phase Managed Collection Strategy
Phase 1: Speaker Recruitment & Screening (Weeks 1–2)
We deployed a recruitment team in Casablanca and Rabat, working with local community organizations and universities to source native Darija speakers. Every candidate underwent a three-step screening process:
- Language verification — A 10-minute conversational interview with a native linguist to confirm Darija fluency and assess code-switching patterns.
- Demographic verification — Age, region of origin, education level, and occupation verified to ensure demographic balance.
- Recording quality test — A 5-minute test recording in our target env
[Full translation coming soon — English version above]

