ブログに戻る

Case Study: How We Collected 2,000 Hours of Moroccan Arabic Conversational Speech Data

ケーススタディ: 72日間でAIのための2,000時間モロッコアラビア語音声データを収集した方法。Smart Language Service。

read time8 min
case studyReal Project
updated2026

Case Study: How We Collected 2,000 Hours of Moroccan Arabic Conversational Speech Data

The Challenge: A Rare Language, a Tight Deadline

When a leading AI research lab approached us in late 2024, they had a problem that most speech data providers couldn't solve: they needed 2,000 hours of conversational Moroccan Arabic (Darija) speech data for training a new multilingual voice recognition model. Not Modern Standard Arabic. Not Gulf Arabic. Darija — a language with no standardized written form, heavy code-switching with French and Spanish, and significant regional variation across Morocco.

Three other providers had already declined the project. The challenges were clear:

  • No existing datasets to bootstrap from — Moroccan Arabic is classified as a "low-resource" language in NLP research.
  • Conversational, not read speech — they needed natural dialogue, not scripted monologues.
  • Demographic diversity — balanced across age (18–65), gender, region (Casablanca, Rabat, Fès, Marrakech, Tangier), and socioeconomic background.
  • 90-day delivery window — their model training cycle was locked to a quarterly release.

Our Approach: A Three-Phase Managed Collection Strategy

Phase 1: Speaker Recruitment & Screening (Weeks 1–2)

We deployed a recruitment team in Casablanca and Rabat, working with local community organizations and universities to source native Darija speakers. Every candidate underwent a three-step screening process:

  1. Language verification — A 10-minute conversational interview with a native linguist to confirm Darija fluency and assess code-switching patterns.
  2. Demographic verification — Age, region of origin, education level, and occupation verified to ensure demographic balance.
  3. Recording quality test — A 5-minute test recording in our target env

    [Full translation coming soon — English version above]

    Contact Smart Language Service →