darija-open-dataset/dataset
darija <-> english dataset
What it solves
It provides a comprehensive, open-source dataset for the Moroccan dialect (Darija), addressing the lack of high-quality linguistic resources for Natural Language Processing (NLP) in this specific dialect. It bridges the gap between Darija and English, supporting both Arabic and Latin script spellings.
How it works
The project compiles human-curated entries including semantic and syntactic categorizations, verb conjugations, and translated sentences. It offers two main data formats:
- Human Dataset: A collection of reviewed human entries and sentences.
- DODa-500K: A large-scale parallel file containing over 500,000 rows, combining human-validated data with synthetic, model-generated expansions for machine translation training.
Access to the human dataset is facilitated via PyDODa, a Python wrapper library that allows developers to programmatically retrieve translations and explore linguistic categories.
Who it’s for
- NLP Practitioners: Those building machine translation or language models for Darija.
- Researchers: Linguists and AI researchers studying the Moroccan dialect.
- Language Enthusiasts: People interested in learning or analyzing Darija.
Highlights
- Large Scale: Includes approximately 150,000 human entries and a 500K parallel dataset for MT training.
- Multiscript Support: Covers both Arabic and Latin (Arabizi) alphabets.
- Linguistic Depth: Includes verb-to-noun correspondences, masculine-to-feminine variations, and detailed verb conjugations.
- Developer Friendly: Provides a dedicated Python library (PyDODa) for easy data integration.
Related
- Project
- Project
- Project
- Project