The tachelhiyt-darija dataset is a parallel dataset for Moroccan Darija and Tachelhiyt:
Paper
If you use this dataset, please cite our paper:
@inproceedings{atouf2025tachelhiyt,
title={Tachelhiyt-Darija: a parallel speech corpus for two underrepresented languages},
author={Atouf, Noureddine and Issa, Elsayed and Ouzbayr, Said},
booktitle={Proceedings of the 8th International Conference on Natural Language and Speech Processing (ICNLSP-2025)},
pages={379--384},
year={2025}
}