Are LLMs Good Text Diacritizers? An Arabic and Yorùbá Case Study
Hawau Olamide Toyin, Samar Magdy, Hanan Aldarmaki
We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation, we introduce a novel multilingual dataset MultiDiac