This repository functions as a centralized, public knowledge base tailored for annotating Oromo language datasets to train and evaluate Natural Language Processing (NLP), speech recognition, and content safety models. It standardizes how linguistic data should be processed by human labeling teams to ensure consistency.