
This dataset contains a manually annotated Yoruba constituency treebank consisting of 1,000 sentences drawn from diverse sources, including grammar books, religious texts, news media, literary works, and spoken Yoruba.
Each sentence is annotated with phrase-structure representations capturing core syntactic categories such as NP, VP, IP, and CP, with particular attention to Yoruba-specific constructions including serial verb constructions, focus structures, and embedded clauses.
The treebank was developed as part of an undergraduate linguistics project and is intended to support syntactic research and computational modeling for Yoruba, an under-resourced Niger-Congo language. The dataset can be used for linguistic analysis, constituency parsing experiments, and the development or evaluation of NLP models.