This dataset contains a manually annotated Yoruba constituency treebank consisting of 1,000 sentences. The treebank was developed as part of an undergraduate linguistics research project focused on Yoruba syntax and computational parsing for under-resourced languages.
The annotations follow a phrase-structure (constituency) framework, including labels such as NP, VP, IP, and CP.
Dataset Contents