This is a ~14 million tokens preprocessed corpus of Asante Twi, realized for the Luminos Fund's Bilingual Boost initiative thanks to funding from the Gates Foundation.
An interactive, searchable and visualization-rich version of this corpus is vailable at
asante-twi-corpus.fly.dev
Corpus composition
Subcorpus
Documents
Tokens
Share
Doctrinal
3,214
12,125,333
86.5%
General
2,246
1,353,049
9.7%
Children
190
535,293
3.8%
Total
5,650
14,013,675
100%
The corpus also includes 34,647 Twi-English aligned sentences
Data Quatlity
Preprocessing accuracy has been evalauted against a single-annotator triple-proofread set of 13 documents for a total of 12,162 proofed tokens (10,372 words + punctuation; 9,460 development / 2,702 sealed held-out, expert-proofed and re-tokenized to the delivered corpus). Accuracy, words-only (punctuation excluded — it is tagged trivially and inflates scores):
Metric · words-only (excl. punctuation)
Dev · 8,028 words
Held-out · 2,344 words
Lemma accuracy (row · aligned micro-F1)
0.915
0.902
— correct / missed
7,342 / 686
2,115 / 229
Lemma F1 (bag-of-tokens)
0.926
0.908
PoS accuracy (token)
0.873
0.827
PoS macro-F1 (per-tag)
0.799
0.788
— all-token (incl. punctuation): lemma / PoS acc
0.927 / 0.892
0.914 / 0.846
Lemma accuracy (~0.91) and PoS macro-F1 (~0.79) hold across development and the sealed held-out set, so the combined 12,162-token gold standard confirms the pipeline generalizes rather than overfits. 5-fold spread (words-only): dev lemma 0.915 ± 0.005 / PoS 0.873 ± 0.008; held-out lemma 0.902 ± 0.011 / PoS 0.827 ± 0.024.
Resources Included
5,650 corpus files
one metadata file detailing genre, title, author... of each corpus document
34,647 Twi-English aligned sentences