This dataset contains a large-scale collection of word bigrams extracted from text across 154 African languages. It is meticulously designed to support general research and analysis of African languages, providing foundational data based on two-word sequences.
Background