The Bayelemabaga dataset is a collection of 46976 aligned machine translation ready Bambara-French lines, originating from Corpus Bambara de Reference. The dataset is constitued of text extracted from 264 text files, varing from periodicals, books, short stories, blog posts, part of the Bible and the Quran.
Lines
46976
French Tokens (spacy)