This is a parallel Twi-English dataset designed for training machine translation models that understand paragraph structure, including empty lines between text blocks and multiple consecutive sentences.
Source Language: English
Target Language: Twi (Akan)
Dataset Structure: Parallel paragraphs with 3 sentences each
Paragraph Patterns:
50% with 1 sentence + blank line + 2 sentences