Logo Lanfrica

Hierarchical Curriculum Transfer for Swahili QA with mT5: SQuAD Pretraining and KenSwQuAD Extractive-to-Abstractive Refinement

Domaine:

natural language processing

Type de record:

paper
Créateur:
JarBen
Éditeur:
Nex
Hôte:
The introduction of the Kencorpus Swahili Question Answering Dataset (KenSwQuAD) presents a compelling opportunity to advance Machine Reading Comprehension (MRC) for the Swahili language. However, the mixed composition of this dataset—comprising 67.6\% extractive and 32.4\% abstractive answers—introduces substantial hurdles for standard training pipelines. In this study, we benchmark the performance of the multilingual T5 (mT5) architecture on KenSwQuAD, deploying a robust Hierarchical Curriculum Learning strategy. We introduce a novel data restructuring technique that algorithmically partitions the corpus into distinct extractive and abstractive subsets, thereby enabling meticulously phased fine-tuning. By sequentially progressing from structural transfer via the English SQuAD, to morphological alignment on the Swahili extractive subset, and culminating in abstractive refinement, we establish that the mT5-base model can achieve a SacreBLEU score of 48.99 on extractive tasks. For abstractive reasoning, whilst the BLEU score is comparatively modest (15.52), the model attains a remarkable BERTScore F1 of 77.21\%, reflecting a profound degree of semantic comprehension. Furthermore, we evaluate the utility of structure-aware input formatting (context scaffolding) for successfully navigating extensive narrative contexts. Ultimately, our findings indicate that whilst modern transformer architectures can successfully internalise Swahili QA logic, a persistent  fluency gap characterises the abstractive domain. This gap underscores the urgent necessity for more expansive, targeted datasets to reconcile semantic accuracy with lexical precision.

Languages

Similaires