Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

osinkolu/Ehugbo-Crosslingual-Retrieval

Domain:

natural language processing

Record type:

dataset
Creator:
osi
Host:
# Ehugbo-QA: A Multimodal Benchmark for Cross-Dialect Information Retrieval (CDIR) **Ehugbo-QA** is the first dedicated multimodal benchmark for evaluating information access in the **Ehugbo** dialect (the Afikpo variety of Igbo, spoken in Ebonyi State, Nigeria). This repository provides the dataset and diagnostic evaluation suite used to identify the **"Alignment Gap"**—a systemic failure in regional African foundation models to bridge the semantic space between high-resource queries and dialectal documents. ## 📊 Final Leaderboard (Zero-Shot) We evaluate models on a **1-vs-312 Global Gallery Retrieval** task, matching English text queries to Ehugbo documents. | Model Class | Model Name | Task | Top-1 Accuracy | MRR | | :--- | :--- | :--- | :--- | :--- | | **Global Bi-Encoder** | **LaBSE** | Text-to-Text | **85.26%** | **0.8990** | | **Regional MLM** | Serengeti-E250 | Text-to-Text | 4.17% | 0.0789 | | **Regional MLM** | AfriBERTa | Text-to-Text | 1.28% | 0.0610 | | **Multimodal** | CLAP | Text-to-Audio | 0.32% | 0.0229 | --- ## 🔍 Key Diagnostic Findings ### 1. The Alignment Gap (Visualized) Our t-SNE analysis reveals a stark contrast in embedding topology. While **LaBSE** achieves language invariance (merging English and Ehugbo), regional models like **Serengeti** exhibit "Disconnected Islands." **Figure 1: Multimodal & Multilingual Alignment Comparison** * **Left (LaBSE):** Integrated semantic space. * **Right (Serengeti):** Segregated linguistic islands with a massive "no-man's land" in the middle. ### 2. The Hubness Phenomenon Qualitative error probing shows that unaligned models (Serengeti/AfriBERTa) suffer from **Hubness**. In our tests, Serengeti incorrectly mapped diverse theological queries to a single "hub" verse regarding baptism with **94.7% confidence**. --- ## 📂 Repository Contents - `data/ehugbodataset.csv`: The parallel corpus of 312 verses (Ehugbo audio links, transcripts, and English translations). - `notebooks/Ehugbo_Experiments.i …

Visit

github.com

Tasks

embeddingsinformation retrieval

Languages

Igbo

Licenses

MIT

Similar

Ehugbo TTS: biblical text to speech dataset in Ehugbo LanguageKachiengineers/Ehugbo-AudioImpact of Monolingual-Crosslingual Data Ratios on Zero-Shot Retrieval Accuracy in Low-Resource XNLIosinkolu/fongbe-hausa-asrosinkolu/yecs-asr-benchmarkosinkolu/Umoja-Hack-Africa-2022

Ehugbo TTS: biblical text to speech dataset in Ehugbo Language

This dataset contains audio recordings of Bible verses in Ehugbo, a dialect of Igbo (a Niger-Congo language spoken in Nigeria). It contains 312 audio recordings of biblical text-to-speech data comprising 1 hour and 30 seconds of speech data.

This dataset

Kachiengineers/Ehugbo-Audio

This is a repository of my Ehugbo project working on the first publicly available Ehugbo audio data

Impact of Monolingual-Crosslingual Data Ratios on Zero-Shot Retrieval Accuracy in Low-Resource XNLI

Information retrieval across different languages is an increasingly important challenge in natural l

osinkolu/fongbe-hausa-asr

# Fongbe ASR Dataset Creator A pipeline for building a unified Automatic Speech Recognition (ASR) d

osinkolu/yecs-asr-benchmark

Finetune Omnilingual ASR on YECS; benchmark vs MMS/Whisper (Yoruba-English code-switch) # YECS ASR

osinkolu/Umoja-Hack-Africa-2022

#Umoja Hacks. This repo contains team solo's best solutions for the umoja hackathon 2022 hosted on