Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

asmaatagelsir/Sudan-MM-2025-Metadata-Architecture

Domain:

natural language processing

Record type:

dataset
Creator:
asm
Host:
Multimodal dataset architecture for Sudanese culture and dialect, featuring 695 image/video pairs with MSA and Sudanese Arabic captions. Developed by Team 4sparks for the Sudan-MM 2025 initiative # Sudan-MM-2025-Metadata-Architecture Multimodal dataset architecture for Sudanese culture and dialect, featuring 695 image/video pairs with MSA and Sudanese Arabic captions. Developed by Team 4sparks for the Sudan-MM 2025 initiative # Sudan-MM 2025: Multimodal Dataset Architecture This project represents the technical documentation and metadata architecture for the **Sudan-MM 2025** dataset, developed by **Team 4sparks**. ## Project Overview Sudan-MM 2025 is a multimodal dataset designed to bridge the linguistic gap in African AI. It features 695 pairs (640 images and 55 videos) representing daily life, culture, and environments across various states in Sudan. ## My Role (Asmaa Tagelsir) As a core member of Team 4sparks, my primary responsibilities included: - **Media Collection:** Sourcing 700+ high-quality images and videos ensuring geographic diversity across Sudan. - **Linguistic Annotation:** Writing and refining authentic captions in both **Modern Standard Arabic (MSA)** and **Sudanese Arabic dialect**. - **Ethical Sourcing:** Managing permissions and licensing from local content creators to ensure a legally sound and ethically compliant dataset. ## Technical Specifications - **Format:** Image/Video (.jpg, .mp4) with synchronized Audio (.mp3). - **Categories:** Agriculture, Public Infrastructure, Marketplaces, and Local Culture. - **Goal:** To support the development of Sovereign AI that understands Sudanese linguistic nuances. ## Team: 4sparks - Nafisa Ahmed - Rafa Abdallah - Miada Abdalmonim - Asmaa Tagelsir ## Dataset Access The multimodal media files (640 images, 55 videos, and 695 voice recordings) are hosted on Google Drive for high-resolution integrity. - **Dataset Root Folder:** [drive.google.com] - **Metadata Inventory:** The `metadata.csv` file in this repository contains the synchronized captions and file IDs.

Visit

github.com

Tasks

computer visionimage-text retrieval

Languages

Arabic, Sudanese Spoken

Similar

A7mdos/sudan-mm-2025-automatorSudan-MMsouthsudanemr/ihs-sudan-metadataMetadata 0200-20141220 - Metadata Documenting Tabaq, a Hill Nubian language of the Sudan, in its sociolinguistic context

A7mdos/sudan-mm-2025-automator

# Sudan-MM-2025 Automator A Streamlit web application for automating multimodal data collection wor

Sudan-MM

Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-r

southsudanemr/ihs-sudan-metadata

# IHS South Sudan Bahmni Metadata ## Description Bahmni metadata which includes the following: * P

Metadata 0200-20141220 - Metadata Documenting Tabaq, a Hill Nubian language of the Sudan, in its sociolinguistic context

Language_Name: Tabaq Language_Region: Africa Language_Country: Sudan Project_Status: Ongoi