Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Bridging the Data Provenance Gap Across Text, Speech and Video

Domaine:

natural language processing

Type de record:

paper

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities--popular text, speech, and video datasets--from their detailed sourcing trends and use restrictions to their geographical and linguistic representation. Our manual analysis covers nearly 4000 public datasets between 1990-2024, spanning 608 languages, 798 sources, 659 organizations, and 67 countries. We find that multimodal machine learning applications have overwhelmingly turned to web-crawled, synthetic, and social media platforms, such as YouTube, for their training sets, eclipsing all other sources since 2019. Secondly, tracing the chain of dataset derivations we find that while less than 33% of datasets are restrictively licensed, over 80% of the source content in widely-used text, speech, and video datasets, carry non-commercial restrictions. Finally, counter to the rising number of languages and geographies represented in public AI training datasets, our audit demonstrates measures of relative geographical and multilingual representation have failed to significantly improve their coverage since 2013. We believe the breadth of our audit enables us to empirically examine trends in data sourcing, restrictions, and Western-centricity at an ecosystem-level, and that visibility into these questions are essential to progress in responsible AI. As a contribution to ongoing improvements in dataset transparency and responsible use, we release our entire multimodal audit, allowing practitioners to trace data provenance across text, speech, and video.

Visit

arxiv.org

Tasks

computer visionspeech processing

Tags

dataset auditanalysismeta-datasetdata governancedataset transparencydataset documentationdata sourcingdata ethicsAI accountabilitydataset analysis+7

Similaires

SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion DetectionBridging the Legal-technical Gap: AI Surveillance and Data Sovereignty in AfricaBridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech RecognitionCultural Narratives: Bridging the gap. Educating Speech-language Pathologists to Work in Multicultural PopulationsCross-lingual Matryoshka Representation Learning across Speech and TextBridging Nigeria's Energy Gap - A Geospatial Data-Driven Approach

SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Detection

We present our shared task on text-based emotion detection, covering more than 30 languages from sev

Bridging the Legal-technical Gap: AI Surveillance and Data Sovereignty in Africa

South African and Nigerian law guarantees rights to privacy and due process, but national security c

Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition

Modern machine learning models for audio tasks often exhibit superior performance on English and oth

Cultural Narratives: Bridging the gap. Educating Speech-language Pathologists to Work in Multicultural Populations

Preparation of the clinician to work in multicultural contexts involves the identification of a rang

Cross-lingual Matryoshka Representation Learning across Speech and Text

Speakers of under-represented languages face both a language barrier, as most online knowledge is in

Bridging Nigeria's Energy Gap - A Geospatial Data-Driven Approach

This report outlines the design and implementation of the REA Geospatial Management System (REA-GMS)