Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ArPod2.0: A Novel Arabic Podcast Dataset for Spoken Topic Identification

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Lou
Éditeur:
HDS
Hôte:avatar
This study introduces a new corpus designed to explore spoken topics in Arabic. Despite the growing interest in building new resources for Arabic, to our knowledge, there is no existing corpus designed specifically for recognizing the topics discussed. We assembled a substantial collection of Arabic podcasts, known as ArPod2.0, covering diverse subjects such as Business, Health, and Technology. This corpus comprises over 30 hours of speech, featuring both standard Arabic and various dialects from regions including Saudi Arabia and Egypt. Careful attention was given to curating this dataset to facilitate computer-based analysis of speech. Our focus was on six main topics, and we employed advanced computer techniques to analyze the data. Our findings highlight the efficacy of one such technique, the Multilayer Perceptron (MLP), which demonstrated exceptional performance in accurately discerning topics, even amidst mixed content. Updated Version

Visit

doi.orgzenodo.org

Tasks

speech processingtext classificationtopic classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Deep Confessions Podcast Arabic Speech DatasetSpoken Arabic Algerian dialect identificationArabic topic identification based on empirical studies of topic modelsA Hierarchical Approach for Topic IdentificationArabic topic identification based on empirical studies of topic models Identification thématique Arabe basée sur des études empiriques des topic modelsA Novel Dataset for Arabic Speech Recognition Recorded by Tamazight Speakers

Deep Confessions Podcast Arabic Speech Dataset

The Deep Confessions Podcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech

Spoken Arabic Algerian dialect identification

Arabic topic identification based on empirical studies of topic models

This paper focuses on the topic identification for the Arabic language based on topic models. We stu

A Hierarchical Approach for Topic Identification

Colloque avec actes et comité de lecture. internationale. International audience This

Arabic topic identification based on empirical studies of topic models Identification thématique Arabe basée sur des études empiriques des topic models

International audience This paper focuses on the topic identification for the Arabic

A Novel Dataset for Arabic Speech Recognition Recorded by Tamazight Speakers

Automatic Speech Recognition (ASR) is an area of research that's constantly evolving, thanks to impo