Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

Domain:

natural language processing

Record type:

papermodel
Creator:
SukCho
Host:avatar
Large pre-trained speech models excel in downstream tasks but their deployment is impractical for resource-limited environments. In this paper, we introduce HArnESS, the first Arabic-centric self-supervised speech model family, designed to capture Arabic speech nuances. Using iterative self-distillation, we train large bilingual HArnESS (HL) SSL models and then distill knowledge into compressed student models (HS, HST), preserving Arabic-specific representations. We use low-rank approximation to further compact the teacher's discrete supervision into shallow, thin models. We evaluate HArnESS on Arabic ASR, Speaker Emotion Recognition (SER), and Dialect Identification (DID), demonstrating effectiveness against HuBERT and XLS-R. With minimal fine-tuning, HArnESS achieves SOTA or comparable performance, making it a lightweight yet powerful alternative for real-world use. We release our distilled models and findings to support responsible research and deployment in low-resource settings. 5 pages, 4 figures

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and Language

Similar

ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion RecognitionCasablanca: Data and Models for Multidialectal Arabic Speech RecognitionMinimum Bayes risk discriminative language models for Arabic speech recognitionDialectal Adaptation of Foundation Models for Low-Resource Speech Synthesis: A Case Study on Adamawa FulfuldeEvaluating New AI Cell Foundation Models on Challenging Kidney Pathology Cases Unaddressed by Previous Foundation ModelsLightweight Deep Learning Models for Brain Tumor Classification

ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition

Speech emotion recognition is vital for human-computer interaction, particularly for low-resource la

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

In spite of the recent progress in speech processing, the majority of world languages and dialects r

Minimum Bayes risk discriminative language models for Arabic speech recognition

Dialectal Adaptation of Foundation Models for Low-Resource Speech Synthesis: A Case Study on Adamawa Fulfulde

The rapid evolution of neural speech synthesis has achieved near-human naturalness for high-resource

Evaluating New AI Cell Foundation Models on Challenging Kidney Pathology Cases Unaddressed by Previous Foundation Models

Accurate cell nuclei segmentation is critical for downstream tasks in kidney pathology and remains a

Lightweight Deep Learning Models for Brain Tumor Classification

Differentiating brain tumours by MRI using computer algorithms remains a huge challenge in clinical