Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning

Domaine:

natural language processing

Type de record:

paper
Créateur:
SamIslHosBhu
Hôte:avatar
The rapid growth of speech synthesis and voice conversion systems has made deepfake audio a major security concern. Bengali deepfake detection remains largely unexplored. In this work, we study automatic detection of Bengali audio deepfakes using the BanglaFake dataset. We evaluate zeroshot inference with several pretrained models. These include Wav2Vec2-XLSR-53, Whisper, PANNsCNN14, WavLM and Audio Spectrogram Transformer. Zero-shot results show limited detection ability. The best model, Wav2Vec2-XLSR-53, achieves 53.80% accuracy, 56.60% AUC and 46.20% EER. We then f ine-tune multiple architectures for Bengali deepfake detection. These include Wav2Vec2-Base, LCNN, LCNN-Attention, ResNet18, ViT-B16 and CNN-BiLSTM. Fine-tuned models show strong performance gains. ResNet18 achieves the highest accuracy of 79.17%, F1 score of 79.12%, AUC of 84.37% and EER of 24.35%. Experimental results confirm that fine-tuning significantly improves performance over zero-shot inference. This study provides the first systematic benchmark of Bengali deepfake audio detection. It highlights the effectiveness of f ine-tuned deep learning models for this low-resource language. Accepted for publication in 2025 28th International Conference on Computer and Information Technology (ICCIT)

Visit

arxiv.org

Tasks

speech processing

Tags

SoundArtificial Intelligence

Similaires

Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned TransformersZero-Shot Transfer Learning using Affix and Correlated Cross-Lingual Embeddings.Exploring Transliteration-Based Zero-Shot Transfer for Amharic ASRIntermediate-Task Training and Few-Shot Learning for Cross-Domain Robustness in Zero-Shot Cross-Lingual TransferContrastive Learning with Synthetic Cross-Lingual Pairs for Zero-Shot Transfer on XTREME-RScaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREME

Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers

Large language models (LLMs) can produce text that closely resembles human writing. This capability

Zero-Shot Transfer Learning using Affix and Correlated Cross-Lingual Embeddings.

Learning morphologically supplemented embedding spaces using cross-lingual models has become an acti

Exploring Transliteration-Based Zero-Shot Transfer for Amharic ASR

Intermediate-Task Training and Few-Shot Learning for Cross-Domain Robustness in Zero-Shot Cross-Lingual Transfer

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Contrastive Learning with Synthetic Cross-Lingual Pairs for Zero-Shot Transfer on XTREME-R

Recently, although pre-trained language models have achieved great success on multilingual NLP (Natu

Scaling Model Size and Zero-Shot Transfer to Low-Resource Languages in XTREME

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni