Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers

Domaine:

natural language processing

Type de record:

paper
Créateur:
IslSamHosZam
Hôte:avatar
Large language models (LLMs) can produce text that closely resembles human writing. This capability raises concerns about misuse, including disinformation and content manipulation. Detecting AI-generated text is essential to maintain authenticity and prevent malicious applications. Existing research has addressed detection in multiple languages, but the Bengali language remains largely unexplored. Bengali's rich vocabulary and complex structure make distinguishing human-written and AI-generated text particularly challenging. This study investigates five transformer-based models: XLMRoBERTa-Large, mDeBERTaV3-Base, BanglaBERT-Base, IndicBERT-Base and MultilingualBERT-Base. Zero-shot evaluation shows that all models perform near chance levels (around 50% accuracy) and highlight the need for task-specific fine-tuning. Fine-tuning significantly improves performance, with XLM-RoBERTa, mDeBERTa and MultilingualBERT achieving around 91% on both accuracy and F1-score. IndicBERT demonstrates comparatively weaker performance, indicating limited effectiveness in fine-tuning for this task. This work advances AI-generated text detection in Bengali and establishes a foundation for building robust systems to counter AI-generated content. Accepted for publication in 2025 28th International Conference on Computer and Information Technology (ICCIT)

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageArtificial Intelligence

Similaires

Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer LearningComparative Performance of Code-Switched Zero-Shot and XNLI-Fine-Tuned Retrieval Models on MIRACL Low-Resource Languagesadedejimakinde/Zero-Shot-vs.-Fine-Tuned-Approaches-for-Yoruba-Sentiment-AnalysisCross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on RomanianLow Resource Word Sense Disambiguation in Oromo with Fine Tuned Small TransformersSequential Fine-Tuning Language Variation in Zero-Shot Euphemism Detection

Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning

The rapid growth of speech synthesis and voice conversion systems has made deepfake audio a major se

Comparative Performance of Code-Switched Zero-Shot and XNLI-Fine-Tuned Retrieval Models on MIRACL Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

adedejimakinde/Zero-Shot-vs.-Fine-Tuned-Approaches-for-Yoruba-Sentiment-Analysis

Zero-Shot vs. Fine-Tuned Approaches for Yoruba Sentiment Analysis: Examining the Role of Sub-word To

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotate

Low Resource Word Sense Disambiguation in Oromo with Fine Tuned Small Transformers

Sequential Fine-Tuning Language Variation in Zero-Shot Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec