Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Counter Turing Test ($CT^2$): Investigating AI-Generated Text Detection for Hindi -- Ranking LLMs based on Hindi AI Detectability Index ($ADI_{hi}$)

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
KavRanChaKum
Hôte:avatar
The widespread adoption of Large Language Models (LLMs) and awareness around multilingual LLMs have raised concerns regarding the potential risks and repercussions linked to the misapplication of AI-generated text, necessitating increased vigilance. While these models are primarily trained for English, their extensive training on vast datasets covering almost the entire web, equips them with capabilities to perform well in numerous other languages. AI-Generated Text Detection (AGTD) has emerged as a topic that has already received immediate attention in research, with some initial methods having been proposed, soon followed by the emergence of techniques to bypass detection. In this paper, we report our investigation on AGTD for an indic language Hindi. Our major contributions are in four folds: i) examined 26 LLMs to evaluate their proficiency in generating Hindi text, ii) introducing the AI-generated news article in Hindi ($AG_{hi}$) dataset, iii) evaluated the effectiveness of five recently proposed AGTD techniques: ConDA, J-Guard, RADAR, RAIDAR and Intrinsic Dimension Estimation for detecting AI-generated Hindi text, iv) proposed Hindi AI Detectability Index ($ADI_{hi}$) which shows a spectrum to understand the evolving landscape of eloquence of AI-generated text in Hindi. The code and dataset is available at github.com Accepted at EMNLP 2024 Findings

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similaires

Hindi Text Summarization DataNigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text DetectionBidirectional Machine Translation for Punjabi-English, Punjabi-Hindi, and Hindi-English Language PairsA Corpus of English-Hindi Code-Mixed Tweets for Sarcasm DetectionHypernymy Detection for Low-resource Languages: A Study for Hindi, Bengali, and AmharicIdentifying Text Classification Failures in Multilingual AI-Generated Content

Hindi Text Summarization Data

The data is crawled from Hindi News Media for the purpose of extractive text summarization.

Nigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text Detection

A curated corpus of pre-AI era (2005–2022) Nigerian academic writing paired with AI-generated equiva

Bidirectional Machine Translation for Punjabi-English, Punjabi-Hindi, and Hindi-English Language Pairs

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Social media platforms like twitter and facebook have be- come two of the largest mediums used by pe

Hypernymy Detection for Low-resource Languages: A Study for Hindi, Bengali, and Amharic

Numerous attempts for hypernymy relation (e.g., dog “is-a” animal) detection have been made for reso

Identifying Text Classification Failures in Multilingual AI-Generated Content

With the rising popularity of generative AI tools, the nature of apparent classification failures by