Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Examining Accuracy Heterogeneities in Classification of Multilingual

Domain:

natural language processing

Record type:

paper
Creator:
Rag
Publisher:
Aca
Host:
Tools for detection of AI-generated texts are used globally, however, the nature of the apparent accuracy disparities between languages must be further observed. This paper aims to examine the nature of these differences through testing OpenAI’s “AI Text Classifier” on a set of various AI and human-generated texts in English, Swahili, German, Arabic, Chinese, and Hindi. Current tools for detecting AI-generated text are already fairly easy to discredit, as misclassifications have shown to be fairly common, but such vulnerabilities often persist in slightly different ways when non-English languages are observed: classification of human-written text as AI-generated and vice versa could occur more frequently in specific language environments than others. Our findings indicate that false positives are more likely to occur in Hindi and Arabic, whereas false negative labelings are more likely to occur in English. Other languages tested had a tendency to not be confidently labeled at all.

Visit

doi.org

Tasks

text classification

Languages

Swahili

Similar

Timera-ctrl/multilingual-news-classificationAssessing the Accuracy of Different Supervised Classification Methods of Satellite ImageAutomatic classification of multilingual occupations using transfer learningSentiment Classification in Swahili Language Using Multilingual BERTSeasonal variation of land cover classification accuracy of Landsat 8 images in Burkina FasoSpatial Heterogeneities of Warming Impacts on Corn Yields in Ghana

Timera-ctrl/multilingual-news-classification

English and isiXhosa news classification using PyTorch, TF-IDF, and multinomial logistic regression.

Assessing the Accuracy of Different Supervised Classification Methods of Satellite Image

Assessing the accuracy of the classification map is an essential area in remote sensing digital imag

Automatic classification of multilingual occupations using transfer learning

Socio-economic research often requires detailed information on individual occupations in order to st

Sentiment Classification in Swahili Language Using Multilingual BERT

The evolution of the Internet has increased the amount of information that is expressed by people on

Seasonal variation of land cover classification accuracy of Landsat 8 images in Burkina Faso

Abstract. In the seasonal tropics, vegetation shows large reflectance variation because of phenology

Spatial Heterogeneities of Warming Impacts on Corn Yields in Ghana

In this study, we utilize a panel of subnational district-level yields for corn matched to weather d