Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
GaiRanZamHom
Hôte:avatar
The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable exceptions, most research on automatic offensive language identification has dealt with English. To address this shortcoming, we introduce MOLD, the Marathi Offensive Language Dataset. MOLD is the first dataset of its kind compiled for Marathi, thus opening a new domain for research in low-resource Indo-Aryan languages. We present results from several machine learning experiments on this dataset, including zero-short and other transfer learning experiments on state-of-the-art cross-lingual transformers from existing data in Bengali, English, and Hindi. Accepted to RANLP 2021

Visit

arxiv.org

Tasks

hate speech detectiontext classificationtransfer learning

Tags

Computation and LanguageArtificial IntelligenceMachine LearningNeural and Evolutionary ComputingSocial and Information Networks

Similaires

Isomorphic Cross-lingual Embeddings for Low-Resource LanguagesCross-lingual metaphor detection for low-resource languagesGlotLID: Language Identification for Low-Resource LanguagesCross-Lingual Retrieval Augmented Prompt for Low-Resource LanguagesZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource LanguagesThe Illusion of Cross-Lingual Safety in Low-Resource Languages

Isomorphic Cross-lingual Embeddings for Low-Resource Languages

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt

Cross-lingual metaphor detection for low-resource languages

State-of-the-art metaphor detection (MD) models achieve human-like performance for English data, whi

GlotLID: Language Identification for Low-Resource Languages

International audience Several recent papers have published good solutions for langua

Cross-Lingual Retrieval Augmented Prompt for Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent em

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank

The Illusion of Cross-Lingual Safety in Low-Resource Languages

Safety alignment in large language models (LLMs) is largely developed in English, assuming these saf