Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-lingual Dataless Classification for Languages with Small Wikipedia Presence

Domaine:

natural language processing

Type de record:

paper
Créateur:
SonMayRot
Hôte:avatar
This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingual Dataless Document Classification (CLDDC) relies on mapping the English labels or short category description into a Wikipedia-based semantic representation, and on the use of the target language Wikipedia. Consequently, performance could suffer when Wikipedia in the target language is small. In this paper, we focus on languages with small Wikipedias, (Small-Wikipedia languages, SWLs). We use a word-level dictionary to convert documents in a SWL to a large-Wikipedia language (LWLs), and then perform CLDDC based on the LWL's Wikipedia. This approach can be applied to thousands of languages, which can be contrasted with machine translation, which is a supervision heavy approach and can be done for about 100 languages. We also develop a ranking algorithm that makes use of language similarity metrics to automatically select a good LWL, and show that this significantly improves classification of SWLs' documents, performing comparably to the best bridge possible.

Visit

arxiv.org

Tasks

text classification

Tags

Computation and Language

Similaires

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource LanguagesUniversal Cross-Lingual Text ClassificationEnhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word AlignmentCross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource LanguagesAdversarial Deep Averaging Networks for Cross-Lingual Sentiment ClassificationZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verificatio

Universal Cross-Lingual Text Classification

Text classification, an integral task in natural language processing, involves the automatic categor

Enhancing Cross-lingual Sentence Embedding for Low-resource Languages with Word Alignment

The field of cross-lingual sentence embeddings has recently experienced significant advancements, bu

Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource Languages

In cross-lingual dependency annotation projection, information is often lost during transfer because

Adversarial Deep Averaging Networks for Cross-Lingual Sentiment Classification

In recent years great success has been achieved in sentiment classification for English, thanks in p

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank