Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A comparison of character neural language model and bootstrapping for language identification in multilingual noisy texts

Domaine:

natural language processing

Type de record:

paper
Créateur:
AdoDobBerSem
Éditeur:
DepDépAss
Éditeur:
CCSDAss
Hôte:avatar
International audience This paper seeks to examine the effect of including background knowledge in the form of character pre-trained neural language model (LM), and data bootstrapping to overcome the problem of unbalanced limited resources. As a test, we explore the task of language identification in mixed-language short non-edited texts with an under-resourced language, namely the case of Algerian Arabic for which both labelled and unlabelled data are limited. We compare the performance of two traditional machine learning methods and a deep neural networks (DNNs) model. The results show that overall DNNs perform better on labelled data for the majority categories and struggle with the minority ones. While the effect of the untokenised and unlabelled data encoded as LM differs for each category, bootstrapping, however, improves the performance of all systems and all categories. These methods are language independent and could be generalized to other under-resourced languages for which a small labelled data and a larger unlabeled data are available.

Visit

cea.hal.science

Tasks

language identification

Languages

Arabic, Algerian Spoken

Tags

Natural Language ProcessingDeep Neural NetworksLanguage IdentificationCharacter Neural Language ModelMultilingual Noisy Texts[INFO]Computer Science [cs][INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]

Licenses

http://hal.archives-ouvertes.fr/licences/publicDomain/info:eu-repo/semantics/OpenAccess

Similaires

Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?Phone Clustering Methods for Multilingual Language IdentificationAfroLID: A Neural Language Identification Tool for African LanguagesAfroLID: A Neural Language Identification Tool for African LanguagesCHaymaK/Language-Identification-Sentiment-Analysis-for-Arabic-and-Tunisian-Texts

Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high- resource languages. Building language mod- els and, more generally, NLP systems for non- standard

Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural la

Phone Clustering Methods for Multilingual Language Identification

This paper proposes phoneme clustering methods for multilingual language identification (LID) on a m

AfroLID: A Neural Language Identification Tool for African Languages

AfroLID is a powerful neural toolkit for African languages identification which covers 517 African languages.

AfroLID: A Neural Language Identification Tool for African Languages

Language identification (LID) is a crucial precursor for NLP, especially for mining web data. Problematically, most of the world's 7000+ languages today are not covered by LID technologies. We address this pressing issue for Africa by introducing AfroLID, a neural

CHaymaK/Language-Identification-Sentiment-Analysis-for-Arabic-and-Tunisian-Texts

This repository contains two example of text processing tasks using NPL. # Language-Identification-