Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
AssOuyVosWan
Éditeur:
Und
Hôte:avatar
Indigenous languages remain largely invisible in commercial language identification systems—a stark reality exemplified by Google Translate's LangID tool, which supports 109 languages but excludes all 150 Indigenous languages of North America. This technological marginalization is particularly acute for Alaska's 20 Native languages, all of which face endangerment despite their rich linguistic heritage. We present GenAlaskan, a framework demonstrating how both large language models and specialized classifiers can effectively identify these languages with minimal data. Working closely with Native Alaskan community members, we created Akutaq-2k, a carefully curated dataset of 2000 sentences spanning all 20 languages—named after the traditional Yup'ik dessert symbolizing the blending of diverse elements. We design few-shot learning on proprietary and open-source LLMs, achieving nearly 100% accuracy with just 40 examples per language. While initial zero-shot attempts showed limited success, our systematic attention head pruning revealed critical architectural components for accurate language differentiation, providing insights into model decision-making for low-resource languages. Our results challenge the notion that effective Indigenous language identification requires massive resources or corporate infrastructure, demonstrating that targeted technological interventions can drive meaningful progress in preserving endangered languages in the digital age.

Visit

doi.orgunderline.io

Tasks

language identification

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

A Portuguese Native Language Identification DatasetCross-language identification of non-native lexical toneEvaluation Mirage: A Layered Evaluation of Large Language Models and Language Identification for African NLPMELAI-1/polarization-identification-nlpOP‐Synthetic: identification of optimal genetic manipulations for the overproduction of native and non‐native metabolitesPreprints as a Voice for Africa: Amplifying Visibility and Inclusion

A Portuguese Native Language Identification Dataset

In this paper we present NLI-PT, the first Portuguese dataset compiled for Native Language Identific

Cross-language identification of non-native lexical tone

We extend to lexical-tone systems a model of second-language perception, the Perceptual Assimilation

Evaluation Mirage: A Layered Evaluation of Large Language Models and Language Identification for African NLP

David Ifeoluwa Adelani (Supervisor) As Large Language Models (LLMs) are increasingly deployed in glo

MELAI-1/polarization-identification-nlp

NLP project for multi-task classification of polarization in English and Swahili social media text.

OP‐Synthetic: identification of optimal genetic manipulations for the overproduction of native and non‐native metabolites

Constraint‐based flux analysis has been widely used in metabolic engineering to predict genetic opti

Preprints as a Voice for Africa: Amplifying Visibility and Inclusion

This presentation was by Dr Lamis Yahia Mohamed Elkheir, Director of Training and Resource Developme