Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Stylistic Transfer from Annotator Communities to Large Language Models

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssCho
Éditeur:
Und
Hôte:avatar
Large language models (LLMs) are post-trained on human feedback collected from annotator communities, yet the linguistic influence of these annotator communities on language models remains poorly understood. We investigated the stylistic transfer from Nigerian annotators to the LLaMA family of models through a natural experiment with LLaMA 2 and LLaMA 3.1, as their release dates are separated by the shutdown of a major data annotation service provider in Nigeria. We generated corpora from both model families and measured linguistic style by computing the difference-in-difference of the Jensen-Shannon distance on the bigram distribution between model outputs and corpora of Nigerian English and US English. We found that, although both pre-trained model variants exhibit similar proximity to both English variants, the LLaMA 2 post-trained model moved toward Nigerian English, while the LLaMA 3.1 post-trained model moved away from Nigerian English. Qualitatively, we found that post-trained LLaMA 2 models used significantly fewer contractions, in line with Nigerian English speakers opting to use a formal register due to its role as an index of knowledgeability. Our findings suggest that annotator communities can imprint linguistic style on large language models, with potential implications such as a disproportionately higher false positive rate in AI plagiarism detection for users who share a linguistic style with annotator communities.

Visit

doi.orgunderline.io

Tasks

language modeling

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similaires

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferCode-Switching In-Context Learning for Cross-Lingual Transfer of Large Language ModelsFrom Facts to Folklore: Evaluating Large Language Models on Bengali Cultural KnowledgeFew-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource LanguagesMultilinguality of Large Language Models From a Structural PerspectiveLeft Behind: Cross-Lingual Transfer as a Bridge for Low-Resource Languages in Large Language Models

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer

Despite remarkable advancements in few-shot generalization in natural language processing, most mode

Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models

While large language models (LLMs) exhibit strong multilingual abilities, their reliance on English

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (L

Few-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource Languages

Large pre-trained language models (PLMs) are at the forefront of advances in Natural Language Proces

Multilinguality of Large Language Models From a Structural Perspective

Large language models (LLMs) have excelled in processing multiple languages through pre- and post-tr

Left Behind: Cross-Lingual Transfer as a Bridge for Low-Resource Languages in Large Language Models

We investigate how large language models perform on low-resource languages by benchmarking eight LLM