Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

How Far Can Multilingual Text Embeddings Be Trained From Scratch? A Compute-Efficient Study of Arabic, English, and Urdu

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
ShaMenSanHam
Éditeur:
Zenodo
Hôte:avatar

We investigate how much useful multilingual embedding capability can be obtained when the student encoder starts from random initialization, under a strict single-GPU compute budget. We present the mentee-embed series — three controlled experiments (v1–v3) using a 41M-parameter Transformer encoder with a custom 50,000-entry BPE tokenizer, trained entirely from scratch on Arabic, English, and Urdu. The student is randomly initialized; a frozen pretrained teacher (multilingual-e5-base) supplies relational distillation signal only during Stage B but transfers no weights in to the student.

Our three-version controlled progression isolates the dominant training factors: batch size and retrieval training data domain matter more than model size in this regime. The v1→v2→v3 ablation demonstrates this strikingly: scaling from 41M to 125M parameters while reducing batch size (v2) degrades Protocol A avg MRR@10 from 0.585 to 0.429, while returning to 41M with batch size 512 and adding MS-MARCO retrieval data (v3) recovers to 0.655 — a 3× Protocol C MS-MARCO improvement (0.215→0.645) and STS-B Spearman ρ = 0.683.

 Model weights: huggingface.co
 Code: github.com

Visit

doi.org

Tasks

embeddings

Languages

Ndasa

Tags

multilingual embeddingsknowledge distillationArabic NLPUrdu NLPsentence embeddingsinformation retrievalfrom scratch training

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

A Comparative Study of From-Scratch Squeezeformer Training for Arabic and English Speech RecognitionHow facilitatory can lexical information be during word recognition? Evidence from Moroccan ArabicMultilingual Jointly Trained Acoustic and Written Word EmbeddingsML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldBULaMU-Dream: The First Text-to-Image Model Trained from Scratch for an African LanguageF2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World

A Comparative Study of From-Scratch Squeezeformer Training for Arabic and English Speech Recognition

While modern speech recognition systems achieve impressive performance through large-scale pre-train

How facilitatory can lexical information be during word recognition? Evidence from Moroccan Arabic

Multilingual Jointly Trained Acoustic and Written Word Embeddings

Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be lear

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

The development of high-quality text embeddings is increasingly drifting toward an exclusionary futu

BULaMU-Dream: The First Text-to-Image Model Trained from Scratch for an African Language

This paper introduces BULaMU Dream, the first text-to-image, conditional diff

F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World

We present F2LLM-v2, a new family of general-purpose, multilingual embedding models in 8 distinct si