Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AlokChedambath64/EfficientCPT-of-LLMs-for-LRLs

Domaine:

natural language processing

Type de record:

software
Créateur:
Alo
Hôte:
Implementation of the two core algorithms in the paper "Efficient Continual Pre-training of LLMs for Low-resource Languages. # EfficientCPT-of-LLMs-for-LRLs Implementation of the two core algorithms in the paper "Efficient Continual Pre-training of LLMs for Low-resource Languages. 1. Vocabulary.ipynb: Algorithm to choose the best tokens to extend your LLMs vocabulary with 2. Corpus.ipynb: Algorithm to choose the best sentence to pick from your corpus Adding tokens to your vocabulary/Adding sentences to your training data increase compute. These algorithms help you choose the most efficient additions of both. # Paper: arxiv.org # Corpus Selection Algorithm # Vocabulary Selection Algorithm

Visit

github.com

Tasks

language modeling