Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Adapting BERT and AgriBERT for agroecology: A small-corpus pretraining approach

Domaine:

natural language processingagriculture

Type de record:

papermodel
Créateur:
MecAuzJonRoc
Éditeur:
AgrTerMatWEB
Éditeur:
CCSDSpringer
Hôte:avatar
Source Agritrop Cirad (agritrop.cirad.fr) * Autres projets (id;sigle;titre): 101081973;IntercropValuES;(EU) Developing Intercropping for agrifood Value chains and Ecosystem Services delivery in Europe and Southern countries// International audience Source variables, or observable properties, used to describe agroecological experiments are often heterogeneous, non-standardized, and multilingual, making them challenging to understand, explain, and utilize in cropping system modeling and multicriteria evaluations of agroecological system performance. A potential solution is data annotation via a controlled vocabulary, known as candidate variables, from the Agroecological Global Information System (AEGIS). However, matching source and candidate variables via their textual descriptions remains a challenging task in agroecology. Domain-general language models, such as BERT, often struggle with domain-specific tasks due to their general-purpose training data. In the literature, these models are adapted to specialized domains through further pretraining, pretraining from scratch, and/or fine-tuning on downstream tasks. However, pretraining a domain-general model on a domain-specific corpus is resource-intensive, requiring substantial time, energy, and computational resources. To the best of our knowledge, no study has further pretrained a domain-general model on a small corpus (less than 100 MB) to adapt it to a domain-specific task and evaluated it on downstream tasks without fine-tuning. To address these shortcomings, this paper proposes further pretraining BERT and AgriBERT on a small agroecology-related corpus. This approach is designed to be both time- and resource-efficient while enhancing domain adaptation. We evaluate the pretrained models on the task of matching source and candidate variable descriptions without fine-tuning. Our results show that our further pretrained AgriBERT (+ Experts + Core) model outperforms all others by more than 8% from P@1 to P@10. These findings showed that small-scale pretraining can significantly improve performance on domain-specific tasks without requiring fine-tuning.

Visit

hal.science

Tags

Word embeddingsSemantic matchingMasked Language Modeling (MLM)Pretrained Language Models (PLMs)Observable properties[SDV]Life Sciences [q-bio]

Licenses

https://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/OpenAccess

Similaires

Parsing with Multilingual BERT, a Small Corpus, and a Small TreebankTunBERT: Pretraining BERT for Tunisian Dialect UnderstandingMoroccan Darija Pretraining Corpus for NanochatAmharic Pretraining CorpusLyte/darija-pretraining-corpusRethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

Parsing with Multilingual BERT, a Small Corpus, and a Small Treebank

Pretrained multilingual contextual representations have shown great success, but due to the limits o

TunBERT: Pretraining BERT for Tunisian Dialect Understanding

Moroccan Darija Pretraining Corpus for Nanochat

Parquet shards prepared for nanochat pretraining. 83 train shards 1 validation shard 82,738,910 tra

Amharic Pretraining Corpus

Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining

Lyte/darija-pretraining-corpus

Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources

Large Language Models (LLMs) exhibit significant disparities in performance across languages, primar