Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ZeroBERTo: Leveraging Zero-Shot Text Classification by Topic Modeling

Domain:

natural language processing

Record type:

papersoftware
Creator:
AlcPalGerBus
Editor:
EscTélUniIns
Publisher:
CCSDSpringer International Publishing
Host:avatar
International audience Traditional text classification approaches often require a good amount of labeled data, which is difficult to obtain, especially in restricted domains or less widespread languages. This lack of labeled data has led to the rise of low-resource methods, that assume low data availability in natural language processing. Among them, zero-shot learning stands out, which consists of learning a classifier without any previously labeled data. The best results reported with this approach use language models such as Transformers, but fall into two problems: high execution time and inability to handle long texts as input. This paper proposes a new model, ZeroBERTo, which leverages an unsupervised clustering step to obtain a compressed data representation before the classification task. We show that ZeroBERTo has better performance for long inputs and shorter execution time, outperforming XLM-R by about 12% in the F1 score in the FolhaUOL dataset.

Visit

telecom-paris.hal.science

Tasks

text classification

Tags

TransformersTopic modelingZero-shot learningUnlabeled dataLow-resource NLPTransformers[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][INFO.INFO-AI]Computer Science [cs]/Artificial Intelligence [cs.AI]

Licenses

info:eu-repo/semantics/OpenAccess