Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BabelBot at AraFinNLP2024: Fine-tuning T5 for Multi-dialect Intent Detection with Synthetic Data and Model Ensembling

Domain:

natural language processing

Record type:

paper
Creator:
AssFarTou
Publisher:
Und
Host:avatar
This paper presents our results for the Arabic Financial NLP (AraFinNLP) shared task at the Second Arabic Natural Language Processing Conference (ArabicNLP 2024). We participated in the first sub-task, Multi-dialect Intent Detection, which focused on cross-dialect intent detection in the banking domain. Our approach involved fine-tuning an encoder-only T5 model, generating synthetic data, and model ensembling. Additionally, we conducted an in-depth analysis of the dataset, addressing annotation errors and problematic translations. Our model was ranked third in the shared task, achieving a F1-score of 0.871.

Visit

doi.orgunderline.io

Tasks

text classification

Tags

Computational LinguisticsNatural Language ProcessingLinguisticsFOS: Languages and literature

Similar

Fine-Tuning Transformers and LLMs for Fake News Detection in Algerian DialectBrain Tumor Segmentation in Sub-Sahara Africa with Advanced Transformer and ConvNet Methods: Fine-Tuning, Data Mixing and EnsemblingSequential Fine-Tuning with Typologically Similar Languages for Yoruba Euphemism DetectionMultilingual Multi-Label Emotion Classification at Scale with Synthetic Data**Scaling Model Size and Fine-Tuning Strategies in Cross-Lingual Euphemism Detection**Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning

Fine-Tuning Transformers and LLMs for Fake News Detection in Algerian Dialect

Brain Tumor Segmentation in Sub-Sahara Africa with Advanced Transformer and ConvNet Methods: Fine-Tuning, Data Mixing and Ensembling

Brain tumors are among the deadliest cancers worldwide, with particularly devastating impact in Sub-

Sequential Fine-Tuning with Typologically Similar Languages for Yoruba Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data

Emotion classification in multilingual settings remains constrained by the scarcity of annotated dat

**Scaling Model Size and Fine-Tuning Strategies in Cross-Lingual Euphemism Detection**

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning

Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic