Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Latent Intent Forecasting for Low-Resource: A JEPA Approach to Amharic Conversational Intelligence Using Dialogue Dataset

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
KinDurMesMoh
Éditeur:
MDP
Hôte:
Intent forecasting in dialogue models remains a challenge for low-resource languages such as Amharic. Amharic is the official language of Ethiopia. More than 57 million people speak it. Amharic lacks large annotated datasets and high-performance computing, which limits model accuracy and slows progress in conversational intelligence models. We present Joint-Embedding Predictive Architecture (JEPA), a lightweight adaptation of the Joint-Embedding Predictive Architecture to text-based dialogue. JEPA operates entirely in latent space: a frozen multilingual encoder extracts 512-dimensional representations of each dialogue turn, and a compact three-layer Transformer predictor learns to forecast the latent embedding of the next turn without generating text. We introduce the Amharic Dialogue Benchmark (ADB-1K), a curated corpus of 1000 context-response pairs spanning five intent categories, augmented with orphological and noisy variants. Trained with an Exponential-Moving-Average target branch and mean-squared-error loss, JEPA reaches a validation Latent Cosine Similarity of 0.9228 at epoch 10 and achieves 60% intent-probe accuracy on the held-out test set, outperforming a fine-tuned GPT-2 baseline (15%) and a random control (20%) while using only 2.1% of GPT-2's parameter count (2.6M versus 124.4M). Morphological robustness degradation is zero (∆mr = 0.000), confirming tolerance to Amharic inflectional variation.

Visit

doi.org

Languages

Amharic

Licenses

http://creativecommons.org/licenses/by/4.0

Similaires

Intent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource LanguagesA latent class approach to understanding patterns of peer victimization in four low-resource settingsLarge Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachMULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented DialogueLow Resource Question Answering: An Amharic Benchmarking DatasetBengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems

Intent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource Languages

Building Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic

A latent class approach to understanding patterns of peer victimization in four low-resource settings

Abstract Background: Peer victimizat

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

MULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented Dialogue

Task-oriented dialogue (TOD) systems have been widely deployed in many industries as they deliver mo

Low Resource Question Answering: An Amharic Benchmarking Dataset

BengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems

This dataset presents a meticulously curated benchmark collection of 1,200 Bengali voice assistant u