Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Latent Intent Forecasting for Low-Resource: A JEPA Approach to Amharic Conversational Intelligence Using Dialogue Dataset

Domain:

natural language processing

Record type:

datasetmodel
Creator:
KinDurMesMoh
Publisher:
MDP
Host:
Intent forecasting in dialogue models remains a challenge for low-resource languages such as Amharic. Amharic is the official language of Ethiopia. More than 57 million people speak it. Amharic lacks large annotated datasets and high-performance computing, which limits model accuracy and slows progress in conversational intelligence models. We present Joint-Embedding Predictive Architecture (JEPA), a lightweight adaptation of the Joint-Embedding Predictive Architecture to text-based dialogue. JEPA operates entirely in latent space: a frozen multilingual encoder extracts 512-dimensional representations of each dialogue turn, and a compact three-layer Transformer predictor learns to forecast the latent embedding of the next turn without generating text. We introduce the Amharic Dialogue Benchmark (ADB-1K), a curated corpus of 1000 context-response pairs spanning five intent categories, augmented with orphological and noisy variants. Trained with an Exponential-Moving-Average target branch and mean-squared-error loss, JEPA reaches a validation Latent Cosine Similarity of 0.9228 at epoch 10 and achieves 60% intent-probe accuracy on the held-out test set, outperforming a fine-tuned GPT-2 baseline (15%) and a random control (20%) while using only 2.1% of GPT-2's parameter count (2.6M versus 124.4M). Morphological robustness degradation is zero (∆mr = 0.000), confirming tolerance to Amharic inflectional variation.

Visit

doi.org

Languages

Amharic

Licenses

http://creativecommons.org/licenses/by/4.0

Similar

Intent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource LanguagesA latent class approach to understanding patterns of peer victimization in four low-resource settingsLarge Scale Speech Recognition for Low Resource Language Amharic, an End-to-End ApproachMULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented DialogueLow Resource Question Answering: An Amharic Benchmarking DatasetBengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems

Intent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource Languages

Building Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic

A latent class approach to understanding patterns of peer victimization in four low-resource settings

Abstract Background: Peer victimizat

Large Scale Speech Recognition for Low Resource Language Amharic, an End-to-End Approach

Speech recognition, or automatic speech recognition (ASR), is a technology designed to convert spoke

MULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented Dialogue

Task-oriented dialogue (TOD) systems have been widely deployed in many industries as they deliver mo

Low Resource Question Answering: An Amharic Benchmarking Dataset

BengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems

This dataset presents a meticulously curated benchmark collection of 1,200 Bengali voice assistant u