Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Empirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-resource Settings

Domain:

natural language processing

Record type:

paperdataset
Creator:
BoiVilBes
Host:avatar
Since Bahdanau et al. [1] first introduced attention for neural machine translation, most sequence-to-sequence models made use of attention mechanisms [2, 3, 4]. While they produce soft-alignment matrices that could be interpreted as alignment between target and source languages, we lack metrics to quantify their quality, being unclear which approach produces the best alignments. This paper presents an empirical evaluation of 3 main sequence-to-sequence models (CNN, RNN and Transformer-based) for word discovery from unsegmented phoneme sequences. This task consists in aligning word sequences in a source language with phoneme sequences in a target language, inferring from it word segmentation on the target side [5]. Evaluating word segmentation quality can be seen as an extrinsic evaluation of the soft-alignment matrices produced during training. Our experiments in a low-resource scenario on Mboshi and English languages (both aligned to French) show that RNNs surprisingly outperform CNNs and Transformer for this task. Our results are confirmed by an intrinsic evaluation of alignment quality through the use of Average Normalized Entropy (ANE). Lastly, we improve our best word discovery model by using an alignment entropy confidence measure that accumulates ANE over all the occurrences of a given alignment pair in the collection. Interspeech 2019

Visit

arxiv.org

Languages

Mbosi

Tags

Computation and Language

Similar

Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated PseudotranscriptsSequence-to-Sequence Models Can Directly Translate Foreign SpeechAfriTeVa: Extending “Small Data” Pretraining Approaches to Sequence-to-Sequence ModelsMultilingual unsupervised sequence segmentation transfers to extremely low-resource languagesAmharic Word Sequence PredictionR2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging

Overcoming Domain Mismatch in Low Resource Sequence-to-Sequence ASR Models using Hybrid Generated Pseudotranscripts

Sequence-to-sequence (seq2seq) models are competitive with hybrid models for automatic speech recogn

Sequence-to-Sequence Models Can Directly Translate Foreign Speech

We present a recurrent encoder-decoder deep neural network architecture that directly translates spe

AfriTeVa: Extending “Small Data” Pretraining Approaches to Sequence-to-Sequence Models

Pretrained language models represent the state of the art in NLP, but the successful construction of

Multilingual unsupervised sequence segmentation transfers to extremely low-resource languages

We show that unsupervised sequence-segmentation performance can be transferred to extremely low-reso

Amharic Word Sequence Prediction

The significance of computers and handheld devices are not deniable in the modern world of today. Texts are entered to these devices using word processing programs as well as other techniques and word prediction is one of the techniques. Word Prediction is the acti

R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging

We introduce the Rule-to-Tag (R2T) framework, a hybrid approach that integrates a multi-tiered syste