Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards Using Machine Translation Techniques to Induce Multilingual Lexica of Discourse Markers

Domaine:

natural language processing

Type de record:

paper
Créateur:
Lopde CabRib
Hôte:avatar
Discourse markers are universal linguistic events subject to language variation. Although an extensive literature has already reported language specific traits of these events, little has been said on their cross-language behavior and on building an inventory of multilingual lexica of discourse markers. This work describes new methods and approaches for the description, classification, and annotation of discourse markers in the specific domain of the Europarl corpus. The study of discourse markers in the context of translation is crucial due to the idiomatic nature of these structures. Multilingual lexica together with the functional analysis of such structures are useful tools for the hard task of translating discourse markers into possible equivalents from one language to another. Using Daniel Marcu's validated discourse markers for English, extracted from the Brown Corpus, our purpose is to build multilingual lexica of discourse markers for other languages, based on machine translation techniques. The major assumption in this study is that the usage of a discourse marker is independent of the language, i.e., the rhetorical function of a discourse marker in a sentence in one language is equivalent to the rhetorical function of the same discourse marker in another language. 6 pages

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageI.2.7

Similaires

Development of Multilingual Corpora in Medical Domain Using Neural Machine TranslationTowards a Deep Understanding of Multilingual End-to-End Speech TranslationAnalyzing Subword Techniques to Improve English to Sinhala Neural Machine TranslationUsing Interlinear Glosses as Pivot in Low-Resource Multilingual Machine TranslationTextual Augmentation Techniques Applied to Low Resource Machine Translation: Case of SwahiliBackdoor Attack on Multilingual Machine Translation

Development of Multilingual Corpora in Medical Domain Using Neural Machine Translation

Development of Multilingual Corpora in Medical Domain Using Neural Machine Translation

Poster presented at the Deep Learning Indaba 2022 by AJAGBE, Sunday Adeola

Towards a Deep Understanding of Multilingual End-to-End Speech Translation

In this paper, we employ Singular Value Canonical Correlation Analysis (SVCCA) to analyze representa

Analyzing Subword Techniques to Improve English to Sinhala Neural Machine Translation

Neural machine translation (NMT) is a remarkable approach which performs much better than the Statis

Using Interlinear Glosses as Pivot in Low-Resource Multilingual Machine Translation

We demonstrate a new approach to Neural Machine Translation (NMT) for low-resource languages using a

Textual Augmentation Techniques Applied to Low Resource Machine Translation: Case of Swahili

In this work we investigate the impact of applying textual data augmentation tasks to low resource machine translation. There has been recent interest in investigating approaches for training systems for languages with limited resources and one popular approach is

Backdoor Attack on Multilingual Machine Translation

While multilingual machine translation (MNMT) systems hold substantial promise, they also have secur