Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Massively multilingual morpho-syntactic analysis using typological resources, universal annotations, and multilingual word embeddings Analyse morpho-syntaxique massivement multilingue à l'aide de ressources typologiques, d'annotations universelles et de plongements de mots multilingues

Domain:

natural language processing

Record type:

paper
Creator:
Sch
Editor:
TraAixAleCar
Publisher:
CCSD
Host:avatar
Data annotation is a major problem in all machine learning tasks. In the field of NLP, this problem is multiplied by the number of existing languages.Many languages do not have any annotations, and are therefore excluded from NLP systems. One possible solution to integrate these languages into the systems is to try to leverage the languages having many annotations, and to try to learn information about these resource-rich languages, and to transfer this knowledge to the low-resources languages.It is possible to rely on initiatives such as Universal Dependencies, which propose a universal annotation scheme between languages. The use of multilingual word embeddings and typological features from resources such as the WALS are solutions allowing knowledge sharing between languages.These tracks are studied in the framework of this thesis, through the prediction of parsing, morphology and parts of speech on 41 languages in total. We show that the impact of the WALS can be positive in a multilingual setting, but that its usefulness is not systematic in a zero-shot learning setting. Other language representations can be learned from the data, and perform better than the WALS, but have the downside of not working in a zero-shot setting. We also highlight the importance of the presence of a nearby language when learning patterns, as well as the problems associated with using a character pattern for isolated languages. L'annotation de données est un problème majeur dans toutes les tâches d'apprentissage automatique. Dans le domaine du TAL, ce problème est multiplié par le nombre de langues existantes.De nombreuses langues se retrouvent sans annotations, et sont alors mises à l'écart des systèmes de TAL. Une solution possible pour intégrer ces langues dans les systèmes est de tenter d'exploiter les langues disposant de nombreuses annotations, d'apprendre des informations sur ces langues bien dotées, et de transférer ce savoir vers les langues peu dotées. Pour cela, il est possible de se reposer sur des initiatives comme les Universal Dependencies, qui proposent un schéma d'annotation universel entre les langues. L'utilisation de plongements de mots multilingues et de traits typologiques issus de ressources comme le WALS sont des solutions permettant un partage de connaissances entre les langues.Ces pistes sont étudiées dans le cadre de cette thèse, à travers la prédiction de l'analyse syntaxique, de la morphologie et des parties du discours sur 41 langues au total. Nous montrons que l'impact du WALS peut être positif dans un cadre multilingue, mais que son utilité n'est pas systématique dans une configuration d'apprentissage zero-shot. D'autres représentations des langues peuvent être apprises sur les données, et donnent de meilleurs résultats que le WALS, mais ont l'inconvénient de ne pas fonctionner dans un cadre de zero-shot. Nous mettons également en évidence l'importance de la présence d'une langue proche lors de l'apprentissage des modèles, ainsi que les problèmes liés à l'utilisation d'un modèle de caractère pour les langues isolées.

Visit

hal.science

Languages

Tal

Tags

zero-shottypological featuresparsingtaggingmultilingualtraits typologiqueszero-shotanalyse syntaxiqueétiquetage morpho-syntaxiquemultilingue+3

Licenses

info:eu-repo/semantics/OpenAccess

Similar

Massively Multilingual Word EmbeddingsDescription morpho-syntaxique de la langue tikarTraitement automatique du dialecte tunisien à l'aide d'outils et de ressources de l'arabe standard : application à l'étiquetage morphosyntaxiqueMorpho-Syntactic Analysis Framework for Tone Language Text-to-Speech Systems إطار التحليل الصرفي النحوي لأنظمة تحويل النص إلى كلام بلغة النغمة Cadre d'analyse morpho-syntaxique pour les systèmes de synthèse vocale en langage tonal Marco de análisis morfo-sintáctico para sistemas de lenguaje tonal de texto a vozLearning Multilingual Word Embeddings Using Image-Text DataMorpho-syntactic Analysis of Negation in Moroccan Arabic.pdf

Massively Multilingual Word Embeddings

We introduce new methods for estimating and evaluating embeddings of words in more than fifty langua

Description morpho-syntaxique de la langue tikar

Traitement automatique du dialecte tunisien à l'aide d'outils et de ressources de l'arabe standard : application à l'étiquetage morphosyntaxique

Le développement d’outils de traitement automatique pour les dialectes de l’arabe se heurte à l’abse

Morpho-Syntactic Analysis Framework for Tone Language Text-to-Speech Systems إطار التحليل الصرفي النحوي لأنظمة تحويل النص إلى كلام بلغة النغمة Cadre d'analyse morpho-syntaxique pour les systèmes de synthèse vocale en langage tonal Marco de análisis morfo-sintáctico para sistemas de lenguaje tonal de texto a voz

This paper presents a morpho-syntactic analysis framework using the data-driven methodology.The prop

Learning Multilingual Word Embeddings Using Image-Text Data

There has been significant interest recently in learning multilingual word embeddings -- in which se

Morpho-syntactic Analysis of Negation in Moroccan Arabic.pdf

This paper has a two-fold goal. First, it attempts to argue that the circumfixal nature of the negat