Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The African Varieties of Portuguese: Compiling Comparable Corpora and Analyzing Data-Derived Lexicon

Domaine:

natural language processing

Type de record:

paper
"Linguistic Resources for the Study of the Portuguese African Varieties" is an ongoing project that aims at the constitution, treatment, analysis and availability of a corpus of the African varieties of Portuguese, with 3 million words of written and spoken texts, constituted by five comparable subcorpora, corresponding to the varieties of Angola, Cape Verde, Guinea-Bissau, Mozambique and Sao Tome and Principe. This material will allow intra and intercorpora comparative studies, which will make visible variations that result from discursive and pragmatic differences of each corpus and aspects of linguistic unity or diversity that characterise the spoken Portuguese of this referred five African countries. The five corpora are comparable in size (600,000 words each), in chronology (the last 30 years) and in types and genres (24,000 spoken words and c. 580,000 written words, the last belonging to newspapers, literature and varia). The corpus is automatically annotated and after the extraction of alphabetical lists of lexical forms, these data will be automatically lemmatised. Five separated lists of vocabulary for each variety will be established. A tool for word extraction and preferential calculus according to predefined indexes in order to achieve lexicon comparison of the African Portuguese Varieties is being developed. Concordances extraction will be also performed.

Visit

aclanthology.orgwww.lrec-conf.org

Tags

acl

Similaires

Editorial. Possession and Location in African Varieties of Portuguese Disambiguating vectors for bilingual lexicon extraction from comparable corpora Razdvoumljanje vektorjev za izboljšanje luščenja dvojezičnih leksikonov iz primerljivih korpusovPortuguese Phonetic Lexicon DatasetAnalysis and evaluation of comparable corpora for under resourced areas of machine translationThe agreement <i>continuum</i> in urban samples of African, Brazilian and European varieties of PortugueseGrouping conversational markers across languages by exploiting large comparable corpora and unsupervised segmentation

Editorial. Possession and Location in African Varieties of Portuguese 

Early versions of the four papers in this special collection were presented at the first workshop of

Disambiguating vectors for bilingual lexicon extraction from comparable corpora Razdvoumljanje vektorjev za izboljšanje luščenja dvojezičnih leksikonov iz primerljivih korpusov

International audience This paper presents an approach to enhance the extraction of t

Portuguese Phonetic Lexicon Dataset

This dataset contains phonetic and morphological information for Portuguese words, collected from th

Analysis and evaluation of comparable corpora for under resourced areas of machine translation

The agreement <i>continuum</i> in urban samples of African, Brazilian and European varieties of Portuguese

Abstract Our aim is to contrast number expression in nominal phrases (NPs) and in verbal phrases (

Grouping conversational markers across languages by exploiting large comparable corpora and unsupervised segmentation

International audience This work approaches Conversational and Discourse Markers (her