Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Spoken Portuguese Corpus

Domain:

natural language processing

Record type:

dataset
Publisher:
ELR
Host:avatar
The Spoken Portuguese corpus was collected among sociolinguistically diverse speakers having Portuguese as mother tongue or as second language. In a total of 86 recordings, the texts exemplify the Portuguese spoken in Portugal (30), in Brazil (20), in the African countries with Portuguese as its official language: Angola, Cape Verde, Guinea-Bissau, Mozambique and Sao Tome and Principe (5 each), in Macao (5), in Goa (3) and in East-Timor (3), corresponding to a total of 8h44m of recording. The corpus was recorded in a situation of spontaneous oral communication, on different themes of everyday life, with speakers of different ages and social and professional backgrounds.The recordings cover a period that goes from 1970 to 2001, and approximately 70% of them fall within the nineties. The corpus contains 153,588 tokens.The corpus consists of audio files in .wav format, aligned transcriptions in XML Exmaralda format and transcriptions in plain text. The plain text files also have automatically assigned POS-tag information. The transcriptions of the corpus are also available in html format. The characters have been encoded in UTF-8.

Visit

catalog.elra.info

Licenses

Rights available for: nonCommercialUse, commercialUse

Similar

Portuguese-Changana Parallel CorpusMultilingual Spoken Words CorpusEuclidesDoRosario/Emakhuwa-to-Portuguese-Corpusxhosa Corpus of spoken isiXhosaML Commons: Multilingual Spoken Words Corpusdrineszribi/Spoken-Tunisian-Arabic-Corpus-STAC-

Portuguese-Changana Parallel Corpus

The first publicly available sentence-level parallel corpus for Portuguese and Changana (Xichangana/

Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.

EuclidesDoRosario/Emakhuwa-to-Portuguese-Corpus

Parallel corpus of Emakhuwa to Portuguese for Machine Translation Emakhuwa to Portuguese Corpus ---

xhosa Corpus of spoken isiXhosa

The Corpus of Spoken isiXhosa The Corpus of Spoken isiXhosa consists of transcribed and annotated

ML Commons: Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.0. The dataset contains more than 340,000 keyw

drineszribi/Spoken-Tunisian-Arabic-Corpus-STAC-

Spoken Tunisian Arabic Corpus (STAC)