Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles

Domain:

natural language processing

Record type:

paper
We present a new major release of the OpenSubtitles collection of parallel corpora. The release is compiled from a large database of movie and TV subtitles and includes a total of 1689 bitexts spanning 2.6 billion sentences across 60 languages. The release also incorporates a number of enhancements in the preprocessing and alignment of the subtitles, such as the automatic correction of OCR errors and the use of meta-data to estimate the quality of each subtitle and score subtitle pairs.

Visit

aclanthology.orgwww.duo.uio.no

Connected records

dataset

Tasks

machine translation

Languages

Afrikaans

Licenses

Similar

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal SupervisionExtracting Signs from Weakly Aligned Sign Language Corpora: A Study on LSF and LSMAutshumato English-Xitsonga Parallel CorporaAutshumato English-Setswana Parallel CorporaAutshumato English-Afrikaans Parallel CorporaAutshumato English-Sepedi Parallel Corpora

Extracting Parallel Sentences from Low-Resource Language Pairs with Minimal Supervision

Abstract At present, machine translation in the market depends on parallel sentence

Extracting Signs from Weakly Aligned Sign Language Corpora: A Study on LSF and LSM

International audience This paper presents a framework for the automatic annotation o

Autshumato English-Xitsonga Parallel Corpora

Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with e

Autshumato English-Setswana Parallel Corpora

Aligned English-Setswana parallel corpus. This set contains data that was translated by professional

Autshumato English-Afrikaans Parallel Corpora

Aligned parallel corpora for the language pair English-Afrikaans. The data is given as two separate

Autshumato English-Sepedi Parallel Corpora

Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF