A copy of the deduplicated the stack Markdown files, annotated with fastText to include language labels (lid column) and probabilities (lid_prob). This was done using the NLLB fastText model facebook/fasttext-language-identification.
The following languages were detected:
abk_Cyrl
ace_Arab
ace_Latn
ady_Cyrl
afr_Latn
aka_Latn
als_Latn
amh_Ethi
arb_Arab
arb_Latn
arn_Latn
asm_Beng
ast_Latn
ayr_Latn
azb_Arab
azj_Latn
bak_Cyrl
bam_Latn
ban_Latn
bel_Cyrl
bem_Latn
ben_Beng
bho_Deva
bis_Latn