Logo Lanfrica

The Xi’an Multi-Language Learner Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
ZhaZhang, LingDanFen
Éditeur:
Lin
Hôte:avatar

Introduction

The Xi’an Multi-Language Learner Corpus was developed by Xi'an International Studies…. It is comprised of 526 argumentative essays in 15 languages by Chinese L1 university students studying second languages, along with student metadata and writing prompts. It was developed to support second language learner research and to provide a database for cross-linguistic comparison of second languages.

Data

The essays were produced by undergraduate students at XISU and Yunnan Minzu University (YM… in response to writing prompts prepared by the corpus development team. Data was collected in 2023 and 2024. Participating students were linguistic majors or studying one of the foreign languages available at XISU and YMU. Off-topic essays and incomplete texts were excluded

All texts were cleaned and formatted. No changes were made to the texts in relation to grammatical tense or turn of phrase accuracy.

Text and token counts by language are as follows:

Languagetextstokens
Arabic81,762
English10732,822
Filipino101,371
French12939,944
German7810,941
Hindi162,972
Indonesian142,630
Korean242,630
Malay365,208
Persian121,751
Russian338,018
Swahili101,840
Thai121,661
Turkish223,719
Urdu153,645

 

LancsBox X 4.0 was used for counting Swahili, Persian, French, Urdu, and Hindi tokens. AntConc 4.2.4 was used for counting tokens in the other languages.

The essays and writing prompts are stored in UTF-8 encoded plain text files. Metadata is presented in .csv files.