Logo Lanfrica

Student Translation Corpus (STC): Decolonial Strategies in Archaeological and Historical Discourse (English-Russian / Russian-English)

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Kal
Éditeur:
Abd
Éditeur:
Zenodo
Hôte:avatar
This dataset contains the Student Translation Corpus (STC), a specialized bilingual learner translation corpus comprising 195 translation solutions (approximately 17,400 tokens, including source texts). The corpus was compiled to analyze cognitive and pragmatic attitudes of novice translators working with postcolonial and decolonial translation strategies within the archaeological and historical discourse of Kazakhstan. The STC features a bidirectional (bivector) architecture: Inward Vector (English to Russian) contains 100 translation texts focusing on the import of Western academic discourse and the identification/translation of colonial markers (e.g., epistemic arrogance, exoticization). Outward Vector (Russian to English) includes 95 translation texts focusing on the export of Kazakhstan's national, sacred, and archaeological heritage (e.g., Saryarka Petroglyphs, Botai Culture, Emir Timur, and Karazhartas Pyramid) to a global audience, analyzing pragmatic adaptation and decolonial hypercorrection. The dataset includes a three-level pragmatic indexing matrix for evaluating the ideological intervention of the translator, shifting away from binary "correct/incorrect" assessments toward analyzing the social agency of the translator in postcolonial contexts.

Languages

Similaires