
Public sample of 500 sentence pairs from a parallel corpus of 4,056 unique Indonesian–Sekar sentence pairs built to support neural machine translation for the endangered Papua Kokas language (Sekar; ISO 639-3: skz; Glottolog: seka1247), an Austronesian language of the North Bomberai subgroup spoken in the Kokas District of Fakfak Regency, West Papua, Indonesia.
The corpus was constructed through remote elicitation with a heritage-community speaker: thematically composed Indonesian stimulus sentences were translated into Sekar by a native speaker of the older generation from Kokas District through iterative WhatsApp correspondence during 2024, with verbal informed consent for research use. The 500 pairs were sampled uniformly at random from the corpus training split only; the held-out validation and test splits are not included, so this sample can be used without contaminating future evaluation on the full corpus.
Governance of the corpus follows the CARE Principles for Indigenous Data Governance. The full corpus (4,056 pairs with train/validation/test splits) is available on reasonable request to the corresponding author of the associated paper, under a data-use agreement that respects the speaker community's authority over the resource. See README.md for collection methodology, format, citation, and contact.