Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

JoshuaPianoGuy/NLP-Reseach--Afrikaans-Compound-Segmentation

Domain:

natural language processing

Record type:

software
Creator:
Jos
Host:
Code for semester long research project that looked at using pre-trained language models for Afrikaans compound segmentation # NLP-Research--Afrikaans-Compound-Segmentation ## Usage Code for semester long research project that looked at using pre-trained language models for Afrikaans compound segmentation. To use this repo, you first need to download the freely available dataset from the Aucopro project at: repo.sadilar.org Next up, remove the duplicates, and run the split_new.py script which groups the data according to stems, removes any morpheme boundaries, replaces compound boundaries with '@' signs and randomly distributes the data into a traing partition of 80% of the total size, and validation and testing sets which are each 10% of the total size. To finish pre-processing, run the remove_empty.py script to remove empty lines from stem groupings. Finally run the fine_tune_val.py script to fine-tune your chosen model. The model with the best f1 score on the validation set is loaded at the end. To test the model and see its predictions, run the improved_testing.py script. ## Acknowledgements This code is adapted from the LLMSegm codebase developed by Pranjić et al. (2024). ## References M. Pranjić, M. Robnik-Šikonja, and S. Pollak, ‘LLMSegm: Surface-level Morphological Segmentation Using Large Language Model’, in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024, pp. 10665–10674.

Visit

github.com

Languages

Afrikaans

Licenses

MIT

Similar

Automatic Compound Processing: Compound Splitting and Semantic Analysis for Afrikaans and DutchBioTech-America-Reseach/BOX-MARADHIThe Development of Dutch and Afrikaans Language Resources for Compound Boundary Analysis.An investigation of a lexical segmentation strategy for Afrikaansclarin-pl/combo-nlp-xlm-roberta-base-afrikaans-afribooms-ud2.17Kinyarwanda-NLP/NLP

Automatic Compound Processing: Compound Splitting and Semantic Analysis for Afrikaans and Dutch

BioTech-America-Reseach/BOX-MARADHI

Mfumo huu umebuniwa kuwa kama kitabu cha kidijitali ambapo kila herufi kutoka A mpaka Z inawakilisha

The Development of Dutch and Afrikaans Language Resources for Compound Boundary Analysis.

An investigation of a lexical segmentation strategy for Afrikaans

clarin-pl/combo-nlp-xlm-roberta-base-afrikaans-afribooms-ud2.17

Kinyarwanda-NLP/NLP

NLP in Kinyarwanda # NLP NLP in Kinyarwanda