Logo Lanfrica

fgaim/tigrinya-nlp-anthology

Domaine:

natural language processing

Type de record:

paper
Créateur:
fga
Hôte:
Tigrinya NLP Anthology of Academic Papers and Computational Resources # Tigrinya NLP Anthology: Papers and Computational Resources This repository contains the official, curated anthology of research papers and computational resources for Natural Language Processing (NLP) in Tigrinya. The collection was compiled as part of the paper: **Natural Language Processing for Tigrinya: Current State and Future Directions**. ## Overview This work presents a comprehensive survey of NLP research for Tigrinya, analyzing over **50 studies** published between **2011 and 2025**. We systematically review the current state of computational resources, models, and applications across multiple downstream tasks, including: - **Morphological Processing** - Rule-based stemmers, finite-state transducers, neural boundary detection - **Machine Translation** - Statistical MT, Neural MT, transfer learning approaches - **Part-of-Speech Tagging** - CRF/SVM models, BiLSTM architectures, transformer fine-tuning - **Named Entity Recognition** - Small-scale to large-scale annotated datasets and models - **Text Classification** - Topic categorization, sentiment analysis, hate speech detection - **Question-Answering** - Factoid QA, reading comprehension benchmarks, educational datasets - **Language Modeling** - Monolingual corpora construction, pre-trained transformer models - **Automatic Speech Recognition** - Corpus design, deep neural networks, hybrid CTC-Attention models - **Text-to-Speech** - Concatenative synthesis, neural TTS (Tacotron), massively multilingual speech models - **Optical Character Recognition** - GLOCR dataset, CRNN-based models for Ge'ez script - **Information Retrieval** - Word stemming for lexical retrieval, bi-encoder models for semantic search - **Language/Dialect Identification** - Mutual intelligibility studies, dialect detection The research spans from early rule-based systems (2011) to modern neural architectures and large language models (2025), highlighting how community-led dataset creation has been the primary catalyst for adva …

Languages