Logo Lanfrica

SaifWiyar/low-resource-nlp-ml-research

Domaine:

natural language processing

Type de record:

project
Créateur:
Sai
Hôte:
Research portfolio for low-resource language NLP, dataset creation, annotation, machine learning, deep learning, model training, benchmarking, and custom AI solutions. # Low-Resource NLP & Machine Learning Research ### Dataset Creation · Model Training · NLP Engineering · Applied AI Research A professional research and development portfolio focused on building datasets, machine-learning models, NLP systems, and intelligent applications for low-resource languages and real-world problems. --- ## Research Overview This repository presents the NLP, Machine Learning, Data Science, and dataset-development work of **Saif Wiyar**, a Data Scientist and AI Engineer from Afghanistan. My research and development work focuses on transforming raw text, speech, documents, and domain-specific information into structured datasets and trained intelligent systems. I specialize in: - Low-resource language datasets - Natural Language Processing - Machine Learning model development - Deep Learning - Text classification - Language-model applications - Dataset collection and annotation - Data cleaning and preprocessing - Model training and evaluation - AI-powered educational systems - Production integration of trained models I design datasets containing **thousands of carefully organized and labeled samples**, depending on the language, research objective, data availability, and client requirements. --- ## Primary Research Focus ### Low-Resource Language Technology Many languages have limited digital resources, labeled datasets, pretrained models, and research benchmarks. My work aims to support low-resource languages through: - Text corpus construction - Speech-data collection - Dataset cleaning - Manual and assisted annotation - Language-specific preprocessing - Tokenization research - Text normalization - Benchmark creation - Machine Learning evaluation - NLP model development - Research documentation Current areas of interest include: - Pashto - Dari - Regional languages - Multilingual NLP - Right-to-left language processing - Low-resource educational technology --- ## Dataset Development I design datasets …