Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MariaYasmeen/NLP-Low-Resource-Paraphrase-detection

Domain:

natural language processing

Record type:

dataset
Creator:
Mar
Host:
# Monolingual Paraphrase Detection - Low Resource Sindhi Lang at Sentence Level This project focuses on **sentence-level paraphrase detection for the Sindhi language**. The system classifies whether two Sindhi sentences convey the same meaning (paraphrase) or not (non-paraphrase). ## Overview Paraphrase detection is important for many NLP applications such as question answering, plagiarism detection, information retrieval, and text similarity systems. Most previous work focuses on English, while Sindhi remains a low-resource language. This project helps fill that gap by creating a dataset and building detection models. ## Objectives * Build a Sindhi paraphrase corpus * Train machine learning and deep learning models * Compare different feature representations * Develop a real-time prediction interface ## Dataset * Collected Sindhi sentence pairs from online news sources * Manually annotated into paraphrase and non-paraphrase * Preprocessed and cleaned the text data * Split into training and testing sets ## Methods Used The following techniques were implemented and compared: * N-gram features * FastText embeddings * Sentence Transformers * Feature fusion approach ## Model Training * Implemented in Python * Used Scikit-learn and HuggingFace Transformers * Experiments conducted on Google Colab * Evaluated using standard classification metrics ## Results * Feature fusion provided the best performance * Sentence Transformers showed strong semantic understanding * The system achieved reliable paraphrase classification for Sindhi text *(You can add your exact accuracy/F1 score here.)* ## Web Interface A simple web-based interface was developed where users can: * Input two Sindhi sentences * Get real-time paraphrase prediction ## How to Run 1. Clone the repository ``` git clone cd ``` 2. Install dependencies ``` pip install -r requirements.txt ``` 3. Run the model or notebook ## Tools & Technologies * Python * Google Colab * Scikit-learn * HuggingFa …

Visit

github.com

Similar

oyinkanchekwas/low-resource-nlp-toolkitToluClassics/Low-Resource-NLP-Tutorialsellis-nlp/low-resource-MTDecolonizing NLP for “Low-resource Languages”SaifWiyar/low-resource-nlp-ml-researchpriscilla-adenuga/low-resource-stress-tests-nlp

oyinkanchekwas/low-resource-nlp-toolkit

Selective language routing and code-switch audits, evaluated on a reproducible AfriSenti benchmark.

ToluClassics/Low-Resource-NLP-Tutorials

Getting started in NLP for low resource languages # Low-Resource-NLP The goal of this repository i

ellis-nlp/low-resource-MT

Work on machine translation in low resource scenarios and for minority and under-resourced languages

Decolonizing NLP for “Low-resource Languages”

Today African languages are spoken by more than a billion people, yet in the world of machine transl

SaifWiyar/low-resource-nlp-ml-research

Research portfolio for low-resource language NLP, dataset creation, annotation, machine learning, de

priscilla-adenuga/low-resource-stress-tests-nlp

Berlin Buzzwords 2026 talk repo on low-resource languages as stress tests for NLP # Low-Resource La