Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Sinhala Similarity and Semantic Plagiarism Detection Using Modern Embedding Models

Domaine:

natural language processing

Type de record:

software
Créateur:
PriRanAbh
Éditeur:
Zenodo
Hôte:avatar
Abstract While plagiarism detection has developed to some degree in English and in some high-resource languages, it has not yet been fully developed in Sinhala language. Plagiarism detection is now well developed in English, and in some high resource languages, but there is little practical support for Sinhala language. This situation is worse in learning environments with presence of manual or tool-based checking of Sinhala submissions that doesn't identify semantic rewriting and translated copying. Current Sinhala research demonstrates success in direct matching and sentence similarity measures, yet these are applied on a limited scale between small corpora and/or rely on very basic similarity measurements or on a sentence level model that cannot be scaled to a full detection pipeline. In this paper, a research design that will tackle 2 related tasks in a single system for plagiarism detection system is presented for Sinhala plagiarism. The first one searches for similarity plagiarism within the Sinhala language, and is able to highlight copied and paraphrased content on a sentence level. It is written in English and then translated into Sinhala and is used without a reference. Semantic plagiarism is where English source material is translated and then used without a source. The proposed solution includes the use of a corpus to build the system, annotating the corpus, generating contextual embeddings, using a Siamese neural network, creating bilingual sentence representations, conducting a focused web crawl, and deploying the system as a service. A common framework enables the ingestion of documents, splitting them into sentences, comparing them, ranking them by ‘thresholder' and generating the report. The paper also dissects the data plan, workflow for implementation, progression of the prototypes and the evaluation plan for both components. The research integrates same language detection and cross language detection technologies within a single framework to improve the natural language processing for Sinhala language and at the same time simultaneously support academia in the low resource scenario by using it as a tool on academic integrity.

Visit

doi.orgzenodo.org

Tasks

embeddings

Tags

Sinhala plagiarism detectionSemantic plagiarismCross-language plagiarism detectionLaBSEMultilingual BERT

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Enhancing semantic question similarity in Arabic using hybrid models and knowledge graph alignmentWord Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English SentencesNeural Models for Detecting Binary Semantic Textual Similarity for Algerian and MSAPredicting Embedding Reliability in Low-Resource Settings Using Corpus Similarity MeasuresIntroducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and RelatednessGATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training

Enhancing semantic question similarity in Arabic using hybrid models and knowledge graph alignment

Semantic question similarity (SQS) plays a crucial role in enhancing question answering (QA) systems

Word Embedding-Based Approaches for Measuring Semantic Similarity of Arabic-English Sentences

International audience Semantic Textual Similarity (STS) is an important component in

Neural Models for Detecting Binary Semantic Textual Similarity for Algerian and MSA

Predicting Embedding Reliability in Low-Resource Settings Using Corpus Similarity Measures

This paper simulates a low-resource setting across 17 languages in order to evaluate embedding simil

Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness

We present two novel datasets for the low-resource language Vietnamese to assess models of semantic

GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training

Semantic textual similarity (STS) is a critical task in natural language processing (NLP), enabling