Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Versatile Corpus for Intrinsic Plagiarism Detection, Text Reuse Analysis, and Author Clustering in Urdu

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Man
Éditeur:
HasFarFarAbi
Éditeur:
Men
Hôte:avatar
Plagiarism detection (PD) is a process of identifying instances where someone has presented another person's work or ideas as their own. Plagiarism detection is categorized into two types (i) Intrinsic plagiarism detection primarily concerns the assessment of authorship consistency within a single document, aiming to identify instances where portions of the text may have been copied or paraphrased from elsewhere within the same document. Author clustering, closely related to intrinsic plagiarism detection, involves grouping documents based on their stylistic and linguistic characteristics to identify common authors or sources within a given dataset. On the other hand, (ii) extrinsic plagiarism detection delves into the comparative analysis of a suspicious document against a set of external source documents, seeking instances of shared phrases, sentences, or paragraphs between them, which is often referred to as text reuse or verbatim copying. Detection of plagiarism from documents is a long-established task in the area of NLP with remarkable contributions in multiple applications. A lot of research has already been conducted in the English and other foreign languages but Urdu language needs a lot of attention especially in intrinsic plagiarism detection domain. The major reason is that Urdu is a low resource language and unfortunately there is no high-quality benchmark corpus available for intrinsic plagiarism detection in Urdu language. This study presents a high-quality benchmark Corpus comprising 10,872 documents. The corpus is structured into two granularity levels: sentence level and paragraph level. This dataset serves multifaceted purposes, facilitating intrinsic plagiarism detection, verbatim text reuse identification, and author clustering in the Urdu language. Also, it holds significance for natural language processing researchers and practitioners as it facilitates the development of specialized plagiarism detection models tailored to the Urdu language. These models can play a vital role in education and publishing by improving the accuracy of plagiarism detection, effectively addressing a gap and enhancing the overall ability to identify copied content in Urdu writing.

Visit

doi.orgdata.mendeley.com

Tasks

text classification

Tags

Authoring SystemNatural Language ProcessingMachine LearningPlagiarismUrdu LanguageText Processing

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

PAN Arabic Intrinsic Plagiarism Detection Shared Task CorpusPAN Arabic External Plagiarism Detection Shared Task Corpus"RoU-AudioTox: Roman Urdu Audio-Text Dataset for Offensive Speech Detection and Transcription"Development of an Algorithm for Plagiarism DetectionFine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed TextHate Speech Detection in Roman Urdu

PAN Arabic Intrinsic Plagiarism Detection Shared Task Corpus

Evaluation corpus for ARAbic INtrinsic plagiarism detection (InAra Corpus) This corpus has been used

PAN Arabic External Plagiarism Detection Shared Task Corpus

Evaluation Corpus for ARAbic EXternal plagiarism detection (ExAra Corpus) This corpus has been used

"RoU-AudioTox: Roman Urdu Audio-Text Dataset for Offensive Speech Detection and Transcription"

"Since online communications are growing, those who speak Roman Urdu experience more problems with o

Development of an Algorithm for Plagiarism Detection

For many years, plagiarism has remained a serious problem in Higher Learning Institutions (HLIs). De

Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

The use of derogatory terms in languages that employ code mixing, such as Roman Urdu, presents chall

Hate Speech Detection in Roman Urdu

Hate speech is a specific type of controversial content that is widely legislated as a crime that mu