Logo Lanfrica

Sindhi Open Lexicon Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Sin
Éditeur:
FaySin
Éditeur:
Har
Hôte:avatar
Sindhi Open Lexicon Dataset (223K+ Entries) for AI, NLP & Computational Linguistics. This project is a large-scale structured lexical dataset for the Sindhi language containing over 223,000 entries including definitions, linguistic metadata, and normalized forms. Sindhi is a historically rich but low-resource language in AI. This dataset aims to support NLP, AI systems, and computational linguistics. Objectives - Provide AI-ready Sindhi dataset - Support NLP research - Enable search engines, chatbots, OCR, and language tools - Preserve linguistic heritage digitally Dataset Features - 223,000+ entries - Definitions in Sindhi - Variants with/without diacritics - Normalized text - Domain classification - Formats: CSV, JSONL, SQLite