Logo Lanfrica

A MEDOID-BASED SEMI-SUPERVISED CLUSTERING APPROACH FOR WORD SENSE DISAMBIGUATION IN TELUGU

Domaine:

natural language processing

Type de record:

paper
Créateur:
DUR
Éditeur:
Lit
Hôte:avatar
Word Sense Disambiguation (WSD) is critical component in all the Natural Language Processing related tasks. Due to limited availability of sense annotated corpus and morphologically richness of Telugu language its’ very challenging to develop WSD systems. This paper proposes a semi-supervised clustering technique for Telugu WSD that minimizes the utilization of large annotated corpora. The proposed method uses IndicBERT-based sentence embeddings to find contextual semantics. A novel seed selection approach based on sum-of-squared error (SSE) is proposed to make sure the initialization of sense clusters, followed by a formula-based medoid selection mechanism and an adaptive similarity thresholding scheme for sense propagation. The approach is evaluated on a manually annotated corpus and also compared with baseline methods Most Frequent Sense, Most Common Sense and simple K-medoid. The performance metrics show that the proposed system surpasses the baseline approaches. This shows the effectiveness of proposed approach for Telugu WSD and applicability to Information retrieval tasks.