Initial release of "Context-Aware Intent Classification in Code-Mixed Kannada: Modeling the Affection–Aggression Paradox."
This repository introduces a framework for intent classification in code-mixed Kannada conversations, where expressions that appear aggressive may convey social closeness depending on context. The work addresses limitations of standard toxicity and hate speech models by incorporating contextual and role-based signals.
Key contributions:
Novel label taxonomy including AFFECTIONATE_AGGRESSION
Annotated dataset of code-mixed Kannada text (JSONL format)
Role-aware and context-aware annotation framework
Baseline transformer-based classification models
Evaluation pipeline with error analysis
Contents:
Annotated dataset and annotation guidelines
Training and evaluation scripts
Baseline models for multi-class intent classification
Configurations and reproducible workflows
This release supports research in multilingual NLP, low-resource language modeling, and more accurate intent detection in socially nuanced communication.
License: MIT