Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CantoTalk: Probing Teacher Expertise From Fine-Tuned Talk Move Representations

Domain:

natural language processingeducation

Record type:

datasetpaper
Creator:
MayXinGor
Editor:
BotT. SinOga
Publisher:
Int
Host:avatar
Classroom discourse profoundly shapes student learning, yet analyzing teacher talk at scale remains challenging in non-Western and low-resource language contexts. This paper introduces CantoTalk, a dataset of 7,518 Cantonese teacher utterances from Hong Kong mathematics classrooms annotated with ten talk-move categories. We investigate whether LLMs can reliably classify these moves and encode systematic differences in teacher expertise. Fine-tuning five open-weight LLMs yields strong performance, with the best model (Qwen3-8B) achieving micro-F1 of 0.81 and macro-F1 of 0.77. Probing utterance-level embeddings reveals that teacher expertise is linearly separable with 0.79 balanced accuracy, well above chance even after controlling for surface features. Clustering analyses uncover three coherent discourse styles differing in pedagogical authority, scaffolding, and dialogic engagement, with qualitative analysis showing systematic differences in how experienced and novice teachers execute similar talk moves. These findings demonstrate that fine-tuned LLM representations capture teacher expertise, offering a new lens for analyzing classroom discourse and informing teacher feedback tools.

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched SpeechAn analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and Englishpmmlv2-fine-tuned-hausapmmlv2-fine-tuned-yoruba0xnu/pmmlv2-fine-tuned-hausardhinaz/BERT-Fine-Tuned-Hausa

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

An analysis of fine-tuned representations for code-switched speech recognition of Yorùbá and English

Poster presented at the Deep Learning Indaba 2022 by Tolúlọpẹ́ Ògúnrẹ̀mí

pmmlv2-fine-tuned-hausa

pmmlv2-fine-tuned-yoruba

0xnu/pmmlv2-fine-tuned-hausa

rdhinaz/BERT-Fine-Tuned-Hausa