Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Akhara: A Study on a Style-Conditioned Latent Diffusion Model for Khmer Handwriting Synthesis

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
PenSin
Éditeur:
Zenodo
Hôte:avatar
Khmer, which is a language spoken by more than 16 million speakers, is considered a low-resource language when it comes to Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) because of the limited availability of annotated handwriting data and the high cost of manually annotating writing systems that consist of consonant stacking, ligatures, and diacritics.Akhara presents the DiffusionPen framework with an approach that utilizes a CANINE-C character-level encoding scheme without tokenization and a writer-specific MobileNetV2-based style extractor to generate labeled synthetic Khmer handwriting from Unicode text. This paper presents results of applying the presented solution to the task of creating synthetic handwriting for the Khmer language. The quality of generated images was estimated based on visual similarity (Fréchet Inception Distance), OCR readability (Tesseract), and human evaluation.The results have indicated that latent diffusion under conditional style is a scalable approach to manual annotation, which can lead to better performance in OCR/HTR systems of Khmer.

Visit

doi.orgzenodo.org

Tasks

computer visionoptical character recognition

Tags

KhmerSynthetic Data GenerationLatent Diffusion ModelsGenerative ModelsLow-Resource Language

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode