SegmOnto
Layout analysis model trained with YALTAi and relying on YOLO models. Data are annotated with the SegmOnto controlled vocabulary. Most ot the training data are French texts, mainly prints but not only, produced by the Gallic(orpor)a, the FoNDUE and the SETAF projects.
If you need to quote the paper:
@inproceedings{solfrini_OCR_2024, author={solfrini, Sonia and Gabay, Simon and Pinche, Ariane and Beaulnes, Pierre-Olivier and Marques Oliveira, Aurélia and Gross, Geneviève and Solfaroli Camillocci, Daniela}, title={Océriser les imprimés du XVIe siècle en langue française : le cas d'un corpus romand en caractères gothiques}, address={Meknes, Morocco}, year={2024}, month={May}, booktitle={Humanistica 2024}, publisher={Association francophone des humanités numériques} }