Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Type de record:

paperdataset
Créateur:
IkuWöhAiz
Hôte:avatar
Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panels differ from natural images, computational systems traditionally had to be designed specifically for manga. Recently, the adaptive nature of modern large multimodal models (LMMs) shows possibilities for more general approaches. To provide an analysis of the current capability of LMMs for manga understanding tasks and identifying areas for their improvement, we design and evaluate MangaUB, a novel manga understanding benchmark for LMMs. MangaUB is designed to assess the recognition and understanding of content shown in a single panel as well as conveyed across multiple panels, allowing for a fine-grained analysis of a model's various capabilities required for manga understanding. Our results show strong performance on the recognition of image content, while understanding the emotion and information conveyed across multiple panels is still challenging, highlighting future work towards LMMs for manga understanding. This work has been submitted to the IEEE for possible publication

Visit

arxiv.org

Tags

Computer Vision and Pattern RecognitionMultimedia

Similaires

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsLet's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of SportsLarge Multimodal Models for Low-Resource Languages: A SurveyKazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCRPolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and DialectsContext-Aware Large Language Models for Multilingual Understanding

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR

Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evalua

Context-Aware Large Language Models for Multilingual Understanding

Multilingual large language models (LLMs) have demonstrated strong performance in cross-lingual task