Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DCAT: A Novel Transformer-Based Approach for Dynamic Context-Aware Image Captioning in the Tamil Language

Domaine:

natural language processing

Type de record:

paper
Créateur:
JotAruManGop
Éditeur:
MDP
Hôte:
The task of image captioning in low-resource languages like Tamil is fraught with challenges due to limited linguistic resources and complex semantic structures. This paper addresses the problem of generating contextually and linguistically coherent captions in Tamil. We introduce the Dynamic Context-Aware Transformer (DCAT), a novel approach that synergizes the Vision Transformer (ViT) with the Generative Pre-trained Transformer (GPT-3), reinforced by a unique Context Embedding Layer. The DCAT model, tailored for Tamil, innovatively employs dynamic attention mechanisms during its Initialization, Training, and Inference phases to focus on pertinent visual and textual elements. Our method distinctively leverages the nuances of Tamil syntax and semantics, a novelty in the realm of low-resource language image captioning. Comparative evaluations against established models on datasets like Flickr8k, Flickr30k, and MSCOCO reveal DCAT’s superiority, with a notable 12% increase in BLEU score (0.7425) and a 15% enhancement in METEOR score (0.4391) over leading models. Despite its computational demands, DCAT sets a new benchmark for image captioning in Tamil, demonstrating potential applicability to other similar languages.

Visit

doi.org

Tasks

image-text retrievalcomputer vision

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse DatasetsContext-Aware Dynamic Chunking for Streaming Tibetan Speech Recognitiongautamiyer31/Image-CaptioningExplainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer ModelsA Context-Aware, Computer-Vision-Based Approach for the Detection of Taxi Street-Hailing Scenes from Video Streamsamanuelbyte/amharic-image-captioning

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by

Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition

In this work, we propose a streaming speech recognition framework for Amdo Tibetan, built upon a hyb

gautamiyer31/Image-Captioning

A Machine Learning image captioning (image-to-text) Model for three languages – Hausa, Kyrgyz, and

Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models

Transformer models achieve state-of-the-art performance across domains and tasks, yet their deeply l

A Context-Aware, Computer-Vision-Based Approach for the Detection of Taxi Street-Hailing Scenes from Video Streams

International audience With the increasing deployment of autonomous taxis in differen

amanuelbyte/amharic-image-captioning