Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DCAT: A Novel Transformer-Based Approach for Dynamic Context-Aware Image Captioning in the Tamil Language

Domain:

natural language processing

Record type:

paper
Creator:
JotAruManGop
Publisher:
MDP
Host:
The task of image captioning in low-resource languages like Tamil is fraught with challenges due to limited linguistic resources and complex semantic structures. This paper addresses the problem of generating contextually and linguistically coherent captions in Tamil. We introduce the Dynamic Context-Aware Transformer (DCAT), a novel approach that synergizes the Vision Transformer (ViT) with the Generative Pre-trained Transformer (GPT-3), reinforced by a unique Context Embedding Layer. The DCAT model, tailored for Tamil, innovatively employs dynamic attention mechanisms during its Initialization, Training, and Inference phases to focus on pertinent visual and textual elements. Our method distinctively leverages the nuances of Tamil syntax and semantics, a novelty in the realm of low-resource language image captioning. Comparative evaluations against established models on datasets like Flickr8k, Flickr30k, and MSCOCO reveal DCAT’s superiority, with a notable 12% increase in BLEU score (0.7425) and a 15% enhancement in METEOR score (0.4391) over leading models. Despite its computational demands, DCAT sets a new benchmark for image captioning in Tamil, demonstrating potential applicability to other similar languages.

Visit

doi.org

Tasks

image-text retrievalcomputer vision

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse DatasetsContext-Aware Dynamic Chunking for Streaming Tibetan Speech Recognitiongautamiyer31/Image-CaptioningExplainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer ModelsA Context-Aware, Computer-Vision-Based Approach for the Detection of Taxi Street-Hailing Scenes from Video Streamsamanuelbyte/amharic-image-captioning

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by

Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition

In this work, we propose a streaming speech recognition framework for Amdo Tibetan, built upon a hyb

gautamiyer31/Image-Captioning

A Machine Learning image captioning (image-to-text) Model for three languages – Hausa, Kyrgyz, and

Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models

Transformer models achieve state-of-the-art performance across domains and tasks, yet their deeply l

A Context-Aware, Computer-Vision-Based Approach for the Detection of Taxi Street-Hailing Scenes from Video Streams

International audience With the increasing deployment of autonomous taxis in differen

amanuelbyte/amharic-image-captioning