Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Neural Fashion Image Captioning : Accounting for Data Diversity

Domain:

natural language processing

Record type:

datasetpaper
Creator:
HacSay
Publisher:
arXiv
Host:avatar
Image captioning has increasingly large domains of application, and fashion is not an exception. Having automatic item descriptions is of great interest for fashion web platforms, sometimes hosting hundreds of thousands of images. This paper is one of the first to tackle image captioning for fashion images. To address dataset diversity issues, we introduced the InFashAIv1 dataset containing almost 16.000 African fashion item images with their titles, prices, and general descriptions. We also used the well-known DeepFashion dataset in addition to InFashAIv1. Captions are generated using the Show and Tell model made of CNN encoder and RNN Decoder. We showed that jointly training the model on both datasets improves captions quality for African style fashion images, suggesting a transfer learning from Western style data. The InFashAIv1 dataset is released on Github to encourage works with more diversity inclusion.

Visit

doi.orgarxiv.org

Tasks

image-text retrievalcomputer vision

Tags

Computer Vision and Pattern Recognition (cs.CV)Artificial Intelligence (cs.AI)FOS: Computer and information sciencesFOS: Computer and information sciencesI.2.10; I.2.7; I.4.10

Licenses

Creative Commons Attribution Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-sa/4.0/legalcode

Similar

Afro SpecDetect A multimodal dataset for African fashion image captioninggautamiyer31/Image-CaptioningImage Captioning in Amharic Language using Deep Convolutional Neural Networks with LSTMamanuelbyte/amharic-image-captioningMuphulusiDzivhani/isiZulu-image-CaptioningCPE-OOU/NIGERIA-IMAGE-CAPTIONING

Afro SpecDetect A multimodal dataset for African fashion image captioning

Afro SpecDetect A multimodal dataset for African fashion image captioning

Poster presented at the Deep Learning Indaba 2023 by Nouréini Sayouti Souleymane

gautamiyer31/Image-Captioning

A Machine Learning image captioning (image-to-text) Model for three languages – Hausa, Kyrgyz, and

Image Captioning in Amharic Language using Deep Convolutional Neural Networks with LSTM

amanuelbyte/amharic-image-captioning

MuphulusiDzivhani/isiZulu-image-Captioning

COS801 Project – isiZulu Image Captioning ## COS 801 Project – Bridging the Visual-Linguistic Divid

CPE-OOU/NIGERIA-IMAGE-CAPTIONING

# Nigeria Image Captioning This repository contains the code and resources for our final year proje