The Yorùbá Flickr Audio Caption Corpus (YFACC) dataset extends the Flickr8k image-text dataset to Yorùbá with three modalities:
1. Yorùbá translations of 6k of the captions.
2. Corresponding spoken recordings of these translations, obtained from a single speaker.
3. Temporal alignments of 67 Yorùbá keywords for a subset of 500 of the captions.