This dataset combines agricultural and satellite images with English captions and their synthetic Wolof translations.
Captions were translated using Oolel.
Wolof lacks multimodal training data. This dataset attempts to address it by pairing real-world agricultural imagery with natural Wolof descriptions, enabling vision-language model training in the language.
Methodology