Discourse coherence is vital for language understanding but underexplored in low-resource languages. Multimodal cues—gestures, prosody, and visual context—can complement scarce text, yet systematic integration is lacking. We present a pilot study on multimodal discourse coherence in Wolof, introducing 45 videos (6.2 hours), 1,096 utterances, and 1,432 annotated relations across six types. Our fusion model combines textual (mBERT), auditory (Wav2Vec + prosody), and visual (ResNet + pose) features, achieving 52.4 macro F1, surpassing text-only (48.7) and audio-only (35.2) baselines. Audio aids causal relations, visuals support contrast, and we discuss annotation challenges and cross-modal alignment. This work provides a foundation for discourse-aware NLP in low-resource languages.