This preliminary work addresses the critical gap in multimodal Large Language Models (LLMs) for African languages, which remain underrepresented despite their rich multimodal communication traditions. We propose a framework that leverages simulated multimodal data and cross-lingual transfer learning to bootstrap multimodal capabilities. Our initial experiments with Swahili demonstrate that proxy multimodal embeddings can be effectively generated using pre-trained encoders, achieving an average cosine similarity of 0.72 for culturally relevant concepts. We further show that simple fusion methods can effectively combine these embeddings, and that transfer learning from high-resource languages yields a 28% improvement in multimodal alignment over zero-shot approaches. These results validate the feasibility of our approach and provide a foundation for culturally-aware multimodal LLMs in low-resource African language contexts.