Federated learning offers a promising paradigm for training machine learning models across distributed creative industry clusters without centralising sensitive cultural or commercial data. However, its deployment in Ugandan creative districts confronts two interrelated statistical challenges: distribution shift across heterogeneous participating clusters and extreme label sparsity in domains where annotation expertise is scarce. We introduce a client-weighted objective that couples a distributionally robust optimisation term with a semi-supervised consistency regulariser, and we prove convergence guarantees under non-identical data partitions and partial label availability. The framework is analysed theoretically through a novel combination of Wasserstein distance-based robustness bounds and graph-based label propagation error estimates. We further propose a practical protocol for differential privacy that accommodates the communication constraints typical of Ugandan creative districts. The analysis demonstrates that robustness to distribution shift and resilience to sparse labels are not competing objectives but can be jointly optimised through a principled weighting scheme.