The West African Power Pool synchronized its member grids in late 2025 and is moving toward a common regional electricity market, motivating collaborative load forecasting among utilities that cannot centralize raw operational data. This paper tests whether federated collaboration improves day-ahead load forecasting relative to each country training on its own data alone. Ten member grids, calibrated to published national statistics and spanning a sixteen-fold range in scale, are used to compare five strategies sharing an identical patch-transformer architecture: local-only training, Federated Averaging, FedProx, centralized pooling, and personalized federated learning. With full data, FedAvg and FedProx are worse than local-only training, at 8.27 and 8.36 percent mean absolute percentage error against 7.45 percent, a gap significant for the largest client (Diebold-Mariano p less than 0.001) and explained by severe heterogeneity in scale, growth, and sectoral composition. Reducing each client to 60 days of data does not reverse this: local-only still wins for eight of ten clients, contrary to the expectation that collaboration helps most under scarcity. Clustering clients by economic scale recovers part of the gap for larger economies only. Personalization and centralized pooling close it entirely, identifying personalization rather than vanilla aggregation as the strategy suited to this heterogeneity.