Voice corpora are fundamental resources for developing speech technologies, such as automatic speech recognition (ASR), speaker identification, and natural language processing. However, the creation of such datasets remains particularly challenging in low-resource settings, where linguistic diversity, limited technological infrastructure, and financial constraints hinder their systematic development. This systematic literature review aims to synthesize existing research on voice corpus dataset creation in low-resource contexts, focusing on the reported challenges, data collection strategies, and opportunities for advancing speech technologies in under-resourced languages and regions.