To digitize existing high-quality text belonging to a certain low-resource language, we are often faced with two major bottlenecks. First, we may have non-machine-readable documents such as dictionaries, stories, grammars, etc. that need to be digitized correctly to create datasets for downstream use. To tackle this, there is a rich body of research that investigates improving optical character recognition (OCR) quality, and in this thesis, we propose post-OCR semi-automatic processing methods that can improve the readability and structure of this output. Second, even if we assume access to machine-readable text, we often lack granular language identification (langID) models that can accurately identify and classify target text into a large number of languages and language varieties. For this, we propose a simple and effective hierarchical solution, based on previous contributions in the langID domain, and propose training script-agnostic word representations to better represent language families with immense script diversity. In summary, this thesis aims to tackle the resource bottleneck for low-resource languages in the text domain using language identification and optical character recognition. In addition, it attempts to study and improve resource creation efforts, which will hopefully positively impact communities that employ such languages when interfacing with technology.