Logo Lanfrica

Robust speech recognition for low-resource languages

Domain:

natural language processing

Record type:

paper
Creator:
Rom
Editor:
UniUni
Publisher:
Uni
Host:avatar
Process of human-machine interaction is an integral part of everyday human life in a modern world. The various interfaces are intended to facilitate this interaction and provide maximum comfort for users. The speech interface is the most universal, and it can be used in a vast majority of systems. Speech is the main way of communication between people, from this point of view speech interface looks completely natural. Nowadays, human-machine speech interfaces can consist of many complex components in order to implement an interaction with a user. However, a primary element of such interfaces is an automatic speech recognition (ASR) system. Automatic recognition of spontaneous speech in a telephone channel can be solved very successfully in the presence of significant amount of speech training data, but the collection and processing of such data is a complex, time-consuming and expensive task. Great amount of data is available only for high-resource languages such as English, Spanish, German, Chinese, French, etc. At the same time, there are many languages with no significant amount of training data for ASR systems preparation. Such languages are commonly referred to as low-resource languages. The number of native speakers for such languages can reach tens, and sometimes hundreds of millions of people. Taking this into account, building ASR systems for such languages is an important task. The main objective main objective of this thesis is the development of universal methodology for robust speech recognition of low-resource languages. To achieve it, we have solved three tasks. First, we have developed a flexible methodology in order to build an ASR system for large vocabulary conversational speech on low-resource languages. Second, we have identified a set of languages representing a wide variety of acoustic and grammatical features that influence on ASR systems. Third, we have conducted experimental studies of the developed universal methodology and compared the accuracy of the resulting ASR systems with the state-of-the-art results. We have contributed to several key components of the ASR system as part of the first task. We have proposed speaker-dependent bottleneck acoustic features trained with a multilingual approach, to improve acoustic modeling. Besides, a new approach of combining acoustic features has been proposed. In terms of language modeling, we have proposed a new method of text data augmentation with the artificial texts generated by the Char-RNN model. Finally, we have created a universal methodology that combines both existing and proposed methods. This methodology represents a complete pipeline for building ASR systems for low-resource languages. As a part of the second task, we have identified the most important language features which affect the building process of ASR systems. We have considered such factors as word order, agglutination, inflection, presence of an accent, dialects, and tone. Based on the analysis, we have created a set of low-resource languages that covers the factors described above and several additional features. This set includes eight languages: Georgian, Turkish, Kazakh, Vietnamese, Swahili, Tagalog, Zulu, and Russian. As a part of the third task, we have conducted large-scale experimental studies. We have identified the most effective basic acoustic features and the training pipeline for initial GMM-HMM models. We have verified proposed multilingual speaker-dependent bottleneck features and confirmed the effectiveness of the delayed combination for acoustic features. We have determined the usefulness of the proposed text data augmentation approach. Moreover, we have found this approach complementary to the addition of filtered web-text data. We have experimented with various neural network-based acoustic models and identified the most effective ones. Finally, we have compared the accuracy of resulting ASR systems with the state-of-the-art results for studied low-resource languages. We believe that this work will accelerate further research of speech recognition in low-resource languages and make speech recognition systems available for a wide variety of world languages in the near future. Diese Dissertation entstand im Rahmen einer Kooperation mit der Universität ITMO in St. Petersburg. / This dissertation was written in the context of a cooperation with ITMO University in St. Petersburg.