Text Normalization for Low-Resource Languages of Africa
Training data for machine learning models can come from many different
sources, which can be of dubious quality. For resource-rich languages like
English, there is a lot of data available, so we can afford to throw out the
dubious data. For low-resource languages w