Translation quality assessment are crucial challenges in computational linguistics. This study explores the use of data compression techniques to evaluate translation accuracy by identifying distinct linguistic patterns. Traditional methods for translation evaluation rely on style metric analysis and machine learning; however, these approaches are often influenced by text length and predefined linguistic features. To address these limitations, we employ an information-theoretic method based on data compression.
Our methodology utilizes compression algorithms to analyze translation of quality assessment. We assess the unconscious stylistic contribution of translators by comparing multiple translations of the same literary works. Additionally, we apply compression-based classification to distinguish between original Amharic texts, human-translated Amharic-to-English texts, and computer-translated texts. In our Experiments were conducted using six original Amharic novels for authorship styles and for translation quality assessment we utilize well-known translated works by human translator and computer translators. Among various lossless data compression algorithms, the following were tested: Prediction by Partial Matching (PPM), Huffman coding, Barrows-Wheeler Transform (BWT) and Lempel-Ziv-Markov Algorithm (LZMA), in order to evaluate their performance. According to the Cramer’s V coefficient calculated from different experiments, the Prediction by Partial Matching (PPM) algorithm showed the highest stability and was therefore selected for all subsequent analyses.
Results indicate that PPM achieves the highest classification accuracy, with a Cramer coefficient (V) of 0.89 for Amharic authorship works, 0.762 and 1 for human-translated English-to-Amharic texts,0.91 for computer based translated Amharic-to-English texts and 0.53 for English Amharic computer translated tasks.
The study demonstrates that data compression techniques provide a viable, language-independent approach for translation quality assessment, particularly for low-resource languages like Amharic. These findings highlight the potential of information-theoretic methods in linguistic analysis and computational translation studies.