Socio-economic research often requires detailed information on individual occupations in order to study aspects such as occupational prestige, hazards, or gender disparities. Traditionally, this data collection has involved manual coding of respondents’ job descriptions into specific classifications. However, for reasons of efficiency and quality, this process is increasingly being automated using machine learning. While these automated approaches have shown success with large, high-quality datasets in single languages, multilingual occupation coding remains a challenge, especially for low-resource languages with limited training data.
This paper investigates the potential of transfer learning to improve multilingual occupation coding. We explore how models can leverage the predictive power of other languages within large multilingual datasets. Using extensive training data from the German Socio-Economic Panel (SOEP) and the Survey of Health Aging and Retirement in Europe (SHARE), we fine-tune several pre-trained language models (DistilBERT) to predict one- and four-digit ISCO08 codes, representing simple and complex classification tasks, respectively.
We compare the prediction accuracy of models trained exclusively on country-specific data against those trained on the full multilingual dataset. In addition, we examine the impact of boosting a specific language, in this case German using SOEP data, on the accuracy for this and other countries. Our results aim to demonstrate the effectiveness of transfer learning in improving multilingual occupation classification. Specifically, for cross-country studies such as SHARE, we provide insights into whether and how we can mitigate the challenge of coding occupations with high accuracy in low-resource languages.