We propose a new approach for learning contextualised cross-lingual word embeddings based on a small parallel corpus (e.g. a few hundred sentence pairs). Our method obtains word embeddings via an LSTM encoder-decoder model that simultaneously translates and reconstructs an input sentence. Through sharing model parameters among different languages, our model jointly trains the word embeddings in a common cross-lingual space. We also propose to combine word and subword embeddings to make use of orthographic similarities across different languages. We base our experiments on real-world data from
Research goal: How does the integration of word alignment techniques impact the performance of cross-lingual sentence embeddings on the XNLI benchmark when evaluated through zero-shot transfer accuracy in low-resource languages compared to high-resource languages?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.9/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.9/10.