Zero-Shot POS Tagging for Under-Resourced Languages
# Cross-Lingual POS Tagging: XLM-R vs. Glot500
## Project Goal
Compare the performance of XLM-R and Glot500 when fine-tuned for POS tagging on a better-resourced language and then applied directly to a low-resource language without further training. Analyze the impact of subword tokenization on cross-lingual transfer.
## Results
### Fragmentation Rate
|High-Resource Language| XLMR Fragment Rate| Glot500 Fragment Rate|
|--------------|----------|----------|
| English|1.30 | 1.19|
| French| 1.44|1.32 |
| Standard Arabic| 1.0| 1.0|
| Russian| 1.67|1.49 |
|Low-Resource Language| XLMR Fragment Rate| Glot500 Fragment Rate|
|--------------|----------|----------|
| Wolof| 1.81 |1.38 |
| Catalan|1.41 |1.30 |
| Urdu| 1.31| 1.28|
| Ukranian| 1.74| 1.57|
### Cross-Lingual Transfer Performance: Mono-Lingual Fine-Tuning Dataset
| Language Pair | Model | Accuracy | F1 Score |
|--------------|---------|----------|-----------|
| English → Wolof | XLM-R | 37% | 34% |
| | Glot500 | 47% | 46% |
| Standard Arabic → Urdu | XLM-R | 17% | 11% |
| | Glot500 | 24% | 10% |
| French → Catalan | XLM-R | 47% | 47% |
| | Glot500 | 68% | 68% |
| Russian → Ukrainian | XLM-R | 54% | 53% |
| | Glot500 | 72% | 69% |
| Welsh → Irish | XLM-R | 47% | 42% |
| | Glot500 | 33% | 19% |
### Cross-Lingual Transfer Performance: Mono-Lingual Fine-Tuning Dataset with Noise Injection
| Language Pair | Model | 25% Noise | | 75% Noise | |
|--------------|--------|------------|------------|------------|------------|
| | | Accuracy | F1 Score | Accuracy | F1 Score |
| English → Wolof | XLM-R | 23% | 20% | 22% | 19% |
| | Glot500 | 42% | 41% |39% | 38% |
| Standard Arabic → Urdu | XLM-R | 24% | 11% | 23% | 11% |
| | Glot500 | 24% | 11% | 24% | 11% |
| French → Catalan | XLM-R | 46% | 46% | 46% | 45% |
| | Glot500 | 68% | 67% | 67% | 66% |
| Russian → Ukrainian | XLM-R | 54% | 53% | 51% | 50% |
| | Glot500 | 66% | 61% | 59% …