Logo Lanfrica

RamzyBakir/roberta-aegyptustranslit-classifier

Domain:

natural language processing

Record type:

model
Creator:
Ram
Host:
A fine-tuned RoBERTa-base model for classifying Ancient Egyptian transliterations into their respective historical time periods ('Predynastic & Early Dynastic', 'Old Kingdom & First Intermediate', 'Middle Kingdom & Second Intermediate', 'New Kingdom & Third Intermediate', 'Late Period & Greco-Roman Egypt'). # roberta-aegyptustranslit-classifier A fine-tuned `RoBERTa-base` model for classifying **Ancient Egyptian transliterations** into their respective historical **time periods** ('Predynastic & Early Dynastic', 'Old Kingdom & First Intermediate', 'Middle Kingdom & Second Intermediate', 'New Kingdom & Third Intermediate', 'Late Period & Greco-Roman Egypt'). View and use the model on HuggingFace via: huggingface.co ## Model Details - **Base model**: `roberta-base` - **Task**: Text classification - **Classes**: 5 historical time periods - **Fine-tuned on**: Custom labeled dataset (transliteration, period) ## Training Results | Metric | Value | |--------|-------| | **F1 Score** | ~0.562 | | **Weighted F1** | ~0.567 | | **Validation Loss** | ~1.5295 | | **Epochs** | 20 | | **Learning Rate** | 2e-5 | | **Batch Size** | 64 | ## Per-class F1 scores: - **Predynastic & Early Dynastic:** F1 = 0.576 - **Old Kingdom & First Intermediate:** F1 = 0.432 - **Middle Kingdom & Second Intermediate:** F1 = 0.468 - **New Kingdom & Third Intermediate:** F1 = 0.713 - **Late Period & Greco-Roman Egypt:** F1 = 0.608 ## Intended Use & Limitations This model is designed for **historical text classification** and intended for exploratory research and as a performance baseline. Current constraints include: - **Data limitations**: ~10k balanced samples may not represent all orthographic variations. - **Period bias**: Middle Kingdom classification (F1=0.47) underperforms due to: - Orthographic overlap with neighboring periods - **Best practices**: Always verify critical classifications with primary sources. **Roadmap**: - v2.0: Expand to ccorpus samples (balanced & unbalanced) ## Data Used For Training Thesaurus Linguae Aegyptiae, Late Egyptian sentences, corpus v19, premium, huggingface.co, v1.0, 1/19/2025 ed. by Tonio Sebastian Richter & Daniel A. Werni …