This release contains the first version of AIULEC, a learner corpus of written English produced by Libyan Arabic (L1) university students majoring in English at Al-asmarya Islamic University (Spring 2024).
Corpus highlights:
272 documents (~20,000 words) from English majors across all semesters
Anonymized learner texts written under 60-minute timed conditions without grammar/spell-check tools
XML files with embedded metadata and CLAWS C5 POS tags
UCLEE error-tagged files for manual error annotation
Plain-text (TXT) versions of all documents
Combined metadata table (CSV) and PDF versions for teacher-friendly reading
Intended use:
Learner corpus research
SLA studies
Error analysis
NLP tasks (tagging, classification, grammatical error correction)
License: Creative Commons Attribution–NonCommercial 4.0 International (CC BY-NC 4.0)
Citation: To be added via Zenodo DOI after publication.
For questions or contributions, please open an issue or contact @alkishir on GitHub.