Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Developing a Bilingual English-Arabic Dataset for Textbook Question Answering: A Hybrid Translation and Validation Approach

Domaine:

natural language processingeducation

Type de record:

dataset
Créateur:
Ama
Éditeur:
Yay
Hôte:
Textbook Question Answering has been a central feature of educational artificial intelligence enabling curriculumaligned machine reading to support personalized learning and diagnostic testing. While there is significant advancement in English-language TQA datasets, there is still a lag in Arabic because of a lack of sufficient highquality domain-specific resources. A new bilingual English-Arabic TQA data is presented in this paper, and it was created using a hybrid translation and validation method. It combines machine translation of CK12-QA dataset with Google sheet translator. Semantic consistency was evaluated using automated metrics based on multilingual sentence embeddings and translation quality scores. Cosine similarity (0.87) and BLEU score (38.5) confirmed strong semantic equivalence and translation reliability across the bilingual dataset. These results demonstrate robust linguistic alignment and completeness. This approach is a balance between conflicting scalability and accuracy in long-standing semantic drift, morphological variation and in context misalignment issues in Arabic education datasets compared to previous efforts to use machine translation or mini-batch annotation only. Output dataset has a parallel format structure of English-Arabic question-answer pair that facilitates simple cross-lingual research in multiple-choice and textbook conditions. By focusing on K-12 science curriculum in specific subject areas, this contribution can enable improved monolingual and cross-lingual educational QA applications model training and testing. This does not only make AI-based learning more inclusive among Arabic students but also provides impetus to creation of cross-lingual transfer learning and benchmarking in TQA. The sources and information are openly published in an attempt to further increase the reproducibility, verifiable peer cooperation and further promote the development of AI in multilingual education

Visit

doi.org

Tasks

machine translationquestion answering

Licenses

https://creativecommons.org/licenses/by-nc-sa/4.0

Similaires

TroubleQA-BE: A Bilingual Troubleshooting Question Answering Dataset for Bangla and EnglishArabicaQA: A Comprehensive Dataset for Arabic Question AnsweringACQAD: A Dataset for Arabic Complex Question AnsweringPre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative StudyA Question-Entailment Approach to Question AnsweringDeep learning-based approach for Arabic open domain question answering

TroubleQA-BE: A Bilingual Troubleshooting Question Answering Dataset for Bangla and English

BTQA is a bilingual troubleshooting question answering dataset created to facilitate research and de

ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources

ACQAD: A Dataset for Arabic Complex Question Answering

International audience In this paper, we tackle the problem of Arabic complex Questio

Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study

Question answering(QA) is one of the most challenging yet widely investigated problems in Natural La

A Question-Entailment Approach to Question Answering

One of the challenges in large-scale information retrieval (IR) is to develop fine-grained and domai

Deep learning-based approach for Arabic open domain question answering

Open-domain question answering (OpenQA) is one of the most challenging yet widely investigated probl