Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TroubleQA-BE: A Bilingual Troubleshooting Question Answering Dataset for Bangla and English

Domaine:

natural language processing

Type de record:

dataset
Créateur:
HanRah
Éditeur:
Daffodil International University
Éditeur:
Men
Hôte:avatar
BTQA is a bilingual troubleshooting question answering dataset created to facilitate research and development in Natural Language Processing (NLP), multilingual artificial intelligence, conversational agents, and automated technical support systems. The dataset contains a total of 1,006 manually curated question-answer pairs, consisting of 503 Bangla QA pairs and 503 English QA pairs. The dataset focuses on troubleshooting and technical assistance scenarios commonly encountered in daily technology usage, including issues related to mobile devices, computers, software applications, internet connectivity, operating systems, hardware functionality, account management, and general technical support. Each question is paired with a concise, accurate, and contextually relevant answer designed to simulate real-world technical support interactions and improve the performance of AI-based support systems. The primary objective of BTQA is to provide a high-quality bilingual resource that can support a variety of downstream NLP and machine learning tasks, including question answering (QA), chatbot development, retrieval-augmented generation (RAG), semantic search, multilingual information retrieval, intent understanding, cross-lingual learning, and large language model (LLM) fine-tuning. The dataset is particularly valuable for Bangla NLP research, as publicly available domain-specific Bangla QA datasets remain limited compared to English resources. By including aligned troubleshooting content in both Bangla and English, BTQA enables comparative multilingual experimentation and supports the development of more inclusive AI systems for low-resource languages. The dataset is distributed in CSV format with UTF-8 encoding to ensure compatibility with multilingual text processing pipelines and modern machine learning frameworks. All entries were manually reviewed to maintain linguistic clarity, consistency, and practical relevance. BTQA is intended for academic research, educational purposes, benchmarking multilingual NLP systems, and building intelligent customer support applications capable of operating across multiple languages. The dataset aims to contribute to the advancement of multilingual AI technologies by providing a structured and domain-focused bilingual troubleshooting corpus that can serve as a valuable resource for researchers, students, and developers working in both academia and industry.

Visit

doi.orgdata.mendeley.com

Tasks

question answering

Languages

Samba Leko

Tags

Natural Language Processing

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Developing a Bilingual English-Arabic Dataset for Textbook Question Answering: A Hybrid Translation and Validation ApproachSwahili Question-Answering Dataset for HorticultureArabicaQA: A Comprehensive Dataset for Arabic Question AnsweringACQAD: A Dataset for Arabic Complex Question AnsweringAmQA: Amharic Question Answering DatasetTigrinya Question-Answering Dataset (TiQuAD)

Developing a Bilingual English-Arabic Dataset for Textbook Question Answering: A Hybrid Translation and Validation Approach

Textbook Question Answering has been a central feature of educational artificial intelligence enabli

Swahili Question-Answering Dataset for Horticulture

The dataset was created to contribute to Swahili language resources for natural language processing

ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources

ACQAD: A Dataset for Arabic Complex Question Answering

International audience In this paper, we tackle the problem of Arabic complex Questio

AmQA: Amharic Question Answering Dataset

Question Answering (QA) returns concise answers or answer lists from natural language text given a c

Tigrinya Question-Answering Dataset (TiQuAD)

TiQuAD is a human-annotated question-answering dataset for the Tigrinya language. The dataset contai