A Bangla Grammar Question Answering Dataset with Prerequisite Knowledge, Multi-Stage Hints, and Hint Convergence Scores

Published: 3 September 2026| Version 1 | DOI: 10.17632/tsdvtzwwpn.1
Contributors:
,
,

Description

This dataset contains 10,000 manually curated Bangla grammar question-answering instances. It was developed for hint-based learning, intelligent tutoring systems, adaptive and curriculum-aware learning, knowledge tracing, educational NLP, and LLM evaluation. All questions, grammar concepts, reference answers, explanations, and instructional content were collected from the official SSC textbook "বাংলা ভাষার ব্যাকরণ ও নির্মিতি" ("Bangla Bhashar Byakaran O Nirmity"), published by the National Curriculum and Textbook Board (NCTB), Bangladesh. The source material was manually curated into structured QA instances. Each record contains a Bangla grammar question, reference answer, explanation, topic, subtopic, difficulty, prerequisite concepts, semantic tags, five progressive hints, and stage-wise hint convergence scores. Hints progress from general contextual guidance to near-explicit support. Each 0–1 convergence score represents intended guidance toward the correct answer: lower scores are broader and higher scores are more specific. Scores reflect pedagogical progression, not empirical probabilities or model-confidence values. The dataset is distributed in UTF-8 JSON Lines (.jsonl): full.jsonl (10,000 instances), train.jsonl (70%), valid.jsonl (15%), test.jsonl (15%), and sample.jsonl (20 instances). Each record includes: id, topic, subtopic, difficulty, question, prerequisites, hints, convergence, candidate_answers, exact_answer, explanation, and topic_tags. Example Record: { "id": "bn_grammar_05900", "topic": "কারক ও বিভক্তি", "subtopic": "বিভক্তির প্রকারভেদ ও প্রয়োগ", "difficulty": "medium", "question": "বাক্যে ব্যবহৃত একটি পদের সাথে অন্য পদের অর্থগত সম্পর্ককে অর্থবহ করতে কোন ব্যাকরণিক উপাদানের ভূমিকা প্রধান?", "prerequisites": ["বিভক্তি", "বাক্যতত্ত্ব"], "hints": [ "শব্দের শেষে যুক্ত হয়ে পদের অন্বয় রক্ষা করে।", "বাক্যের কারক সম্পর্কের মূল ভিত্তি এটি।", "বিভক্তি নির্দেশক ব্যাকরণিক উপাদান।", "বিভক্তি রূপ নির্দেশক বিকল্প।", "বিভক্তি নির্দেশক সঠিক উত্তর।" ], "convergence": { "1": 0.35, "2": 0.55, "3": 0.74, "4": 0.88, "5": 0.97 }, "candidate_answers": ["বিভক্তি", "উপসর্গ", "সন্ধি", "সমাস"], "exact_answer": "বিভক্তি", "explanation": "বাক্যস্থিত পদসমূহের পারস্পরিক অন্বয় বা অর্থগত সম্বন্ধ স্পষ্ট ও অর্থপূর্ণ করতে শব্দবিভক্তি ও ক্রিয়াবিভক্তির ভূমিকাই প্রধান।", "topic_tags": ["Bangla Grammar", "কারক ও বিভক্তি", "বিভক্তি"] } Records were reviewed for completeness, annotation consistency, factual accuracy, and valid JSON formatting before release. The dataset supports controlled studies of how progressively revealed hints influence learning and answer selection. It is suitable for evaluating AI models that generate, rank, or select pedagogically useful hints at different levels of instructional support. This dataset is intended for research and educational purposes, particularly for building progressive, adaptive, and concept-aware hint-based systems for Bangla grammar learning.

Files

Steps to reproduce

1. Collect Bangla grammar concepts, examples, and instructional content from the NCTB SSC textbook *বাংলা ভাষার ব্যাকরণ ও নির্মিতি*. 2. Create and manually curate grammar questions covering diverse topics, subtopics, and difficulty levels. 3. Assign prerequisite concepts, topic tags, candidate answers, correct answers, and explanations for each question. 4. Develop five progressively informative hints for every question, ordered from general guidance to near-explicit support. 5. Assign stage-wise hint convergence scores on a 0–1 scale and validate all questions, answers, hints, and annotations through consistency checking and error correction. 6. Structure the final dataset in UTF-8 encoded JSON Lines (`.jsonl`) format, including training, validation, test, full, and sample files.

Institutions

Categories

Linguistics, Computer Science, Artificial Intelligence, Computational Linguistics, Data Science, Natural Language Processing

Licence