Please use this identifier to cite or link to this item:
http://repository.iiitd.edu.in/xmlui/handle/123456789/2096| Title: | Assistive technologies for education |
| Authors: | Goel, Ansh Vashishtha, Aryan Goyal, Vikram (Advisor) Mohania, Mukesh (Advisor) |
| Keywords: | Large Language Models Automated Question Generation Pedagogical Scaffolding |
| Issue Date: | 2-Dec-2024 |
| Publisher: | IIIT-Delhi |
| Abstract: | In Indian physics classrooms, instructors often identify when a student is struggling, guessing, or applying a concept incorrectly. Rather than relying solely on final answers, teachers pose targeted sub-questions to diagnose conceptual gaps and evaluate reasoning. This diagnostic style of ques tioning is central to JEE preparation but is largely inaccessible to self-learners. Instruction-tuned open-source Large Language Models (LLMs)—Llama3-8B-Instruct, Gemma2-9B-IT, Mistral-7B Instruct-v0.3, and Qwen-2.5-7B-Instruct—offer a potential way to automate such instructor-style probing, though their pedagogical reliability remains untested. This study systematically bench marks four LLMs on a diversified JEE Physics dataset using strict prompting constraints. Our hybrid evaluation pipeline combines embedding-based redundancy detection, taxonomy-based rele vancy scoring, LLM-as-judge assessments, and extensive human annotation to measure redundancy, coverage of key concept, and relevance. Our results show a consistent performance hierarchy. Based on manual scores across 180 problems, Qwen-2.5-7B outperforms all models, achieving the highest relevancy (87.35%), best redundancy score (0.9722), and strongest conceptual coverage (2.5603 on a 0–3 scale). Gemma follows with moderate relevancy (75.42%) and coverage (1.9922), while Mistral and Llama lag behind in both semantic accuracy and coverage depth. A deeper meta-evaluation of the judges reveals a Complexity–Reliability Gap: small models are nearly perfect judges for simple tasks (e.g., Mistral MAE = 0.015 for redundancy) but fail dramatically on semantic tasks, showing high error (≥0.24) and systematic negative bias as complexity increases. Overall, this benchmark offers the first evidence that (1) Qwen is the strongest generator for JEE-style diagnostic ques tions, and (2) judge selection must be task-specific, as no single open-source LLM reliably evaluates complex pedagogical outputs. |
| URI: | http://repository.iiitd.edu.in/xmlui/handle/123456789/2096 |
| Appears in Collections: | Year-2025 |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| IIITD_BTP_Report - Aryan Vashishtha.pdf Restricted Access | 1.25 MB | Adobe PDF | View/Open Request a copy |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.