Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2096
Title: Assistive technologies for education
Authors: Goel, Ansh
Vashishtha, Aryan
Goyal, Vikram (Advisor)
Mohania, Mukesh (Advisor)
Keywords: Large Language Models
Automated Question Generation
Pedagogical Scaffolding
Issue Date: 2-Dec-2024
Publisher: IIIT-Delhi
Abstract: In Indian physics classrooms, instructors often identify when a student is struggling, guessing, or applying a concept incorrectly. Rather than relying solely on final answers, teachers pose targeted sub-questions to diagnose conceptual gaps and evaluate reasoning. This diagnostic style of ques tioning is central to JEE preparation but is largely inaccessible to self-learners. Instruction-tuned open-source Large Language Models (LLMs)—Llama3-8B-Instruct, Gemma2-9B-IT, Mistral-7B Instruct-v0.3, and Qwen-2.5-7B-Instruct—offer a potential way to automate such instructor-style probing, though their pedagogical reliability remains untested. This study systematically bench marks four LLMs on a diversified JEE Physics dataset using strict prompting constraints. Our hybrid evaluation pipeline combines embedding-based redundancy detection, taxonomy-based rele vancy scoring, LLM-as-judge assessments, and extensive human annotation to measure redundancy, coverage of key concept, and relevance. Our results show a consistent performance hierarchy. Based on manual scores across 180 problems, Qwen-2.5-7B outperforms all models, achieving the highest relevancy (87.35%), best redundancy score (0.9722), and strongest conceptual coverage (2.5603 on a 0–3 scale). Gemma follows with moderate relevancy (75.42%) and coverage (1.9922), while Mistral and Llama lag behind in both semantic accuracy and coverage depth. A deeper meta-evaluation of the judges reveals a Complexity–Reliability Gap: small models are nearly perfect judges for simple tasks (e.g., Mistral MAE = 0.015 for redundancy) but fail dramatically on semantic tasks, showing high error (≥0.24) and systematic negative bias as complexity increases. Overall, this benchmark offers the first evidence that (1) Qwen is the strongest generator for JEE-style diagnostic ques tions, and (2) judge selection must be task-specific, as no single open-source LLM reliably evaluates complex pedagogical outputs.
URI: http://repository.iiitd.edu.in/xmlui/handle/123456789/2096
Appears in Collections:Year-2025

Files in This Item:
File Description SizeFormat 
IIITD_BTP_Report - Aryan Vashishtha.pdf
  Restricted Access
1.25 MBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.