Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2096
Full metadata record
DC FieldValueLanguage
dc.contributor.authorGoel, Ansh-
dc.contributor.authorVashishtha, Aryan-
dc.contributor.authorGoyal, Vikram (Advisor)-
dc.contributor.authorMohania, Mukesh (Advisor)-
dc.date.accessioned2026-09-05T09:37:20Z-
dc.date.available2026-09-05T09:37:20Z-
dc.date.issued2024-12-02-
dc.identifier.urihttp://repository.iiitd.edu.in/xmlui/handle/123456789/2096-
dc.description.abstractIn Indian physics classrooms, instructors often identify when a student is struggling, guessing, or applying a concept incorrectly. Rather than relying solely on final answers, teachers pose targeted sub-questions to diagnose conceptual gaps and evaluate reasoning. This diagnostic style of ques tioning is central to JEE preparation but is largely inaccessible to self-learners. Instruction-tuned open-source Large Language Models (LLMs)—Llama3-8B-Instruct, Gemma2-9B-IT, Mistral-7B Instruct-v0.3, and Qwen-2.5-7B-Instruct—offer a potential way to automate such instructor-style probing, though their pedagogical reliability remains untested. This study systematically bench marks four LLMs on a diversified JEE Physics dataset using strict prompting constraints. Our hybrid evaluation pipeline combines embedding-based redundancy detection, taxonomy-based rele vancy scoring, LLM-as-judge assessments, and extensive human annotation to measure redundancy, coverage of key concept, and relevance. Our results show a consistent performance hierarchy. Based on manual scores across 180 problems, Qwen-2.5-7B outperforms all models, achieving the highest relevancy (87.35%), best redundancy score (0.9722), and strongest conceptual coverage (2.5603 on a 0–3 scale). Gemma follows with moderate relevancy (75.42%) and coverage (1.9922), while Mistral and Llama lag behind in both semantic accuracy and coverage depth. A deeper meta-evaluation of the judges reveals a Complexity–Reliability Gap: small models are nearly perfect judges for simple tasks (e.g., Mistral MAE = 0.015 for redundancy) but fail dramatically on semantic tasks, showing high error (≥0.24) and systematic negative bias as complexity increases. Overall, this benchmark offers the first evidence that (1) Qwen is the strongest generator for JEE-style diagnostic ques tions, and (2) judge selection must be task-specific, as no single open-source LLM reliably evaluates complex pedagogical outputs.en_US
dc.language.isoen_USen_US
dc.publisherIIIT-Delhien_US
dc.subjectLarge Language Modelsen_US
dc.subjectAutomated Question Generationen_US
dc.subjectPedagogical Scaffoldingen_US
dc.titleAssistive technologies for educationen_US
dc.typeOtheren_US
Appears in Collections:Year-2025

Files in This Item:
File Description SizeFormat 
IIITD_BTP_Report - Aryan Vashishtha.pdf
  Restricted Access
1.25 MBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.