Evaluating the Instructional Quality of LLM Across Prompting Strategies for Explaining Computer Science Concepts

Authors

  • Meghna Gomatam Wakeland High School, Frisco, TX
  • Achyut Nuli Bridgewater-Raritan Regional High School, Bridgewater, NJ
  • Neel Bhaskar Thomas Jefferson High School for Science and Technology, Alexandria, VA
  • Mihai Boicu Department of Information Sciences and Technology, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5629

Abstract

Large language models (LLMs) offer immediate programming assistance, but overreliance risks reducing critical thinking and problem-solving skills. While research highlights AI’s educational benefits and coding accuracy, little work evaluates how prompting strategies impact the instructional quality of AI-generated explanations. This study compares the explanation quality of four LLMs (ChatGPT 5.5, Gemini 3.1 Pro, Claude Opus 5, and Grok 4.5) across six prompting strategies, three single turn: direct answers, step-by-step teaching, analogy-based, and three multi-turn: Socratic, hint-based, and interactive tutoring, for three learning objectives (binary search, linear search and depth first search). The responses were evaluated by three independent reviewers and scored using a weighted rubric assessing pedagogical quality through technical accuracy, completeness, concept coverage, clarity, organization, test cases, and time and space complexity. Statistical analysis examined the impact of prompting strategy, model type, and their interaction on quality scores, showing that Step by Step Teaching achieved the highest mean score (85.94), substantially outperforming Direct Answer prompting (61.38) while multi-turn prompting consistently improved response quality across all evaluated models. Across models, ChatGPT 5.5 attained the highest overall average score (78.39), closely followed by Claude Opus 5 (78.02). By establishing a standardized evaluation framework, this study emphasizes how prompting guidelines can transform AI-assistants into effective instructional tools. Additional research will extend this framework to other STEM disciplines, directly assess student learning outcomes, and incorporate newer LLM models.

Published

2026-09-24

Issue

Section

College of Engineering and Computing: Department of Information Sciences and Technology