Evaluating Socratic Prompting Strategies to Improve Student Python Code Understanding

Authors

  • Timothy Guo Oakton High School, Oakton, VA
  • Linda Li Thomas Jefferson High School for Science and Technology, Alexandria, VA
  • Nihanth Tatikonda Bothell High School, Bothell, WA
  • Avani Thakur Saint Francis High School, CA
  • Harihar Vadrevu Tom Glenn High School, Leander, TX
  • Mihai Boicu Personalized Learning with Artificial Intelligence Technologies (PLAIT) Lab, Department of Information Sciences and Technology, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5626

Abstract

Large language models (LLMs) are increasingly used as tutors for programming education, but their value depends on how effectively they guide learners' reasoning rather than providing coding solutions. Socratic prompting uses questioning to support step-by-step discovery, however there is limited research comparing existing Socratic prompting techniques on how they improve learners’ code understanding. Our mini-experiment compared five Socratic tutoring strategies—guided discovery, assumption checking, comparative questioning, sequential scaffolding, and clarification questions—on two CRUXEval Python problems representing beginner and intermediate difficulty across three AI Assistants (AIAs)—ChatGPT 5.5 Instant, Gemini 3.6 Flash, and Claude Sonnet 5. Each strategy was tested using a mixed learner profile, and the resulting conversations were graded with a common rubric, evaluating code behavior correctness, explanation correctness, completeness, conciseness, and specificity. Problem ID significantly affected rubric scores (p = .014), indicating that performance varied depending on the Python problem. The prompting strategy also had a significant overall effect (p = .015). Comparative Questioning achieved the highest mean score but exceeded the next-highest strategy by only 0.066 points, and no pairwise differences remained statistically significant after correction for multiple comparisons. Neither AI assistant (p=0.161) nor the prompting-strategy-by-assistant interaction (p=0.477) significantly affected scores. Additionally, inter-rater reliability was low (ICC = 0.133), so the experiment is being revised using standardized learner-response scripts and clearer problem-specific scoring criteria. Our findings suggest that standardized learner interactions and more reliable evaluation methods are necessary to determine which Socratic prompting technique most effectively supports Python code understanding.

Published

2026-09-24

Issue

Section

College of Engineering and Computing: Department of Information Sciences and Technology