
HeuristicEdu: New Training Method Aligns Qwen2.5-7B AI to Teach Like Socrates Instead of Giving Answers
Researchers developed HeuristicEdu, a two-phase pipeline using supervised warm-up and GRPO to train the Qwen2.5-7B model as a Socratic guide. The method uses 797 multi-turn Chinese children's science dialogues to reward cognitive engagement over direct answers, addressing the problem of LLMs acting as direct answerers in education.