
New Study Reveals AI Struggles to Grade Human Math Reasoning Like Teachers
A new benchmark called RealMath-Eval shows that even the best AI models can't reliably grade real student math work. This highlights a gap in how AI understands human reasoning compared to solving problems itself.
