Why Reinforcement Learning Beats Fine-Tuning in Math Problem-Solving
A new arXiv study finds that AI models trained with reinforcement learning (RL) develop superior internal representations for mathematical reasoning compared to supervised fine-tuning (SFT) models, explaining their better performance on math problems.