PRO-Step: Step-Level Process Reward Optimization Improves Multi-Hop Reasoning in RAG Models
Researchers introduced PRO-Step (Step-level Process Reward Optimization), a method that improves multi-hop reasoning in Retrieval-Augmented Generation (RAG) models by rewarding each intermediate reasoning step rather than just the final answer, reducing error propagation.