DeepSeek R3 Achieves State-of-the-Art Reasoning on Math and Code Benchmarks
DeepSeek's latest model, R3, has set new records on AIME and HumanEval, outperforming GPT-5 and Gemini 2.5 Pro on multi-step reasoning tasks while remaining fully open-weights.