Research

Research News

1035 stories curated by AInformed · page 43 of 44

Rethinking Generalization in Reasoning SFT
research

Rethinking Generalization in Reasoning SFT

Researchers challenge the notion that supervised finetuning memorizes while reinforcement learning generalizes. They find that cross-domain generalization is conditional, influenced by optimization, data, and model capability. This challenges prevailing narratives in LLM post-training.

research

Pramana Fine-Tunes LLMs

Pramana is a novel approach that teaches large language models explicit epistemological methods to improve their reasoning. This approach aims to address the epistemic gap in AI, where models struggle with systematic reasoning and often produce unfounded claims.