
New Benchmark HypoArena Tests LLMs on Prospective Hypothesis Discovery from Incomplete Data
Researchers introduce Prospective Hypothesis Discovery (PHD) and the HypoArena benchmark (988 cases) to evaluate how well large language models can autonomously generate grounded, discriminative, and testable hypotheses from anomalous observations and fragmented records.