
UnpredictaBench: Testing AI's Ability to Capture Real-World Distributional Randomness
Researchers introduced a new benchmark, UnpredictaBench, to evaluate whether large language models (LLMs) can capture true underlying distributions rather than collapsing to a single plausible answer. This is critical as AI is increasingly used as a substitute for real entities in economic simulations and other modeling tasks.
