
KWBench: Evaluating LLMs' Ability to Recognize Professional Scenarios
Researchers introduce KWBench, a new benchmark for assessing whether large language models can identify professional scenarios without explicit prompting. This focuses on a critical yet often overlooked step in knowledge work: recognizing the structure of a situation before attempting to solve it.
