
LLMs Struggle with Abstract Meaning Comprehension More Than Expected
A new study reveals that large language models, including GPT-4o, perform poorly on abstract meaning comprehension tasks. The findings highlight significant challenges in interpreting non-concrete, high-level semantics.