
DLR Framework Boosts VLM Reasoning by Preserving Visual Latents
Researchers introduce Decompose, Look, and Reason (DLR), a new framework that solves visual information loss in Vision-Language Models by using continuous visual latents instead of textual chains of thought. This approach dynamically decomposes queries and grounds reasoning in visual data, outperforming existing patch-based methods.