What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit
Researchers have identified that performance benchmarks for vision-language models are significantly influenced by how the models are prompted to output their answers. While these tests are often assumed to be neutral, the study shows that a model's perceived capability limit frequently reflects its ability to process specific formatting conventions rather than its actual visual reasoning skill.
Covered by 1 source
- AarXiv CS.AI↗Alfredo F. Frontera Del Valle10h ago