Show, Don’t Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Published in arXiv, 2026
ProVisE changes the response interface of existing spatial benchmarks without changing their task semantics or metrics. A task-aware router selects a fixed visual protocol, the model expresses its answer in pixel space, and a deterministic parser converts the result back into the benchmark’s structured answer format.
SpatialGen-Bench contains 470 curated samples spanning perception, understanding, reasoning, and interaction.
Recommended citation: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, and Xuhong Zhang. (2026). Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text. arXiv:2607.21072.
Download Paper
