StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

Authors

  • Michelle Lin Department of Computer Science, George Mason University, Fairfax, VA
  • Rui Yu Department of Computer Science, George Mason University, Fairfax, VA
  • Mingrui Liu Department of Computer Science, George Mason University, Fairfax, VA

DOI:

https://doi.org/10.13021/jssr2026.5577

Abstract

Vision-language models are often evaluated on broad multimodal benchmarks, but their ability to recover latent spatial state from clean visual inputs remains undercharacterized. We introduce StateSight, a synthetic diagnostic benchmark for spatial-state reconstruction across three single-image tasks: cube-net opposite-face prediction, occlusion-aware cube-tower counting, and 4-neighbor connected-component counting. Each task contains 300 procedurally generated examples with deterministic ground truth, text-free images, and exact-match answer parsing. Under direct answer-only prompting, GPT-5.5 achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%, with zero output-format failures. On sampled items, a 30-participant human baseline achieved 80.8%, 68.8%, and 64.3%, outperforming both models on every task and revealing human-model gaps of up to 57.0 percentage points. Visible-derivation runs and conservative response-level coding indicate recurring failures in visual-state reconstruction, reasoning-strategy selection, and occasional fabricated visual facts. StateSight provides a controlled stress test for VLM spatial inference beyond object recognition and formatting compliance.

 

Published

2026-09-24

Issue

Section

College of Engineering and Computing: Department of Computer Science