AM-Bench Evaluates 2D Spatial Reasoning in Text-Only Models
September 1, 2026
The Autoregressive Mosaics benchmark differentiates between a model's ability to translate spatial descriptions into code and its actual internal 2D spatial reasoning. Tests on eight open-weight models show that layout composition performance varies significantly even when code-generation abilities are equal.
HOW THIS AFFECTS YOU
●
researcherYou should use this benchmark to distinguish between syntactic code proficiency and true spatial cognitive capabilities in LLMs.