Paper deep dive
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, Hokin Deng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/13/2026, 12:30:50 AM
Summary
The paper introduces ConservationBench, a benchmark designed to evaluate whether Vision Language Models (VLMs) can reason about physical transformations and conservation principles. Testing 112 VLMs across 23,040 trials, the study reveals that models fail to maintain transformation-invariant representations, often relying on superficial heuristics or textual priors rather than genuine physical reasoning. Performance remains near chance, and models exhibit systematic biases that fail to distinguish between conserving and non-conserving scenarios.
Entities (4)
Relation Signals (2)
ConservationBench â evaluates â Vision-Language Models
confidence 100% ¡ We introduce ConservationBench evaluating conservation... across 112 VLMs.
Vision-Language Models â failstoreasonabout â Physical Transformation
confidence 95% ¡ These findings show that current VLMs fail to maintain transformation-invariant representations of physical properties across dynamic scenes.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains unclear. We introduce ConservationBench evaluating conservation -- whether physical quantities remain invariant under transformations. Spanning four properties with paired conserving/non-conserving scenarios, we generate 23,040 questions across 112 VLMs. Results reveal systematic failure: performance remains near chance with improvements on conservation tasks accompanied by drops on controls. Control experiments show strong textual priors favoring invariance, yet models perform worse with visual content. Neither temporal resolution, prompting, nor curated sampling helps. These findings show that current VLMs fail to maintain transformation-invariant representations of physical properties across dynamic scenes.
Tags
Links
- Source: https://arxiv.org/abs/2603.07109v1
- Canonical: https://arxiv.org/abs/2603.07109v1
Trouble viewing inline? Open PDF directly â
Full Text
78,650 characters extracted from source content.
Expand or collapse full text
Vision Language Models Cannot Reason About Physical Transformation Dezhi Luo * 1 Yijiang Li * 2 Maijunxian Wang 3 Tianwei Zhao 4 Bingyang Wang 5 Siheng Wang 6 Pinyuan Feng 7 Pooyan Rahmanzadehgervi 8 Ziqiao Ma 1 Hokin Deng 9 Abstract Understanding physical transformations is funda- mental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains unclear. We introduce Conservation- Bench evaluating conservationâwhether phys- ical quantities remain invariant under transfor- mations. Spanning four properties with paired conserving/non-conserving scenarios, we gener- ate 23,040 questions across 112 VLMs. Results reveal systematic failure: performance remains near chance with improvements on conservation tasks accompanied by drops on controls. Control experiments show strong textual priors favoring invariance, yet models perform worse with visual content. Neither temporal resolution, prompting, nor curated sampling helps. These findings show that current VLMs fail to maintain transformation- invariant representations of physical properties across dynamic scenes. 1. Introduction RecentadvancesinVisionLanguageModels (VLMs) (Zhang et al., 2024c; Radford et al., 2021; Alayrac et al., 2022; Li et al., 2023) have demonstrated remarkable capabilities of perception (Wang et al., 2024; Chen et al., 2025; Jiang et al., 2025; Team et al., 2025; Cheng et al., 2024b), reasoning (Zhang et al., 2024b; Xu et al., 2024; Cheng et al., 2024a), and visual commonsense understanding (Zellers et al., 2019; Park et al., 2020). These capabilities hold promise for real-world applications (Brohan et al., 2023), particularly in embodied tasks (Driess et al., 2023; Nasiriany et al., 2024) that demand * Equal contribution 1 University of Michigan 2 University of California San Diego 3 University of California Berkeley 4 Johns Hopkins University 5 Emory University 6 University of Toronto 7 Brown University 8 Auburn University 9 Carnegie Mellon Univer- sity. Correspondence to: Dezhi Luo<ihzedoul@umich.edu>, Yijiang Li <yijiangli@ucsd.edu>. Preprint. a genuine understanding of the physical world and its underlying properties (Chow et al., 2025b; Gao et al., 2024a). Yet it remains unclear whether VLMs possess a true understanding of physical principles or the capacity to operate reliably in embodied physical environments. A key factor in human intelligence that enables successful navigation in an embodied, physically grounded world is the ability to understand and reason about physical transfor- mations (Piaget, 1950; 1952; 1965; Baillargeon et al., 1985; 1990; Baillargeon, 1987; 1986; Spelke et al., 1992; Bail- largeon & Carey, 2012; Bear et al., 2021; Piloto et al., 2022). This capacity includes tracking objects over time (Spelke et al., 1994; 1995), managing occlusions (Gredeb Ě ack & von Hofsten, 2004), and adapting to dynamic environ- ments (Allen et al., 2020). While there are benchmarks evaluating physically plausible video generation (Motamed et al., 2025; Meng et al., 2024; Yang et al., 2025; Liu et al., 2025; Shi et al., 2024) and physical understanding in VLMs, spanning from everyday scenes (Zheng et al., 2024; Chow et al., 2025a) to high-school physics questions (Wang et al., 2025) and Olympiad-level problems (Qiu et al., 2025; Wang et al., 2025), these efforts focus either on video generation or physical properties in static scenes, leaving underexplored whether VLMs can genuinely reason about physical transfor- mationsâwhere specific properties may or may not remain invariant. To bridge this gap, we evaluate conservation in VLMsâthe understanding that physical quantities remain invariant un- der transformation despite changes in appearance. Here, physical quantity refers to the measurable magnitude of objects along certain dimensions, while spatial transforma- tion denotes the continuous process through which objects change in appearance or position. For example, an agent demonstrating conservation would recognize that pouring water into a differently shaped glass does not alter its vol- ume, despite the change in visible form. Achieving con- servation thus requires more than linguistic knowledge of quantity: it demands a systematic understanding that is both reversible and grounded in visual as well as conceptual representations. We introduce ConservationBench, a cogni- tively grounded benchmark for evaluating whether VLMs can reason about physical transformations. The benchmark consists of 192 video-based tasks across four core quanti- 1 arXiv:2603.07109v1 [cs.AI] 7 Mar 2026 Vision Language Models Cannot Reason About Physical Transformation Is the number of coins in the upper row the same as in the lower row in the final image? Number Is the length of the upper straw the same as the length of the lower in the final image? Length Is the size of the playdough in the first image the same as in the final image? Size Is the amount of liquid in the left glass in the first image the same as in the right in the final image? Volume ConservingQuestion Task Non-conserving Figure 1. Illustrative Tasks and Frame Selection Pipeline in Conservation Bench. tative properties (number, length, volume, and size), each requiring models to judge whether a quantity is conserved despite visual transformations. To control for shortcut ex- ploitation, we include 192 matched non-conserving controls where the target quantity changes while irrelevant features remain constant. We systematically vary frame extraction method, temporal resolution, and prompting strategy, yield- ing 60 conditions and 23,040 total trials. Evaluating 112 VLMs, we find that models consistently fail to integrate temporal information to track conserved properties across dynamic scenes. High accuracy on con- servation tasks is often driven by default heuristics, which reverse in non-conserving scenarios, revealing brittle, non- generalizable reasoning. Furthermore, prompting with cues encouraging transformation reasoning or providing higher temporal resolution does not help. These findings expose a fundamental limitation in current VLMs and underscore the need for more grounded, temporally-aware models capable of systematic physical inference. 2. Related Works Evaluating VLMs.Early benchmarking efforts relied on single-task benchmarks such as VQA (Antol et al., 2015), OK-VQA (Marino et al., 2019), and OCR (Liu et al., 2023). However, with the emergence of VLMs that claim broader perceptual and reasoning abilities, evaluation has shifted to- ward holistic benchmarks such as MMMU (Yue et al., 2024), SEED-Bench (Li et al., 2024), and MMBench (Liu et al., 2024). A growing line of benchmarks focuses specifically on quantity understanding (Rane et al., 2024; Paiss et al., 2023; Rahmanzadehgervi et al., 2024; Yuksekgonul et al., 2022). These tasks typically assess a modelâs ability to in- dividuate and count discrete objects in static scenes. While useful, such evaluations largely reduce to surface-level enu- meration and do not test whether models encode numerical invarianceâinvariance of quantity across transformations. In contrast, our work examines whether VLMs go beyond perceptual counting to represent quantity as a conserved property. Physical Understanding and Conservation. Insights from cognitive science underscore conservation as a critical benchmark for systematic physical reasoning. First pro- posed by Piaget, success on conservation tasks has long been viewed as evidence of emerging mental operations (Piaget & Inhelder, 1969). Developmental studies show that solving these tasks requires constructing transformation-invariant representations while suppressing misleading perceptual cues (Goldin-Meadow & Beilock, 2010; Houd Ě e et al., 2011; Poirel et al., 2012). Behavioral and neurocognitive research further demonstrates that conservation performance depends on sensorimotor grounding and inhibitory control, high- lighting the embodied nature of transformation understand- ing (Beilock & Goldin-Meadow, 2010; Lozada & Carro, 2016). Conservation also builds on more rudimentary abili- ties such as object permanence and individuation, revealed through studies exploiting the tunnel effect and violation-of- expectation paradigms (Burke, 1952; Flombaum & Scholl, 2006; Noles et al., 2005; Scholl, 2007), which themselves provide essential foundations for robust physical reasoning. In this light, conservation is widely recognized as a fun- damental cognitive capacity for the higher-level physical reasoning needed to navigate dynamic, embodied environ- ments (Fodor, 1975; Baillargeon & Carey, 2012; Barsalou, 2020; Luo et al., 2025b). Recent studies have examined modelsâ abilities to reason about physical properties, causal interactions, and material dynamics (Chow et al., 2025a; Patel et al., 2022; Zheng et al., 2024; Li et al., 2025a). Growing evidence suggests that VLMs struggle with fundamental aspects of visual rea- soning and physical understanding (Campbell et al., 2024; Gao et al., 2024b; Sun et al., 2024; 2025; Gao et al., 2025; Schulze Buschoff et al., 2025; Buschoff et al., 2025), with 2 Vision Language Models Cannot Reason About Physical Transformation some work exploring how modular frameworks or synthetic training data might address these limitations (Balazadeh et al., 2024; 2025; Luo et al., 2025b). However, these efforts largely emphasize outcome prediction or descriptive infer- ence, without testing whether models recognize that certain properties remain invariant under transformation. In many cases, success appears to stem from outcome-based heuris- tics rather than structured mental operations (Newman et al., 2024; Isola et al., 2015). Consequently, it remains unclear whether current VLMs can genuinely integrate sequential evidence to track physical transformations while maintain- ing stable representations of underlying propertiesâa core cognitive capacity directly targeted by conservation tasks (Mitchell & Krakauer, 2023). 3. Experimental Design 3.1. Conservation Tasks To systematically measure the conservation ability of VLMsâ the understanding that specific physical properties remain in- variant under transformations despite changes in appearance, we construct a suite of conservation tasks in the form of videos that visually depict physical transformations across four fundamental quantitative properties. We illustrate con- servation tasks across each property in Figure 2, with full descriptions provided in Appendix B. Although the four conservation types probe distinct physical properties, the tasks follow a unified structure: a transition from an initial to a final state mediated by an observable transformation. Each video begins with an initial state, proceeds through a continuous transformation (e.g., pouring, spreading, flattening), and ends with a new state where the surface appearance of the object of interest is altered. This design mirrors real-world scenarios where physical reasoning depends on integrating perceptual evidence across time. Generalization across Task-irrelevant Features.To en- sure the robustness and generalizability of the conclusions drawn from our benchmark, we systematically vary key vi- sual parameters in each conservation task (Table 4). These parameters include object count, size, color, layout, con- tainer shape, and the direction of transformation. Each conservation property consists of 48 unique video instances of different configurations, resulting in a total of 192 videos. This controlled variation guarantees that the core conser- vation principle is preserved across a wide range of visual contexts, thus preventing models from relying on memo- rized templates or superficial cues. Transformation-mandatory vs. -helpful. Notably, con- servation tasks differ in how strongly they depend on observ- ing the transformation. We classify them into two categories: transformation-mandatory and transformation-helpful. In mandatory tasks (volume and size), witnessing the transfor- mation is essentialâfor instance, in volume conservation, seeing the liquid poured is necessary, since the final height alone is insufficient for judging quantity. In helpful tasks (number and length), correct judgments can still be made from the initial and final states, as the relevant quantity re- mains visually accessible despite superficial changes. This distinction enables a more diagnostic evaluation: models that excel on helpful but not on mandatory tasks may rely on static cues rather than forming internal representations of the process. To this end, we further curated a set of 96 tasks derived solely from the final frame of transformation-helpful tasks. Here, models are prompted to compare numbers and lengths directly based on simple counting and intuitive judgments of spatial extent. This design isolates pre-conceptual, rudi- mentary forms of quantitative assessmentâsuch as item enumeration and perceptual matchingâfrom the broader representational demands of transformation-based reason- ing. By contrasting performance on these static tasks with temporal conservation trials, we can reveal how basic quan- titative sensitivity relates to the more systematic representa- tions of quantity that underlie conservation reasoning. 3.2. Non-conserving Tasks A key limitation of applying conservation tasks to model evaluation is the uniformity of ground-truth labels: since all standard tasks involve quantity preservation, models can appear accurate simply by defaulting their responses to indi- cate invariance, due to biases from either visual contexts or linguistic patterns in the prompts, without genuinely reason- ing about the physical transformation itself (Li et al., 2025b). To address this, we create non-conserving counterfactuals as a set of controlled experiments where the quantity of interest is explicitly altered during the transformation without chang- ing the task-irrelevant features. That is to say, these manipu- lations are performed within the same environments, using identical object sets and visual contexts, thereby ensuring a controlled comparison. This design enables fine-grained as- sessment of model sensitivity to actual changes in quantity, rather than reliance on superficial heuristics or distributional priors. Details regarding control tasks across each property are available in Figure 2, with full descriptions provided in Appendix B. Following this design, we curated a control set in which each non-conserving control is paired with a conservation task under matched configurations, yielding an additional 192 videos. 3.3. Adaptation to Multi-frame Input Temporal Resolution The ability to understand physi- cal transformations critically depends on comprehending 3 Vision Language Models Cannot Reason About Physical Transformation dynamic processes over time. Unlike static snapshot reason- ing, robust comprehension requires recognizing continuity across successive observations. Human perception benefits from high frame rates (e.g.âź30-60 frames per second) that convey rich temporal information, while the architec- tural and computational limitations of VLMs restrict them to inferring such dynamics from discrete and often sparse inputs. To investigate the impact of temporal resolution on conservation understanding, we vary the number of frames extracted from each video: â˘3-frame condition: Only three frames are pro- videdâthe first, the last, and one intermediate frame. This condition presents minimal temporal information while retaining just enough cues for humans to solve the task. â˘5-, 7-, and 9-frame condition: More frames are sam- pled to offer moderate temporal granularity. This con- dition is designed to contrast qualitatively with the 3-frame condition by enabling multi-frame representa- tions of the temporally continuous scene. â˘16-frame condition: Sixteen frames are sampled to provide finer-grained temporal information, offering a more detailed depiction of the transformation process, contrasting quantitatively with the 8-frame condition. All conditions include initial and final states. This design tests whether models can leverage higher temporal reso- lution to extract transformation-relevant information for conservation reasoning. Sampling StrategyIn studying physical transformations, the sequence and selection of visual inputs are crucial. This raises an important question: do different frame selection strategies influence the modelâs understanding of dynamic scenes? Additionally, do humans and models rely on differ- ent criteria when identifying informative visual moments? To examine this, we implement and compare three frame extraction strategies, each reflecting distinct assumptions about what defines a ârepresentativeâ moment in a physical event. â˘Uniform Sampling: Frames are sampled uniformly across the timeline, serving as a baseline approach commonly used in prior work, based on the assumption that temporal regularity sufficiently represents infor- mational diversity. â˘Human-based: To obtain a baseline for human intu- ition in frame extraction, we recruited N = 18 annota- tors. Each annotator was randomly assigned a subset of the dataset and asked to manually select the interme- diate frames that captured the essential stages of the transformation. â˘Model-based: We adopt SEVILA (Yu et al., 2023) and leverage a BLIP-2-based Localizer to identify language-aware keyframes. Prompted with the same instruction assigned to humans (âextract the most com- plete set of frames that capture the entire processâ), the Localizer module selects frames with high relevance scores, which are then passed to the Answerer mod- ule for inference. This method formalizes a strategy akin to semantic salience: choosing frames that are maximally informative given a specific query. This design allows us to test whether different frame se- lection strategies affect model performance on physical transformation reasoning. We hypothesize that optimiz- ing frame selection, rather than merely increasing frame quantity, leads to more effective representations of dynamic events. We detail our data curation process in Appendix A and prompting strategies in Appendix C, and provide example input in Appendix D. Table 1. Overview of Multi-image Task Conditions and Evaluation Scale ComponentCount Core Dataset Conservation Tasks192 Non-conserving Control192 Total Videos384 Multi-frame Conditions (Factorial) Extraction Method3 Frame Count5 Prompting4 Factorial Combinations3Ă 5Ă 4 = 60 Total Evaluation Trials384Ă 60 = 23, 040 4. Experiments 4.1. Inference and Evaluation Inference. We evaluate 112 VLMs spanning diverse model architectures, training data, and parameter scales, cov- ering both mainstream commercial systems and advanced open-source models. To ensure fidelity, comparability, and reproducibility, we strictly adhere to reference configura- tions and implementations from the official codebases. Re- fer to Appendix E for further details. Evaluation. To evaluate free-form outputs of VLMs on multiple-choice questions (MCQs), we follow the two-stage scoring method of Li et al. (2025a). In Stage 1, each VLM output is mapped to a unique choice from the provided op- tions or labeled FAIL when no unambiguous mapping is possible. Mapping follows a hybrid strategy: deterministic template matching is applied first, and unresolved cases are adjudicated by an LLM-as-a-Judge constrained to the option set. Models exhibiting persistently high FAIL rates are ex- 4 Vision Language Models Cannot Reason About Physical Transformation Figure 2. Overall Performance on ConservationBench. A. Accuracy averaged across conservation tasks and non-conserving control compared to strict pairwise calculation (Top 30 models; full results available in Appendix H; B. Performance on non-conserving control in relation to conservation tasks. cluded from further analyses to avoid bias from nonsensical outputs. In Stage 2, the mapped option is compared against the ground-truth answer, with all FAILs scored as incorrect. Details are provided in Appendix F. 4.2. Human Baseline Given the large number of questions and the cost of human annotation, we curated a representative subset by randomly selecting one out of every eight task configurations for each quantitative property, counterbalanced across conservation tasks and non-conserving controls. This resulted in a total of 864 questions. We hypothesize that reduced variation in task-irrelevant features is unlikely to compromise the bench- markâs validity or generalizability given the robustness of human reasoning. Participants received the same stimuli and three-choice questions as the VLMs, with the excep- tion that they directly selected answers rather than requir- ing LLM judge parsing. The aggregated human accuracy reaches 98.35%, consistent with decades of developmental research showing that humans from late childhood reliably solve conservation tasks with near-perfect accuracy (Piaget, 1965; Houd Ě e, 1997; Pezzulo et al., 2013; Viarouge et al., 2019). These results validate our benchmark design and its adaptation for evaluating VLMs. Detailed breakdown and comparison with models are highlighted in Appendix H. 4.3. Main Results As shown in Figure 2A, model accuracy across 112 VLMs ranges from 20% to 69%, with most performing only marginally above the 33.3% chance level. In contrast, hu- man participants exceed 98% accuracy (Section 4.2), high- lighting a clear gap between VLMs and intuitive human reasoning. Collectively, these results reveal a core limita- tion: VLMs struggle to integrate temporal cues or track invariant properties through dynamic transformations, a key requirement for grounded physical reasoning. We report the performance of all models and human baseline in Appendix H. Non-Conserving Control Reveal Systematic Bias By comparing model performance on non-conserving control tasks against conservation tasks, we observe a moderate negative correlation (r = â0.510,n = 112): models that perform better on conservation tasks tend to perform worse on the corresponding control tasks, and vice versa (Fig- ure 2B). Most models cluster in the lower-right quadrant, exhibiting moderately high conservation accuracy (40â80%) but low non-conserving accuracy (10â40%), revealing a sys- tematic bias toward quantity invariance regardless of actual transformation evidence. Only a small subset of models ap- proaches balanced performance near the diagonal (y = x), while virtually none achieve high accuracy on both task types simultaneously. This pattern generalizes across four quantitative domains (see Appendix I for details). Crucially, this pattern reveals a diagnostic failure: models are not sim- ply underperforming but exhibiting asymmetric reliance on default heuristics that systematically reverse across matched task conditions, demonstrating an inability to flexibly adjust reasoning based on transformation evidence. We further validate this pattern using a strict pairwise eval- uation across the full set matched conservation and non- conserving control tasks (Figure 2A; labeled in purple). In this analysis, a model is marked correct only if it answers 5 Vision Language Models Cannot Reason About Physical Transformation both tasks in a pair correctlyâcapturing whether it can jointly recognize quantity preservation and detect meaning- ful violations under matched visual conditions. We find that most models (82/112, 73.2%) perform well below chance, achieving strict accuracy rates under 10%. Only three top- performing modelsâGEMINI-2.5-PRO, DOUBAO-SEED- 1.6-VISION, and CLAUDE-SONNET-4-5âexceed chance level (33.3%). This indicates that models are unable to re- liably distinguish between conserving and non-conserving scenarios. The gap between average accuracy and strick evaluations suggests that modelsâ success are driven largely by bias toward quantity variance or invariance rather than genuine reasoning about physical transformations. This finding further supports the conclusion that models fail to internalize structured physical reasoning and instead rely on brittle default strategies for quantity assessment. Dissociating Sources of Bias. To dissociate the source of biasâwhether it arises from visual features or textual priorsâwe conducted two control experiments on 62 VLMs that support both image and text-only inputs. First, we reran the same experiments using fully white, content-free im- ages while keeping all text input constant (Empty Image Control). Second, we removed visual input entirely, pre- senting only text prompts (Text Control). We used the 7-frame condition for all comparisons to enable direct pair- wise evaluation. Model responses were evaluated as if they were answering standard conservation tasks. If performance were driven purely by visual cues, models should operate at chance when visual content is removed. Conversely, sys- tematic deviations from chance would indicate reliance on textual biasesâfavoring either conservation (bias toward invariance) or non-conservation (bias toward perceptual change). ConserveNot Conserve (More) Not Conserve (Less) Empty Image Control Conserve Not Conserve (More) Not Conserve (Less) Conservation Task 85.7% n=65K 8.1% n=6K 6.3% n=5K 71.5% n=30K 17.1% n=7K 11.4% n=5K 75.4% n=18K 13.7% n=3K 10.9% n=3K ConserveNot Conserve (More) Not Conserve (Less) Text Control Conserve Not Conserve (More) Not Conserve (Less) Conservation Task 73.7% n=57K 11.3% n=9K 15.0% n=11K 69.1% n=29K 14.0% n=6K 16.9% n=7K 68.2% n=16K 13.8% n=3K 18.0% n=4K AB High Low Maintained Changed Figure 3. Response patterns under Empty Image and Text-only conditions compared to standard one. We report a distribution change in prediction between Empty Image and standard condition (A) to the left; Text-only control condition and standard condition (B) to the right. Figure 3 reveals striking patterns in how models respond when visual information is degraded or removed. In the Empty Image Control (Panel A), when the actual conserva- tion task required a âConserveâ answer, 85.7% of responses remained âConserveâ under empty images. Critically, when the actual task required âNot Conserveâ answers, models overwhelmingly switched to âConserveâ responsesâ71.5% for âMoreâ scenarios and 75.4% for âLessâ scenarios. The Text Control (Panel B) shows a similar but slightly attenu- ated pattern: 73.7% maintain âConserveâ answers, while 69.1% and 68.2% of non-conserving scenarios shift to âCon- serveâ responses. Further details regarding model-level trends in control condition responses compared to conserva- tion task performance are provided in Appendix J. These results reveal that textual priors strongly favor quan- tity invarianceâa bias that is correct for conservation tasks but incorrect for non-conserving controls, explaining the inverse correlation we observe between conservation and non-conserving task performance. Notably, removing vi- sual content while maintaining the visual modality (Empty Image) yields stronger conservation responses (85.7%) than removing the visual modality entirely (Text Control: 73.7%), suggesting that the presence of the visual channel amplifies textual biases even without meaningful visual information. Critically, however, models perform worse on actual con- servation tasks with real visual content (average accuracy âź60%) than with empty images (85.7%), indicating that visual content actively interferes with the correct textual prior. Rather than enhancing transformation reasoning, vi- sual information causes models to override their correct de- fault bias with faulty visual processing. This demonstrates that the core deficit lies in visual transformation reasoning: models cannot reliably extract and integrate transformation- relevant information from sequential visual evidence, lead- ing them to incorrectly reject quantity invariance even when visual content should confirm it. The combination of correct textual priors but impaired visual processing accounts for both the moderate success on conservation tasks and the systematic inverse failures on non-conserving controls. 4.4. Different Prompting Strategies, Frame Numbers, and Sampling Methods We further analyzed model performance across three experi- mental factorsâprompt type, frame count, and frame sam- pling methodâevaluated separately for Number & Length versus Volume & Size conservation tasks. To properly account for the hierarchical structure of our data (multi- ple observations nested within 112 models), we employed repeated-measures ANOVA with models as the unit of anal- ysis. For each model, we first averaged accuracy across the irrelevant experimental conditions (e.g., when testing frame number effects, we averaged across all prompt types and extraction methods), then conducted repeated-measures ANOVA to test main effects, followed by Bonferroni- corrected pairwise comparisons for significant factors. This approach accounts for the dependency structure in our data while avoiding inflated Type I error from treating non- independent observations as separate trials. We highlight the main conclusions below. 6 Vision Language Models Cannot Reason About Physical Transformation Figure 4. Model performance showing main effects by (A) prompt type, (B) number of frames, and (C) frame sampling method. Each panel averages across the other 2 factors from the full factorial design (4 prompts Ă 5 frame counts Ă 3 extraction methods). Continuous cues aid performance, CoT makes it worse (Figure 4A). For Number & Length tasks, prompt type shows a highly significant main effect (F(3, 333) = 18.28, p < 0.001). Bonferroni-corrected pairwise comparisons re- veal that CoT prompting performs significantly worse than all other prompt types: Continuous (p < 0.001), Sequential (p < 0.001), and Direct (p = 0.0022). Additionally, Con- tinuous promptsâwhich explicitly frame transformations as continuous processesâsignificantly outperform Direct questions (p = 0.0191). These results indicate that concep- tual cues emphasizing continuity can provide modest benefit for transformation-helpful tasks, while forcing step-by-step verbalization consistently impairs performance, likely by amplifying reliance on brittle heuristics. However, for Vol- ume & Size tasks, prompt type shows no significant effect (F(3, 333) = 2.00,p = 0.114), suggesting that linguistic scaffolding provides no benefit when transformation reason- ing demands are higher. Temporal resolution shows no reliable benefit across task types (Figure 4B). For Number & Length tasks, frame count shows no significant main effect (F(4, 444) = 0.98, p = 0.416), indicating that additional frames do not reliably improve performance even when transformation cues are helpful. For Volume & Size tasks, frame count shows a modest significant effect (F(4, 444) = 2.66,p = 0.032), but Bonferroni-corrected pairwise comparisons reveal only one significant difference: 7 frames outperform 9 frames (p corrected = 0.0329). This lack of consistent improvement with increased temporal information demonstrates that cur- rent VLMs are unable to effectively integrate sequential visual evidence for transformation reasoning. Additional frames do not enable models to track continuous physical changes, even when such tracking is essential for task suc- cess. Frame extraction shows significant task-dependent ef- fects (Figure 4C). For Number & Length tasks, extrac- tion method shows no significant effect (F(2, 222) = 1.36, p = 0.258), suggesting that different sampling methods perform comparably when transformation reasoning is help- ful but not mandatory. However, for Volume & Size tasks, extraction method shows a highly significant effect (F(2, 222) = 8.75,p = 0.0002). Bonferroni-corrected pair- wise comparisons reveal that uniform sampling significantly outperforms both human-selected (p corrected = 0.0006) and SeViLA-selected frames (p corrected = 0.0014), with no dif- ference between the two curated methods (p corrected = 1.0). These findings suggest that frame selection strategies in- teract weakly with task demands. For transformation- mandatory tasks (Volume & Size), curated frame selection even inadvertently emphasize misleading static features, suggesting that models are unable to utilize task-relevant bi- ases for reasoning, further demonstrating the lack of ability to understand physical transformations. 4.5. Does Scaling of Model Size Help? The advancement of LLMs has been closely tied to the em- pirical scaling lawâpredictable power-law improvements in performance with increased compute, parameters, and train- ing data (Kaplan et al., 2020; Henighan et al., 2020; Zhai et al., 2022)âas well as emergence, the abrupt appearance of qualitatively new abilities as models grow larger (Wei et al., 2022; Aghajanyan et al., 2023; Bubeck et al., 2023; Berti et al., 2025). This raises a natural question: Does the capacity to understand physical transformations and conservation similarly emerge with scale? To this end, we examine the performance v.s. model size (measured in log-scale parameters) across 112 VLMs, rang- ing from 1B to 76B parameters. We hold the frame number constant at 7 to control for potential confounding effects of multi-frame processing capacity. We find strikingly di- vergent patterns (Figure 7). For conservation tasks, model size exhibits virtually no predictive power (R 2 = 0.019). In contrast, non-conserving task accuracy exhibits a mod- erate positive relationship with model size (R 2 = 0.239, 7 Vision Language Models Cannot Reason About Physical Transformation Figure 5. Conservation reasoning does not emerge with model scale. Model performance on (left) conservation tasks shows no relationship with parameter count (R 2 = 0.019), while (right) non-conserving task accuracy exhibits only modest scaling effects (R 2 = 0.239), both evaluated at 7-frame condition across 112 VLMs. y = 10.48x + 17.81), indicating that larger models tend to do better on non-conserving controls. However, even this relationship accounts for less than 24% of the variance, with substantial scatter persisting across all model scales. These results demonstrate that conservation reasoning emerges only at a small scale in current VLMs. 5. Discussions We introduce a cognitively grounded benchmark evalu- ating whether VLMs can reason about physical transfor- mations through conservation tasks and non-conserving controls. Our findings reveal that current models consis- tently fail to integrate sequential visual evidence to maintain transformation-invariant representations of physical proper- ties across dynamic scenes. Control experiments reveal that models possess strong textual priors favoring quantity invari- ance yet perform worse with actual visual content, revealing reliance on brittle heuristics. Neither increased temporal resolution, targeted prompting, nor human-curated frame sampling induces robust transformation reasoning. These results expose a fundamental deficit in structured physical understanding and highlight critical challenges for develop- ing grounded AI systems capable of systematic inference in dynamic environments. Beyond documenting these failures, our benchmark provides an enduring diagnostic test as the field advances. The tasks curated in this study can serve as sanity checks for foundation modelsâ transformation rea- soning, particularly as models achieve breakthroughs on high-level physical reasoning benchmarks while robustness challenges persist. The failures documented in this work have direct implica- tions for understanding VLMsâ capacity for physical rea- soning and multi-frame understanding more broadly. Con- servation reflects a foundational cognitive substrate that scaffolds higher-level physical reasoning in humans. The failure modes exposed here resonate with recent work ar- guing that without such core representations, higher-level abilities remain brittle and fail to generalize beyond curated benchmarks. Models lacking transformation-invariant rea- soning about basic quantities pose risks for robust physical understanding in complex, real-world scenarios (Li et al., 2025a; Cai et al., 2025). Moreover, our results suggest that limitations in encoding and processing sequential visual- spatial information likely constitute a primary bottleneck for transformation reasoning. This speaks to broader con- cerns regarding reliance on coarse-grained visual encodings in VLMs, which may be mechanistically unsuitable for con- ducting structured physical reasoning (Zhang et al., 2024a; Luo et al., 2025a; Fu et al., 2025). While our behavioral findings do not establish the precise architectural causes, we highlight the need for mechanistic investigations to iden- tify the specific representational constraints prevent current models from constructing transformation-invariant object representations. 6. Conclusion and Limitation We introduce a comprehensive benchmark on physically grounded inference property conservation, the principle that physical quantities remain invariant under transformations despite appearance changes. We show that current VLMs consistently fail to maintain transformation-invariant rep- resentations of physical properties across dynamic scenes. These failures indicate fundamental deficits in systematic physical reasoning, posing risks for tasks such as embodied AI. Our benchmark provides an enduring diagnostic test for transformation reasoning as the field advances. We acknowledge several limitations in our study. First, our evaluation focuses on conservation tasks across four quantitative properties under controlled laboratory condi- tions. While targeting foundational cognitive capacities, more complex scenarios involving occlusions, deformable objects, noisy observations, and ambiguous transformations are planned for future work. Second, our study is primar- ily behavioral, documenting failure modes without estab- lishing precise mechanistic causes. While we hypothesize that coarse-grained visual encodings prevent transformation- invariant representations, mechanistic interpretability stud- ies to validate these claims are planned. Finally, our eval- uation focuses on perceptual judgment rather than goal- directed applications. Whether conservation deficits impair downstream tasks such as planning, tool use, or robotic manipulation remains for future investigation. Acknowledgments This work was supported by a MiraclePlus Com- pute Grant to Growing AI Like A Child (https:// growing-ai-like-a-child.github.io/).. 8 Vision Language Models Cannot Reason About Physical Transformation References Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Ham- bardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L. Scaling laws for generative mixed-modal language models. In International Confer- ence on Machine Learning, p. 265â279. PMLR, 2023. Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few- shot learning. Advances in neural information processing systems, 35:23716â23736, 2022. Allen, K. R., Smith, K. A., and Tenenbaum, J. B. Rapid trial- and-error learning with simulation supports flexible tool use and physical reasoning. Proceedings of the National Academy of Sciences, 117(47):29302â29310, 2020. Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, p. 2425â2433, 2015. Baillargeon, R. Representing the existence and the location of hidden objects: Object permanence in 6-and 8-month- old infants. Cognition, 23(1):21â41, 1986. Baillargeon, R. Young infantsâ reasoning about the physi- cal and spatial properties of a hidden object. Cognitive development, 2(3):179â200, 1987. Baillargeon, R. and Carey, S. Core cognition and beyond: The acquisition of physical and numerical knowledge. Early childhood development and later outcome, 1, 2012. Baillargeon, R., Spelke, E. S., and Wasserman, S. Object permanence in five-month-old infants. Cognition, 20(3): 191â208, 1985. Baillargeon, R., Graber, M., Devos, J., and Black, J. Why do young infants fail to search for hidden objects? Cognition, 36(3):255â284, 1990. Balazadeh, V., Ataei, M., Cheong, H., Hosein Khasahmadi, A., and Krishnan, R. G. Synthetic vision: Training vision- language models to understand physics. arXiv e-prints, p. arXivâ2412, 2024. Balazadeh, V., Ataei, M., Cheong, H., Khasahmadi, A. H., and Krishnan, R. G. Physics context builders: A modu- lar framework for physical reasoning in vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 7318â7328, 2025. Barsalou, L. W. Challenges and opportunities for grounding cognition. Journal of Cognition, 3(1), 2020. Bear, D. M., Wang, E., Mrowca, D., Binder, F. J., Tung, H.-Y. F., Pramod, R., Holdaway, C., Tao, S., Smith, K., Sun, F.-Y., et al. Physion: Evaluating physical prediction from vision in humans and machines. arXiv preprint arXiv:2106.08261, 2021. Beilock, S. and Goldin-Meadow, S.Gesture changes thought by grounding it in action. Psychological Sci- ence, 21(11):1605â1610, 2010. Berti, L., Giorgi, F., and Kasneci, G. Emergent abilities in large language models: A survey. arXiv preprint arXiv:2503.05788, 2025. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., Florence, P., Fu, C., Arenas, M. G., Gopalakr- ishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Leal, I., Lee, L., Lee, T.-W. E., Levine, S., Lu, Y., Michalewski, H., Mordatch, I., Pertsch, K., Rao, K., Reymann, K., Ryoo, M., Salazar, G., Sanketi, P., Sermanet, P., Singh, J., Singh, A., Soricut, R., Tran, H., Vanhoucke, V., Vuong, Q., Wahid, A., Welker, S., Wohlhart, P., Wu, J., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023. URL https://arxiv.org/abs/2307.15818. Bubeck, S., Chadrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lund- berg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023. Burke, L. On the tunnel effect. Quarterly Journal of Exper- imental Psychology, 4(3):121â138, 1952. Buschoff, L. M. S., Voudouris, K., Akata, E., Bethge, M., Tenenbaum, J. B., and Schulz, E. Testing the limits of fine- tuning to improve reasoning in vision language models. arXiv preprint arXiv:2502.15678, 2025. Cai, Z., Wang, Y., Sun, Q., Wang, R., Gu, C., Yin, W., Lin, Z., Yang, Z., Wei, C., Qian, O., et al. Holistic evaluation of multimodal llms on spatial intelligence. arXiv preprint arXiv:2508.13142, 2025. Campbell, D., Rane, S., Giallanza, T., De Sabbata, C. N., Ghods, K., Joshi, A., Ku, A., Frankland, S., Griffiths, T., Cohen, J. D., et al. Understanding the limits of vision language models through the lens of the binding problem. Advances in Neural Information Processing Systems, 37: 113436â113460, 2024. Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., Gu, L., Wang, X., Li, Q., Ren, Y., Chen, Z., Luo, J., Wang, J., Jiang, T., Wang, B., 9 Vision Language Models Cannot Reason About Physical Transformation He, C., Shi, B., Zhang, X., Lv, H., Wang, Y., Shao, W., Chu, P., Tu, Z., He, T., Wu, Z., Deng, H., Ge, J., Chen, K., Zhang, K., Wang, L., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W. Expanding performance boundaries of open-source multimodal mod- els with model, data, and test-time scaling, 2025. URL https://arxiv.org/abs/2412.05271. Cheng, K., Li, Y., Xu, F., Zhang, J., Zhou, H., and Liu, Y. Vision-language models can self-improve reasoning via reflection. arXiv preprint arXiv:2411.00855, 2024a. Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., Zhao, D., and Bing, L. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024b. URL https://arxiv.org/abs/2406.07476. Chow, W., Mao, J., Li, B., Seita, D., Guizilini, V., and Wang, Y. Physbench: Benchmarking and enhancing vision- language models for physical world understanding. arXiv preprint arXiv:2501.16411, 2025a. Chow, W., Mao, J., Li, B., Seita, D., Guizilini, V., and Wang, Y.Physbench: Benchmarking and enhanc- ing vision-language models for physical world under- standing, 2025b. URLhttps://arxiv.org/abs/ 2501.16411. Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P. Palm-e: An embodied multimodal language model, 2023. URL https://arxiv.org/abs/2303.03378. Flombaum, J. I. and Scholl, B. J. A temporal same-object ad- vantage in the tunnel effect: facilitated change detection for persisting objects. Journal of Experimental Psychol- ogy: Human Perception and Performance, 32(4):840, 2006. Fodor, J. A. The Language of Thought. MIT Press, 1975. Fu, S., Bonnen, T., Guillory, D., and Darrell, T. Hidden in plain sight: Vlms overlook their visual representations. arXiv preprint arXiv:2506.08008, 2025. Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Ma- jumdar, A., and Sadigh, D. Physically grounded vision- language models for robotic manipulation, 2024a. URL https://arxiv.org/abs/2309.02561. Gao, Q., Li, Y., Lyu, H., Sun, H., Luo, D., and Deng, H. Vision language models see what you want but not what you see. arXiv preprint arXiv:2410.00324, 2024b. Gao, Q., Pi, X., Liu, K., Chen, J., Yang, R., Huang, X., Fang, X., Sun, L., Kishore, G., Ai, B., et al. Do vision-language models have internal world models? towards an atomic evaluation. arXiv preprint arXiv:2506.21876, 2025. Goldin-Meadow, S. and Beilock, S. Actionâs influence on thought: the case of gesture. Perspectives on psychologi- cal science : a journal of the Association for Psychologi- cal Science, 5(6):664â674, 2010. Gredeb Ě ack, G. and von Hofsten, C. Infantsâ evolving repre- sentations of object motion during occlusion: A longitu- dinal study of 6- to 12-month-old infants. Infancy, 6(2): 165â184, 2004. doi: 10.1207/s15327078in06022. Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701, 2020. Houd Ě e, O. Numerical development: From the infant to the child. Cognitive Development, 12(3):373â391, 1997. Houd Ě e, O., Pineau, A., Leroux, G., Poirel, N., Perchey, G., Lano Ě e, C., Lubin, A., Turbelin, M.-R., Rossi, S., Simon, G., Delcroix, N., Lamberton, F., Vigneau, M., Wisniewski, G., Vicet, J.-R., and Mazoyer, B. Func- tional magnetic resonance imaging study of piagetâs conservation-of-number task in preschool and school-age children: a neo-piagetian approach. Journal of experi- mental child psychology, 110(3):332â346, 2011. Isola, P., Lim, J. J., and Adelson, E. H. Discovering states and transformations in image collections. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1383â1391, 2015. Jiang, Y., Wang, Y., Zhao, R., Parag, T., Chen, Z., Liao, Z., and Unnikrishnan, J.Videop2r: Video under- standing from perception to reasoning. arXiv preprint arXiv:2511.11113, 2025. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., and Shan, Y. Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13299â13308, 2024. Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. CONFERENCE, 2023. 10 Vision Language Models Cannot Reason About Physical Transformation Li, Y., Gao, Q., Zhao, T., Wang, B., Sun, H., Lyu, H., Hawkins, R. D., Vasconcelos, N., Golan, T., Luo, D., and Deng, H. Core knowledge deficits in multi-modal lan- guage models. arXiv preprint arXiv:2410.10855, 2025a. Li, Y., Wang, B., Zhao, T., Gao, Q., Deng, H., and Luo, D. Evaluating multi-modal language models through con- cept hacking. In Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions, 2025b. Liu, D., Zhang, J., Dinh, A.-D., Park, E., Zhang, S., Mian, A., Shah, M., and Xu, C. Generative physical ai in vision: A survey, 2025. URLhttps://arxiv.org/abs/ 2501.10928. Liu, Y., Li, Z., Yang, B., Li, C., Yin, X., Liu, C.-l., Jin, L., and Bai, X. On the hidden mystery of ocr in large multimodal models. arXiv preprint arXiv:2305.07895, 2023. Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player?In European conference on computer vision, p. 216â233. Springer, 2024. Lozada, M. and Carro, N. Embodied action improves cog- nition in children: Evidence from a study based on piage- tian conservation tasks. Frontiers in psychology, 7(393), 2016. Luo, D., Gao, Q., and Deng, H. Rethinking the simulation vs. rendering dichotomy: No free lunch in spatial world modelling. arXiv preprint arXiv:2510.20835, 2025a. Luo, D., Li, Y., and Deng, H. The philosophical foun- dations of growing ai like a child. arXiv preprint arXiv:2502.10742, 2025b. Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. Ok-vqa: A visual question answering benchmark requir- ing external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, p. 3195â3204, 2019. Meng, F., Liao, J., Tan, X., Shao, W., Lu, Q., Zhang, K., Cheng, Y., Li, D., Qiao, Y., and Luo, P. Towards world simulator: Crafting physical commonsense-based benchmark for video generation, 2024. URLhttps: //arxiv.org/abs/2410.05363. Mitchell, M. and Krakauer, D. C. The debate over under- standing in aiâs large language models. Proceedings of the National Academy of Sciences, 120(13):e2215907120, 2023. Motamed, S., Culp, L., Swersky, K., Jaini, P., and Geirhos, R. Do generative video models understand physical principles?, 2025. URLhttps://arxiv.org/abs/ 2501.09038. Nasiriany, S., Xia, F., Yu, W., Xiao, T., Liang, J., Dasgupta, I., Xie, A., Driess, D., Wahid, A., Xu, Z., Vuong, Q., Zhang, T., Lee, T.-W. E., Lee, K.-H., Xu, P., Kirmani, S., Zhu, Y., Zeng, A., Hausman, K., Heess, N., Finn, C., Levine, S., and Ichter, B. Pivot: Iterative visual prompting elicits actionable knowledge for vlms, 2024. URL https://arxiv.org/abs/2402.07872. Newman, K., Wang, S., Zang, Y., Heffren, D., and Sun, C. Do pre-trained vision-language models encode object states? arXiv preprint arXiv:2409.10488, 2024. Noles, N. S., Scholl, B. J., and Mitroff, S. R. The persis- tence of object file representations. Perception & Psy- chophysics, 67(2):324â334, 2005. Paiss, R., Ephrat, A., Tov, O., Zada, S., Mosseri, I., Irani, M., and Dekel, T. Teaching clip to count to ten. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 3170â3180, 2023. Park, J. S., Bhagavatula, C., Mottaghi, R., Farhadi, A., and Choi, Y. Visualcomet: Reasoning about the dynamic context of a still image, 2020. URL https://arxiv. org/abs/2004.10796. Patel, M., Gokhale, T., Baral, C., and Yang, Y. Cripp- vqa: Counterfactual reasoning about implicit physical properties via video question answering. arXiv preprint arXiv:2211.03779, 2022. Pezzulo, G., Barsalou, L. W., Cangelosi, A., Fischer, M. H., McRae, K., and Spivey, M. J. Computational grounded cognition: a new alliance between grounded cognition and computational modeling. Frontiers in psychology, 3: 612, 2013. Piaget, J. The Psychology of Intelligence. Harcourt, Brace, 1950. Piaget, J. The Origins of Intelligence in Children. Interna- tional Universities Press, 1952. Piaget, J. The Childâs Conception of Number. W.W. Norton and Company, 1965. Piaget, J. and Inhelder, B. The Psychology of the Child. Basic Books, New York, 1969. Piloto, L. S., Weinstein, A., Battaglia, P., and Botvinick, M. Intuitive physics learning in a deep-learning model inspired by developmental psychology. Nature human behaviour, 6(9):1257â1267, 2022. 11 Vision Language Models Cannot Reason About Physical Transformation Poirel, N., Borst, G., Simon, G., Rossi, S., Cassotti, M., Pineau, A., and Houd Ě e, O. Number conservation is related to childrenâs prefrontal inhibitory control: an fmri study of a piagetian task. PloS one, 7(7):e40802, 2012. Qiu, S., Guo, S., Song, Z.-Y., Sun, Y., Cai, Z., Wei, J., Luo, T., Yin, Y., Zhang, H., Hu, Y., et al. Phybench: Holis- tic evaluation of physical perception and reasoning in large language models. arXiv preprint arXiv:2504.16074, 2025. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv preprint arXiv: 2103.00020, 2021. Rahmanzadehgervi, P., Bolton, L., Taesiri, M. R., and Nguyen, A. T. Vision language models are blind. arXiv preprint arXiv:2407.06581, 2024. Rane, S., Ku, A., Baldridge, J., Tenney, I., Griffiths, T., and Kim, B. Can generative multimodal models count to ten? Proceeding of the Annual Meeting of the Cognitive Science Society, 46:1235â1241, 2024. Scholl, B. J. Object persistence in philosophy and psychol- ogy. Mind & Language, 22(5):563â591, 2007. Schulze Buschoff, L. M., Akata, E., Bethge, M., and Schulz, E. Visual cognition in multimodal large language models. Nature Machine Intelligence, 7(1):96â106, 2025. Shi, X., Huang, Z., Wang, F.-Y., Bian, W., Li, D., Zhang, Y., Zhang, M., Cheung, K. C., See, S., Qin, H., Dai, J., and Li, H. Motion-i2v: Consistent and controllable image- to-video generation with explicit motion modeling, 2024. URL https://arxiv.org/abs/2401.15977. Spelke, E. S., Breinlinger, K., Macomber, J., and Jacobson, K. Origins of knowledge. Psychological review, 99(4): 605, 1992. Spelke, E. S., Katz, G., Purcell, S. E., Ehrlich, S. M., and Breinlinger, K. Early knowledge of object motion: Conti- nuity and inertia. Cognition, 51(2):131â176, 1994. Spelke, E. S., Kestenbaum, R., Simons, D. J., and Wein, D. Spatiotemporal continuity, smoothness of motion and ob- ject identity in infancy. British journal of developmental psychology, 13(2):113â142, 1995. Sun, H., Gao, Q., Lyu, H., Luo, D., Li, Y., and Deng, H. Probing mechanical reasoning in large vision language models. arXiv preprint arXiv:2410.00318, 2024. Sun, H., Yu, S., Li, Y., Gao, Q., Lyu, H., Deng, H., and Luo, D. Probing perceptual constancy in large vision language models. arXiv preprint arXiv:2502.10273, 2025. Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ram Ě e, A., Rivi ` ere, M., Rouillard, L., Mesnard, T., Cideron, G., bastien Grill, J., Ramos, S., Yvinec, E., Casbon, M., Pot, E., Penchev, I., Liu, G., Visin, F., Kenealy, K., Beyer, L., Zhai, X., Tsitsulin, A., Busa-Fekete, R., Feng, A., Sachdeva, N., Coleman, B., Gao, Y., Mustafa, B., Barr, I., Parisotto, E., Tian, D., Eyal, M., Cherry, C., Peter, J.-T., Sinopalnikov, D., Bhupatiraju, S., Agarwal, R., Kazemi, M., Malkin, D., Kumar, R., Vilar, D., Brusilovsky, I., Luo, J., Steiner, A., Friesen, A., Sharma, A., Sharma, A., Gilady, A. M., Goedeckemeyer, A., Saade, A., Feng, A., Kolesnikov, A., Bendebury, A., Abdagic, A., Vadi, A., Gy Ě orgy, A., Pinto, A. S., Das, A., Bapna, A., Miech, A., Yang, A., Paterson, A., Shenoy, A., Chakrabarti, A., Piot, B., Wu, B., Shahriari, B., Petrini, B., Chen, C., Lan, C. L., Choquette-Choo, C. A., Carey, C., Brick, C., Deutsch, D., Eisenbud, D., Cattle, D., Cheng, D., Paparas, D., Sreepathihalli, D. S., Reid, D., Tran, D., Zelle, D., Noland, E., Huizenga, E., Kharitonov, E., Liu, F., Amirkhanyan, G., Cameron, G., Hashemi, H., Klimczak-Pluci Ě nska, H., Singh, H., Mehta, H., Lehri, H. T., Hazimeh, H., Ballantyne, I., Szpektor, I., Nardini, I., Pouget-Abadie, J., Chan, J., Stanton, J., Wieting, J., Lai, J., Orbay, J., Fernandez, J., Newlan, J., yeong Ji, J., Singh, J., Black, K., Yu, K., Hui, K., Vodrahalli, K., Greff, K., Qiu, L., Valentine, M., Coelho, M., Ritter, M., Hoffman, M., Watson, M., Chaturvedi, M., Moyni- han, M., Ma, M., Babar, N., Noy, N., Byrd, N., Roy, N., Momchev, N., Chauhan, N., Sachdeva, N., Bunyan, O., Botarda, P., Caron, P., Rubenstein, P. K., Culliton, P., Schmid, P., Sessa, P. G., Xu, P., Stanczyk, P., Tafti, P., Shivanna, R., Wu, R., Pan, R., Rokni, R., Willoughby, R., Vallu, R., Mullins, R., Jerome, S., Smoot, S., Gir- gin, S., Iqbal, S., Reddy, S., Sheth, S., P Ě oder, S., Bhat- nagar, S., Panyam, S. R., Eiger, S., Zhang, S., Liu, T., Yacovone, T., Liechty, T., Kalra, U., Evci, U., Misra, V., Roseberry, V., Feinberg, V., Kolesnikov, V., Han, W., Kwon, W., Chen, X., Chow, Y., Zhu, Y., Wei, Z., Egyed, Z., Cotruta, V., Giang, M., Kirk, P., Rao, A., Black, K., Babar, N., Lo, J., Moreira, E., Martins, L. G., Sanseviero, O., Gonzalez, L., Gleicher, Z., Warkentin, T., Mirrokni, V., Senter, E., Collins, E., Barral, J., Ghahra- mani, Z., Hadsell, R., Matias, Y., Sculley, D., Petrov, S., Fiedel, N., Shazeer, N., Vinyals, O., Dean, J., Hass- abis, D., Kavukcuoglu, K., Farabet, C., Buchatskaya, E., Alayrac, J.-B., Anil, R., Dmitry, Lepikhin, Borgeaud, S., Bachem, O., Joulin, A., Andreev, A., Hardin, C., Dadashi, R., and Hussenot, L. Gemma 3 technical report, 2025. URL https://arxiv.org/abs/2503.19786. Viarouge, A., Houd Ě e, O., and Borst, G. The progressive 6- year-old conserver: Numerical saliency and sensitivity as core mechanisms of numerical abstraction in a piaget-like 12 Vision Language Models Cannot Reason About Physical Transformation estimation task. Cognition, 190:137â142, 2019. Wang, L., Su, E., Liu, J., Li, P., Xia, P., Xiao, J., Zhang, W., Dai, X., Chen, X., Meng, Y., et al. Physunibench: An undergraduate-level physics reasoning benchmark for multimodal models. arXiv preprint arXiv:2506.17667, 2025. Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J. Qwen2-vl: Enhancing vision-language modelâs perception of the world at any resolution, 2024. URL https://arxiv.org/abs/2409.12191. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Met- zler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022. Xu, G., Jin, P., Hao, L., Song, Y., Sun, L., and Yuan, L. Llava-o1: Let vision language models reason step-by- step. arXiv preprint arXiv:2411.10440, 2024. Yang, X., Li, B., Zhang, Y., Yin, Z., Bai, L., Ma, L., Wang, Z., Cai, J., Wong, T.-T., Lu, H., and Jia, X. Vlipp: Towards physically plausible video generation with vi- sion and language informed physical prior, 2025. URL https://arxiv.org/abs/2503.23368. Yu, S., Cho, J., Yadav, P., and Bansal, M. Self-chained image-language model for video localization and question answering. Advances in Neural Information Processing Systems, 36:76749â76771, 2023. Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 9556â9567, 2024. Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., and Zou, J. When and why vision-language models behave like bags-of-words, and what to do about it? arXiv preprint arXiv:2210.01936, 2022. Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. From recogni- tion to cognition: Visual commonsense reasoning, 2019. URL https://arxiv.org/abs/1811.10830. Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 12104â12113, 2022. Zhang, J., Hu, J., Khayatkhoei, M., Ilievski, F., and Sun, M. Exploring perceptual limitation of multimodal large lan- guage models. arXiv preprint arXiv:2402.07384, 2024a. Zhang, R., Zhang, B., Li, Y., Zhang, H., Sun, Z., Gan, Z., Yang, Y., Pang, R., and Yang, Y. Improve vision lan- guage model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198, 2024b. Zhang, Y., Wu, J., Li, W., Li, B., Ma, Z., Liu, Z., and Li, C. Video instruction tuning with synthetic data, 2024c. URL https://arxiv.org/abs/2410.02713. Zheng, Z., Yan, X., Chen, Z., Wang, J., Lim, Q. Z. E., Tenen- baum, J. B., and Gan, C. Contphy: Continuum physi- cal concept learning and reasoning from videos. arXiv preprint arXiv:2402.06119, 2024. 13 Vision Language Models Cannot Reason About Physical Transformation A. Data Curation Curation and Quality Control. ConservationBench was curated by three annotators with college-level training in cognitive science or computer science. Each video underwent two independent cross-review passes; items failing to meet design criteria were removed or revised. Data Acquisition. All videos were captured under standardized recording conditions using a fixed camera setup, with consistent lighting and background held constant within each property category. Each transformation was carefully scripted to ensure visual clarity, reproducibility, and minimal ambiguity. Design Principles. To ensure conceptual integrity and interdisciplinary rigor, we adopt three design criteria for each item: (i) Discriminativenessâtasks are constructed such that models lacking the targeted knowledge are systematically driven toward incorrect responses; (i) Minimal confoundingâinstances are designed to minimize reliance on ancillary skills (e.g., object recognition); and (i) Minimal textual shortcutsâtasks cannot be solved using textual cues alone and instead require genuine multimodal reasoning. B. Task Design Table 2 presents paired descriptions of conservation tasks and their matched non-conserving controls across all four quantitative properties, with corresponding illustrations in Figure 2. Table 2. Task descriptions for conservation and non-conserving control scenarios across four quantitative properties. PropertyConservation TaskNon-conserving Control NumberTwo rows of identical coins are presented in an initial configuration. One row is then spread apart without adding or removing any coins. Two rows of identical coins are presented in an initial configuration. One row is then spread apart, with one coin added to said row. LengthTwo straws of identical length are shown in an initial configuration. One of the straws is then repositioned without altering its actual length. Two straws of identical length are shown in an initial configuration. One of the straws is then repositioned while its actual length is altered (extendable straws are used). VolumeA fixed volume of liquid is poured from one con- tainer into another of a different shape. Although the height changes significantly, the volume remains constant throughout the transformation. A fixed volume of liquid is poured from one con- tainer into another of a different shape. A significant portion of water is left in the original container in- stead of being completely poured in. SizeA lump of playdough is reshaped from one form (e.g., a ball) into another (e.g., a flattened disc). While the shape and surface features change, the total mass remains the same across both states. A lump of playdough is reshaped from one form (e.g., a ball) into another (e.g., a flattened disc). A significant portion of playdough is left in the experi- menterâs hand without being integrated into the new shape. C. Prompting Strategy Reasoning about conservation often requires interpreting the transformation as a continuous process across the videos or sequence of frames. To examine how prompts influence temporal integration and transformation-based reasoning, we design four prompt types, each progressively enhancing the modelâs awareness of the underlying continuous process, as summarized in Table 3. Together, these prompting strategies enable us to evaluate how different forms of linguistic scaffolding shape model engagement with visual dynamics. The âSequentialâ and CoT prompts encourage frame-by-frame perception with step-by- step reasoning, directing attention to frame-wise visual evidence. In contrast, the âContinuousâ prompt explicitly presents the multi-frame input as a continuous process, offering a conceptual cue to support conservation reasoning. 14 Vision Language Models Cannot Reason About Physical Transformation Table 3. Four different prompt formats used in our benchmark. Prompt TypePrompt Example Direct QuestionIs the number of coins in the upper row the same as in the lower row in the final image? âSequentialâ PromptPlease process the images below sequentially, and then answer: Is the number of coins in the upper row the same as in the lower row in the final image? CoT PromptPlease process the images below sequentially. First describe what happens across the images, then answer: Is the number of coins in the upper row the same as in the lower row in the final image? âContinuousâ PromptThe above images represent a continuous process. Please answer: Is the number of coins in the upper row the same as in the lower row in the final image? D. Example Input To provide clarity on the exact format of inputs provided to models, we present a complete example task below, including both the visual frames and the full textual prompt. Task: Conservation of Number (Conserving condition) Task Configuration: This example demonstrates a Number conservation task using Uniform extraction method with 7 frames and Direct Question prompt format. Visual Input: The model receives a sequence of frames extracted from the video in temporal order (Frame 1 through Frame 7), ensuring that the transformation process is presented chronologically without any frame order disruption. Figure 6 shows an example with 7 frames, where frames are sampled uniformly across the video timeline. The first frame shows the initial state (two rows of coins with equal numbers), intermediate frames capture the transformation process (spreading one row), and the final frame shows the end state (one row spread out while maintaining the same number of coins). Frame 1 Frame 2 Frame 3 Frame 4 Frame 5 Frame 6 Frame 7 Figure 6. Example visual input: A sequence of 7 frames from a number conservation task, showing the initial state, transformation process, and final state. Textual Input: Below is the structure of the prompt provided to the model (using the âDirect Questionâ format). The [Image] placeholders indicate where the corresponding frames from Figure 6 are embedded in the actual input: Frame 1: [Image] Frame 2: [Image] Frame 3: [Image] Frame 4: [Image] Frame 5: [Image] Frame 6: [Image] 15 Vision Language Models Cannot Reason About Physical Transformation Frame 7: [Image] Is the number of coins in the upper row the same as in the lower row in the final image? Please choose one of the following options: (A) No, the lower row has more coins. (B) No, the upper row has more coins. (C) Yes, they are the same. Ground Truth: Option (C) - Yes, they are the same. Alternative Prompt Formats: For other prompt types, the question is prefixed with additional instructions. For example, the âSequentialâ format would begin with âPlease process the images below sequentially, and then answer: [question]â, while the CoT format would include âPlease process the images below sequentially. First describe what happens across the images, then answer: [question]â. See Table 3 for details on all four prompt formats. E. Model Inference We evaluate 112 VLMs spanning diverse architectures, training regimes, and parameter scales, including mainstream proprietary models as well as advanced open-source models ranging from 1B to 76B parameters. Inference is conducted on a cluster equipped with 8Ă NVIDIA H100 (80 GB) GPUs. As a practical policy, models of 1â13B parameters typically run on a single GPU; 13â32B on two GPUs; 32â70B on four GPUs; and >70B on all eight GPUs. To preserve fidelity and reproducibility, we adhere to configurations and reference implementations from the official codebases, avoiding unnecessary modifications. We build a scalable evaluation framework supporting parallel execution and compartmentalized environments. Inference jobs are distributed across GPUs via a dynamic scheduler that maximizes utilization and minimizes idle time. We additionally develop a lightweight modality-verification suite that prompts each model to summarize the media information it receives, and then the responses are checked by human reviewers to verify correct input routing and modality handling in our inference pipelines. F. Evaluation Rule-based template matching degrades with complex model outputs, yielding elevated false positives/negatives and requiring continual template optimization to cover corner cases. LLM-based matching better identifies intended choices within free-form text but can hallucinate, especially when brief answers are embedded in extensive context. To balance these trade-offs, we introduce Hybrid Matching, which prioritizes deterministic template matching and, on failure, falls back to an ensemble of four LLM judges (Qwen2.5-72B-Instruct, Mixtral-8x7B-Instruct-v0.1, DeepSeek-R1-Distill-Llama-70B, and Llama-3.1-70B). The ensemble decision is accepted only if at least three three of four models return a consistent extraction; otherwise, the mapping is deemed unsuccessful. By coupling the precision of template extraction with the semantic flexibility of LLM adjudication, Hybrid Matching delivers more reliable mappings across diverse response styles. 16 Vision Language Models Cannot Reason About Physical Transformation G. Counterbalancing Conditions Complete counterbalancing parameters and their factorial combinations for all four quantitative property domains are provided in Table 4. Table 4. Counterbalanced variations of task-irrelevant features. Each unique combination of parameter values yields 48 distinct task instances per domain. DomainParameterVariations Number P1: Object Type2 variants (Uniform, Mixed) P2: Mapping Shift2 variants (Lower vs. Upper row moved) P3: Distance Spread2 variants (Near, Far) P4: Number of Objects6 variants (3â8 coins) Total combinations:2Ă 2Ă 2Ă 6 = 48 Length P1: Object Type2 variants (Uniform, Mixed) P2: Mapping Shift2 variants (Lower vs. Upper straw moved) P3: Distance Moved2 variants (Near, Far) P4: Direction2 variants (Left, Right) P5: Transformation Action3 variants (Slide, Rotate, Vertical) Total combinations:2Ă 2Ă 2Ă 2Ă 3 = 48 Volume P1: Liquid Color8 variants P2: Glass Transaction2 variants (Tallâ Short, Shortâ Tall) P3: Liquid Volume3 variants (Small, Medium, Large) Total combinations:2Ă 8Ă 3 = 48 Size P1: Object Color8 variants P2: Shape Transformation6 variants (Crossing Sphere, Cylinder, Plane) Total combinations:6Ă 8 = 48 17 Vision Language Models Cannot Reason About Physical Transformation H. Complete Model Results Aggregated Across Domains and Conditions Complete performance metrics for all 112 evaluated VLMs, ranked by average accuracy, are provided in Tables 5 and 6. Table 5. Complete Model Rankings by Average Accuracy (Ranks 1-56) RankModelConserve (%)Non-Conserve (%)Average (%)Strict (%) 0Human Baseline98.5598.1598.3596.72 1gemini-2.5-pro94.3343.8869.1141.20 2doubao-seed-1.6-vision82.9946.4264.7138.22 3gpt-591.4835.8263.6533.14 4claude-sonnet-4-575.4345.3960.4137.23 5Qwen-3-VL-32B-Instruct64.5752.0658.3131.15 6Qwen-2.5-VL-72B-Instruct89.6921.9855.8317.32 7hunyuan-t1-vision74.0536.0555.0526.80 8InternVL-3-38B-Instruct61.2946.3653.8329.74 9InternVL-3.5-38B-Instruct75.4931.8053.6419.32 10InternVL-3.5-38B-MPO72.7533.9853.3620.85 11InternVL-3-38B57.4748.4252.9529.41 12InternVL-2.5-78B60.7144.4352.5719.27 13Qwen-2.5-VL-32B-Instruct81.8222.8352.3316.66 14InternVL-3.5-8B72.5329.9351.2315.62 15Qwen-2-VL-72B-Instruct69.4431.5250.4815.24 16InternVL-3.5-8B-Instruct74.5326.0650.3015.45 17InternVL-2.5-38B64.9735.5750.2719.12 18InternVL-2.5-78B-MPO55.9744.3950.1823.42 19InternVL-3.5-8B-MPO73.8726.2550.0615.28 20InternVL-3.5-1B-MPO96.392.4649.421.60 21InternVL-3.5-1B-Instruct93.453.4548.451.96 22InternVL-2-2B77.5618.1347.858.16 23Ovis1.6-Gemma2-9B74.2420.2347.247.26 24InternVL-3-14B-Instruct74.8619.2447.058.94 25Qwen-3-VL-32B-Thinking61.4432.4446.9420.23 26InternVL-3.5-30B-A3B-MPO75.1317.2946.219.15 27InternVL-3-14B69.0722.9045.9911.00 28Qwen-2-VL-7B-Instruct66.5825.1445.866.20 29InternVL-3.5-30B-A3B-Instruct77.1914.4745.836.84 30InternVL-2.5-38B-MPO50.5241.0945.8120.24 31InternVL-2-40B43.3547.9345.6414.50 32Qwen-2.5-VL-3B-Instruct76.9314.2545.592.89 33InternVL-2.5-1B-MPO76.4714.5245.495.07 34InternVL-2-26B73.8516.2645.055.15 35InternVL-3-2B-Instruct81.517.9344.721.50 36R1-Onevision-7B67.6621.7444.7013.49 37InternVL-3.5-1B81.936.7444.333.06 38Phi-3.5-vision-instruct71.5317.1244.325.35 39InternVL-3-2B76.0112.0544.032.75 40InternVL-2.5-2B59.5928.2543.9210.89 41InternVL-2.5-4B-MPO77.0310.6343.836.55 42llava-onevision-qwen2-7b-si-hf58.7627.9343.343.40 43InternVL-2-8B74.0912.0743.084.12 44Qwen-2.5-VL-7B-Instruct64.1221.5542.837.39 45llava-next-interleave-qwen-7b-dpo50.1335.2042.666.17 46InternVL-2-Llama3-76B55.6729.2442.4611.90 47Qwen-3-VL-8B-Instruct53.1231.5242.328.59 48InternVL-3-8B-Instruct70.9513.6142.287.97 49InternVL-2.5-1B74.869.4142.143.69 50InternVL-3-9B-Instruct51.4132.7242.077.41 51Mini-InternVL-Chat-4B-V1-579.774.2442.010.75 52InternVL-3.5-4B-Instruct57.0726.5141.798.13 53InternVL-3-9B50.4632.8841.677.90 54Qwen-2.5-Omni-3B80.223.0641.641.48 55VLAA-Thinker-Qwen2VL-2B48.6034.2041.405.31 56InternVL-2-4B78.264.2241.241.76 18 Vision Language Models Cannot Reason About Physical Transformation Table 6. Complete Model Rankings by Average Accuracy (Ranks 57-112, cont.) RankModelConserve (%)Non-Conserve (%)Average (%)Strict (%) 57InternVL-3-8B65.5516.7441.148.12 58InternVL-3.5-4B59.0422.4340.738.73 59Phi-3-vision-128k-instruct57.6123.5440.586.28 60InternVL-2.5-4B72.668.0640.364.66 61Qwen-2.5-Omni-7B66.4814.1140.307.84 62VLAA-Thinker-Qwen2.5VL-7B54.7725.5640.1610.56 63InternVL-2.5-26B47.6632.4240.049.60 64Qwen-2-VL-2B-Instruct44.7334.6839.704.33 65InternVL-3-78B-Instruct45.2733.9239.6015.14 66InternVL-3-78B43.0435.9039.4715.97 67VLAA-Thinker-Qwen2VL-7B52.6825.6939.187.84 68InternVL-3.5-14B-Instruct42.5835.3138.959.46 69InternVL-3.5-38B58.6919.1638.9211.28 70InternVL-3.5-4B-MPO52.6025.0838.847.67 71Qwen-3-VL-2B-Instruct50.8225.9938.405.62 72llava-onevision-qwen2-7b-ov-hf47.3529.0638.212.07 73llava-onevision-qwen2-7b-ov-chat-hf47.6428.4038.022.01 74InternVL-Chat-V1-552.1023.6437.876.76 75InternVL-3.5-2B63.9711.6137.794.23 76InternVL-3.5-2B-Instruct62.8512.3237.584.46 77xgen-m-phi3-mini-instruct-interleave-r-v1.557.6117.3537.482.33 78InternVL-3.5-2B-MPO61.4312.7637.104.31 79InternVL-2.5-26B-MPO34.8038.0736.449.70 80Mini-InternVL-Chat-2B-V1-542.3229.7536.032.67 81Phi-4-multimodal-instruct33.7936.9435.365.45 82InternVL-2.5-8B47.2823.0935.195.82 83Qwen-3-VL-4B-Instruct49.0520.8334.946.77 84Qwen-3-VL-8B-Thinking52.9216.3134.617.42 85InternVL35-GPT-OSS-20B-A4B-Preview57.0111.6134.314.64 86VLAA-Thinker-Qwen2.5VL-3B36.4431.6634.059.13 87Mantis-llava-7b31.4636.2133.834.23 88InternVL-2.5-8B-MPO37.5329.1333.337.42 89InternVL-2.5-2B-MPO32.2034.2433.224.92 90llava-v1.6-mistral-7b-hf33.3332.8533.090.00 91InternVL-3-1B52.1713.9833.070.47 92InternVL-3-1B-Instruct52.9313.0933.010.25 93InternVL-Chat-V1-123.5442.3432.942.45 94Mantis-8B-Idefics229.5736.3032.931.48 95llama3-llava-next-8b-hf32.4432.4732.450.00 96llava-onevision-qwen2-0.5b-si-hf19.0641.8630.461.10 97llava-onevision-qwen2-0.5b-ov-hf3.7054.8429.270.54 98Qwen-3-VL-4B-Thinking40.7717.6629.216.73 99Mantis-8B-clip-llama318.6539.7129.182.79 100Mantis-8B-siglip-llama312.3343.7828.051.47 101InternVL-3.5-30B-A3B47.538.4728.004.92 102Xinyuan-VL-2B11.8244.1127.973.20 103Mantis-8B-Fuyu20.3135.0727.690.36 104InternVL-Chat-V1-211.1143.8527.482.93 105InternVL-Chat-V1-2-Plus0.7151.5326.120.28 106Mantis-bakllava-7b17.1034.4425.771.15 107InternVL-2-1B3.0447.5525.300.34 108JanusFlow-1.3B29.3920.2524.821.76 109Janus-Pro-7B28.6620.8924.782.74 110Janus-1.3B29.0517.6123.330.66 111Janus-Pro-1B19.5123.6221.571.42 112Qwen-3-VL-2B-Thinking30.109.5219.812.51 19 Vision Language Models Cannot Reason About Physical Transformation I. Combined Model Performance By Quantitative properties NumberLengthSizeVolume Domain 0% 20% 40% 60% 80% 100% Accuracy 0.681 0.577 0.538 0.408 0.180 0.297 0.232 0.355 Average Accuracy by Domain: Conservation Task vs Non-Conserving Control Chance Level (33.3%) Conservation Task Non-Conserving Control Figure 7. Average model performance by quantitative domains. Models consistently perform worse on non-conserving controls compared to conservation tasks. J. Model response under empty image and text control 0%20%40%60%80%100% Conservation Task Accuracy 0% 20% 40% 60% 80% 100% Empty Image Control Accuracy r = 0.678 p < 0.0001 Conservation Task vs Empty Image Control (n=62) Linear Fit y=1-x Chance Level (33.3%) 0%20%40%60%80%100% Conservation Task Accuracy 0% 20% 40% 60% 80% 100% Text Control Accuracy r = 0.475 p < 0.0001 Conservation Task vs Text Control (n=62) Linear Fit y=1-x Chance Level (33.3%) Figure 8. Model-level correlations between conservation task performance and control condition biases. (Left) Conservation task accuracy versus Empty Image Control accuracy shows strong positive correlation (r = 0.578,p < 0.0001), indicating models performing better on conservation tasks exhibit stronger textual priors favoring quantity invariance when visual content is removed. (Right) Conservation task accuracy versus Text Control accuracy shows similar but slightly weaker correlation (r = 0.475,p < 0.0001), with both patterns demonstrating that success on conservation tasks is driven primarily by textual biases rather than visual transformation reasoning. Evaluated on 62 VLMs supporting both empty image and text-only inputs under 7-frame conditions. 20