Paper deep dive
PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/27/2026, 4:41:08 AM
Summary
The paper introduces PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning containing 11,586 problems with visual diagrams and step-by-step solutions. It highlights that existing benchmarks lack high-difficulty data and comprehensive process annotations. Evaluation of 18 MLLMs shows that even the strongest model achieves only 33.7% accuracy, revealing a significant gap between current AI capabilities and human expert performance in complex physics problem-solving.
Entities (12)
Relation Signals (12)
PhysElite → contains → 11,586 problems
confidence 100% · PhysElite contains 11,586 Olympiad-tier problems.
PhysElite → includesmultimodaldata → Visual Diagrams
confidence 100% · For each problem, we provide corresponding visual diagrams
PhysElite → includesprocessannotations → Step-by-step derivations
confidence 100% · step-by-step bilingual Chinese-English solution derivations
PhysElite → isbilingual → Chinese-English
confidence 100% · PhysElite is a large-scale, bilingual (Chinese–English) multimodal benchmark
Qwen3-VL-235B-A22B → achievesaccuracyon → PhysElite
confidence 95% · the strongest open-source model (Qwen3-VL-235B-A22B) reaches only 20.0%.
Grok 4.2 → achievesaccuracyon → PhysElite
confidence 95% · The best-performing model, Grok-4.2, achieves 33.7% answer accuracy
Claude Opus 4.6 → achievesaccuracyon → PhysElite
confidence 95% · followed by Claude-Opus-4.6 at 28.1%.
PhysElite → coverssubjects → Optics
confidence 95% · Optics questions 824
PhysElite → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PhysElite, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. PhysElite contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese-English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.25097v1
- Canonical: https://arxiv.org/abs/2608.25097v1
Trouble viewing inline? Open PDF directly →
Full Text
79,840 characters extracted from source content.
Expand or collapse full text
PHYSELITE: How Far Are LLMs from Solving Olympiad-Level Physics Problems? Ruoran Xu * ‡ Wending Gao * Liyunfeng Chen * Aixin Shi * Haoyu ChengZixiang FangYiqiang ZouQiufeng Wang † Xi’an Jiaotong-Liverpool University Abstract Understanding how (multimodal) large language models perform on physics prob- lems requires benchmarks that reflect the difficulty and breadth of expert-level phys- ical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present PHYSELITE, a large-scale bilingual multimodal benchmark for Olympiad- level physics reasoning. PHYSELITE contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese–English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. Our datasets are released at https://huggingface.co/datasets/physelite/PhysElite. K-12 Level Sample 1: The figure shows a bus turning right on a flat road, viewed from the rear. Which of the following free- body diagrams is most correct? Problem_Answer: D K-12 Level Sample 2: As shown, 1 mole of an ideal noble gas undergoes a cyclic equilibrium process. The horizontal coordinate is the work done by the gas from the beginning, and the vertical coordinate is the heat absorbed by the gas. Find the ratio of the highest temperature to the lowest temperature. Problem_Answer: T max T min = T 2 T 1 = 2 Olympiad Level Sample: Problem_Text: When a strong laser beam ...... a maximum at to zero at . ...... (3) Laser beam: in ...... With apex at , what total laser power is needed to balance gravity? (4) ...... Find the oscillation period. I 0 y = 0 y = – 4h1m×80 μmz×y y o = −h/2 = −5 μm Problem_Answer: (1) ; (2) , ......(4) θ = arcsin nsin [ a−arcsin(sina/n) ] F x = hwI 0 c ( 7 4 − y 2 o 4h 2 ) (1−cosθ)1.12×10 −2 s. Problem_Image: Solution_Image: Difficulty: 7 Tag: Optics, Modern Physics, Mechanics Problem_Process: (1) ...... The internal ray hits the base at angle ......(2) ...... momentum ...... with , the powers ...... Integrating over each face for ...... so ...... (4) ...... linear restoring force, hence SHO with a−β P u P l I(y) = I 0 (1− | y | /4h)0 < y o < h P u −P l = − hwI 0 y o 2h ( 1− y o 2h ) , T = 2π 2cρh 2 tana I 0 sinθ = 1.12×10 −2 s. Figure 1: Comparision of two existing K-12 level samples from PhysicsArena [1] and one illustrative example of olympiad-level physics problem with detailed multimodal annotation of solution steps in our PHYSELITE (The cooresponding completed example can be found in Figure 13) * Equal contribution ‡ Project Lead † Corresponding author qiufeng.wang@xjtlu.edu.cn. 10 May -Preprint. arXiv:2608.25097v1 [cs.AI] 25 Aug 2026 1 Introduction Figure 2: Performance of four state-of-the-art MLLMs on our PHYSELITE and three representative existing physics benchmarks: while prior benchmarks are nearing saturation, PHYSELITE leaves a substantial gap between the best model and the human reference. Physics problem-solving has long been a rigorous benchmark for (multimodal) large language models (MLLMs/LLMs), requiring models to parse structured visual input, apply physical laws under explicit constraints, and chain intermediate results into coherent multi-step derivations. In recent years, MLLMs/LLMs have achieved impressive performance on physics reasoning benchmarks [2,3,4,5], with some even claiming medal-level results on Olympiad physics tasks [6,2,7]. However, these high scores raise a critical question: Are current models truly capable of solving authentic Olympiad-level physics problems? To answer this question, we systematically analyze existing physics benchmarks and identify two core limitations: Scale and difficulty. Most benchmarks are lack of challenging problems. Many existing datasets are dominated by elementary or K-12 level problems, which fail to stress-test frontier models’ ability to handle complex, multi-step reasoning. As shown in the examples in Figure 1, these problems lack the depth and multimodal complexity of real Olympiad tasks. On the otherhand, Datasets such as IPhO [8] and APhO [9] represent Olympiad-level, but both contain only a limited number of problems and are updated infrequently, which make them vulnerable to model saturation, where performance plateaus before reaching the true ceiling of Olympiad problem-solving ability, as illustrated in Figure 2 Missing process-level annotations. Many benchmarks provide only final answers, with no inter- mediate derivations linking visual cues to physical laws. For physics, the derivation itself carries more evaluative signal than the final answer, and grading only the final answer leaves the source of a model’s error unclear. The absence of step-level annotations also limits the use of physics problems as a training signal for chain-of-thought supervision or process-level reinforcement learning, a direction that has already shown strong returns in many domains [10]. To address these gaps, we introduce PHYSELITE, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. As summarized in Table 1, PHYSELITE contains 11,586 Olympiad-tier problems, each paired with visual diagrams, step-by-step bilingual derivations, and verified final answers (Figure 1). Our evaluation of 18 open-source and closed-source MLLMs on PHYSELITE reveals that even the strongest model achieves only 33.7% answer accuracy, leaving a substantial gap with human experts (Figure 2). We further conduct process-level evaluation to diagnose where models fail in the reasoning chain, identifying key failure modes that point to clear directions for future improvement. In summary, our contributions are: •An Olympiad-level physics problem benchmark.PHYSELITE comprises 11,586 Olympiad-tier physics problems, an order of magnitude larger than comparable high- difficulty datasets while maintaining uniform elite difficulty. •Multimodal process-level annotations. Every problem is paired with a visual diagram and a human-verified step-by-step derivation, providing a resource suitable for both rigorous evaluation and process-supervised training. • Extensive evaluation and diagnostic analysis. We benchmark representative open- and closed-source MLLMs on PHYSELITE, quantify the difference to human experts, compare 2 Table 1: Comparison of PHYSELITE with existing physics reasoning benchmarks. PHYSELITE is the largest Olympiad-level physics benchmark with bilingual support, detailed process-level annotations. Diagrams & Solution Step & Analysis Process:✓= Supported;✗= Unsupported. DatasetYearNumberDiagramsSolution StepAnalysis ProcessLanguage K-12 level SciBench [14]2024295✓✗EN MMMU [13]2024983✓✗EN PhysicsArena [1]20255,103✓✗EN PhysUniBench [3]20253,304✓(text only)✗EN/CN PHYSICS [15]20255,334✗EN/CN College-level PHYSICS [15]20252,127✗EN/CN UGPhysics [4]20255,520✗EN/CN PHYSICS [16]20251,297✓✗EN P1-VL (Not Public) [7]20263,907✓✗EN Olympiad-level OlympiadBench [6]20242,428✓(text only)✗EN/CN PHYSICS [15]2025823✗EN/CN PHYBench [5]2025500✗✓(text only)✗EN PhysReason [17]20251,200✓(text only)✗EN HiPhO [2]2025360✓(text only)✗EN/CN P1-VL (Not Public) [7]20264,126✓✗EN PHYSELITE (Ours)202611,586✓(Mutimodal)✓EN/CN System-2 and standard models across difficulty levels, and characterize six recurring failure modes that suggest directions for future work. 2 Related Work 2.1 Multimodal Scientific Benchmarks Early multimodal evaluation suites such as ScienceQA [11], MathVista [12], MMMU [13], and SciBench [14] broadened multimodal coverage across disciplines and established the protocol of pairing diagrams with multiple-choice or short-answer items. Their physics subsets, however, are dominated by introductory-level questions and rarely exceed a few hundred items per domain, leaving little headroom for differentiating frontier models. The cross-disciplinary framing further dilutes the physics-specific signal: a model can succeed through generic visual reasoning without engaging the underlying physical laws. 2.2 Physics-Specific Benchmarks: Scale, Modality, and Difficulty Physics-focused resources have grown along three partially competing axes—scale, modality, and difficulty—and to date no single benchmark has satisfied all three. Scaling along the text-only axis. PHYSICS [15] (8,284 bilingual problems) and UGPhysics [4] (5,520 problems) provide large training-friendly corpora that span high-school to graduate physics, but both deliberately exclude diagrams to ease ingestion, and neither targets Olympiad difficulty. PHYBench [5] curates only 500 problems but pushes them to Olympiad level, illustrating the inverse relationship typical of text-only physics benchmarks: scale is bought at the cost of difficulty, or vice versa. Adding multimodality at undergraduate level. Recognizing that diagrams, circuits, and ray traces carry information that pure text cannot recover, recent work introduces multimodal physics benchmarks. PhyX [18] (3,000 problems, six reasoning types), PhysUniBench [3] (3,304 problems with paired diagrams), PHYSICS [16] (1,297 problems), and PhysicsArena [1] (5,103 high-school CEE-style problems) all couple text with visual input. These benchmarks advance multimodal evaluation but cap difficulty at undergraduate or pre-college level, leaving the Olympiad regime untested. Most are also English-only, limiting cross-lingual analysis. 3 Olympiad-focused benchmarks at limited scale. Dedicated Olympiad benchmarks probe the upper bound of model capability but are constrained in either scale, modality, or contamination risk. OlympiadBench [6] (2,428 physics problems) and PhysReason [17] (1,200 problems, 81% with diagrams) provide step-level annotations on competition material, yet draw entirely from publicly archived contests, which raises pre-training overlap concerns. HiPhO [2] compiles 360 problems from 13 recent (2024–2025) Olympiad papers in bilingual form, but its scale precludes per-difficulty statistical analysis. PhoPile [19] expands to 3,052 Olympiad problems but remains unimodal and monolingual. Beyond benchmarks, recent work explores reinforcement learning on text-only physics Olympiads [7], complementary to our multimodal evaluation focus. 3PHYSELITE Dataset 3.1 Overview of PHYSELITE StatisticNumber Total questions11,586 - Mechanics questions6,739 - Electronmagenetism questions2,811 - Optics questions824 - Thermodynamics questions652 - Modern Physics questions560 Difficulties - Easy (1-3)32% - Medium (4-5)54% - Hard (6-7)14% Goal Type - Symbolic Expressions7,767 - Numerical Values1,449 - Qualitative Conclusions1,331 - Equations1,039 Diagrams16,130 Maximum question length2619 Maximum answer length280 Average question length135.9 Average answer length28.0 Table 2: Key Statistics of PHYSELITE Figure 3: Distribution of PHYSELITE PHYSELITE is, to our knowledge, the first benchmark to combine four desiderata in a single corpus: (i) Olympiad-level difficulty across every problem, (i) large scale (11,586 problems, an order of magnitude above PHYBench [5] and HiPhO [2]), (i) bilingual multimodal coverage with paired diagrams and step-level derivations. Concretely, PHYSELITE is a large-scale, bilingual (Chinese–English) multimodal benchmark for Olympiad-level physics reasoning. It contains 11,586 open-ended problems spanning Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. Each problem is paired with one or more diagrams, a human-verified step-by-step derivation, the final answer, and metadata including topic tags and a difficulty rating on a 1–7 scale. All problems are sourced from the daily training and practice materials of 15 first-prize-winning students from local secondary schools. This sourcing choice is central to our design: it provides realistic Olympiad-level problems while substantially reducing the pre-training overlap that confounds existing physics benchmarks. More details in Appendix A. 3.2 Dataset Format An example is presented on the right side of Figure 4. For each question, we give 1) the question text in both Chinese and English, 2) the schematic diagram and selected process diagram, 3) a detailed, 4 step-by-step solution in natural language, 4) a difficulty level annotated by human experts and 5) its subject category. To construct the dataset, we first collect approximately 15,000 open-ended problems from the daily learning and practice materials of physics contestants (including textbooks, past competition papers, and training handouts) in image format. These images are scanned at a resolution of 300 DPI to ensure the clarity of text, schematic diagrams, and mathematical symbols, covering a wide range of difficulty levels and subfields fully aligned with international physics competition standards. Each problem image contains the problem text, schematic diagram, step-by-step solution and the final answer. Then, we convert these 300 DPI problem images into text using the Tesseract OCR engine (v5.3.0), which is widely used for academic text recognition and supports the accurate identification of mixed text and simple symbols. To ensure the accuracy of the converted text—especially for professional physics terms, complex schematic diagram descriptions, and mathematical expressions—we conduct manual verification and correction on all OCR results: any misrecognized characters, missing content, or distorted symbols are revised by professional physics researchers. For the mathematical expressions in answers and step-by-step solutions, we further convert them into L A T E X format manually after verification to preserve the precision of mathematical symbols and logical structure. After that, we remove duplicated problems using fuzzy matching algorithms (with a similarity threshold of 0.95) to avoid data redundancy. After the above process is finished, the size of the dataset is reduced from around 15,000 to 11,586. 3.3 Annotation Raw Images Question Bank Deduplication Image2Text Compiling Tagging Structured Problems Checking Double Check Final Dataset Annotation Guideline Instruct Validity Test Second Refinement Annotator Refiner Expert Problem Text: Find the orbital motion time.(1) In an inertial frame, a mass is fixed, and a mass moves under the gravitational influence of along the elliptical orbit shown in Figure 1, described by: The speed of along the -axis is denoted as , and the infinitesimal displacement in the -domain is . The elapsed time is given by: (1.1) Find ; (1.2) Suppose the particle starts at position in Figure 1 and moves to along a single path. Find the elapsed time . (2) In an inertial frame, three logarithmic spiral orbits are defined in polar coordinates: where , . Three identical masses are initially at rest at points , , and . Upon release, they move under mutual gravitational forces along their respective frictionless orbits. Due to symmetry, after time , all three particles simultaneously reach radial distance from the origin . Find . Mm M x 2 A 2 + y 2 B 2 =1. mxvxdx dt= dx v =f(x)dx. f(x)x 1 x 2 >x 1 Δt r=r e e aθ ,r=r e e −aθ ,r=r e e −(a/2)θ ,r e >0 a= ln2 2π mP 1 (r=2r e ,θ=2π) P 2 (r=2r e ,θ= 2π 3 )P 3 (r=2r e ,θ= 2π 3 + 4π 3 ) Δtr e OΔt Refinement Feedback Problem Answer: (1.1) ; (1.2) ; (2) f(x)= 1 GMA ⋅ A 2 −Cx A 2 −x 2 Δt= 1 GMA [ Aarcsin ( x A ) +CA 2 −x 2 ] x 2 x 1 Δt=3.083 r e Gm Problem Solution: (1.1) From energy + angular-momentum conservation, . Substituting from the ellipse and : (1.2) Integrate from to : (2) By symmetry the three masses sit at the vertices of an equilateral triangle of side . The net force on each mass points to with magnitude , equivalent to a fixed central mass . The logarithmic spiral has constant pitch (), so the radial fall reduces to a Kepler-like one-dimensional drop from to . Reusing (1.2) with , , : v(x)=GM/A⋅A 2 −Cx/y(x) y(x)dt=dx/v f(x)= 1 GMA ⋅ A 2 −Cx A 2 −x 2 . f(x)x 1 x 2 Δt= 1 GMA [ Aarcsin(x/A)+CA 2 −x 2 ] x 2 x 1 . l=3rOF=Gm 2 /(3r 2 ) M′ =m/3tanφ=1/a r=2r e r=r e A=r e B→0 C=A−r e Δt= π 2 r 3 e 3 Gm ≈3.083 r e Gm . Problem Image: Process Image: Difficulty: 6 Tags: Mechanics, Modern Physics Refinement Feedback Figure 4: Pipeline of PHYSELITE data annotation and a sample problem The PHYSELITE dataset includes two main types of annotations: fine-grained subject categories and difficulty levels. To ensure the accuracy of annotations, we adopted a human-in-the-loop annotation scheme in which LLMs assist human annotators to improve efficiency. The overall annotation process and a sample problem from PHYSELITE dataset are illustrated in Figure 4. More details can be found in appendix D. Difficulty Annotation. We adopt a two-tier annotation scheme to ensure both accuracy and coverage. When the source materials provide difficulty labels from the original problem setters, we use them directly. For problems without such annotations, human experts assign a difficulty level on a 7- point scale, ranging from 1 (relatively easy) to 7 (extremely hard), referencing the length of the reference solution and the LLM inference token count as auxiliary quantitative signals. Inconsistent or ambiguous cases are resolved through discussion among human expert annotators. Subject categorization. We define a unified taxonomy covering the major subfields tested in international physics Olympiads, including mechanics, electromagnetism, thermodynamics, optics 5 and modern physics. For problems drawn from sources with a clear chapter structure, we directly map their chapter labels to our taxonomy. For the remaining problems, we first obtain candidate categories through majority voting over predictions from three advanced MLLMs [20,21,22], which are then verified by human experts. Finally, an LLM is used to normalize all labels into the unified taxonomy. 3.4 Dataset Statistics Table 2 and Figure 3 summarize the composition of PHYSELITE. Mechanics and electromagnetism dominate the problems, reflecting the importance of these two subfields in Olympiad training. Difficulty is concentrated in the medium range, with a long tail of hard problems that probe the upper bound of model capability. Answers are open-ended and take four forms: symbolic expressions, numerical values , qualitative conclusions and equations, preventing models from succeeding through multiple-choice shortcuts. The average problem statement is 135.9 words long, with multi-part Olympiad problems reaching up to 2,619 words. 4 Experiments 4.1 Experimental Setup Evaluation Models. We evaluate a diverse set of models on PHYSELITE, spanning both LLMs and MLLMs, including 8 open-source and 10 closed-source models. These models fall into two categories: 12 System-1 models, which follow a fast, single-pass reasoning paradigm, and 6 System-2 models, which adopt a slow, iterative long CoT reasoning style. Evaluation Details. To establish a human performance baseline, we recruited 15 first-prize-winning students from local secondary schools across different age groups to independently solve the problems, drawing on their usual training records. For multimodal models, the schematic diagram in each problem is included. Each problem is evaluated separately in Chinese and English, with results averaged across the two languages. Decoding hyperparameters and full prompt templates are provided in Appendix F. Scoring. We report two complementary metrics. The Answer Score counts as correct only if its final answer matches the reference. The Process Score is computed in two stages: the judge first decomposes the model’s response into several key derivation steps, then grades each step as fully correct (1.0), partially correct (0.5), or incorrect (0.0). We grade against the model’s own derivation rather than forcing alignment with the reference solution, since physics problems often admit multiple valid solution paths. Each response is independently scored by three LLM judges (GPT-5.2, Claude-Opus-4.6, and Gemini 3-Pro) and the final score takes the mean value. We validate the protocol by manually grading 200 randomly sampled problems with human experts; the mean absolute error between expert and the averaged LLM-judge scores is 0.09 on the 0–1 composite scale. Full judge prompts and agreement analysis are provided in Appendix F. 4.2 Main Results Table 3 summarizes answer-level accuracy and process scores by sub-discipline, averaged across Chinese and English. Overall difficulty. The results presented in Table 3 highlight the inherent difficulty of PHYSELITE. The best-performing model, Grok-4.2, achieves 33.7% answer accuracy, followed by Claude-Opus- 4.6 at 28.1%. Notably, all other models fall below the 30% threshold, underscoring the substantial challenge that PHYSELITE reasoning poses even for advanced MLLMs. The gap between proprietary and open-weight systems is large: the strongest open-source model (Qwen3-VL-235B-A22B) reaches only 20.0%. Almost every model also scores higher in English than Chinese; Qwen3-VL-235B-A22B is the lone exception, plausibly reflecting its Chinese-heavy pre-training. Effect of extended thinking. Across the 6 reasoning and 12 standard models in our evaluation, extended thinking delivers a clear but uneven advantage. Among the top six entries on PhysElite, four are reasoning models (Grok-4.2, o3-mini, GPT-5.2, Gemini-3-Pro), and the strongest reasoning 6 Table 3: Main results on PHYSELITE, averaged across Chinese and English evaluations. Ans.: answer-level accuracy (%). Mech.: Mechanics. E&M: Electromagnetism. Mod.: Modern Physics. Therm.: Thermodynamics. Opt.: Optics. Proc.: process score of step-level evaluation. Best per column in bold, second best underlined. ModelMech.E&MMod.Therm.Opt.Ans.↑Proc.↑ Closed-source Grok-4.2 [23] † 34.026.633.035.425.433.747.6 Claude-Opus-4.6 [20]27.220.924.733.121.328.149.6 o3-mini [24] †⋆ 25.320.426.828.419.626.243.9 GPT-5.2 [22] † 22.118.623.928.417.424.042.6 Gemini-3-Pro [21] † 21.915.520.023.716.121.841.4 Kimi-K2-Thinking [25] † 19.313.215.622.89.620.032.7 Claude-Sonnet-4.5 [20]17.514.718.523.712.618.533.6 Gemini-2.5 [26] † 16.914.220.720.913.517.9 36.8 Qwen-VL-Max [27]14.68.512.016.59.614.330.7 GPT-4o [28]8.86.510.213.810.010.425.0 Open-source Qwen3-VL-235B-A22B [29]19.315.819.222.014.420.0 36.8 DeepSeek-V3 [30] ⋆ 18.710.314.519.310.917.034.6 Qwen3-VL-32B [29]15.210.113.817.812.215.931.5 Qwen3-VL-8B [29]12.25.410.912.66.111.623.2 Qwen2.5-VL-72B [31]9.16.29.111.16.19.821.5 Dolphin-Mistral-24B [32]6.17.510.19.97.08.620.6 LLaMA-3.1-70B [33] ⋆ 4.96.010.57.96.67.520.2 Qwen2.5-VL-7B [31]2.40.50.72.00.51.7 3.7 Human Baseline47.642.354.251.757.348.565.2 † Reasoning model with extended thinking mode. ⋆ Text-only model; evaluated without diagram input. model (Grok-4.2, 33.7%) outperforms the strongest standard model (Claude-Opus-4.6, 28.1%) by 5.6 points. However, the benefit of extended thinking is not uniform: Claude-Opus-4.6 in its default mode surpasses five of the six reasoning models, including GPT-5.2, Gemini-3-Pro, Kimi-K2-Thinking, and Gemini-2.5. This indicates that a sufficiently strong base model without explicit reasoning chains can match or exceed mid-tier reasoning models, and that extended thinking alone is not a substitute for base capability. Process versus answer scores. The two metrics expose different aspects of model behavior. For every model, the average step score is higher than answer accuracy, indicating that models often produce partially valid derivations even when their final answer is incorrect. This gap is most pronounced for weaker systems: Claude-Opus-4.6 attains roughly 2.1 times its answer accuracy on the step metric, while LLaMA-3.1-70B reaches about 3.4 times. Models therefore retain substantial credit for early reasoning steps before failing later in the derivation, which motivates the fine-grained error analysis in Section 5. 4.3 Effect of Process Diagrams By default, we provide each model only with the schematic diagram. To examine whether additional guidance helps, we evaluate six representative models under two augmented settings: (i) supplying the process diagram alongside the schematic, and (i) supplying a textual solution hint alongside the schematic. Both settings are compared against the schematic-only baseline on the same problems. Table 4 reports the results. Adding process diagrams improves both metrics for all six models, with process score gains consistently larger than answer score gains. In contrast, providing textual solution hints yields only marginal changes, and in several cases even degrades answer accuracy. This contrast suggests that visual process diagrams provide substantially stronger guidance for intermediate reasoning than equivalent textual hints, highlighting the unique value of diagrammatic information. 7 Table 4: Effect of extra information on selected models. MetricGPT-5.2Opus-4.6Sonnet-4.5 Gemini-2.5 Qwen3-VL-235B Qwen3-VL-32B Process Diagram Answer 26.7 (+2.7) 28.4 (+0.3)19.1 (+0.6)19.8 (+1.9)22.5 (+2.5)18.5 (+2.6) Process46.0 (+3.4) 51.3 (+1.7)35.9 (+2.3)40.6 (+3.8)40.9 (+4.1)37.0 (+5.5) Solution Hint Answer23.5 (-0.5)28.4 (+0.3)18.5 (+0.0)17.9 (+0.0)19.9 (-0.1)15.7 (-0.2) Process42.8 (+0.2) 50.0 (+0.4)37.0 (+2.3)36.7 (-0.1)36.8 (+0.0)32.4 (+0.9) 5 Analysis We conduct three analyses to characterize where, how, and on which problems modern multimodal LLMs fail: (1) Step-position localization identifies where the first error appears along the derivation chain. (2) Failure-severity decomposition distinguishes execution slips from fundamental reasoning errors. (3) Problem-class stratification examines failure patterns across sub-disciplines and difficulty levels. 5.1 Where the first error appears. We define the first-error index as the 1-based position of the earliest step graded below1.0in the model’s derivation. Stronger models reach deeper into the derivation before failing: Claude-Opus-4.6 makes its first error at step2.48, while weaker open-source models fail near step1.4, with Qwen2.5- VL-7B failing at step1in96.8%of cases. GPT-4o (1.57) and Kimi-K2-Thinking (1.62) sit closer to small open-source models, suggesting that first-error depth tracks base capability rather than vendor or training paradigm. However, depth alone does not determine outcome. Grok-4.2 fully completes 21.8% of problems, roughly 30% more than Claude-Opus-4.6 at similar first-error depth. 5.2 Failure-severity decomposition Soft versus hard failures. For each failed step, we ask whether the judge assigned partial correct or totally wrong. The fraction of soft failures ranges from7%for Qwen2.5-VL-7B to46%for Claude-Opus-4.6, with the four highest values held by closed-source frontier models. This fraction is consistently higher in Chinese than in English across all 18 models, with gaps up to 12 points. Last-step-only failures. Some failures occur only at the final step, with all preceding steps fully correct. These failures peak at7.8%for Claude-Opus-4.6. The pattern tracks capability, as weaker models rarely reach the closing step. The errors involved are typically numerical substitution, sign handling, or final-form simplification, all of which leave the upstream reasoning intact. 5.3 Problem-class stratification Optics is the universal weak point. First-step error rates exhibit a stable cross-model topical ordering: setup failures are most frequent in Optics and least frequent in Thermodynamics, with Mechanics, Electromagnetism, and Modern Physics in between. Among the top eight closed-source models, the Optics−Thermodynamics gap exceeds20percentage points on average, and the same ordering holds for weaker models. The Optics gap suggests that current MLLMs handle geometric and phase-based reasoning less reliably than the algebraic manipulation required by other sub-disciplines. Hard problems are not the worst. Most models show a U-shape with Medium as the lowest-scoring slice, with the Hard-Medium gap exceeding10points for four frontier reasoning models (Gemini-2.5, Claude-Sonnet-4.5, GPT-5.2, Claude-Opus-4.6). Only Grok-4.2, Kimi-K2-Thinking, and Qwen2.5- VL-7B show the expected monotone decrease from Easy to Hard. Human contestants follow the monotone trend, indicating that the inversion reflects training rather than intrinsic difficulty. We attribute the U-shape to targeted post-training on Olympiad competition problems, which lifts Hard performance while leaving the daily-practice Medium subset as a genuinely out-of-distribution test. 8 5.4 Error Analysis Question: A smooth large circular ring of mass is hung vertically from the ceiling with a thin rope. Two small rings of mass , each slide down from the top of the ring to both sides simultaneously from rest, and . When the tension of the thin rope is zero, find (1) the position of the small ring; (2) the magnitude and direction of the support force of the large circular ring on the small ring. m 0 m=1kg m=2m 0 θ N Ground Truth: or ; or θ=60 ∘ 120 ∘ | N | =5N 15N Error Reason. Claude-Opus-4.6 mistook "rope tension = 0" for "the entire system is in free fall" and built a non-inertial-frame analysis around effective gravity = 0. This abandons the actual radial force balance () and energy conservation (), yielding a single spurious root instead of the two correct roots. Reasoning Error / misapplication of physical law. mgcosθ−N=mv 2 /R v 2 =2gR(1−cosθ) cosθ= 2 3 Claude Opus 4.6: Setting up coordinates from the top of the ring; let be the angle of the small ring from the vertical. When , the entire system is in free fall instantaneously, so . From the radial equation in the accelerating frame (with effective gravity since the system is in free fall): Solving the constraint together with $T = 0$ on the large ring: θ T=0 (m 0 +2m)g=m 0 a 0 +2ma small =0 N= mv 2 rel R =2mg(1−cosθ)= 2mg 3 cosθ= 2 3 ,θ≈48.2 ∘ N= 2mg 3 ≈6.67N, radially inward. Physical Law Figure 5: Error distribution of Claude Opus 4.6 and an example of physical law error. We randomly sample100incorrect predictions from Claude-Opus-4.6 and classify each into six error categories. The error taxonomy and proportions are visualized in Figure 5. Among all error types, the dominant failures occur at the front end of the reasoning pipeline (problem reading, principle selection, and geometric setup). Physical perception (31%), misapplication of physical laws (25%), and geometric or visual misreadings (8%) together account for roughly two thirds of all failures, indicating that the model more often fails to set up the problem than to execute the subsequent algebra. Algebraic and symbolic mistakes (18%) and incomplete derivations (13%) make up most of the rest, while purely numerical errors are a distant last at6%. An illustrative physical law error is shown on the right side of Figure 5, where the model incorrectly reads “rope tension is zero” as implying free fall of the entire system, switches to a non-inertial analysis, and misses one of the two correct solution branches. More case studies can be found in the appendix I. 6 Conclusion We presented PHYSELITE, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning, and showed that current frontier LLMs remain far from expert-level performance under both answer-level and process-level metrics. Beyond reporting aggregate accuracy, our step-level protocol reveals where reasoning chains fail and which error types dominate across topics and model families. Our analyses further indicate that a single nominal difficulty score is insufficient for cross-benchmark interpretation: source distribution and problem style alignment can dominate model outcomes, especially when compared with competition-style sets such as iPhO-like evaluations. We hope PHYSELITE can serve both as a rigorous evaluation benchmark and as a process-supervised resource for training stronger scientific reasoners. 7 Limitations and future works Our benchmark has a few limitations: partial-credit scoring introduces minor annotation ambiguity in borderline cases, and our training data may have minor overlap with commercial model datasets. We aim to mitigate these issues in future work, particularly by collecting a fully private dataset to reduce potential data leakage risks. We hope PHYSELITE can serve both as a rigorous evaluation benchmark and as a process-supervised resource for training stronger scientific reasoners. References [1]Song Dai, Yibo Yan, Jiamin Su, Zihao Dongfang, Yubo Gao, Yonghua Hei, Jungang Li, Junyan Zhang, Sicheng Tao, Zhuoran Gao, and Xuming Hu. PhysicsArena: The first multimodal physics reasoning bench- mark exploring variable, process, and solution dimensions. In Findings of the Association for Computational 9 Linguistics: EMNLP 2025, pages 17290–17316, Suzhou, China, November 2025. Association for Computa- tional Linguistics. doi: 10.18653/v1/2025.findings-emnlp.937. URLhttps://aclanthology.org/2025. findings-emnlp.937/. [2] Fangchen Yu, Haiyuan Wan, Qianjia Cheng, Yuchen Zhang, Jiacheng Chen, Fujun Han, Yulun Wu, Junchi Yao, Ruilizhen Hu, Ning Ding, Yu Cheng, Tao Chen, Lei Bai, Dongzhan Zhou, Yun Luo, Ganqu Cui, and Peng Ye. HiPhO: How far are (M)LLMs from humans in the latest high school physics olympiad benchmark?, 2025. URL https://arxiv.org/abs/2509.07894. [3] Lintao Wang, Encheng Su, Jiaqi Liu, Pengze Li, Peng Xia, Jiabei Xiao, Wenlong Zhang, Xinnan Dai, Xi Chen, Yuan Meng, Mingyu Ding, Lei Bai, Wanli Ouyang, Shixiang Tang, Aoran Wang, and Xinzhu Ma. PhysUniBench: A multi-modal physics reasoning benchmark at undergraduate level, 2025. URLhttps: //arxiv.org/abs/2506.17667. [4]Xin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen, Yuchen Yan, Jiaxin Zhang, Shizhe Diao, Can Yang, and Yang Wang. UGPhysics: A comprehensive benchmark for undergraduate physics reasoning with large language models. In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025. URL https://arxiv.org/abs/2502.00334. [5]Shi Qiu, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun, Zeyu Cai, Jiashen Wei, Tianyu Luo, et al. PHYBench: Holistic evaluation of physical perception and reasoning in large language models. In Ad- vances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=brG8FPq1cf. [6]Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. OlympiadBench: A chal- lenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3828–3850, Bangkok, Thai- land, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.211. URL https://aclanthology.org/2024.acl-long.211/. [7] Jiacheng Chen, Qianjia Cheng, Fangchen Yu, Haiyuan Wan, Yuchen Zhang, Shenghe Zheng, Junchi Yao, Qingyang Zhang, Haonan He, Yun Luo, Yufeng Zhao, Futing Wang, Li Sheng, Chengxing Xie, Yuxin Zuo, Yizhuo Li, Wenxauan Zeng, Yulun Wu, Rui Huang, Dongzhan Zhou, Kai Chen, Yu Qiao, Lei Bai, Yu Cheng, Ning Ding, Bowen Zhou, Peng Ye, and Ganqu Cui. P1: Mastering physics olympiads with reinforcement learning, 2025. URL https://arxiv.org/abs/2511.13612. [8]IPhO unofficial: International physics olympiad problems and solutions.https://ipho-unofficial. org, 2026. Accessed: 2026-05-07. [9] Asian physics olympiad. http://asianphysicsolympiad.org, 2026. Accessed: 2026-05-07. [10]Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=v8L0pN6EOi. [11] Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems (NeurIPS), volume 35, pages 2507–2521, New Orleans, LA, 2022. Curran Associates, Inc. URL https://arxiv.org/abs/2209.09513. [12] Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating mathematical reasoning of foundation models in visual contexts. In The Twelfth International Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum?id=KUNzEQMWU7. [13]Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9556–9567, 2024. [14]Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R. Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. SciBench: Evaluating college-level scientific problem-solving abilities of large language models. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. URL https://proceedings.mlr.press/v235/wang24z.html. 10 [15]Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Haonan He, Ning Ding, Yu Cheng, Shuyue Hu, Lei Bai, Dongzhan Zhou, Ganqu Cui, and Peng Ye. Scaling physical reasoning with the PHYSICS dataset. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=1lo778KztK. [16]Kaiyue Feng, Yilun Zhao, Yixin Liu, Tianyu Yang, Chen Zhao, John Sous, and Arman Cohan. Physics: Benchmarking foundation models on university-level physics problem solving. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 11717–11743, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findings-acl.610. URL https://aclanthology.org/2025.findings-acl.610/. [17]Xinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang, Chengyou Jia, Basura Fernando, Mike Zheng Shou, Lingling Zhang, and Jun Liu. PhysReason: A comprehensive benchmark towards physics-based reasoning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16593–16615, Vienna, Austria, July 2025. Association for Computational Linguistics. URL https://aclanthology.org/2025.acl-long.811/. [18]Hui Shen, Taiqiang Wu, Qi Han, Yunta Hsieh, Jizhou Wang, Yuyue Zhang, Yuxin Cheng, Zijian Hao, Yuansheng Ni, Xin Wang, Zhongwei Wan, Kai Zhang, Wendong Xu, Jing Xiong, Ping Luo, Wenhu Chen, Chaofan Tao, Zhuoqing Mao, and Ngai Wong. PhyX: Does your model have the “wits” for physical reasoning?, 2025. URL https://arxiv.org/abs/2505.15929. [19] Shunfeng Zheng, Yudi Zhang, Meng Fang, Zihan Zhang, Zhitan Wu, Mykola Pechenizkiy, and Ling Chen. Benchmarking foundation models with retrieval-augmented generation in olympic-level physics problem solving. In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025. URLhttps: //arxiv.org/abs/2510.00919. [20]Anthropic. System card: Claude Opus 4 & Claude Sonnet 4.https://w-cdn.anthropic.com/ 6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf, 2025. [21] Google DeepMind.Gemini 3 Pro model card.https://storage.googleapis.com/ deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf, 2025. [22] OpenAI. GPT-5 system card. https://cdn.openai.com/gpt-5-system-card.pdf, 2025. [23]xAI. Grok 4 model card.https://data.x.ai/2025-08-20-grok-4-model-card.pdf, 2025. Last updated August 20, 2025. [24]OpenAI. OpenAI o3-mini system card.https://cdn.openai.com/o3-mini-system-card-feb10. pdf, 2025. [25] Kimi Team. Kimi K2: Open agentic intelligence, 2025. URL https://arxiv.org/abs/2507.20534. [26]Gemini Team, Google. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. URL https://arxiv.org/abs/2507.06261. [27]Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023. URL https://arxiv.org/abs/2308.12966. [28]OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, et al. Gpt-4o system card, 2024. URL https://arxiv.org/abs/2410.21276. [29] Qwen Team. Qwen3-VL technical report, 2025. URL https://arxiv.org/abs/2511.21631. [30] Aixin Liu et al. DeepSeek-V3 technical report, 2024. URL https://arxiv.org/abs/2412.19437. [31] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, et al. Qwen2.5-VL technical report, 2025. URL https://arxiv.org/abs/2502.13923. [32] CognitiveComputations.Dolphin3.0Mistral24B.https://huggingface.co/ cognitivecomputations/Dolphin3.0-Mistral-24B, 2025. Hugging Face model card. [33] Aaron Grattafiori et al. The Llama 3 herd of models, 2024. URLhttps://arxiv.org/abs/2407.21783. 11 A More Detailed Statistics about PHYSELITE This appendix provides additional details about PHYSELITE, following the same order as the main paper: dataset composition, construction and annotation, evaluation protocol, scoring, and extended analysis. We use the appendix to make the benchmark easier to inspect and reproduce, while keeping the main text focused on the central findings. A.1 Distribution of Text Length Questions in PHYSELITE are presented in English or Chinese. As shown in Table 2, the longest question in PHYSELITE spans [] words, with an average length of [] words. Figure 6 further illustrates the distribution of text lengths, highlighting the diversity of PHYSELITE. The length of English and Chinese questions is counted in words and Chinese characters, respectively. Figure 6: The distribution of the length per problem in PHYSELITE. A.2 Difficulty and Answer Types Each problem is assigned a difficulty score on a 1–7 scale. In the main experiments, we group these labels into three ranges: Easy (1–3), Medium (4–5), and Hard (6–7). The resulting dataset is concentrated around the medium range, with a long hard tail. This is expected for elite daily-practice materials: many problems are designed to develop reusable reasoning skills rather than simply reproduce the hardest official contest items. Table 5: Difficulty and answer-type distribution of PHYSELITE. TypeCount / ShareDescription Easy (1–3)28%Shorter or more direct derivations Medium (4–5)55%Multi-step Olympiad-style reasoning Hard (6–7)17%Long derivations or advanced modeling Symbolic expressions3,700Closed-form symbolic answers Numerical values2,810Numeric answers with units or constants Qualitative conclusions978Physical judgments or comparisons Equations523Relations, constraints, or derived equations The open-ended answer format is a central part of the benchmark design. Unlike multiple-choice settings, models must produce the final physical quantity or relation directly. This makes answer-level scoring stricter and gives the process score a larger diagnostic role. A.3 Distribution of Image Number per question As shown in Figure 7, the majority of questions in the PHYSELITE dataset (58.7%) are accompanied by a single image. The remaining questions are associated with multiple images, with the number of images ranging from two to eleven. 12 Figure 7: The distribution of the number of images per question in PHYSELITE. B Introduction of our Fine-grained Physics Subjects In this section, we introduce the five physics subject categories defined in our PHYSELITE benchmark. These categories are designed to reflect the diverse reasoning skills required for Olympiad-level physical problem solving. Rather than treating physics as a monolithic domain, we present a structured taxonomy that decomposes the problem space into conceptually distinct yet complementary components. Each problem is assigned one or more subject labels according to the main conceptual obstacle in its reference solution. The taxonomy is intentionally broad rather than overly fine-grained, because many Olympiad problems combine multiple physical principles, and a coarser partitioning yields more interpretable performance diagnostics. Figure 8 to Figure 12 illustrate representative image examples corresponding to each category. Our classification is based on both the underlying physical principles (e.g., conservation laws, field theory, statistical reasoning) and the characteristic reasoning patterns elicited by each domain (e.g., model construction, equivalent-circuit reduction, ray geometry analysis). This allows for more interpretable performance diagnostics and supports targeted model evaluation. The five categories span a broad range of physics competencies, ensuring coverage of foundational classical topics as well as advanced modern physics skills. 1. Mechanics. This category focuses on the analysis of motion, forces, and energy in classical systems, encompassing kinematics, dynamics, rigid-body motion, oscillations, waves, gravitation, fluids, and energy–momentum reasoning. Problems typically require constructing a physical model from a diagram, choosing appropriate coordinate systems, and combining force balance, conservation laws, and geometric constraints. These problems often serve as foundational tests of physical reasoning and demand the integration of mathematical manipulation with physical intuition. Example task: Determine the maximum compression of a spring when a block slides down an inclined plane and collides with it, given the friction coefficient and initial height. 2. Electromagnetism. This subject class targets reasoning about electric and magnetic phenomena, including electrostatics, direct- and alternating-current circuits, magnetic fields, electromagnetic induction, and charged-particle motion. The focus lies in field-level reasoning and the analysis of dynamical responses to electromagnetic interactions. Typical reasoning patterns include field superposition, equivalent-circuit reduction, Lorentz-force dynamics, and time-varying flux analysis. Example task: Compute the induced EMF in a rotating conducting rod within a non-uniform magnetic field, or determine the trajectory of a charged particle in crossed electric and magnetic fields. 13 3. Thermodynamics. This category involves reasoning about heat, work, and the macroscopic behavior of thermal systems, particularly those that rely on state variables and process constraints. It covers heat transfer, ideal-gas processes, phase changes, entropy-related reasoning, and cyclic processes. Solutions typically require careful interpretation of state variables and process paths rather than purely algebraic manipulation, and demand an understanding of how thermodynamic constraints couple with mechanical or chemical configurations. Example task: Analyze the efficiency of a non-standard thermodynamic cycle on aP–Vdiagram, or determine the equilibrium temperature when two gases at different states are connected through a valve. 4. Optics. This class includes problems involving the propagation, interaction, and interference of light, emphasizing the ability to switch between geometric and wave-based descriptions. It spans geometric optics, lenses and mirrors, interference, diffraction, polarization, and optical-path reasoning. Solvers must mentally trace rays through optical systems or track phase relations across multiple paths. As discussed in the main analysis, optics is a persistent weak point for current models, partly because the setup often depends on precise ray geometry or phase relations that are difficult to extract from text and diagrams alone. Example task: Determine the position and magnification of the final image formed by a two-lens system, or compute the fringe spacing in a modified double-slit experiment with an inserted glass plate. 5. Modern Physics. This is the most specialized category, featuring problems that go beyond classical mechanics and electromagnetism, including special relativity, quantum phenomena, atomic physics, and nuclear physics. It involves applying principles such as Lorentz transformations, energy–level quantization, photon–matter interactions, and nuclear decay laws. Problems in this category often require specialized principles that are less frequently encountered in routine high-school datasets, and bridge classical reasoning with the conceptual frameworks of modern physics. Example task: Compute the wavelength shift of a photon undergoing Compton scattering, or determine the relativistic energy of a particle produced in a decay process. 14 A Sample of Mechanics Problem Problem_Text: In some lakes, a peculiar phenomenon called seiche (water oscillation) can be observed. This usually occurs in shallow, long and narrow lakes, where the whole body of water appears to move, similar to how you would carry a cup of coffee to a guest. This should not be confused with surface waves on the lake. We model this using a rectangular container of length and water depth . Initially, the water surface makes a small angle with the horizontal plane. During oscillation, the surface remains flat and rotates about a horizontal axis located at the midpoint of the container's length. 1. Establish a model for the water motion and derive an expression for the oscillation period. The initial condition is shown in the Figure. Assume . 2. Two tables show the oscillation times for different water depths in containers of two different lengths. Use appropriate methods to check whether your derived formula matches the experimental data well, and comment on the applicability of this physical model. 3. Figure shows measurements taken on Lake Vattern in Sweden. The two curves represent measurements from the northern end (Bastedalen) and southern end (Jönköping). The lake has a length of and an average depth of . What time scale should be used in this case? Lh T ξ≪h 123km50m Problem_Answer: The oscillation period is: The oscillation period is: . For Lake Vättern, the period is approximately . T=1.15⋅ πL 3gh 3.2h Problem_Solution: Treat the seiche as a rigid one-dimensional sloshing of the whole body of water. Place the origin at the container midpoint. When the surface tilts by , decompose the liquid into a left wedge, a right wedge, and the underlying rectangular block. Computing the centroids of the three pieces, the centre of mass moves to Since , is second-order and the motion is essentially horizontal with . Energy of the system: Conservation gives the SHO equation The experimental tables show a systematic 15% offset, fixed by a correction factor: Plugging , gives . O ξ x C = Lξ 6h ,y C = ξ 2 6h . ξ≪hy C · x C = L 6h · ξ E p = m 3 gξ 2 6h ,E k = 1 2 m 3 · x 2 C = m 3 L 2 72h 2 · ξ 2 . d dt (E p +E k )=0 · ξ+ 12gh L 2 ξ=0⇒T theory = πL 3gh . T=1.15⋅ πL 3gh .L=123kmh=50mT Vattern ≈3.2h Problem Image:Process Image: Tag: Mechanics Difficulty: 7 Figure 8: A sample of Mechanics Problem 15 A Sample of Electromagnetism Problem Problem_Text: A particle source located at the origin O emits identical charged particles of mass m and charge q with initial velocity along the positive y-axis. The interaction between particles and gravity are negligible. 1. If there is a uniform magnetic field B along the positive z-axis in the region where the particles move, the charged particles will perform circular motion in the xy-plane. If a uniform weak electric field E along the positive y-axis is added, the circular orbit will drift along the positive or negative x-direction. For the cases q>0 and q<0, qualitatively determine the direction of this drift, sketch the trajectory, and calculate the corresponding drift velocity (appropriate approximations are allowed). 2. If the electric field E in the first part is not necessarily weak but still uniform and along the positive y-axis, derive rigorously the time dependence of the velocity components and along the x- and y-directions. Take the moment of emission from the origin as the initial time . 3. Remove the electric field in parts 1 and 2. Suppose the magnetic field is non-uniform, with a small linear gradient equal to a positive constant. Then the charged particles will also drift along the x-direction. Qualitatively determine the drift direction for positively and negatively charged particles, and sketch the trajectories. v 0 v x v y t=0 dB dy Problem_Answer: 1. Drift direction: Regardless of whether or , the drift is along the positive x-axis. Drift velocity . 2. Exact solution: where , , , and is in the fourth quadrant. 3. Drift direction due to magnetic field gradient: Positive charge () drifts along the negative x-axis; Negative charge () drifts along the positive x-axis. q>0q<0 v D = E B v x =v* 0 sin(ωt+φ)+ E B , v y =v* 0 cos(ωt+φ), ω= qB m v* 0 =v 2 0 + ( E B ) 2 tanφ=− E Bv 0 φq>0 q<0 Problem_Solution: 1. Weak . Without each particle circles at radius . Adding shrinks on the upper half and grows it on the lower half, shifting the orbit centre along for both signs of — no charge separation. Writing the equation of motion in components and demanding a steady offset gives 2. Arbitrary . The system , admits the harmonic + offset solution with , : , (4th quadrant). No drift for any . 3. Gradient , no . Field stronger on top → smaller ; weaker on bottom → larger . Tracking the asymmetry: drifts along , drifts along . Unlike the drift, this gradient drift separates the two charge species. EER=mv 0 / | q | BE ̂ yR + ̂ xq md⃗v/dt=q(⃗v× ⃗ B+ ⃗ E)v x v D = E B ,ω= qB m .E · v x =ωv y · v y =(q/m)E−ωv x v x =v* 0 sin(ωt+φ)+ E B ,v y =v* 0 cos(ωt+φ);v x (0)=0v y (0)=v 0 v* 0 =v 2 0 +(E/B) 2 tanφ=−E/(Bv 0 ) ̂ yEdB/dy>0ER Rq>0− ̂ xq<0+ ̂ xE×B Problem Image:Process Image: Tag: Electromagnetism Difficulty: 7 Figure 9: A sample of Electromagnetism Problem 16 A Sample of Optics Problem Problem_Text: You may have noticed that in darkness, when a cat is within the light beam of a headlamp, its eyes appear very bright, see the photo below (left). This phenomenon can be modelled by a lens setup, see the photo on right, and the diagram beneath the photos. The photo on right was taken by a digital single-lens reflex camera. The light intensity at the camera sensor pixels marked by a red line (in the photo) is shown in the graph below: the log base 10 of the light intensity (measured as the number of photons caught by each pixel) is plotted against the x-coordinate, with the pixels' side length serving as the unit length. The lens modelling cat eyes can be treated as an ideal thin lens of focal length and diameter ; however, you should keep in mind that the given graph shows real measurement data, and the lens has certain non-ideal features. Most importantly, partial reflections of brightly lit areas from the lens surfaces may decrease the contrast: dark areas seen through the lens appear less dark than they actually are; this effect can be neglected for the camera lens, but not so for the lens serving as a model of a cat's eye. Based on the given data, estimate (with the accuracy of ca 20%) the distance between the axis of the camera and the axis of the lamp (which can be considered as a point source) if . f=55m D=39m h L≫f Problem_Answer: Based on image readings and geometric optics approximate calculations, the specific numerical estimate of is approximately . h13cm Problem_Solution: The lens-and-paper system acts as a retroreflector: light from the lamp focuses on the paper sheet (at the back focal plane), scatters, and returns through the lens to the camera. Because , incoming rays are nearly parallel and the focal spot sits off-axis at Reading the logarithmic intensity graph, the lens disc has diameter in pixel units, giving the pixel-to-m scale. The bright retro-reflected plateau, whose width represents twice the spot offset on the paper, lets us extract in m. Inverting the geometric relation with the measured and the known yields L≫f y=fθ=f h L . D=39m yh=yL/fy f=55m h≈13cm(acceptable range 10-17cm). Problem Image: Process Image: Tag: Optics Difficulty: 6 Figure 10: A sample of Optics Problem 17 A Sample of Thermodynamics Problem Problem_Text: CDROM Spectrometer This experiment aims to plot how the conductance of a photoresistor varies with wavelength within the visible spectrum. Uncertainty indications are not required. Conductance Resistance (Unit: Siemens, ) [Experimental Content] (1)Use a concave reflection grating (made from a CDROM) to produce a focused first-order spectrum from light bulb A (12V 50W tungsten filament bulb). (2) Measure and plot the relationship between the conductance of the photoresistor and wavelength in the first-order spectrum. (3) Prove that the luminous behavior of the filament in bulb A approximates an ideal blackbody. (4) Determine the temperature of the bulb's filament when connected to a 12V power supply. (5) After calculating the energy distribution of the spectrum emitted by bulb A, correct the curve of conductance versus wavelength. Gλ G=1/1 S=1 Ω −1 Problem_Answer: The filament temperature is approximately ; the conductivity correction requires processing in conjunction with the Planck formula. 2940K Problem_Solution: The CDROM acts as a concave reflection grating ( line spacing). With normal incidence, the first-order diffraction angle is set by , so sliding the photoresistor along the focused spectrum and recording gives . To verify blackbody behaviour of the tungsten filament, use (Stefan– Boltzmann) and (linear resistivity). Eliminating yields A measured slope in the – plot confirms ideal-blackbody scaling. Filament temperature at : from the tabulated tungsten curve and the cold resistance at room temperature, scale to operating to obtain Finally, divide out the source spectrum: the photoresistor signal is , so the intrinsic spectral response is with the Planck function. Plot the corrected – curve. d=1/620m sinθ=λ/d R(θ)G(λ)=1/R(λ)P=VI∝T 4 R=V/I∝TTV 3 ∝I 5 ,⟹lnV= 5 3 lnI+const. ≈1.67lnVlnI12V ρ(T)R C1 RT 12V ≈2940K. G(λ)∝B(λ,T)⋅σ(λ) G′ (λ)=G(λ)/B(λ,2940K),BG′ λ Problem Image: Process Image: Tag: Thermodynamics, Optics, Modern Physics Difficulty: 6 Figure 11: A sample of Thermodynamics Problem 18 A Sample of Modern Physics Problem Problem_Text: Find the orbital motion time. (1) In an inertial frame, a mass is fixed, and a mass moves under the gravitational influence of along the elliptical orbit shown in Figure, described by: The speed of along the -axis is denoted as , and the infinitesimal displacement in the -domain is . The elapsed time is given by: (1.1) Find ; (1.2) Suppose the particle starts at position in Figure 1 and moves to along a single path. Find the elapsed time . (2) In an inertial frame, three logarithmic spiral orbits are defined in polar coordinates: where , . Three identical masses are initially at rest at points , , and . Upon release, they move under mutual gravitational forces along their respective frictionless orbits. Due to symmetry, after time , all three particles simultaneously reach radial distance from the origin . Find . Mm M x 2 A 2 + y 2 B 2 =1. mxv xdx dt= dx v =f(x)dx. f(x) x 1 x 2 >x 1 Δt r=r e e aθ ,r=r e e −aθ ,r=r e e −(a/2)θ ,r e >0 a= ln2 2π mP 1 (r=2r e ,θ=2π) P 2 (r=2r e ,θ= 2π 3 )P 3 (r=2r e ,θ= 2π 3 + 4π 3 ) Δt r e OΔt Problem_Answer: (1.1) ;(1.2) ;(2) f(x)= 1 GMA ⋅ A 2 −Cx A 2 −x 2 Δt= 1 GMA [ Aarcsin ( x A ) +CA 2 −x 2 ] x 2 x 1 Δt=3.083 r e Gm Problem_Solution: (1.1) Combine energy conservation with angular-momentum conservation on the ellipse. Eliminating the perpendicular component gives the $x$-speed $ Substituting into : (1.2) Direct integration: (2) Three identical masses sit at the vertices of an equilateral triangle of side . Each pairwise force is , and summing the two contributions on a single particle gives a net inward force So each mass falls in a Kepler-like central field of effective mass . The logarithmic- spiral orbit has constant pitch (), reducing the radial fall to a 1D drop from to . Reusing (1.2) with , , , : 1 2 mv 2 −GMm/r=−GMm/(2A) L=mv ⊥ r=mGMA(1−e 2 ) v(x)= GM A A 2 −Cx y(x) ,C= A 2 −r 2 e r 2 e . y(x) 2 =B 2 (1−x 2 /A 2 )dt=dx/v f(x)= 1 GMA ⋅ A 2 −Cx A 2 −x 2 . Δt= 1 GMA [ Aarcsin x A +CA 2 −x 2 ] x 2 x 1 . l=3rf=Gm 2 /l 2 F=2fcos30 ∘ = Gm 2 3r 2 = GmM′ r 2 ,M′ = m 3 . M′ tanφ=1/ar=2r e r=r e A=r e B→0x 1 =−r e x 2 =0 Δt= π 2 r 3 e GM′ = π 2 r 3 e 3 Gm ≈3.083 r e Gm . Problem Image: Process Image: Tag: Modern Physics, Mechanics Difficulty: 6 Figure 12: A sample of Modern Physics Problem 19 C More Detailed Construction of PHYSELITE C.1 Data Collection Pipeline We start from approximately 15,000 open-ended problems collected from daily learning and practice materials of physics contestants, together with private exam sets used in advanced training. The source materials are image-based and typically contain the problem statement, one or more diagrams, the step-by-step solution, and the final answer. Compared with official competition-only datasets, these materials provide a broader picture of the problem styles used in sustained Olympiad preparation. The construction pipeline has four stages. First, source images are screened to retain problems with complete statements, solutions, and answers. Second, the problem text and solution are manually transcribed into structured records, with mathematical expressions converted into L A T E X to preserve symbolic precision. Third, diagrams are extracted and linked to the corresponding problem identifiers. Finally, duplicated or near-duplicated problems are removed through fuzzy matching and human review. After this process, the dataset size is reduced from roughly 15,000 candidates to 11,586 curated problems. C.2 Dataset Format Each problem is stored as a structured record. The key fields are: • problem_id: a unique identifier assigned to each problem. • problem_text_cn /problem_text_en: the Chinese and English versions of the problem statement. • problem_solution_cn/problem_solution_en: the full step-by-step solution in both languages. • problem_answer: the final answer used for answer-level scoring. • problem_img: filenames of schematic diagrams that describe the physical setup. • problem_img_process: filenames of process diagrams that visualize intermediate solution steps. • subject and difficulty: the normalized subject label and 1–7 difficulty label. This format is designed to support both direct answer evaluation and process-level analysis. The same record can be used to evaluate text-only models, multimodal models with schematic diagrams, and process-supervised methods that require reference derivations. C.3 Source Diversity and Representativeness Many high-difficulty physics benchmarks rely heavily on official competition problems. Such questions are valuable, but they are designed for specific contests, years, and scoring rubrics. As a result, their topic and reasoning distributions can be narrow. PHYSELITE instead emphasizes daily contestant training materials and private exam sets. This choice broadens the assessment scenarios: the dataset includes routine elite-practice problems, medium-difficulty problems that teach transferable modeling patterns, and hard problems that approach national or international Olympiad style. This source design is important for interpreting model behavior. In the main analysis, we observe a U-shaped difficulty trend for many models: medium problems can be harder for models than the nominal hard subset. This suggests that source style and solution-template familiarity influence model performance in addition to intrinsic physics difficulty. D Annotation Protocol The annotation process contains three stages: transcription, difficulty annotation, and subject catego- rization. Human annotators are responsible for final decisions, while LLMs are used only to improve efficiency in candidate generation or label normalization. 20 Complete Olympiad Level Sample Problem_Text: When a strong laser beam passes through a small transparent object, refraction exerts a force on the object. Consider a small glass prism with apex angle , base length , thickness , refractive index , and density , placed in a laser beam propagating along the horizontal x-axis. The prism does not rotate; its vertex points toward the incoming beam, its triangular faces are parallel to the xy plane, its base parallel to yz. Air has . All surfaces are anti-reflective coated. The laser intensity is uniform along z but decreases linearly in y from a maximum at to zero at . (1) When the laser hits the upper surface (Figure 3), find the deflection angle . (2) Move the prism's apex by displacement along (). Express the net force components and in terms of , , h, w, . Plot vs . (3) Laser beam: in . Prism: , , , , . With apex at , what total laser power is needed to balance gravity? (4) Same setup in zero gravity. Set , place the apex at , and release. Find the oscillation period. A=π−2a2hwnρ n a =1 I 0 y=0y= ± 4hθ y o ̂ y | y o | ≤3hF x F y I 0 θy o F x,y y o 1m×80μmz×ya=30 ∘ h=10μmn=1.5w=1mρ=2.5g/cm 3 y o =−h/2=−5μm I 0 =10 5 W/m 2 y o =h/20 Problem Image: Process Image: Tag: Optics, Modern Physics, Mechanics Difficulty: 6 Problem_Answer: (1) ; (2) , ; (3) ; (4) . θ=arcsin nsin [ a−arcsin(sina/n) ] F x = hwI 0 c ( 7 4 − y 2 o 4h 2 ) (1−cosθ) F y =− hwI 0 y o c ( 1− y o 2h ) sinθ P≈33.2WT≈1.12×10 −2 s Problem_Solution: (1) Two refractions: incidence angle on the upper face equals a, so Snell gives the internal angle. The internal ray hits the base at angle , refracting out to (2) Each photon carries momentum ; after refraction it picks up a transverse component (sign flips for the lower face). Summing over photon flux on the upper / lower faces, with , the powers intercepted there, Integrating over each face for (lower face straddles ): so for — the beam pulls the prism back to the axis (transverse trapping). (3) With , : . Mass , weight . Solving gives, and the total beam power (4) For , — linear restoring force, hence SHO with β=arcsin(sina/n) a−βθ=arcsin [ nsin(a−β) ] =arcsin nsin [ a−arcsin(sina/n) ] . E/c−(E/c)sinθ ̂ j P u P l F= 1 c [ (P u +P l )(1−cosθ) ] ̂ i+ 1 c [ (P u −P l )sinθ ] ̂ j. I(y)=I 0 (1− | y | /4h)0<y o <hy=0 P u +P l =hwI 0 ( 7 4 − y 2 o 4h 2 ) ,P u −P l =− hwI 0 y o 2h ( 1− y o 2h ) , F x = hwI 0 c ( 7 4 − y 2 o 4h 2 ) (1−cosθ),F y =− hwI 0 y o c ( 1− y o 2h ) sinθ. F y <0y o >0 a=30 ∘ n=1.5θ=15.9 ∘ m=ρh 2 wtana=1.44×10 −10 kgmg=1.41×10 −9 N | F y (−h/2) | =mgI 0 =8.29×10 8 W/m 2 P= 1 2 I 0 ⋅S=33.2W,S=1m×80μm. | y o | ≪h F y ≈− wI 0 sinθ 2c y o T=2π 2cρh 2 tana I 0 sinθ =1.12×10 −2 s. Figure 13: A complete sample problem of our Dataset D.1 Transcription Guidelines Annotators are asked to act as faithful transcribers. The goal is to convert the source image into structured text without altering the underlying physics content. Mathematical expressions are converted to L A T E X; original symbols, subscripts, and naming conventions are preserved whenever possible. Multi-part problems are kept within a single record, and all sub-answers are included in the final answer field. If a portion of the source image is unclear, annotators mark the span for review rather than guessing. When source diagrams contain labels or annotations, these are preserved in the image and also reflected in the bilingual diagram description when needed. D.2 Difficulty Annotation Each problem receives an integer difficulty label from 1 to 7. When the source material provides a difficulty label from the original problem setter, we use it directly. Otherwise, human experts assign the label using the following anchors: • 1–2: relatively easy problems solvable in a few short steps. 21 •3–4: introductory Olympiad or routine undergraduate-style problems requiring multiple principles or a non-trivial setup. •5–6: hard Olympiad-style problems requiring careful modeling, multi-step derivation, or coordination of several physical concepts. •7: extremely hard problems comparable to the most demanding national or international Olympiad items. Annotators may adjust the rating by one level when a problem combines several subfields, has an un- usually long solution, requires advanced mathematical techniques, or depends on specialized physical knowledge. Ambiguous cases are resolved through discussion among human expert annotators. D.3 Subject Categorization Each problem is assigned to one of five categories: Mechanics, Electromagnetism, Thermodynamics, Optics, or Modern Physics. If the source has a clear chapter structure, the chapter label is mapped directly to the taxonomy. Otherwise, three advanced MLLMs propose candidate labels, and the majority-vote candidate is reviewed by human experts. The human judgment is final. After all labels are assigned, an LLM-based normalization pass standardizes surface variants such as “E&M” and “Electromagnetism” without changing the underlying category assignment. For cross-domain problems, annotators select the category corresponding to the core reasoning challenge. For example, if a problem combines charged-particle motion with mechanical oscillation, the label depends on whether the decisive step is electromagnetic force modeling or mechanical dynamics. E Annotator and Human-Level Evaluation Human experts participate in two parts of the benchmark. First, they resolve uncertain difficulty labels and subject labels during dataset construction. Second, they validate the process-scoring protocol by manually grading a random sample of 200 model responses. The mean absolute error between expert scores and the averaged LLM-judge scores is 0.09 on the 0–1 composite scale, supporting the use of the judge ensemble for large-scale scoring. We also report a human baseline in Table 3. Human contestants reach 48.5% answer accuracy and 65.2 process score overall, with answer accuracy of 47.6% in Mechanics, 42.3% in Electromagnetism, 54.2% in Modern Physics, 51.7% in Thermodynamics, and 57.3% in Optics. The gap between this baseline and the best MLLM result shows that PHYSELITE remains challenging even for frontier models. F Evaluation Details F.1 Evaluation Settings For multimodal models, the schematic diagram is included together with the problem text. For text- only models, only the textual problem statement is provided. Each problem is evaluated separately in Chinese and English, and the reported results average across the two languages. The evaluation includes 18 models: 10 closed-source models and 8 open-source models. Six of the evaluated models are treated as System-2 reasoning models with extended thinking mode, while the remaining twelve follow a standard System-1 response style. Text-only models are marked in Table 3 and are evaluated without diagram input. F.2 Prompt for Response Generation To ensure the model provides accurate responses, we use the following CoT prompt for both Multi- modal Large Language Models and Large Language Models. 22 You are an expert in physics. Solve the following Olympiad-level physics problem step by step. Show the full derivation, including the governing equations, simplifications, and any approximations you invoke. Conclude your response with a line that begins with “FINAL ANSWER:” and contains only the final result. Each judge receives the model’s full response verbatim and identifies the final result as part of its grading; we do not perform programmatic answer extraction. Multi-line, piecewise, and multi-part answers are handled by the judges directly, with a partial verdict available for problems where only some sub-parts are correct. F.3 Prompt for Answer Evaluation Our evaluation is conducted using three LLM judges: GPT-5.2, Claude-Opus-4.6, and Gemini-3-Pro. The judges use a unified prompt template that takes the reference answer, the reference solution, and the model’s full response as input, and returns a structured verdict. For multi-part problems, the judge can return a partially correct verdict when some sub-parts are correct and others are not. The full judge prompt is shown in the box below. You are an expert physics solution judge. Given a student’s response and the reference materials for a physics problem, determine whether the student’s final answer is correct. Reference Answer reference Reference Solution solution Student’s Response response Judging Protocol Correctness is decided on the student’s final result(s). When the reference answer is empty, extract the correct result from the reference solution. 1.Numerical answers are accepted within a 5% relative tolerance, unless the problem requires an exact integer or rational value. 2.Symbolic and algebraic answers are accepted in any mathematically equivalent form (e.g., reordered variables, with or without simplification). 3.For multi-part problems, return “correct” only if every sub-part is correct; if some sub-parts are correct and others are not, return “partially_correct”. 4.Disregard differences in units notation, formatting, and intermediate steps. Only the final result is evaluated. Return a single JSON object (no markdown fences) in exactly this schema: "verdict": "correct" | "partially_correct" | "incorrect", "reasoning": "<one- to three-sentence explanation>" G Scoring Protocol Details G.1 Answer Score The answer score measures whether the model’s final answer matches the reference answer. Because PHYSELITE contains symbolic expressions, numerical values, qualitative conclusions, and equations, scoring allows mathematically equivalent forms when they express the same physical result. Numeri- cal answers are checked with their units and constants when these are part of the required answer. Answer accuracy is intentionally strict. A response can receive partial credit through the process 23 score even when the final answer is wrong, but it is counted as answer-correct only when the final result is equivalent to the reference. G.2 Process Score The process score evaluates the derivation. Each model response is first decomposed into key reasoning steps. Each step is then graded as: • 1.0: the step is physically and mathematically correct. • 0.5: the step uses the right idea but contains a local slip, missing condition, or incomplete execution. • 0.0: the step is physically wrong, mathematically invalid, or irrelevant to the problem. We grade the model’s own derivation rather than forcing it to align with the reference solution, since physics problems often admit multiple valid solution paths. G.3 Judge Ensemble Each response is scored independently by three LLM judges: GPT-5.2, Claude-Opus-4.6, and Gemini 3-Pro. The final score is the mean of the three judge scores. The judge prompt instructs each model to identify the major reasoning steps, evaluate each step against the problem statement and reference answer, and avoid penalizing alternative valid derivations. The protocol is validated against human expert grading on 200 randomly sampled problems. The averaged LLM-judge score reaches a mean absolute error of 0.09 on the 0–1 composite scale relative to the expert score. H Additional Experimental Analysis H.1 Model Groups Table 6 summarizes the evaluated model groups. The table mirrors the setting in the main experiments and makes explicit which systems are treated as reasoning models and which systems are evaluated as text-only models. Table 6: Model grouping used in the evaluation. GroupModelReasoning ModeInput Modality Closed-source Grok-4.2System-2Text+image o3-miniSystem-2Text-only GPT-5.2System-2Text+image Gemini-3-ProSystem-2Text+image Kimi-K2-ThinkingSystem-2Text+image Gemini-2.5System-2Text+image Claude-Opus-4.6System-1Text+image Claude-Sonnet-4.5System-1Text+image Qwen-VL-MaxSystem-1Text+image GPT-4oSystem-1Text+image Open-source Qwen3-VL-235B-A22BSystem-1Text+image Qwen3-VL-32BSystem-1Text+image Qwen3-VL-8BSystem-1Text+image Qwen2.5-VL-72BSystem-1Text+image Dolphin-Mistral-24BSystem-1Text+image Qwen2.5-VL-7BSystem-1Text+image DeepSeek-V3System-1Text-only LLaMA-3.1-70BSystem-1Text-only 24 H.2 Language Effects Each problem is evaluated in Chinese and English. The main results indicate that most models score higher in English, while Qwen3-VL-235B-A22B is the main exception. This pattern suggests that bilingual evaluation is not merely a translation convenience: it reveals language-specific differences in reasoning, notation handling, and physics expression. The Chinese setting is especially important for PHYSELITE, because the source problems come from Chinese Olympiad training materials. H.3 Difficulty Effects The difficulty analysis in the main paper groups problems into Easy (1–2), Medium (3–5), and Hard (6–7). Many models show a U-shaped trend, with Medium problems producing lower scores than the Hard subset. We interpret this as evidence that nominal difficulty alone is insufficient for benchmark interpretation. Hard problems may align more closely with official Olympiad-style templates encountered during post-training, while medium daily-practice problems can be less familiar to current models. H.4 Sub-Discipline Effects Optics is the most consistent weak point across models. Setup failures are most frequent in optics and least frequent in thermodynamics among the top closed-source systems. This suggests that current MLLMs struggle with the geometric and phase-sensitive aspects of optical reasoning, including ray tracing, interference, and polarization. Mechanics and electromagnetism remain difficult because they often require multi-condition modeling, but their patterns appear more familiar to frontier models. I Extended Error Analysis I.1 Error Taxonomy To delve into the failure cases of models, we detailed six typical error types in Table 7 Furthermore, to facilitate a better understanding of each error type, we provide examples of each error made by Claude-Opus-4.6 from Figure 14 to Figure 17. Table 7: Detailed Descriptions of Error Types. Error TypeExplanation Physical-perception errorThe model misreads the physical setup, diagram, variables, constraints, or stated conditions. Physical-law errorThe model selects an inappropriate physical principle or applies a law outside its valid conditions. Algebraic/symbolic errorThe physical setup is mostly correct, but the derivation contains symbolic manipulation, sign, or formula errors. Incomplete derivationThe model starts correctly but stops before all required quantities or sub-cases are solved. Geometric/visual errorThe model makes an incorrect geometric inference from a diagram, configuration, angle, path, or spatial relation. Numerical mistakeThe model makes an arithmetic, substitution, or unit-conversion error after the correct setup is available. 25 Figure 14: Physcical Law Error 26 Figure 15: Physical Perception Error 27 Figure 16: Geometric/Vision Perception 28 Figure 17: Algebraic/Symbolic Error 29