Paper deep dive
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
Rongsheng Wang, Minghao Wu, Hongru Zhou, Zhihan Yu, Zhenyang Cai, Junying Chen, Benyou Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 5:11:08 AM
Summary
The paper introduces MicroVerse, a video generation model for microscale simulation, addressing the gap in simulating microscopic phenomena compared to macroscopic ones. It presents MicroWorldBench, a benchmark with 459 expert-annotated criteria for evaluating scientific fidelity, visual quality, and instruction following in organ, cellular, and subcellular simulations. The authors also introduce MicroSim-10K, a dataset of 9,601 expert-verified microscale simulation videos, and demonstrate that current SOTA models fail in microscale simulation due to violations of physical laws and temporal inconsistencies.
Entities (10)
Relation Signals (8)
MicroVerse â basedon â Wan2.1
confidence 95% ¡ MicroVerse is built on Wan2.1 Wan et al. (2025) model
MicroWorldBench â containscriteriafor â Subcellular molecular interactions
confidence 95% ¡ MicroWorldBench enables systematic, rubric-based evaluation through 459 unique expert-annotated criteria spanning multiple microscale simulation task (e.g., ... subcellular molecular interactions)
MicroWorldBench â containscriteriafor â Organ-level simulation
confidence 95% ¡ MicroWorldBench enables systematic, rubric-based evaluation through 459 unique expert-annotated criteria spanning multiple microscale simulation task (e.g., organ-level processes...)
MicroWorldBench â containscriteriafor â Cellular dynamics
confidence 95% ¡ MicroWorldBench enables systematic, rubric-based evaluation through 459 unique expert-annotated criteria spanning multiple microscale simulation task (e.g., ... cellular dynamics...)
MicroVerse â trainedon â MicroSim-10K
confidence 95% ¡ Leveraging this dataset, we train MicroVerse... MicroVerse is built on Wan2.1... and trained with MicroSim-10K
MicroWorldBench â evaluates â MicroVerse
confidence 90% ¡ On MicroWorldBench, MicroVerse surpasses original model... Our work first introduce the concept of Micro-World Simulation and present a proof of concept, which includes a clear objective, a dedicated benchmark...
MicroWorldBench â evaluates â Sora
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds great promise for biomedical applications such as drug discovery, organ-on-chip systems, and disease mechanism studies, while also showing potential in education and interactive visualization. In this work, we introduce MicroWorldBench, a multi-level rubric-based benchmark for microscale simulation tasks. MicroWorldBench enables systematic, rubric-based evaluation through 459 unique expert-annotated criteria spanning multiple microscale simulation task (e.g., organ-level processes, cellular dynamics, and subcellular molecular interactions) and evaluation dimensions (e.g., scientific fidelity, visual quality, instruction following). MicroWorldBench reveals that current SOTA video generation models fail in microscale simulation, showing violations of physical laws, temporal inconsistency, and misalignment with expert criteria. To address these limitations, we construct MicroSim-10K, a high-quality, expert-verified simulation dataset. Leveraging this dataset, we train MicroVerse, a video generation model tailored for microscale simulation. MicroVerse can accurately reproduce complex microscale mechanism. Our work first introduce the concept of Micro-World Simulation and present a proof of concept, paving the way for applications in biology, education, and scientific visualization. Our work demonstrates the potential of educational microscale simulations of biological mechanisms. Our data and code are publicly available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2603.00585v1
- Canonical: https://arxiv.org/abs/2603.00585v1
Trouble viewing inline? Open PDF directly â
Full Text
66,242 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at ICLR 2026 MICROVERSE: A PRELIMINARY EXPLORATION TO- WARD A MICRO-WORLD SIMULATION Rongsheng Wang 1,3â Minghao Wu 1â Hongru Zhou 2 Zhihan Yu 1 Zhenyang Cai 1 Junying Chen 1 Benyou Wang 1,3â 1 The Chinese University of Hong Kong, Shenzhen 2 Peking Union Medical College Hospital 3 Shenzhen Loop Area Institute ABSTRACT Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phe- nomena remains largely unexplored. Microscale simulation holds great promise for biomedical applications such as drug discovery, organ-on-chip systems, and disease mechanism studies, while also showing potential in education and inter- active visualization. In this work, we introduce MicroWorldBench, a multi-level rubric-based benchmark for microscale simulation tasks. MicroWorldBench en- ables systematic, rubric-based evaluation through 459 unique expert-annotated criteria spanning multiple microscale simulation task (e.g., organ-level processes, cellular dynamics, and subcellular molecular interactions) and evaluation dimen- sions (e.g., scientific fidelity, visual quality, instruction following). MicroWorld- Bench reveals that current SOTA video generation models fail in microscale simu- lation, showing violations of physical laws, temporal inconsistency, and misalign- ment with expert criteria. To address these limitations, we construct MicroSim- 10K, a high-quality, expert-verified simulation dataset. Leveraging this dataset, we train MicroVerse, a video generation model tailored for microscale simula- tion. MicroVerse can accurately reproduce complex microscale mechanism. Our work first introduce the concept of Micro-World Simulation and present a proof of concept, paving the way for applications in biology, education, and scientific visualization. Our work demonstrates the potential of educational microscale sim- ulations of biological mechanisms. Our data and code are publicly available at https://github.com/FreedomIntelligence/MicroVerse 1INTRODUCTION World models LeCun (2022); Bruce et al. (2024); Lu et al. (2024) have been extensively studied for their ability to simulate environments and agent interactions. They offer a unified computational framework for perceiving surroundings, controlling actions, and predicting outcomes, thereby re- ducing reliance on real-world trials. This not only robotics engines Luo & Du (2024); Lu et al. (2024) engines and reinforcement learning planners Hafner et al. (2020); Agarwal et al. (2025), but also enhances decision-making, supports safe exploration, and enables scalable learning. Recently, video generative models have demonstrated strong potential to acquire commonsense knowledge directly from raw video data, ranging from physical laws in the real world to embod- ied behavioral patterns Brooks et al. (2024), laying the foundation for their use as real-world sim- ulators. For example, prior work Luo & Du (2024) employs video-guided goal-conditioned explo- ration, grounding large-scale video generation model priors into continuous action spaces through self-supervision, enabling robots to master complex manipulation skills without explicit actions or rewards; and other works Lu et al. (2024) leverage video generation models for embodied decision- â Equal Contribution. â Corresponding author. 1 arXiv:2603.00585v1 [cs.AI] 28 Feb 2026 Published as a conference paper at ICLR 2026 Create a simple animation of DNA turning into an RNA strand through polymerase. Create a medical animation of human lung alveoli with detailed capillaries and blood flow. Prompt Create a simplified video of cell division, from one to two to four Real Video Sora Veo3 Alveoli Demonstration Cell Division DNA Replication Figure 1: Failure cases of Sora and Veo3 on Microscale Simulation. Although Sora and Veo3 generate results that appear visually correct, their violations of physical laws are particularly evident. making, allowing agents to imaginatively explore their environment with high generative quality and consistent exploration. Despite tremendous progress in video generation for natural scenes and human-centered do- mains OpenAI (2024); Google DeepMind (2025); Kong et al. (2024); Wan et al. (2025); Yang et al. (2024), research efforts have remained predominantly focused on the macroscopic scale. This suc- cess has not translated effectively to the microscopic scale, where current state-of-the-art models fail to produce physically plausible or biologically meaningful dynamics, as shown in Figure 1. Microscopic simulation, which tracks the interactions of atoms, molecules, and cells to uncover underlying mechanisms, is crucial for applications in materials science, biomedical research Dario et al. (2000), education Romme (2002), and interactive visualization White (1992). The failure of existing models, primarily due to a lack of incorporated biomedical knowledge, highlights a critical gap despite the strong potential of microscale simulation for generating clinically realistic dynamics in fields like drug discovery and disease modeling. To address this, we aim to explore the potential of educational microscale simulations of biological mechanisms. In this work, we introduce MicroWorldBench, a multi-level rubric-based benchmark for microscale simulation tasks comprising 459 real-world tasks that span organ-level, cellular, and subcellular pro- cesses. These tasks were jointly selected from a large candidate pool by LLMs and domain experts for their diversity and relevance, with each task paired with self-contained, objective evaluation criteria specifying the essentials for valid simulation. Our extensive experiments across a broad spectrum of video generation models reveal that while most maintain superficial visual coherence and adhere to prompts, they perform poorly in microscale settings, consistently failing to generate biologically plausible dynamics. These failures indicate that current models, trained predominantly on human-scale videos, lack grounding in microphysical principles and knowledge. To mitigate the gap, we introduce MicroVerse, a video generation model tailored for microscale simulation. MicroVerse is built on Wan2.1 Wan et al. (2025) model and trained with MicroSim-10K, the first microscale dataset containing 9,601 expert-verified scenarios. Unlike human-scale datasets, MicroSim-10K emphasizes physical plausibility and biological fidelity across diverse microscale mechanisms. On MicroWorldBench, MicroVerse surpasses original model by more than +2.7 in scientific fidelity, highlighting the importance of domain-specific data. Our contributions are summarized as follows: (i) We introduce the concept of Micro-World Sim- ulation and present a proof of concept, which includes a clear objective, a dedicated benchmark, a training dataset, and a tailored model. (i) We propose MicroWorldBench, the first rubric-based benchmark specifically designed for evaluating microscale simulation in video generation; (i) we construct MicroSim-10K, a large-scale, expert-verified dataset of microscale simulation videos; (iv) We introduce MicroVerse , a fine-tuned video generation model built upon MicroSim-10K, achiev- 2 Published as a conference paper at ICLR 2026 Create a microscopic blood flow scene with red blood cells, glucose molecules, and vessel structures. Task Show semi-transparent biconcave red blood cells and crystalline glucose flowing through smooth, elastic vessels. Use realistic lighting to highlight microscopic dynamics, and ensure the visuals clearly demonstrate insulin- driven glucose uptake for better comprehension. Requirement Task DescriptionGenerated Video Yes +1 Yes +1 Yes -0.5 Erythrocytes are biconcave; glucose molecules are nanoscale smaller + 1.0 The blood vessel lining and walls look smooth and unbroken + 1.0 Glucose is misrepresented as crystals, not hydrated cyclic molecules - 0.5 Rubric Criteria & Weights 0 2.5 max Actual Score = 1.5 2.5 =60% Glucose uptake under insulin regulation is shown accurately + 0.5 No 0 1.5 Criterion WeightsPresent Yes -0.5 ... ... Sora Veo3 Wan2.2 CogVideo Hunyun ... Models GPT-5 Figure 2: Illustration of MicroWorldBench Evaluation Process. A MicroWorldBench example con- sists of a generated microscopic video and a set of task-specific evaluation criteria written by experts. A MLLM-based scoring system rates responses according to each criterion. ing competitive performance on MicroWorldBench by reducing violations of scientific constraints and improving temporal and spatial consistency. 2MICROWORLDBENCH: A RUBRIC-BASED BENCHMARK FOR MICROSCALE SIMULATION Generic Evaluation Fails to Capture Microscale Simulation Dynamics Existing evaluation meth- ods for video models often rely on generic scoring rules or high-level principles Huang et al. (2024); Zheng et al. (2025); He et al. (2024), which are insufficient for microscale simulation. Such meth- ods overlook the need for fine-grained microscopic simulations, resulting in misaligned outcomes and failing to capture deficiencies in physical plausibility and biological fidelity. In this work, the proposed Rubric evaluation addresses this gap by introducing task-specific criteria with differenti- ated weights. Rubrics highlight the most critical dimensions identified by experts and ensure that evaluations emphasize substantive shortcomings rather than being diluted by aggregate scoring. In this section, we introduction the core structure of the rubric-based benchmark, covering task selection (Sec. 2.1), prompt design (Sec. 2.2), and rubric construction (Sec. 2.3), and describe the methodology for model evaluation (Sec. 2.4). 2.1TASK CHOICE Biological systems are inherently hierarchical, encompassing levels from society, body, organ, and tissue to cell, organelle, protein, and gene Qu et al. (2011). Given constraints of practicality impact and data availability, in this work we focus on three representative levels as a principled sampling of this hierarchy. Importantly, this choice does not discard existing scientific frameworks, but rather reflects a consensus-based selection of the most representative and tractable scales. 1. Organ-level simulations are essential because they connect microscale behaviors with macroscopic physiological functions. Dynamic processes such as cardiac contraction or vascular deformation are directly related to medical diagnosis, surgical planning, and edu- cation. A benchmark that evaluates these dynamics provides a direct path toward clinically relevant applications. 2. Cellular-level simulations are central to biology and medicine, as cell migration, pro- liferation, and interaction underpin processes such as tissue growth, wound healing, and immune response. Accurate modeling at this level enables researchers and students to vi- sualize and understand the driving forces of health and disease, creating opportunities for both discovery and pedagogy. 3 Published as a conference paper at ICLR 2026 3. Subcellular-level simulations present the most fine-grained view, capturing biochemical and biophysical mechanisms that govern life at its foundationâfusion, apoptosis, signaling cascades. Evaluating generative models at this level is particularly important, as these processes are both visually subtle and mechanistically complex, requiring high fidelity and physical plausibility. 2.2PROMPT SUITE Both the sampling process of diffusion-based video generation models and the development of expert-driven evaluation rubrics are computationally expensive. To ensure efficiency, we control the number of tasks while maintaining diversity and coverage. The construction follows a two-stage pipeline: (1) collecting tasks related to microscale simulation from YouTube; and (2) expert filtering to retain only scientifically meaningful tasks. The final suite contains 459 tasks: 238 at the organ level, 189 at the cellular level, and 32 at the subcellular level. The proportion of tasks is consistent with the distribution of levels in the collected videos. Collecting and Generating Prompts We retrieved over 8,000 YouTube videos using topic-specific queries related to organ-level, cellular-level, and subcellular-level simulations. For each video, we collected metadata including titles and descriptions. This information was then provided to GPT-4o, which generated tasks describing the microscale mechanism. Finally, we generated 8,162 tasks. The prompts used to instruct GPT-4o refer to Appendix J. Expert Filtering We filtered the generated tasks based on two criteria: (1) the diversity of the tasks, and (2) the practical relevance of the tasks. For diversity, we asked GPT-4o to classify each task into one of the following categories: Organ-level simulations, Cellular-level simulations, or Subcellular- level simulations. For practical relevance, we invited three biology experts, and each task had to receive agreement from at least two of the three experts. A task was retained in MicroWorldBench only if it satisfied both criteria. Classification prompts are in Appendix K. 2.3RUBRIC CRITERIA As shown in Figure 2, each MicroWorldBench example includes a task instruction and rubric cri- teria, drafted by LLMs and refined by experts. These criteria evaluate scientific fidelity, visual quality, and instruction following. Scientific fidelity emphasizes mechanistic accuracy rather than visual realism. An LLM-based grader then scores the output, providing a standardized, interpretable assessment. Due to limited expert availability and efficiency concerns, we adopt a collaborative approach where LLMs generate initial rubric drafts and experts perform revision and validation. This method not only improves the efficiency of rubric construction but also ensures broader coverage and more comprehensive consideration despite the small number of experts. Stage 1: Rubric Drafts Generation For each task, GPT-5 generates a set of fine-grained criteria: P = (a i , d i , s i , w i ) N i=1 , where a i denotes the evaluation dimension, d i is the description of the i-th criterion, s i â +1,â1 is the polarity indicating whether the point contributes (+1) or deducts (â1), and w i â (0, 1] is the weight reflecting its importance (e.g., w i = 1.0 for core scientific require- ments, w i = 0.5 for key but secondary requirements, and w i = 0.2 for auxiliary or presentational). The score for each task is defined as: S = P N i=1 s i ¡ w i . To ensure comparability across tasks, we normalize it: S norm = S P N i=1 w + i Ă 100 where P w + i is the maximum score from positive criteria, ensuring a maximum of 100 and preventing minor positives from offsetting severe scientific errors.â Stage 2: Expert Revision and Validation Domain experts refine the LLM-generated rubric through the following actions: ⢠Deleting or filtering criteria: Experts refine the criteria by modifying or removing d i that are redundant, irrelevant, or scientifically trivial. ⢠Adjusting weights: When the weight of certain criteria does not align with the scientific validity of the task, experts modify the corresponding weight w i . 4 Published as a conference paper at ICLR 2026 Table 1: Performance comparison of different video generation models on MicroWorldBench. Bold indicates the best performance. ModelAverageâOrgan-levelâCellular-levelâSubcellular-level â Open-Source Video Generation Models HunyuanVideo23.223.123.819.4 CogVideoX-5B43.539.947.038.6 Wan2.1-T2V-1.3B49.445.951.752.4 Wan2.2-TI2V-5B 51.646.653.949.5 Wan2.1-T2V-14B54.855.754.452.8 Wan2.2-T2V-A14B53.856.352.053.3 MicroVerse-1.3B (Ours)50.247.651.753.3 Commercial Video Generation Models Sora50.755.946.155.0 Veo377.277.576.978.2 Table 2: Performance comparison of different video generation models on MicroWorldBench (dimension-wise scores). Bold indicates the best performance. ModelAverageâScientific Fidelity â Visual QualityâInstruction Followingâ Open-Source Video Generation Models HunyuanVideo23.215.648.223.4 CogVideoX-5B 43.537.464.138.6 Wan2.1-T2V-1.3B49.440.371.850.1 Wan2.2-TI2V-5B51.640.782.747.0 Wan2.1-T2V-14B54.842.786.053.8 Wan2.2-T2V-A14B 53.837.892.855.4 MicroVerse-1.3B (Ours) 50.243.068.549.3 Commercial Video Generation Models Sora50.735.396.437.9 Veo3 77.265.797.077.0 ⢠Supplementing criteria: If the automatically generated criteria fail to cover essential scien- tific dimensions, experts can introduce new tuples (a j , d j , s j , w j ). We invited three experts to participate in the revision and validation process. Each expert first in- dependently reviewed and modified the evaluation criteria, including adjusting weights, removing redundant items, and supplementing any missing dimensions. All modifications were documented with clear rationale to ensure transparency. The proposed changes from all experts were then aggre- gated, and conflicts were resolved through discussion, majority voting. For more analysis on expert revision and validation, refer to the Appendix C. 2.4EVALUATION RESULTS AND ANALYSIS Settings We evaluated video generation models on microscopic simulation tasks using MicroWorld- Bench, including open-source models (e.g., Wan2.1 Wan et al. (2025), HunyuanVideo Kong et al. (2024)) and commercial models (e.g., Sora OpenAI (2024), Veo3 Google DeepMind (2025)). Infer- ence was conducted once per model under default settings to ensure fairness and consistent resolu- tion. Rubric evaluation employed LLM-as-a-Judge Zheng et al. (2023), with GPT-5 serving as the Judge. The configurations and sampling details in the Appendix E. Overall Results As shown in Table 1, the performance of different models varies significantly across organ-level, cellular-level, and subcellular-level tasks. Although commercial closed-source models, such as Veo3, substantially outperform open-source models in overall scores, their advantage is mainly confined to the visual quality dimension rather than scientific fidelity. 5 Published as a conference paper at ICLR 2026 Made at SankeyMATIC.com Original Clips 67,853 Step1 Remaining Clips 33,535 FilteredClassifier 34,318 Step2 Remaining Clips 26,841 Black Borders Filtered 6,694 Step3 Remaining Clips 12,194 Subtitles Filtered 14,647 Final Clips 9,601 Expert Filtered 2,593 Figure 3: Overview of our data filtering pipeline. Each stage applies specific filters and shows the volume of data removed and retained. Visual Quality vs. Scientific Fidelity Table 2 shows that nearly all models achieve high scores in visual quality (80â97), yet their scientific fidelity lags far behind (most open-source models score only 15â43). This result demonstrates that current models often generate videos that âlook rightâ but fail to strictly adhere to physical and biological laws. Performance Differences Across Hierarchical Tasks Both advanced open-source models (e.g., Wan2.2-T2V-A14B) and top commercial models (Sora, Veo3) exhibit lower performance on cellu- lar and subcellular tasks compared to organ-level simulations. This may be attributed to the higher requirements for physical and biological consistency in these tasks, as well as the scarcity of mi- croscale training data that can capture complex dynamics. Scale Effects in Open-Source Models Within the Wan series, increasing model size from 1.3B to 14B mainly improves visual quality, while scientific fidelity shows little significant growth. This suggests that expanding model parameters alone is not sufficient to solve the core scientific fidelity challenges in microscale simulation. 3MICROVERSE: TOWARD MICROSCALE SIMULATION VIA A EXPERT-VERIFIED DATASET The results of MicroWorldBench indicate that current models remain limited in their ability to model microscale mechanism governed by physical and biological principles. Most large-scale video datasetsâsuch as InternVid Wang et al. (2023b), UCF101 Soomro et al. (2012), and OpenVid- 1M Nan et al. (2024)âprimarily consist of natural scenes or human activities, offering little rele- vance to microscopic processes. To address this challenge, we propose a new microscale simula- tion models, termed MicroVerse, which explicitly incorporate physical grounding and fine-grained biological dynamics. A key prerequisite for developing such models is the availability of domain- specific data that accurately capture microscopic processes with physical fidelity. 3.1DATA CONSTRUCTION: MICROSIM-10K Collecting videos from YouTube We used the official YouTube API to search for videos related to microsimulation and filtered them based on the following criteria: (1) resolution of at least 720p; and (2) licensed under Creative Commons. These requirements ensure that the collected videos are suitable and freely available for training. In total, we obtained 12,848 relevant videos. Splitting videos After obtaining the videos, we segmented them into multiple semantically consis- tent and short clips. We used OpenCLIP Ilharco et al. (2021) for video segmentation: whenever the similarity between adjacent frames fell below 0.85, a split was made. In total, 67,853 clips were generated. Since not all clips were related to microsimulation, we trained a classifier based on VideoMAE Tong et al. (2022) to filter them. The model achieved an accuracy of over 92%, signifi- cantly improving the quality of the dataset. With the help of the classifier, 34,318 clips were filtered out. For details of the clip classification model related to microsimulation, refer to the Appendix L. 6 Published as a conference paper at ICLR 2026 Automatic and expert filtering To improve the quality and physical consistency of the clips, we first applied OpenCV 1 to detect black borders and used EasyOCR 2 to detect subtitles in order to filter out those affecting semantic representation, retaining 12,194 clips. Experts then reviewed the data, removing meaningless or physically inconsistent clips, resulting in 9,601 clips. Generating captions We leverage a multimodal LLM (GPT-4o) to generate detailed captions. Due to context limits, we uniformly sampled 8 frames per clip as visual input. To minimize hallucina- tions, we supply the video title and description. Prompt MLLM to Generate Video Caption The provided images are sampled from a video clip (8 evenly spaced frames). This clip is taken from a video with the following metadata: Video Title: Video Title; Video Description: Video Description Using the visual content of the clip, together with the title and description, please generate a clear, detailed, and accurate description of what is shown. Focus on the subject, explains the scene and actions, and emphasizes visible details, textures, and fine structures. 3.2DATA STATISTICS 3.2.1FUNDAMENTAL ATTRIBUTES 59.1% (5672) 22.4% (2148) 18.5% (1778) Organ LevelCellular LevelSubcellular Level (a) Distribution of Categories 10203040506070 Seconds 0 200 400 600 Frequency (b) Distribution of Video Duration 2060100140180220260 Words 0 100 200 300 400 500 600 Frequency (c) Distribution of Caption Length Figure 4: Distributions of fundamental video attributes in the MicroSim-10K. MicroSim-10K is the first large-scale dataset dedicated to microscale simulation, comprising 9,601 high-quality video clips. As shown in Figure 4, all clips have a resolution of at least 720p and a duration of 5â60 seconds, ensuring that each captures a complete and coherent microscopic process. The dataset spans diverse biological mechanisms across organ, cellular, and subcellular levels, of- fering broad coverage of key scenarios. Each clip is paired with a detailed caption generated by a multimodal LLM and validated by experts, with an average length of around 150 words, providing precise semantic alignment for model training. 3.2.2POPULARITY AND RELEVANCE To capture the educational and communicative value of microscale simulations, MicroSim-10K re- tains metadata such as views, likes, and comments. As shown in Figure 5, the videos in MicroSim- 10K have been widely viewed, with many reaching hundreds of thousands of views, and they have received substantial likes and comments, reflecting strong popularity and broad accessibility across both scientific and public communities. 3.2.3REALISM AND DISTRIBUTION We compare its distribution with real-world microscopy videos using Fr Ě echet Video Distance (FVD). Using the method described in Section 3.1, we collected 377 real biological videos from YouTube and obtained 643 video clips after preprocessing. As shown in Table 3, the FVD between MicroSim- 10K and real biological videos is 123.9. This result indicates that our expert-verified MicroSim-10K 1 https://github.com/opencv/opencv-python 2 https://github.com/JaidedAI/EasyOCR 7 Published as a conference paper at ICLR 2026 1234567 Numbers (million) 0 1000 2000 3000 4000 5000 Frequency (a) Distribution of Video Views 2000400060008000 Numbers 0 400 800 1200 Frequency (b) Distribution of Video Likes 100200300400500600700 Numbers 0 1000 2000 3000 4000 Frequency (c) Distribution of Comments Figure 5: Distributions of video popularity indicators in the MicroSim-10K. already lies remarkably close to the real microscopy distribution in terms of visual statistics and structural descriptors, effectively bridging the gap between simulated and real experimental data. Table 3: FVD comparison across models (lower is better). The FVD between MicroSim-10K and real biological videos is 123.9, indicating a close distributional alignment. DataFVD vs. MicroSim-10KâFVD vs. 643-Real-Biological-Clipsâ MicroSim-10K0123.9 643-Real-Biological-Clips123.90 Commercial Models Veo342.6118.1 Sora116.9136.3 Open-Source Models Wan2.1-T2V-1.3B83.0158.9 Wan2.2-TI2V-5B77.6153.3 Wan2.1-T2V-14B53.3137.6 Wan2.2-T2V-A14B65.8132.2 3.3TRAINING MICROVERSE For training, we fine-tune the Wan2.1 model. A text prompt P is encoded as a sequence: P = (p 0 , p 1 , . . . , p m ), while the target video V is decomposed into T frames. Each frame is mapped into the latent space via a VAE Kingma & Welling (2013) encoder, yielding the sequence: L = (l 0 , l 1 , . . . , l T ). The text input P is transformed into embeddings E using CLIP text encoder, and the latent sequence L is processed by a Diffusion Transformer (DiT) Peebles & Xie (2023). The training objective is to predict the latent representation of the video through a denoising diffu- sion process. At timestep t, the loss function is defined as: L =E h âĽÎľâ Îľ θ (L t , t, E)⼠2 i ,(1) where L t is the noisy latent representation at timestep t, Îľ denotes the injected noise, Îľ θ is the modelâs noise prediction, t is the current diffusion timestep, and E is the text embedding. During fine-tuning, with probability defined by the 10%, the text conditioning is entirely masked, enabling Classifier-Free Guidance (CFG) Ho & Salimans (2022) training. This mixture of uncondi- tional and conditional training improves the generation quality of the model during inference. 4EXPERIMENTS Experiment Settings We train MicroVerse using 8 NVIDIA H200 GPUs, fully fine-tuning all pa- rameters of Wan2.1-T2V-1.3B Wan et al. (2025) with a learning rate of 1e-5 and a batch size of 8. The training process is designed to improve the modelâs capability to generate microscopic sim- ulation videos conditioned on text prompts. We conducted a comparative with other models on MicroWorldBench. Additional training details are provided in the Appendix D. 8 Published as a conference paper at ICLR 2026 Human Evaluation To evaluate alignment with human preferences, we conducted a human study comparing MicroVerse with Sora and Veo3. The evaluation included 60 samples across three levels of microsimulation (20 samples per level), all sourced from the 20 most popular microsimulation videos on YouTube. Model outputs were randomly shuffled, and three evaluators independently selected the preferred result based on instruction fidelity and visual clarity, or marked a tie. The final results were reported as preference ratios. 4.1RESULTS OF OUR MICROVERSE Improvement in Scientific Fidelity Table 2 shows that MicroVerse achieves a significant improve- ment in Scientific Fidelity, reaching a score of 43.0 and outperforming all open-source models. This enhancement is attributed to the training on the physics-grounded MicroSim-10K dataset, which en- ables the model to better adhere to biological and physical laws. Although there is a slight decrease in Visual Quality (68.5) and Instruction Following (49.3), this does not affect our core objective: advancing scientific fidelity. Breakthrough in Subcellular-Level Tasks According to Table 1, on the highly challenging subcellular-level tasks, MicroVerse achieves a score of 53.3, surpassing all open-source models. This demonstrates that our dataset enables MicroVerse to make notable progress on microscale sim- ulation tasks where existing models typically struggle. 4.2ANALYSIS Scaling Results We identify two main limitations in the performance of the 1.3B model: first, the improvement in scientific fidelity is relatively modest when fine-tuning a small-parameter model; second, there is a slight decline in visual quality and instruction-following capabilities. To address these issues, we scale the model parameters to 14B and employ a mixed-domain training strategy, combining MicroSim-10K with an equivalent amount of high-quality general-domain data randomly sampled from OpenVid Nan et al. (2024). As shown in Table 4, this dual scaling of both model capacity and data diversity significantly enhances performance across all dimensions, achieving state-of-the-art results among open-source models. For more ablation studies on dataset filtering, dataset size, and training recipes, please refer to Appendix B. Table 4: Impact of Data Scale and Mixed-Domain Training on MicroVerse Performance. Bold indicates the best performance. FT indicates fine-tuning. Data SizeScientific FidelityâVisual QualityâInstruction Followingâ Base Model: Wan2.1-1.3B Baseline040.371.850.1 FT on MicroSim-10K9,60143.068.549.3 FT on General + MicroSim-10K 19,20244.1 (+3.8)74.4 (+2.6)53.8 (+3.7) Base Model: Wan2.1-14B (Full-Blooded Model) Baseline042.786.053.8 FT on MicroSim-10K9,60145.482.751.4 FT on General + MicroSim-10K19,20248.3 (+5.6)87.7 (+1.7)56.9 (+3.1) Human Evaluation Results Figure 6 shows the results of human evaluation. Compared with Wan2.1-1.3B models, MicroVerse performs excellently in the dimension of Scientific Fidelity. Its outstanding performance in Scientific Fidelity further validates the effectiveness of MicroSim-10K. In addition, the Cohenâs Kappa coefficient among the three independent experts was above 0.80, indicating strong interrater agreement and confirming the reliability of the scoring process. More details on Cohenâs Kappa coefficient can be found in Appendix G. Consistency among Judgers in MicroWorldBench To ensure that MicroWorldBenchâs evaluation aligns closely with human judgment across all dimensions, we conducted human preference labeling on a large set of generated videos. Specifically, we computed the consistency of evaluation tasks across different models as well as between the models and humans. Figure 6 shows the consistency relationships among different models and between the models and humans. For more analysis results on evaluation consistency, please refer to Appendix C. 9 Published as a conference paper at ICLR 2026 WinsTieLoses Scientific Fidelity Visual Quality Instruction Following (a) Human Evaluation Result of MicroVerse and Wan2.1-1.3B ScientificVisualFollowing Dimension 0 20 40 60 80 100 Score Human GPT-5 (b) Consistency between LLMs and Humans ScientificVisualFollowing Dimension 0 20 40 60 80 100 Score Gemini-2.5 GPT-5 (c) Consistency across different LLMs Figure 6: Human Evaluation and Consistency Results. 5RELATED WORK World Model World models LeCun (2022); Bruce et al. (2024); Lu et al. (2024) have garnered significant attention. They simulate dynamic environments by predicting future states and estimat- ing rewards based on current observations and actions. Their ability to model state transitions has been extended to real-world scenarios through joint learning of policies and world models, improv- ing sample efficiency in simulated robotics Seo et al. (2023), real-world robots Wu et al. (2022), clinical decision Yang et al. (2025), and autonomous driving Wang et al. (2023a). For example, some work Du et al. (2023) explores long-horizon video planning by combining visionâlanguage and text-to-video models. Others Luo & Du (2024) focus on linking video models to continuous actions through goal-conditioned exploration. Recent works Lu et al. (2024) also use video genera- tive models to let agents explore environments more effectively. MeWM Yang et al. (2025) applies world modeling to medical image analysis and clinical decision-making. Video Generation Video generation has seen rapid progress in the past two years. The release of Sora OpenAI (2024) has ignited strong research interest in text-to-video generation, leading to breakthroughs in quality, coherence, and controllability Blattmann et al. (2023). Other commercial systems such as Veo3, Kling, HunyuanVideo Kong et al. (2024), and Hailuo HailuoAI (2024) have achieved impressive performance and are widely applied in video production, advertising, and edu- cation. With the technology maturing, domain-specific models are emerging to address specialized needs. For instance, MedGen Wang et al. (2025) generates accurate, high-quality medical videos for health education, while AniSora Jiang et al. (2025) focuses on producing detailed and stylistically rich animated content. Despite these advances, the use of video generation for microscale simulation remains largely unexplored. Rubric Evaluation Rubric-based evaluation has become a standard approach for assessing LLMs on open-ended tasks, offering task-specific and interpretable criteria that improve grading consis- tency. HealthBench Arora et al. (2025) scales this paradigm to 5,000 multi-turn conversations with 48k clinician-authored rubrics covering accuracy, safety, and communication. Building on this, Baichuan-M2 Team (2025) dynamically generates case-specific rubrics as verifiable reward signals for reinforcement learning, enabling adaptive and context-aware supervision. Rubrics as Rewards (RaR) Gunjal et al. (2025) further formalizes rubric-based RL and shows significant gains over Likert-style scoring. These efforts highlight rubric-guided evaluation and training as a promising methodology for developing reliable, aligned, and LLMs. 6CONCLUSIONS Video generation excel at natural and human-centered macroscopic scenes but fail to capture faithful microscale dynamics. This work introduces MicroWorldBench, the first rubric-based benchmark for microscale video generation with 459 expert-curated tasks and well-defined rubric criteria. In addi- tion, we build MicroSim-10K and develop MicroVerse which demonstrate remarkable performance on microscale simulation tasks. By integrating physical constraints and expert supervision, Micro- Verse not only improves visual fidelity but also advances toward biologically meaningful dynamics, enabling applications in biomedical research, education, and interactive scientific visualization. 10 Published as a conference paper at ICLR 2026 LIMITATION Our work aims to explore the potential of educational microscale simulations of biological mecha- nisms, rather than the reproduction of results observed in wet lab experiments. However, our current approach does not explicitly incorporate the underlying physical laws that govern biomedical mi- croscale dynamics, such as fluid mechanics in blood flow, diffusionâreaction equations in molecular transport, or biomechanical constraints in cellular processes. This limitation restricts the applicabil- ity of the model in scenarios that require high-precision scientific simulation and prediction. ETHICS STATEMENT All data are publicly available, compliant with YouTubeâs terms, and we exclude personal/sensitive content. Captions were auto-generated (MLLMs) and manually verified to remove inappropriate/i- dentifiable material. The dataset is intended solely and strictly for research purposes and should not be used for non-research settings. We do not own the copyright of these data and will only publicly release the URLs linked to the data instead of the raw data. ACKNOWLEDGEMENTS This work was supported by Major Frontier Exploration Program (Grant No. C10120250085) from the Shenzhen Medical Academy of Research and Translation (SMART), Shen- zhen Medical Research Fund (B2503005), the Shenzhen Science and Technology Pro- gram (JCYJ20220818103001002), NSFC grant 72495131, Shenzhen Doctoral Startup Funding (RCBS20221008093330065), Tianyuan Fund for Mathematics of National Natural Science Foun- dation of China (NSFC) (12326608), Shenzhen Science and Technology Program (Shenzhen Key Laboratory Grant No. ZDSYS20230626091302006), the 1+1+1 CUHK-CUHK(SZ)-GDSTC Joint Collaboration Fund, Guangdong Provincial Key Laboratory of Mathematical Foundations for Artifi- cial Intelligence (2023B1212010001), the International Science and Technology Cooperation Cen- ter, Ministry of Science and Technology of China (under grant 2024YFE0203000), and Shenzhen Stability Science Program 2023. REFERENCES Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chat- topadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025. Rahul K. Arora, Jason Wei, et al. Healthbench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775, 2025. Andreas Blattmann, Sergey Frolov, Robin Rombach, and Patrick Esser. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. https://openai.com/research/ video-generation-models-as-world-simulators, 2024. OpenAI Research. Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative inter- active environments. In Forty-first International Conference on Machine Learning, 2024. Paolo Dario, Maria Chiara Carrozza, Antonella Benvenuto, and Arianna Menciassi. Micro-systems in biomedical applications. Journal of Micromechanics and Microengineering, 10(2):235, 2000. Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B. Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson. Video language planning, 2023. URL https://arxiv.org/abs/2310.10625. 11 Published as a conference paper at ICLR 2026 Google DeepMind. Veo 3, 2025. URL https://deepmind.google/models/veo/. Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Bing Liu, and Sean Hendryx. Rubrics as rewards: Reinforcement learning beyond verifiable domains. arXiv preprint arXiv:2507.17746, 2025. Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination, 2020. URL https://arxiv.org/abs/1912.01603. HailuoAI. Hailuoai, 2024. URL https://hailuoai.video/. Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, et al. Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation. arXiv preprint arXiv:2406.15252, 2024. Jonathan Ho and Tim Salimans.Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022. Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianx- ing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 21807â21818, 2024. Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip, July 2021. URL https://doi.org/10.5281/ zenodo.5143773. If you use this software, please cite it as below. Yudong Jiang, Baohan Xu, Siqian Yang, Mingyu Yin, Jing Liu, Chao Xu, Siqi Wang, Yidi Wu, Bingwen Zhu, Xinwen Zhang, Xingyu Zheng, Jixuan Xu, Yue Zhang, Jinlong Hou, and Huyang Sun. Anisora: Exploring the frontiers of animation video generation in the sora era, 2025. URL https://arxiv.org/abs/2412.10255. Diederik P Kingma and Max Welling.Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114, 2013. Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. Yann LeCun. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. Open Review, 62(1):1â62, 2022. Taiming Lu, Tianmin Shu, Alan Yuille, Daniel Khashabi, and Jieneng Chen. Generative world explorer. arXiv preprint arXiv:2411.11844, 2024. Yunhao Luo and Yilun Du. Grounding video models to actions through goal conditioned exploration. arXiv preprint arXiv:2411.07223, 2024. Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371, 2024. OpenAI. Video generation models as world simulators, February 2024. URL https://openai. com/index/video-generation-models-as-world-simulators/. William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4195â4205, 2023. Zhilin Qu, Alan Garfinkel, James N Weiss, and Melissa Nivala. Multi-scale modeling in biology: how to bridge the gaps between scales? Progress in biophysics and molecular biology, 107(1): 21â31, 2011. Georges Romme. The educational value of microworld simulation. Tilburg Univerisity, Netherlands, 2002. 12 Published as a conference paper at ICLR 2026 Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel. Masked world models for visual control, 2023. URL https://arxiv.org/abs/ 2206.14244. Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012. Baichuan-M2 Team. Baichuan-m2: Scaling medical capability with large verifier system. arXiv preprint arXiv:2509.02208, 2025. Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data- efficient learners for self-supervised video pre-training. Advances in neural information process- ing systems, 35:10078â10093, 2022. Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. Rongsheng Wang, Junying Chen, Ke Ji, Zhenyang Cai, Shunian Chen, Yunjin Yang, and Benyou Wang. Medgen: Unlocking medical video generation by scaling granularly-annotated medical videos, 2025. URL https://arxiv.org/abs/2507.05675. Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drive- dreamer: Towards real-world-driven world models for autonomous driving, 2023a. URL https: //arxiv.org/abs/2309.09777. Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understand- ing and generation. arXiv preprint arXiv:2307.06942, 2023b. Barbara Y White. A microworld-based approach to science education. In New directions in educa- tional technology, p. 227â242. Springer, 1992. Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, and Pieter Abbeel. Daydreamer: World models for physical robot learning, 2022. URL https://arxiv.org/abs/2206. 14176. Yijun Yang, Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang, Rama Chellappa, Zongwei Zhou, Alan Yuille, Lei Zhu, Yu-Dong Zhang, et al. Medical world model: Generative simulation of tumor evolution for treatment planning. arXiv preprint arXiv:2506.02327, 2025. Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, et al. Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness. arXiv preprint arXiv:2503.21755, 2025. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/ abs/2306.05685. 13 Published as a conference paper at ICLR 2026 ADATA FILTERING PIPELINE As illustrated in Figure 7, our data construction process is designed to ensure high fidelity and se- mantic richness. We start with an initial collection of 128K microsimulation videos. To guarantee data quality, we implement a rigorous cleaning protocol that filters out non-microsimulation sam- ples and eliminates temporal inconsistencies. Furthermore, we utilize OpenCV to crop black borders and employ EasyOCR to detect and remove hard-coded subtitles, thereby reducing visual noise. The preprocessed videos undergo a human-in-the-loop verification stage, where experts assess the con- tent for meaningfulness and consistency. In the final stage, we leverage Multimodal Large Language Models (MLLMs) to automatically generate descriptive captions, resulting in a high-quality dataset suitable for downstream representation learning tasks. Figure 7: Overview of the data construction process. To ensure absolute data purity and avoid potential data leakage, we implemented a rigorous three- level deduplication pipeline. As shown in Table 5, we found zero overlap at high similarity thresh- olds across source, text, and vision levels. Table 5: Three-level deduplication pipeline results. Deduplication TypeMethodMetricResult Source-levelVideo ID matching# overlapping IDs0 Text-levelCaption embedding# pairs > 0.900 Vision-levelFrame embedding# pairs > 0.950 Furthermore, we lowered the text-level and vision-level thresholds to 0.60 to help remove poten- tially overlapping tasks and ensure absolute data purity. Table 6 shows the task distribution after deduplication. Table 6: Task distribution after deduplication. Biological LevelOriginalPost-DeduplicationRemoved Organ-level238233-5 Cellular-level189182-7 Subcellular-level3230-2 Total459445-14 We re-evaluated the model performance on the clean test set (445 tasks), as shown in Table 7. The performance scores remained highly stable, with only negligible variations (decimal-level drops), therefore we ultimately adopted the 459 tasks. Regarding private models like Veo3, we acknowledge they may have seen YouTube data. However, MicroWorldBench evaluates scientific fidelity, not just generalization. Even if a model has seen similar data, generating a simulation that strictly adheres to our expert-weighted physical rubrics re- mains a valid measure of capability. We plan to expand the benchmark with specialized microscopy databases in the future to broaden the domain distribution. 14 Published as a conference paper at ICLR 2026 Table 7: Performance comparison on original vs. clean test set. Test SetAverageScientific FidelityVisual QualityInstruction Following Original (459 tasks)50.243.068.549.3 Clean (445 tasks)50.042.768.549.1 BABLATION STUDY In this section, we present a series of ablation studies to evaluate the impact of different components and hyperparameters on the modelâs performance. Ablation Study on dataset filtering. Filtered data (MicroSim-10K) yields more balanced and reliable improvements than using raw, uncleaned data, boosting scientific fidelity while avoiding major drops in visual quality and instruction following. Table 8 summarizes the results. Table 8: Ablation study on dataset filtering. VariantData SizeSci. FidelityVis. QualityInst. Following Baseline Model (Wan2.1-1.3B)040.371.850.1 FT on raw (uncleaned) data34,31842.167.444.5 FT on MicroSim-10K9,60143.068.549.3 Ablation Study on dataset size. Increasing the size of high-quality training data yields steady gains in scientific fidelity and instruction following with only minor visual-quality tradeoffs. Table 9 summarizes the results. Table 9: Ablation study on dataset size. VariantData SizeSci. FidelityVis. QualityInst. Following Baseline Model (Wan2.1-1.3B)040.371.850.1 FT on MicroSim-10K (50% sample)4,80041.969.148.4 FT on MicroSim-10K9,60143.068.549.3 Ablation Study on training recipes (CFG rate). Excessive CFG sharply degrades scientific fi- delity and visual quality. Table 10 summarizes the results. Ablation Study on training recipes (training steps). As shown in the training curves, the model converges rapidly and stabilizes around 5,000 steps. Therefore, we adopt this number of training steps as the default setting for reporting efficiency. Ablation Study on training recipes (number of frames). We selected 81 frames as a balanced trade-off: ⢠It provides sufficient temporal duration to capture complete distinct biological events (e.g., cell division) while maintaining high temporal coherence and manageable computational costs. ⢠Additionally, 81 frames is a native, optimal temporal window for the Wan2.1 architec- ture Wan et al. (2025). CANALYSIS OF THE PROCESS OF EXPERT REVISION AND VALIDATION All experts involved in our evaluation hold doctoral degrees (Ph.D.) and possess extensive research experience in cellular and molecular biology. Table 11 provides the background information for each expert. 15 Published as a conference paper at ICLR 2026 Table 10: Ablation study on CFG rate. VariantData SizeSci. FidelityVis. QualityInst. Following Baseline Model (Wan2.1-1.3B)040.371.850.1 FT on MicroSim-10K (CFG=5%)9,60141.570.249.8 FT on MicroSim-10K (CFG=10%)9,60143.068.549.3 FT on MicroSim-10K (CFG=15%)9,60138.263.151.4 Table 11: Background information of the experts involved in the evaluation. Expert IDField of ExpertiseYears of Experience 1Cell Biology/Genetics12 years 2Biochemistry / Metabolic Pathway Modeling10 years 3Biochemistry / Metabolic Pathway Modeling8 years We analyzed the frequency with which three experts employed the four types of rubric operations when handling different tasks. As shown in Figure 8, all experts tended to favor Adjust Weights, while Supplement was used relatively infrequently. Follow-up interviews with the three experts revealed that the Supplement operation is more cumbersome, as it requires identifying additional evaluation criteria beyond those automatically generated by the LLM, which can introduce extra burden. As detailed in Table 12, experts performed three specific types of interventions to ensure biological and physical processes were captured: (1) Filtering scientifically trivial criteria; (2) Adjusting weights (e.g., setting w i = 1.0) to prioritize underlying scientific mechanisms over superficial visual quality; and (3) Supplementing missing physical constraints. As shown in Table 12, only about 7.6% of the 459 tasks required experts to add missing scientific criteria. This indicates that GPT-5âs initial rubric drafts already captured the essential mechanisms, while experts played a crucial role in refining the final outputs. Furthermore, Table 13 shows that experts intervened more frequently in subcellular-level tasks to enforce strict physical constraints where GPT-5 is less reliable, ensuring scientific accuracy. Furthermore, we conducted a correlation analysis between human experts and LLMs to validate the reliability of automated evaluation. Three domain experts independently evaluated a subset of videos, and we quantified the scoring agreement between a human-only panel and a panel including GPT-5. As shown in Table 14, the inclusion of GPT-5 maintains or even slightly improves the inter-rater agreement (Fleissâ Kappa). We also report the agreement across different biological scales in Table 15. Overall, GPT-5 shows agreement with expert scoring at a level comparable to human evaluators, and in some cases its consistency is even slightly higher. This suggests that GPT-5 is a reliable evaluator rather than a source of additional variance. Adjust Weights Delete and Filter No Action Supplement Actions 0 50 100 150 200 250 300 Frequency (a) Expert 1 Adjust Weights Delete and Filter No Action Supplement Actions 0 50 100 150 200 250 300 Frequency (b) Expert 2 Adjust Weights Delete and Filter No Action Supplement Actions 0 50 100 150 200 250 300 Frequency (c) Expert 3 Figure 8: Actions Log of Three Independent Experts. 16 Published as a conference paper at ICLR 2026 Table 12: Detailed expert actions during the rubric revision process. ActionExpert 1Expert 2Expert 3 Adjusting weights180221189 Filtering137128146 No Action948295 Supplementing482829 Table 13: Expert intervention frequency across different biological scales. Biological LevelNumber of TasksTotal Expert ActionsAvg. Actions per Task Organ-level2383611.52 Cellular-level1895673.00 Subcellular-level321765.50 DTRAINING SETTINGS ON MICROVERSE We implemented our training pipeline utilizing a high-performance computational node equipped with 8 NVIDIA H200 GPUs. To ensure the model fully captures the intricate dynamics and visual nuances specific to the microscopic domain, we adopted a full parameter fine-tuning strategy on the pre-trained Wan2.1-T2V-1.3B model. We utilized a global batch size of 8 and set the learning rate to 1Ă 10 â5 with a consistent weight decay of 0.01 to prevent overfitting. To maximize computational efficiency and memory utilization without compromising numerical precision, we employed bfloat16 (bf16) mixed-precision training. Furthermore, full gradient checkpointing was enabled to significantly reduce the memory footprint during the backpropagation phase. For the visual data configuration, the model was trained to generate high-fidelity video se- quences with a resolution of 480Ă 832 pixels and a temporal duration of 81 frames. A training ControlNet/Classifier-Free Guidance (CFG) rate of 0.1 was applied to randomly drop text condition- ing, thereby reinforcing the modelâs capability for unconditional generation and improving overall robustness. Detailed training configurations are enumerated in Table 16. EINFERENCE SETTINGS ON MICROVERSE To ensure a rigorous and unbiased assessment of generation quality, we standardized the inference protocol across all evaluated models. The inference configuration strictly mirrors the training reso- lution settings (480Ă 832, 81 frames) to avoid potential domain shifts during the generation phase. We employed a standard sampling strategy with 50 inference steps, striking an optimal balance between generation latency and visual fidelity. A guidance scale of 5.0 was selected to ensure strong adherence to the text prompts while maintaining natural visual diversity and avoiding artifacts associated with excessive guidance. Crucially, to uphold the integrity of our evaluation benchmark, we enforced a strict single-shot in- ference policy. For each test prompt, only a single video sample was generated using a fixed random seed. This approach eliminates the possibility of cherry-pickingâselecting the best output from multiple attemptsâthereby providing an honest reflection of the modelâs stability and generalized performance. The comprehensive parameter settings for inference are listed in Table 17. Table 14: Scoring agreement between human experts and LLMs. Evaluation GroupAgreement Score Human-only (Fleissâ Kappa)0.733 Human + GPT-5 (Fleissâ Kappa)0.771 GPT-5 vs Gemini-2.5-pro (Spearman)0.835 GPT-5 vs GPT-5 (Spearman)0.912 17 Published as a conference paper at ICLR 2026 Table 15: Agreement across biological scales. LevelAgreement (Human-only)Agreement (Human + GPT-5) Organ-level0.7280.742 Cellular-level0.7400.739 Subcellular-level0.6910.727 Table 16: Detailed hyperparameters used during the training phase of MicroVerse. ParameterValue --trainbatchsize8 --maxtrainsteps5000 --learning rate1e-5 --mixedprecisionbf16 --trainingcfgrate0.1 --numheight480 --numwidth832 --numframes81 --weightdecay0.01 --ditprecisionfp32 --enablegradientcheckpointingtypefull Table 17: Standardized inference parameter settings for model evaluation. ParameterValue --height480 --width832 --numframes81 --guidancescale5.0 --num inferencesteps50 18 Published as a conference paper at ICLR 2026 FPROMPT GPT-5 TO GENERATE RUBRIC CRITERIA Prompt GPT-5 to Generate Rubric Criteria Task: You are a biology expert. Your task is to design a set of rubrics to evaluate the completion of a given task based on the provided Prompt. The rubric should consist of multiple triplets in the form: a i , d i , s i , w i ⢠a i : the evaluation aspect, restricted to one of the following three categories: â Scientific Fidelity: Accurate representation of organs, cells, and subcellular structures in scale, morphology, and spatial relationships, with dynamic processes consistent with biological and physical laws. â Visual Quality: Emphasis on clarity, detail, and aesthetics, including model precision, rendering, lighting, and color balance. â Instruction Following: Generated videos strictly follow the prompt description. ⢠d i : description of the i-th evaluation criterion. ⢠s i : polarity of the criterion, either +1 (contributes positively) orâ1 (deducts points); leave this field empty. ⢠w i : weight of importance in the range (0, 1]: â 1.0â core scientific requirements â 0.5â important but secondary requirements â 0.2â auxiliary or presentational requirements Output example: "a1": "Scientific Fidelity", "d1": "Key cell structures are clearly defined and proportionally accurate", "s1": "+1", "w1": "1.0" Constraint: 1. Do not give any explanation, output directly. 2. Please describe the evaluation criterion (d i ) in as much detail as possible. 3. Directly describe the key rubrics in the evaluation criterion, and do not use words such as âwhetherâ. 4. Only English output is allowed. Given prompt: prompt GINTER-RATER RELIABILITY AMONG HUMAN EVALUATORS We used Cohenâs Kappa coefficient to measure agreement among the three experts. Table 18 indicate strong agreement among the experts, confirming the reliability of the scoring process. Table 18: Comparison of Cohenâs Kappa values between experts. Comparison PairCohenâs Kappa Expert1 vs. Expert20.87 Expert1 vs. Expert30.83 Expert2 vs. Expert30.81 19 Published as a conference paper at ICLR 2026 HA RUBRIC EXAMPLE FROM MICROWORLDBENCH Figure 9: Example of a cell mitosis simulation video frame, taken from an excellent example video on YouTube. 20 Published as a conference paper at ICLR 2026 Table 19: Rubric Example for Cell Mitosis Evaluation DimensionCriteriaWeight Scientific Fidelity The sequence of stages, segregation patterns, and ploidy changes in mitosis and meiosis are accurately represented; mitosis produces two genetically identical diploid daughter cells, whereas meiosis involves two successive divisions resulting in four genetically diverse haploid gametes. + 1.0 The alignment of chromosomes at the metaphase plate and their segregation from metaphase to anaphase occur correctly; sister chromatids are distinctly differentiated from homologous chromosomes; the attachment of kinetochores to spindle microtubules and the direction of their tension conform to biological principles. + 1.0 Structural details and dynamic coordination between the spindle apparatus and the centrosome (centriole) are accurate; spindle pole positioning, microtubule polarity, and force distribution are appropriate; the relationship between the microtubule-organizing center and cell polarity is correctly established. + 1.0 The timing and mechanisms of DNA replication and genetic recombination are correctly presented; DNA replication occurs during the pre-mitotic S phase, homologous pairing and crossing over take place in prophase I of meiosis, and no DNA replication occurs between meiosis I and I. + 1.0 Visual Quality The image demonstrates high clarity and fine presentation of microstructural details, with sharp edges of subcellular structures, well-defined layer separation, and absence of wax artifact noise. + 0.5 The animation exhibits coherent motion with stable temporal rhythm, smooth phase transitions, and natural movement trajectories, without any stuttering or tearing. + 0.5 The 3D modeling and material texture are credible, with consistent form proportions and scale hierarchy; the textures are detailed, and surface microstructures are discernible. + 0.2 Coordination of lighting, shadows, and depth of field; controlled volumetric scattering and highlights without excess; clear subject contours with well-defined micro-scale detailing. + 0.2 Instruction Following Accurately present the key stages of mitosis in a single somatic cell, ensuring a clear transition and coherent progression between meiotic divisions I and I in gonadal germ cells. + 0.5 Accurate reproduction of subcellular elements and dynamics: chromosome separation following metaphase plate alignment, coordinated movement of spindle fibers and centrioles, and continuous changes and details of the cell membrane/cytokinesis. + 0.5 Accurately describe genetic outcomes and differences: mitosis produces two genetically identical diploid daughter cells, while meiosis results in four genetically diverse haploid gametes, highlighting the mechanistic distinctions. + 0.2 Compliance with technical specifications and viewing angle requirements: within an approximate total duration of 5 seconds, information is organized clearly; microscopic close-up focuses on the single-cell subject; the camera remains stable, transitions are clear, and the subject is unobstructed. + 0.2 Presence of fundamental conceptual and procedural errors: confusion between mitosis and meiosis, incorrect sequencing of stages, inaccurate ploidy descriptions, omission of the two meiotic divisions, or failure to represent the single-cell focus. - 0.5 21 Published as a conference paper at ICLR 2026 IEXAMPLE OF REAL BIOLOGICAL VIDEO CLIPS ... ... Caption: A high-magnification microscopic view reveals a tightly packed layer of rounded, polygonal cells with finely granular cytoplasm and dark, well-defined nuclei, their thin bright borders forming a mosaic-like tissue texture. At the center, one prominent cell is caught in the midst of mitosis, distinguished by its luminous, ring-shaped outline and a nucleus that appears elongated or partially split as condensed chromatin masses align and separateâvisual cues of an active division stage. Surrounding cells remain in interphase, showing intact round nuclei with small dark spots suggestive of nucleoli, providing contrast to the dynamic structural rearrangements within the dividing cell. The overall scene captures the subtle shifts in texture, contrast, and intracellular organization characteristic of living cells undergoing cell division. Figure 10: Example of real biological video clips. JGENERATE PROMPT FROM THE VIDEO TITLE AND DESCRIPTION Generate Prompt from the Video Title and Description. Your task: Based on the âvideo titleâ and âvideo descriptionâ I provide, craft an extremely detailed, professional, and keyword-rich English prompt specifically for generating breathtaking microscopic- world videos. Do not give any explanation, output directly. Donât use bullet points. Write it as a single, complete paragraph. Generate a complete, ready-to-use video-generation prompt. This prompt must include all of the following sections: 1. Main Subject: Clearly describe the central object within the microscopic scene. 2. Scene & Action: Describe what is happening. 3. Details & Textures: Emphasize the details that should be visible. The video title is: video title The video description is: videodescription KPROMPT LLM CLASSIFIES BASED ON TASK DESCRIPTIONS Prompt LLM Classifies Based on Task Descriptions. Your task: You are an expert in scientific video classification. Given a task description, classify it into one of the following categories for diversity: 1. Organ-level simulations â tasks focusing on the behavior, dynamics, or interactions at the scale of whole organs or organ systems. 2. Cellular-level simulations â tasks focusing on the behaviors and interactions of single cells or collections of cells, such as cell division, cell fusion, cell migration, or cell signaling. 3. Subcellular-level simulations â tasks focusing on molecular, genetic, or biochemical pro- cesses within cells, such as protein folding, gene regulation, or intracellular signaling. The task description is: Task Description Please output only the most appropriate category label based on the task description provided. 22 Published as a conference paper at ICLR 2026 LVIDEOMAE-BASED MICROSIMULATION CLASSIFIER To filter out video clips related to microsimulation, we trained a classifier using 2,580 manually annotated samples based on the VideoMAE model. The training was implemented within the Trans- formers, with a learning rate of 5e-5 and a total of 10 epochs, enabling the model to effectively capture video features and achieve accurate classification. Finally, our classifier achieved an accu- racy of 92% on the test set. Table 20: Dataset statistics for microsimulation classification. CategoryTotalTrain (80%)Test (20%) Microsimulation-related1107885222 Non-microsimulation 14731178295 Total25802063517 23