Paper deep dive
TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang, Zukai Chen, Lei Yang, Quan Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 3:59:52 AM
Summary
The paper introduces TransPhy, a framework for physically grounded visual in-context learning (VICL) that infers physical transformation rules from source-target exemplars and adapts them to query images. It proposes PhysVICL-74, a benchmark with 74 rules and 75K contexts, and a method using explicit rule induction and token-wise mixture-of-experts low-rank adaptation (MoE-LoRA) guided by a State-Transition Capturer (STC) for precise rendering.
Entities (7)
Relation Signals (6)
PhysVICL-74 â contains â 74 transformation rules
confidence 98% ¡ PhysVICL-74, comprising 74 physically grounded transformation rules
TransPhy â uses â MoE-LoRA
confidence 95% ¡ TransPhy first predicts the demonstrated rule... and then synthesizes the target image through token-wise mixture-of-experts adaptation
TransPhy â uses â State-Transition Capturer
confidence 93% ¡ We further develop a training-only State-Transition Capturer (STC)... These representations regularize expert routing
State-Transition Capturer â guides â Expert Routing
confidence 91% ¡ These representations regularize expert routing toward transition-sensitive generation behaviors
TransPhy â improves â physical-rule adherence
confidence 90% ¡ Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization
TransPhy â isbasedon â BAGEL
confidence 88% ¡ TransPhy follows a coarse-to-fine path... We instantiate it on BAGEL.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 source--target image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods.
Tags
Links
- Source: https://arxiv.org/abs/2608.24119v1
- Canonical: https://arxiv.org/abs/2608.24119v1
Trouble viewing inline? Open PDF directly â
Full Text
46,192 characters extracted from source content.
Expand or collapse full text
TransPhy: Visual In-Context Learning for Physically Grounded Image Editing Siyi Xie Xuanke Shi Jinsheng Quan Haoran Tang Zukai Chen Lei Yang Quan Wang Abstract Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a sourceâtarget exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 sourceâtarget image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods. 1Peking University 2SenseTime Research 3Zhejiang University 1 Introduction Figure 1: Physically grounded VICL. Top: limitations of prior methods (1) and our TransPhy framework (2). Bottom: representative PhysVICL-74 transformations. Text prompts provide an intuitive interface for image editing (Brooks et al. 2023; Feng et al. 2025; Samadi et al. 2025; Feng et al. 2024; Mao et al. 2026; Ma et al. 2026). However, language often underspecifies transformations involving multiple coupled effects and complex spatial dependencies. Melting, for example, may require coordinated changes in material state, object shape, solid volume, and interaction with the supporting surface. Exhaustively describing these consequences in text is cumbersome and ambiguous. Visual in-context learning (VICL) provides a natural alternative: given a sourceâtarget exemplar pair, a model transfers the demonstrated transformation to a new query image. By showing both changed and preserved content, visual exemplars can specify transformations that are difficult to express precisely through language alone. Existing VICL methods effectively transfer perceptual and semantic relations across stylization, visual effects, attribute editing, representation conversion, and image restoration (Li et al. 2025; Black Forest Labs et al. 2025; Song et al. 2026b; Zhang et al. 2025; Chen et al. 2025a; Gong et al. 2025; Manor et al. 2026; Rajagopalan and Patel 2025). These tasks are largely evaluated by reproducing the demonstrated relation. Physically grounded transformations additionally require query-dependent adaptation: the same rule may yield different outcomes according to material, geometry, interactions, and environment. For example, melting ice and wax produces different deformations, material states, and secondary effects despite sharing a high-level rule. Successful transfer must therefore infer the transformation and adapt it to the query rather than copy the exemplar difference. We call a transformation physically grounded when its observable outcome depends on these query-specific physical factors. Given a sourceâtarget exemplar pair (A,Aâ˛)(A,A ) and a query image B, physically grounded VICL synthesizes a context-appropriate result Bâ˛B while preserving unaffected content. It poses two coupled challenges: First, transformation interpretation must separate the transferable rule from incidental exemplar properties such as identity, texture, color, and background. Second, query-conditioned realization must adapt that rule to the queryâs structure and context. Because these effects are spatially heterogeneous, transformed objects, interaction regions, secondary effects, and preserved areas may require distinct generation behaviors. Existing methods often encode exemplar relations globally or use image- or layer-level adaptation, capturing only salient effects. As Figure 1 shows, a model may transfer melting by generating a liquid puddle while leaving the solid volume unchangedâa visually suggestive but incomplete physical transition. Despite growing interest in physically plausible image editing (Zhao et al. 2025; Lin et al. 2026; Ding et al. 2026; Sheng et al. 2026; Xu et al. 2026), prior work mainly studies text-specified transformations. Whether VICL models can infer such transformations from visual exemplars, adapt them to new instances, and generalize to unseen types remains unclear. We introduce PhysVICL-74, a training-and-evaluation benchmark comprising 74 transformation rules, 5,240 sourceâtarget image pairs, and nearly 75K exemplarâquery contexts. It supports two complementary protocols. Novel-instance transfer applies seen rules to new objects and scenes. Unseen-rule generalization instead holds out entire rules during training. Together, they evaluate whether a model follows the demonstrated rule, adapts it to the query, and preserves rule-irrelevant content. Building on these observations, we propose TransPhy, a physical-transformation induction and transition-aligned rendering framework that decomposes physically grounded VICL into transformation interpretation and query-conditioned realization. The former captures the transferable rule, while the latter predicts how that rule should manifest in the query, reducing reliance on exemplar-specific appearance. Its understanding pathway predicts textual descriptions of the demonstrated rule and query-specific target state, specifying what transfers and how it manifests while reducing exemplar-specific copying. To realize the predicted transformation, token-wise mixture-of-experts low-rank adaptation (MoE-LoRA) then allows spatial tokens to activate specialized rendering experts. We further develop a training-only State-Transition Capturer (STC), which extracts localized transition representations from the vision transformer (ViT) feature differences between B and Bâ˛B . These representations regularize expert routing toward transition-sensitive generation behaviors, encouraging different regions to handle primary changes, secondary effects, and content preservation appropriately. Thus, textual interpretation determines what transformation occurs, while transition-aligned routing controls how it is realized across the query image. Experiments show that TransPhy improves both novel-instance transfer and unseen-rule generalization, producing more complete, physically plausible transformations with competitive perceptual quality and broader visual-relation performance. These results indicate that explicit transformation induction and fine-grained alignment strengthen physically grounded editing and general VICL rule transfer. Our contributions are summarized as follows: ⢠We formulate physically grounded visual in-context learning, a setting that requires models to infer a transformation from a visual exemplar and adapt its consequences to the physical properties and context of the query. ⢠We introduce PhysVICL-74, a dataset and benchmark containing 74 physically grounded transformation rules, 5,240 sourceâtarget image pairs, and nearly 75K exemplarâquery contexts, with separate evaluation of novel-instance transfer and unseen-rule generalization. ⢠We propose TransPhy, combining textual rule interpretation with transition-aligned token-wise expert routing to improve physical plausibility and generalization. 2 Related Work Visual In-Context Learning. In image editing, visual in-context learning (VICL) transfers the transformation demonstrated by an exemplar pair (A,Aâ˛)(A,A ) to a query image B to synthesize Bâ˛B (Sun et al. 2023; Wang et al. 2023). Existing approaches realize this paradigm through pair-specific optimization (Nguyen et al. 2023; Jones et al. 2024; Lu et al. 2025), training-free attention or feature manipulation (Gu et al. 2024; Srivastava et al. 2024; Das Biswas et al. 2025), or learned transfer mechanisms. Representative training-based methods formulate diverse visual tasks as image infilling (Li et al. 2025), encode relational features with lightweight adapters (Gong et al. 2025; Chen et al. 2026), or adapt diffusion transformers through dynamically generated, composed, or routed LoRA modules (Song et al. 2024; Li et al. 2026; Manor et al. 2026). Despite substantial progress in visual analogy, existing methods are predominantly developed and evaluated on appearance-, geometry-, or semantics-driven transformations. They do not explicitly address physical rule induction, where successful transfer requires identifying latent state changes and adapting their spatially coupled, scene-dependent consequences to a new query. Physics-Aware Image Editing. Recent studies show that physical plausibility remains challenging for multimodal models. PhysBench evaluates reasoning about physical properties, relations, and dynamics (Chow et al. 2025), while RISEBench and KRIS-Bench examine broader visual and knowledge-based reasoning (Zhao et al. 2025; Wu et al. 2025b). Physics-aware editing methods and benchmarks further consider latent physical transitions, environmental constraints, causal consistency, physical interactions, and temporal processes (Zhao et al. 2026; Ding et al. 2026; Sheng et al. 2026; Han et al. 2025; Lin et al. 2026; Pu et al. 2025; Wu et al. 2025a; Guo et al. 2026). However, these settings typically provide the desired effect through a textual instruction or a predefined editing task. The model is asked to execute a known physical transformation, rather than discover it from visual evidence. PhysVICL-74 instead evaluates whether a model can infer the latent physical rule from (A,Aâ˛)(A,A ) and transfer its scene-dependent consequences to a new query. Unified Multimodal Models. Unified multimodal models integrate visual understanding and generation within a shared framework. Janus and Janus-Pro separate understanding and generation encoders (Wu et al. 2024a; Chen et al. 2025b); BAGEL uses shared self-attention with specialized transformer experts (Deng et al. 2025); and Show-o2 and JoyAI-Image combine language modeling with generative visual objectives (Xie et al. 2025; Song et al. 2026a). Such models provide a basis for physically grounded VICL, as they can jointly compare exemplars, reason about transformations, and generate edited images. Nevertheless, direct generation may still rely on superficial exemplar differences or transfer incidental visual content. This motivates TransPhy, which explicitly interprets the demonstrated physical transformation and guides query-specific generation through transition-aligned token-wise expert routing. 3 Task and Dataset Construction Task and Scope. Given an exemplar (A,Aâ˛)(A,A ) and a query image B, physically grounded VICL requires a model to infer the implicit transformation explaining AâAâ˛A\!â\!A and synthesize its query-specific realization Bâ˛B while preserving rule-irrelevant content. We use physically grounded to describe visually observable transformations whose qualitative outcomes depend on material properties, geometry, object interactions, or environmental conditions. PhysVICL-74 evaluates whether a model transfers the causal direction and query-appropriate visual consequences of such transformations. Because multiple outcomes may be physically plausible, each reference target represents one human-validated plausible realization rather than a unique physical ground truth. The benchmark therefore does not assess numerical dynamics or simulator-level physical accuracy. Transformation Taxonomy. We group the 74 rules by their dominant scale and mechanism into three families: Scene-Condition Transformations, comprising Geometric, Optical, and Temporal Scene Variation; Mechanically Induced Transformations, comprising Action-Induced and Mechanical Response; and Material-State Transformations, comprising Thermophysical and Reactive Evolution. These seven categories span changes in scene observation, object configuration or deformation, and material state. For compact reporting, we refer to the Scene-Condition, Mechanically Induced, and Material-State families as Scene-Level, Object-Level, and Matter-Level, respectively. Rule Mining, Construction, and Splits. We draw transformation concepts and selected seed images from RISEBench (Zhao et al. 2025), RE-Edit (Ding et al. 2026), InEdit-Bench (Sheng et al. 2026), KRIS-Bench (Wu et al. 2025b), UniREditBench (Han et al. 2025), and WorldEdit (Lin et al. 2026), rather than directly aggregating their original image pairs. After merging semantic duplicates and removing ambiguous or instance-specific edits, we obtain 74 visually observable and transferable rules. GPT Image completes the corresponding targets and generates additional rule instances where needed. The resulting benchmark contains 5,240 sourceâtarget image pairs and approximately 75K exemplarâquery contexts. Each AABB context (A,Aâ˛,B,Bâ˛)(A,A ,B,B ) combines two distinct sourceâtarget instances governed by the same rule. We split base image pairs before constructing contexts, ensuring that no test pair or context containing it appears in training. The protocols evaluate novel-instance transfer for seen rules and generalization to rules held out from training. Target Completion and Verification. To mitigate potential generatorâreviewer circularity, we do not treat model-based screening as sufficient. Dedicated human annotators additionally review rule correctness, physical plausibility, and preservation of rule-irrelevant content; failed samples are regenerated. Detailed annotation criteria, review protocols, image provenance, rule mappings, and representative failure cases are provided in the supplementary material. 4 Methodology 4.1 Overall Framework Given an exemplar transformation pair (A,Aâ˛)(A,A ), a query image B, and a fixed, rule-agnostic prompt p (see the supplementary material), the model must infer the physical rule demonstrated by AâAâ˛A\!â\!A and synthesize its query-specific realization Bâ˛B . The output should realize the transformation while preserving the identity, spatial layout, and rule-irrelevant content of B. This task involves two levels of reasoning. At the coarse level, the model must extract a transformation rule that transfers across objects and scenes. At the fine level, it must adapt the rule to the object properties, spatial structure, and interactions in the query image. Directly generating Bâ˛B from the visual exemplar may conflate these two processes, causing the model to copy the appearance of Aâ˛A without transferring the demonstrated rule. To address this challenge, we propose TransPhy, which progressively translates a compact physical rule prior into fine-grained, spatially adaptive rendering decisions. At the coarse level, the understanding pathway predicts an explicit rule R R and a query-specific target-state description d^BⲠd_B . Together, they specify the transformation demonstrated by the exemplar and its expected effect on the query. At the fine level, the generation pathway employs token-wise MoE-LoRA, allowing different spatial tokens to invoke different low-rank rendering experts. We further introduce a training-only State-Transition Capturer (STC), which derives token-level transition targets from ViT feature differences between (B,Bâ˛)(B,B ) and uses them to supervise expert routing. TransPhy follows a coarse-to-fine path from physical-rule induction to transition-aligned expert rendering. We instantiate it on BAGEL. Figure 2: Overview of TransPhy on BAGEL. The understanding and generation pathways encode the exemplarâquery context; ViT-derived transition evidence aligns token-level routing, while MoE-LoRA experts render the inferred physical rule. 4.2 Progressive Physical Rule Induction and Rendering Coarse-Grained Physical Rule Prior. BAGEL employs a shared multimodal transformer that supports autoregressive understanding and DiT-based visual generation, providing a unified interface between explicit rule inference and image synthesis. Given the interleaved context (A,Aâ˛,B,p)(A,A ,B,p), the understanding pathway first predicts the demonstrated rule R R and a query-specific target-state description d^BⲠd_B . We then combine these textual intermediates with the visual representations to construct the generation context: c=[cA,cAâ˛,cB,Tpâ(R^,d^Bâ˛)],v^=vĎâ(zt,t,c),c=[c_A,c_A ,c_B,T_p( R, d_B )], v=v_Ď(z_t,t,c), (1) where cIc_I denotes BAGELâs multimodal token representation of image I, TpT_p serializes the task prompt and predicted intermediates into text tokens, ztz_t is the noisy visual latent at timestep t, and vĎv_Ď predicts the visual training target. Together, R R and d^BⲠd_B form a compact, coarse-grained physical rule prior. This representation explicitly summarizes the transformation to be transferred and describes its expected realization on the query. Its primary role is to provide coarse-grained semantic constraints. Conditioning generation on this explicit interpretation reduces the shortcut of copying the appearance of Aâ˛A and provides a stable semantic condition for subsequent fine-grained rendering. Fine-Grained Token-Wise Expert Rendering. The physical rule prior describes the global semantics of the target state, but state-transition editing typically exhibits substantial spatial heterogeneity. The transformed object, interaction areas, secondary effects, and preserved content may require different rendering behaviors. We freeze the BAGEL backbone and apply standard LoRA (Hu et al. 2022) to the understanding pathway and most generation projections. However, standard LoRA applies the same low-rank update to every spatial token and therefore cannot adapt its computation according to the spatial role of each region. Motivated by mixtures of LoRA experts (Wu et al. 2024b), we replace the down-projection of the generation MLP with token-wise MoE-LoRA: yi=Wâxi+Îąrââe=1Egi,eâUeâVeâxi,y_i=Wx_i+ Îąr _e=1^Eg_i,eU_eV_ex_i, (2) where xix_i is the i-th generation token, W is the frozen pretrained projection, UeâVeU_eV_e is the rank-r update of expert e, Îą controls the adapter scale, and E is the number of experts. The router independently computes the expert weights for each token using sparse top-k routing: gi=TopKSoftmaxâĄ(Wrâxi/Ď,k),g_i=TopKSoftmax (W_ rx_i/Ď,k ), (3) where WrW_ r is the router projection, Ď is the routing temperature, and k is the number of activated experts. Because gig_i is predicted separately for every token, different spatial positions can invoke different low-rank updates. The coarse-grained physical rule prior can therefore be progressively realized through fine-grained, spatially adaptive expert rendering. Nevertheless, token-wise MoE-LoRA provides only the capacity for fine-grained rendering; it does not guarantee that the router will organize experts according to the spatial effects of the transformation. When trained only with the image-generation objective, the router may instead specialize according to generic appearance cues such as color, texture, or object category. We therefore introduce fine-grained transition alignment. 4.3 Fine-Grained Transition Alignment During training, MoE-LoRA experts learn their rendering behavior under the guidance of the token-wise router. To provide token-level supervision, STC aligns the routerâs spatial perception with localized transition evidence extracted by a frozen ViT. ViT-Derived Transition Targets. We extract semantic features from (B,Bâ˛)(B,B ) using BAGELâs frozen ViT encoder, initialized from SigLIP2-so400m/14 (Deng et al. 2025). Their differences localize the visible effects of the transformation. Let uju_j and ujâ˛u _j denote the frozen ViT embeddings of (B,Bâ˛)(B,B ) at spatial position j. We define the initial fine-grained local transition response as sj=1âujâ¤âujâ˛âujâ2ââujâ˛â2.s_j=1- u_j u _j\|u_j\|_2\|u _j\|_2. (4) A larger sjs_j indicates a stronger semantic change between the source and target states. We retain the top Ď%Ď\% transition-responsive tokens and refine these sparse candidates through connected-component merging and feature-similarity refinement (Shang et al. 2026; Liu et al. 2024). Details of token selection and refinement are provided in the supplementary material. The resulting transition target q is resized onto the VAE token grid of Bâ˛B using nearest-neighbor interpolation. Compared with raw pixel differences, ViT feature differences are less sensitive to color shifts, generation noise, and local texture variations, providing a more stable token-level supervision signal. Router Alignment. Let aiââEa_i ^E denote the pre-softmax router logits for the i-th generation token of Bâ˛B . A two-layer prediction head fSTC:âEââf_ STC:R^E converts the E expert-routing logits into a scalar transition score. We align its sigmoid response with the resized ViT-derived target: âSTC=1|ΊBâ˛|ââiâΊBâ˛(ĎâĄ(fSTCâ(ai))âqi)2,L_ STC= 1| _B | _iâ _B (Ď(f_ STC(a_i))-q_i )^2, (5) where ΊBⲠ_B indexes the VAE token positions of Bâ˛B . This objective encourages the internal routing representation to distinguish transition-responsive positions from relatively stable regions, without assigning a predefined semantic identity to any expert. The rendering behavior of each expert remains learned through the image-generation objective. STC therefore aligns where expert responses should vary, rather than prescribing which expert must represent a particular transformation. To reduce expert collapse, we additionally introduce a standard MoE load-balancing objective: âbal=Eââe=1EpÂŻeââÂŻe,L_ bal=E _e=1^E p_e _e, (6) where pÂŻe p_e is the mean routing probability of expert e and âÂŻe _e is the fraction of tokens assigned to it. The generation objective learns the rendering behavior of the experts, âSTCL_ STC makes the router sensitive to the spatial effects of the transformation, and âbalL_ bal encourages non-collapsed expert specialization. 4.4 Staged Training Objective We train TransPhy in two stages: coarse-grained physical-rule understanding and fine-grained rendering. In Stage 1, we freeze the BAGEL backbone and optimize a rank-16 LoRA with autoregressive supervision y=[R,dBâ˛],y=[R,d_B ], (7) where R is the annotated transformation rule and dBâ˛d_B is the query-specific target-state description. This stage constrains the understanding pathway to a stable âruleâtarget descriptionâ output format, while learning to extract the demonstrated rule and adapt its expected effect to the query. In Stage 2, we freeze Stage 1 and condition generation on its predicted textual intermediates. We employ rank-32 generation adapters and four MoE-LoRA experts, and jointly optimize visual generation, transition alignment, and expert load balancing: ârender=tâ[âvâvĎâ(zt,t,c)â22]+ÎťSTCââSTC+Îťbalââbal,L_ render=E_t [\|v-v_Ď(z_t,t,c)\|_2^2 ]+ _ STCL_ STC+ _ balL_ bal, (8) where v is the training target, and ÎťSTC _ STC and Îťbal _ bal control the transition-alignment and load-balancing terms, respectively. This staged optimization establishes a coarse-to-fine learning process. Stage 1 constructs a stable physical rule prior, while Stage 2 translates this prior into spatially adaptive token-wise expert rendering. By assigning different rendering experts according to the query content and transition evidence, TransPhy improves the spatial precision and controllability of physically grounded state-transition editing. 5 Experiments Figure 3: Qualitative comparison on seen and unseen transformations. Left: new instances of seen transformations. Right: unseen transformations. Table 1: Seen-rule transfer to novel instances across three rule families. Scene-Level, Object-Level, and Matter-Level denote Scene-Condition, Mechanically Induced, and Material-State transformations, respectively. TA, CP, and RP are scored by GPT-5.6; CLIP-D and LPIPS are objective metrics. Results are macro-averaged over rules within each family. FLUX.1-Fill-dev BAGEL-MoE TransPhy Metric Scene-Level Object-Level Matter-Level Scene-Level Object-Level Matter-Level Scene-Level Object-Level Matter-Level TAâ 3.45 3.12 3.33 3.22 3.03 3.63 3.57 3.39 3.73 CPâ 3.23 3.18 3.41 3.12 3.57 3.23 3.43 3.78 3.67 RPâ 3.20 3.67 3.09 3.33 3.21 3.28 3.47 3.36 3.25 CLIP-Dâ 0.20 0.25 0.25 0.21 0.27 0.29 0.23 0.26 0.27 LPIPSâ 0.39 0.36 0.26 0.40 0.36 0.26 0.31 0.31 0.24 Table 2: Generalization to transformations unseen during TransPhy training. Results are reported as PhysVICL-74/Relation252K over 17 held-out PhysVICL-74 rules and 20 sampled Relation252K rules, respectively. Method TAâ CPâ RPâ CLIP-Dâ LPIPSâ FLUX.1-Fill-dev 2.63/1.60 3.45/3.71 3.22/1.79 0.24/0.26 0.51/0.56 BAGEL-MoE 2.55/2.16 3.60/2.96 3.24/1.90 0.24/0.23 0.54/0.61 TransPhy 2.91/2.34 3.75/3.05 3.28/2.90 0.31/0.28 0.48/0.50 RelationAdapter 2.60/3.21 3.17/3.88 3.34/2.82 0.20/0.37 0.52/0.52 VisualCloze 2.76/2.10 3.15/3.41 3.04/2.45 0.25/0.22 0.51/0.54 LoRWeB 1.41/2.16 3.71/2.63 3.44/2.43 0.12/0.26 0.58/0.53 5.1 Settings We instantiate TransPhy on BAGEL-7B-MoT and freeze the backbone. The understanding and generation stages use LoRA ranks 16 and 32, respectively; the generation adapter contains four MoE-LoRA experts with top-1 routing, and STC retains the top 15% transition-responsive ViT tokens. We optimize both stages with AdamW and bfloat16 precision at a learning rate of 1Ă10â41Ă 10^-4 with 100 warmup steps. The understanding stage is trained for 8,000 steps and the generation stage for 20,000 steps. Training uses full-shard FSDP on four A100 GPUs, a per-GPU batch size of 1, and two-step gradient accumulation, yielding an effective batch size of 8. Further implementation details are provided in the supplementary material. 5.2 Benchmark PhysVICL-74 contains 74 transition rules, 5,240 sourceâtarget image pairs, and approximately 75K AABB contexts. The evaluation benchmark combines three disjoint sources: novel instances of 57 PhysVICL-74 rules represented in training, 17 PhysVICL-74 rules entirely unseen during training, and 20 sampled Relation252K rules not included in our training, with four sampled pairs per rule in each source. For the first split, both the exemplar pairs and queries are image-disjoint from training; for the latter two, the rules and all associated images are held out. We split base image pairs before constructing AABB contexts, ensuring that no test pair or context containing it appears in training. 5.3 Baseline Methods For novel instances of seen rules, we compare TransPhy with FLUX.1-Fill-dev and BAGEL-MoE. FLUX.1-Fill-dev is trained on the same PhysVICL-74 split, whereas BAGEL-MoE shares our backbone, adapters, training data, and inference settings but removes STC-based router alignment. For unseen rules, we additionally compare with RelationAdapter (Gong et al. 2025), VisualCloze (Li et al. 2025), and LoRWeB (Manor et al. 2026), using their released checkpoints, official inference settings, and native input formats. All methods receive the same exemplarâquery semantics. Because RelationAdapter and LoRWeB do not publicly disclose detailed training-data splits, exact training-data alignment cannot be verified. 5.4 Evaluation Metrics We evaluate with five complementary metrics. We adopt CLIP Directional Similarity (CLIP-D) (Manor et al. 2026) using OpenAI CLIP ViT-L/14 and Learned Perceptual Image Patch Similarity (LPIPS) (Zhang et al. 2018). Following VLM-based protocols for visual-analogy and knowledge-aware editing (Manor et al. 2026; Lin et al. 2026), GPT-5.6 scores Transition Accuracy (TA), Content Preservation (CP), and Rule Plausibility (RP) on a 0â4 scale, with higher scores being better. These three complementary metrics assess transition fidelity, query preservation, and physical consistency, respectively. Full definitions, prompts, and aggregation details are provided in the supplementary material. We additionally conduct an anonymized, criterion-specific two-alternative forced-choice (2AFC) user study; details are provided in the supplementary material. Quantitative Evaluation. As shown in Table 1, TransPhy achieves clear overall gains over BAGEL-MoE on seen transformations, increasing Object-Level TA from 3.03 to 3.39 and Matter-Level CP from 3.23 to 3.67. On unseen transformations (Table 2), it outperforms the matched BAGEL-MoE baseline on every metric in both settings and ranks first in six of ten metricâsetting combinations among all compared methods. These results demonstrate strong and balanced generalization across transition fidelity, query preservation, physical plausibility, and perceptual fidelity. Qualitative Evaluation. Figure 3 compares seen and unseen transformations. On seen rules (left), TransPhy faithfully realizes diverse physical changes while preserving query geometry; FLUX.1-Fill-dev captures coarse target appearances, whereas BAGEL-MoE spreads or under-applies edits. On unseen rules (right), our method transfers material, representation, and lighting effects with limited background changes. RelationAdapter often under-applies the effect, VisualCloze alters background structures, and LoRWeB leaves queries nearly unchanged, showing that TransPhy better balances complete transfer and query preservation. 6 Ablation Studies To investigate the impact of different components of our framework, we conduct the following ablation studies. (1) Staged Rule Understanding and Expert Alignment. We ablate the two central training components in TransPhy. W/o Stage 1 skips dedicated rule-understanding training and conditions the generation pathway on a single fixed instruction shared across instances. W/o STC alignment retains Stage 1 but removes ÎťSTCââSTC _ STCL_ STC from generation training while keeping the generation and load-balancing objectives unchanged. This directly tests whether explicit rule internalization and STC-guided expert routing are necessary for generalizable transformation transfer. (2) Number of Rendering Experts. We vary the number of MoE-LoRA rendering experts as Eâ4,8,16Eâ\4,8,16\ and use E=4E=4 in the main experiments. This ablation examines whether a larger expert pool provides additional capacity for modeling diverse visual transformations. Table 3: Ablation of training design and expert count. The full model uses staged training, STC alignment, and four experts. The first two variants modify the training design; the last two change only the expert count. Variant TAâ CPâ RPâ CLIP-Dâ LPIPSâ Full model (E=4E=4) 3.25 3.46 3.36 0.27 0.31 W/o Stage 1 1.78 3.08 1.87 0.21 0.39 W/o STC align. 3.18 3.24 2.95 0.26 0.37 E=8E=8 3.28 3.46 3.66 0.29 0.29 E=16E=16 3.34 3.52 3.74 0.29 0.30 (3) Transition-Token Selection Sensitivity. To evaluate STC token localization, GPT-5.6 annotates the transformation-affected ViT-grid tokens M for a fixed subset of five sourceâtarget pairs per transformation rule. The masks are audited by Qwen3-VL-32B, with stratified manual spot checks and corrections; criteria are provided in the supplement. We retain the top Ďâ5,15,30%Ďâ\5,15,30\\% candidates by ViT cosine difference and apply the same merging and refinement steps to obtain SĎS_Ď. We report token-level precision, recall, and F1 between SĎS_Ď and M, averaged across samples. Table 4: Transition-token selection quality. Each Ď is evaluated on the fixed subset after connected-component merging and feature-similarity refinement. Metric Ď=5%Ď=5\% Ď=15%Ď=15\% Ď=30%Ď=30\% Token precisionâ 84.18 69.19 41.65 Token recallâ 28.25 65.03 72.57 Token F1â 42.30 67.05 52.92 Figure 4: Transition-token visualization across Ď. Figure 5: Expert routing visualization. Left: layer-wise expert-routing heat map. Right: results obtained by hard-routing all tokens through experts 0â3. (4) Expert Intervention and Interpretability. Holding all other settings fixed, we route all generation tokens through a single expert eâ0,1,2,3eâ\0,1,2,3\ and compare the outputs with natural layer-wise routing in Figure 5. Results & Analysis. (1) Table 3 shows that removing Stage 1 causes the largest drops in TA (3.25 to 1.78) and RP (3.36 to 1.87), confirming that rule internalization is central to transfer. Removing STC alignment reduces RP to 2.95 and increases LPIPS from 0.31 to 0.37, indicating that expert alignment improves plausible and spatially faithful rendering. (2) Increasing E from 4 to 16 yields modest gains: E=8E=8 gives the lowest LPIPS, whereas E=16E=16 performs best on TA, CP, and RP. This suggests the bottleneck may lie in the base model rather than expert capacity. (3) Table 4 and Figure 4 show that Ď=5%Ď=5\% favors precision but has low transition-region recall, whereas Ď=30%Ď=30\% improves recall at the cost of noisy selections. Ď=15%Ď=15\% achieves the highest F1 of 67.05, validating a moderate selection ratio. (4) Figure 5 shows that routing varies across layers and that forced experts produce distinct localized effects. This provides evidence of emergent, non-redundant expert specialization rather than fixed semantic roles. Taken together, these complementary results provide consistent evidence for the effectiveness of staged induction and transition-aligned expert routing. 7 Conclusion We formulated physically grounded VICL, where models infer a transformation rule from an exemplar and apply it to a new query. We introduced PhysVICL-74, covering 74 rules and approximately 75K exemplarâquery contexts, with protocols for novel-instance transfer and unseen-rule generalization. We proposed TransPhy, combining physical-rule induction with transition-aligned token-wise expert adaptation for fine-grained and physically plausible synthesis. Extensive experiments show strong performance on physically grounded transformations and improved generalization to unseen rules over existing VICL methods. We hope these contributions support future study of visual rule induction and context-aware image editing. References Black Forest Labs et al. (2025) Black Forest Labs, S. Batifol, A. Blattmann, F. Boesel, S. Consul, C. Diagne, T. Dockhorn, J. English, Z. English, P. Esser, S. Kulal, K. Lacey, Y. Levi, C. Li, D. Lorenz, J. MĂźller, D. Podell, R. Rombach, H. Saini, A. Sauer, and L. Smith FLUX.1 kontext: flow matching for in-context image generation and editing in latent space. External Links: 2506.15742 Cited by: §1. Brooks et al. (2023) T. Brooks, A. Holynski, and A. A. Efros InstructPix2Pix: learning to follow image editing instructions. External Links: 2211.09800, Link Cited by: §1. Chen et al. (2026) J. Chen, S. Li, H. Fu, B. Zhao, W. Liu, Y. Liang, L. Qing, and X. Mao Delta-adapter: scalable exemplar-based image editing with single-pair supervision. External Links: 2605.07940 Cited by: §2. Chen et al. (2025a) L. Chen, Q. Mao, Y. Gu, and M. Z. Shou Edit transfer: learning image editing via vision in-context relations. External Links: 2503.13327 Cited by: §1. Chen et al. (2025b) X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan Janus-pro: unified multimodal understanding and generation with data and model scaling. External Links: 2501.17811 Cited by: §2. Chow et al. (2025) W. Chow, J. Mao, B. Li, D. Seita, V. Guizilini, and Y. Wang PhysBench: benchmarking and enhancing vision-language models for physical world understanding. External Links: 2501.16411 Cited by: §2. Das Biswas et al. (2025) S. Das Biswas, M. Shreve, X. Li, P. Singhal, and K. Roy PIXELS: progressive image xemplar-based editing with latent surgery. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 2663â2671. External Links: Document Cited by: §2. Deng et al. (2025) C. Deng, D. Zhu, K. Li, C. Gou, F. Li, Z. Wang, S. Zhong, W. Yu, X. Nie, Z. Song, G. Shi, and H. Fan Emerging properties in unified multimodal pretraining. External Links: 2505.14683 Cited by: §2, §4.3. Ding et al. (2026) Y. Ding, W. Huang, R. Quan, X. Qi, and Y. Yang Is this edit correct? a multi-dimensional benchmark for reasoning-aware image editing. External Links: 2606.05172 Cited by: §1, §2, §3. Feng et al. (2025) A. Feng, W. Qiu, J. Bai, Z. Dong, K. Zhou, X. Zhang, R. Ying, and L. Tassiulas An item is worth a prompt: versatile image editing with disentangled control. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 16559â16567. External Links: Document Cited by: §1. Feng et al. (2024) K. Feng, Y. Ma, B. Wang, C. Qi, H. Chen, Q. Chen, and Z. Wang DiT4Edit: diffusion transformer for image editing. External Links: 2411.03286, Link Cited by: §1. Gong et al. (2025) Y. Gong, Y. Song, Y. Li, C. Li, and Y. Zhang RelationAdapter: learning and transferring visual relation with diffusion transformers. External Links: 2506.02528 Cited by: §1, §2, §5.3. Gu et al. (2024) Z. Gu, S. Yang, J. Liao, J. Huo, and Y. Gao Analogist: out-of-the-box visual in-context learning with image diffusion model. External Links: 2405.10316 Cited by: §2. Guo et al. (2026) S. Guo, S. He, C. Meng, S. Xiao, X. Xiang, S. Zhang, and Q. Fan PhyEditBench: a real-world multi-stage benchmark for physics-aware image editing. External Links: 2606.26551 Cited by: §2. Han et al. (2025) F. Han, Y. Wang, C. Li, Z. Liang, D. Wang, Y. Jiao, Z. Wei, C. Gong, C. Jin, J. Chen, and J. Wang UniREditBench: a unified reasoning-based image editing benchmark. External Links: 2511.01295 Cited by: §2, §3. Hu et al. (2022) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, Cited by: §4.2. Jones et al. (2024) M. Jones, S. Wang, N. Kumari, D. Bau, and J. Zhu Customizing text-to-image models with a single image pair. External Links: 2405.01536, Link Cited by: §2. Li et al. (2026) Z. Li, Z. Duan, J. Ye, C. Chen, D. Chen, Y. Li, and Y. Chen VIRAL: visual in-context reasoning via analogy in diffusion transformers. External Links: 2602.03210 Cited by: §2. Li et al. (2025) Z. Li, R. Du, J. Yan, L. Zhuo, Z. Li, P. Gao, Z. Ma, and M. Cheng VisualCloze: a universal image generation framework via visual in-context learning. External Links: 2504.07960 Cited by: §1, §2, §5.3. Lin et al. (2026) W. Lin, F. Wang, M. Zhang, W. Hu, T. Jin, Z. Zhao, F. Wu, J. Chen, A. Yuille, and S. Ren WorldEdit: towards open-world image editing with a knowledge-informed benchmark. External Links: 2602.07095 Cited by: §1, §2, §3, §5.4. Liu et al. (2024) Y. Liu, F. Wu, R. Li, Z. Tang, and K. Li PAR: prompt-aware token reduction method for efficient large multimodal models. External Links: 2410.07278, Link Cited by: §4.3. Lu et al. (2025) H. Lu, J. Chen, Z. Yang, A. T. Gnanha, F. L. Wang, L. Qing, and X. Mao PairEdit: learning semantic variations for exemplar-based image editing. External Links: 2506.07992 Cited by: §2. Ma et al. (2026) J. Ma, X. Zhu, Z. Pan, Q. Peng, X. Guo, C. Chen, and H. Lu X2Edit: revisiting arbitrary-instruction image editing through self-constructed data and task-aware representation learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 7764â7772. External Links: Document Cited by: §1. Manor et al. (2026) H. Manor, R. Gal, H. Maron, T. Michaeli, and G. Chechik Spanning the visual analogy space with a weight basis of LoRAs. External Links: 2602.15727 Cited by: §1, §2, §5.3, §5.4. Mao et al. (2026) J. Mao, K. Wang, Y. Xiang, and K. Chen TweezeEdit: consistent and efficient image editing with path regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 7936â7944. External Links: Document Cited by: §1. Nguyen et al. (2023) T. Nguyen, Y. Li, U. Ojha, and Y. J. Lee Visual instruction inversion: image editing via visual prompting. External Links: 2307.14331, Link Cited by: §2. Pu et al. (2025) Y. Pu, L. Zhuo, S. Han, J. Xing, K. Zhu, S. Cao, B. Fu, S. Liu, H. Li, Y. Qiao, W. Zhang, X. Chen, and Y. Liu PICABench: how far are we from physically realistic image editing?. External Links: 2510.17681 Cited by: §2. Rajagopalan and Patel (2025) S. Rajagopalan and V. M. Patel AWRaCLe: all-weather image restoration using visual in-context learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 6675â6683. External Links: Document Cited by: §1. Samadi et al. (2025) M. Samadi, F. X. Han, M. Salameh, H. Wu, F. Sun, C. Zhou, and D. Niu FunEditor: achieving complex image edits via function aggregation with diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 6758â6766. External Links: Document Cited by: §1. Shang et al. (2026) Y. Shang, M. Cai, B. Xu, Y. J. Lee, and Y. Yan LLaVA-prumerge: adaptive token reduction for efficient large multimodal models. External Links: 2403.15388, Link Cited by: §4.3. Sheng et al. (2026) Z. Sheng, X. Han, Z. Zhang, Z. Xiong, Y. Ding, A. Ping, X. Li, T. Guo, and Y. Mao InEdit-Bench: benchmarking intermediate logical pathways for intelligent image editing models. External Links: 2603.03657 Cited by: §1, §2, §3. Song et al. (2026a) L. Song, W. Li, G. Ma, W. Tang, B. Wang, Y. Zhang, Y. Yang, Y. Xiao, J. Liu, Y. Zhang, G. Zhang, W. Zhang, H. Xu, N. Jiang, X. Han, H. Sun, M. Zhang, H. Huang, and N. Duan Awaking spatial intelligence in unified multimodal understanding and generation. External Links: 2605.04128 Cited by: §2. Song et al. (2026b) W. Song, H. Jiang, Z. Yang, Z. Cheng, R. Quan, and Y. Yang Insert anything: image insertion via in-context editing in dit. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 9097â9105. External Links: Document Cited by: §1. Song et al. (2024) X. Song, J. Cui, H. Zhang, J. Shi, J. Chen, C. Zhang, and Y. Jiang LoRA of change: learning to generate LoRA for the editing instruction from a single before-after image pair. External Links: 2411.19156 Cited by: §2. Srivastava et al. (2024) A. Srivastava, T. R. Menta, A. Java, A. Jadhav, S. Singh, S. Jandial, and B. Krishnamurthy ReEdit: multimodal exemplar-based image editing with diffusion models. External Links: 2411.03982, Link Cited by: §2. Sun et al. (2023) Y. Sun, Y. Yang, H. Peng, Y. Shen, Y. Yang, H. Hu, L. Qiu, and H. Koike ImageBrush: learning visual in-context instructions for exemplar-based image manipulation. External Links: 2308.00906, Link Cited by: §2. Wang et al. (2023) Z. Wang, Y. Jiang, Y. Lu, Y. Shen, P. He, W. Chen, Z. Wang, and M. Zhou In-context learning unlocked for diffusion models. External Links: 2305.01115, Link Cited by: §2. Wu et al. (2024a) C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan, and P. Luo Janus: decoupling visual encoding for unified multimodal understanding and generation. External Links: 2410.13848 Cited by: §2. Wu et al. (2025a) J. Z. Wu, X. Ren, T. Shen, T. Cao, K. He, Y. Lu, R. Gao, E. Xie, S. Lan, J. M. Alvarez, J. Gao, S. Fidler, Z. Wang, and H. Ling ChronoEdit: towards temporal reasoning for image editing and world simulation. External Links: 2510.04290 Cited by: §2. Wu et al. (2024b) X. Wu, S. Huang, and F. Wei Mixture of LoRA experts. In International Conference on Learning Representations, Cited by: §4.2. Wu et al. (2025b) Y. Wu, Z. Li, X. Hu, X. Ye, X. Zeng, G. Yu, W. Zhu, B. Schiele, M. Yang, and X. Yang KRIS-Bench: benchmarking next-level intelligent image editing models. External Links: 2505.16707 Cited by: §2, §3. Xie et al. (2025) J. Xie, Z. Yang, and M. Z. Shou Show-o2: improved native unified multimodal models. External Links: 2506.15564 Cited by: §2. Xu et al. (2026) R. Xu, D. Zhou, X. Shen, F. Ma, and Y. Yang PhyEdit: towards real-world object manipulation via physically-grounded image editing. External Links: 2604.07230, Link Cited by: §1. Zhang et al. (2018) R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §5.4. Zhang et al. (2025) Z. Zhang, J. Xie, Y. Lu, Z. Yang, and Y. Yang In-context edit: enabling instructional image editing with in-context generation in large scale diffusion transformer. External Links: 2504.20690 Cited by: §1. Zhao et al. (2026) L. Zhao, L. Zhuo, S. Paul, H. Li, and M. Elhoseiny From statics to dynamics: physics-aware image editing with latent transition priors. External Links: 2602.21778, Link Cited by: §2. Zhao et al. (2025) X. Zhao, P. Zhang, K. Tang, H. Li, Z. Zhang, G. Zhai, J. Yan, H. Yang, X. Yang, and H. Duan Envisioning beyond the pixels: benchmarking reasoning-informed visual editing. External Links: 2504.02826 Cited by: §1, §2, §3.