Paper deep dive
AI models of unstable flow exhibit hallucination
Ramdhan Wibawa, Birendra Jha
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/26/2026, 5:42:06 PM
Summary
The paper reports the first systematic evidence of 'hallucination' in AI models of fluid dynamics, specifically in the context of viscous fingering. The authors demonstrate that models like Vision Transformers (ViT) and DAE-LSTM produce visually coherent but physically implausible results, such as spurious fluid interfaces and reverse diffusion, due to spectral bias. To address this, they introduce 'DeepFingers', a framework combining the Fourier Neural Operator (FNO) with a Deep Operator Network (DeepONet) and U-FNO layers. DeepFingers achieves balanced learning across spatial modes and accurately captures complex phenomena like tip splitting and finger merging, outperforming baseline models in both physical consistency and global mixing metrics.
Entities (9)
Relation Signals (4)
DeepFingers ā combines ā Fourier Neural Operator
confidence 100% Ā· combining the Fourier Neural Operator with a Deep Operator Network
Vision Transformer ā exhibits ā Hallucination
confidence 100% Ā· ViT model struggles particularly at later time steps, producing unrealistic structures... Such unrealistic flow structures can be regarded as hallucinations
Viscous Fingering ā ismodeledby ā DeepFingers
confidence 100% Ā· DeepFingers... to predict the spatiotemporal evolution of viscous fingers.
DeepFingers ā mitigates ā Spectral Bias
confidence 90% Ā· DeepFingers achieves a balanced distribution across all modes... avoids the spectral biases observed in ViT and DAEāLSTM
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We report the first systematic evidence of hallucination in AI models of fluid dynamics, demonstrated in the canonical problem of hydrodynamically unstable transport known as viscous fingering. AI-based modeling of flow with instabilities remains challenging because rapidly evolving, multiscale fingering patterns are difficult to resolve accurately. We identify solutions that appear visually realistic yet are physically implausible, analogous to hallucinations in large language models. These hallucinations manifest as spurious fluid interfaces and reverse diffusion that violate conservation laws. We show that their origin lies in the spectral bias of AI models, which becomes dominant at high flow rates and viscosity contrasts. Guided by this insight, we introduce DeepFingers, a new framework for AI-driven fluid dynamics that enforces balanced learning across the full spectrum of spatial modes by combining the Fourier Neural Operator with a Deep Operator Network to predict the spatiotemporal evolution of viscous fingers. By conditioning on both time and viscosity contrast, DeepFingers learns mappings between successive concentration fields across regimes. The framework accurately captures tip splitting, finger merging, and channel formation while preserving global metrics of mixing. The results open a new research direction to investigate fundamental limitations in AI models of physical systems.
Tags
Links
- Source: https://arxiv.org/abs/2604.20372v1
- Canonical: https://arxiv.org/abs/2604.20372v1
Trouble viewing inline? Open PDF directly ā
Full Text
48,553 characters extracted from source content.
Expand or collapse full text
AI models of unstable flow exhibit hallucination Ramdhan Wibawa and Birendra Jha Department of Chemical Engineering and Materials Science, University of Southern California, 925 Bloom Walk, Los Angeles, 90089, CA, USA. *Corresponding author(s). E-mail(s):bjha@usc.edu; Abstract We report the first systematic evidence of hallucination in AI models of fluid dynamics, demonstrated in the canonical problem of hydrodynamically unstable transport known as viscous fingering. AI-based modeling of flow with instabili- ties remains challenging because rapidly evolving, multiscale fingering patterns are diļ¬icult to resolve accurately. We identify solutions that appear visually real- istic yet are physically implausible, analogous to hallucinations in large language models. These hallucinations manifest as spurious fluid interfaces and reverse diffusion that violate conservation laws. We show that their origin lies in the spec- tral bias of AI models, which becomes dominant at high flow rates and viscosity contrasts. Guided by this insight, we introduce DeepFingers, a new framework for AI-driven fluid dynamics that enforces balanced learning across the full spec- trum of spatial modes by combining the Fourier Neural Operator with a Deep Operator Network to predict the spatiotemporal evolution of viscous fingers. By conditioning on both time and viscosity contrast, DeepFingers learns mappings between successive concentration fields across regimes. The framework accurately captures tip splitting, finger merging, and channel formation while preserving global metrics of mixing. The results open a new research direction to investigate fundamental limitations in AI models of physical systems. Keywords:Viscous fingering, Neural Operator, Spectral bias, AI hallucination, Fluid dynamics Introduction Hydrodynamic instabilities such as viscous fingering influence many natural and engi- neered systems, including chemical [ 1], pharmaceutical [2,3], food processing [4], mantle convection [5], and groundwater systems [6]. Viscous fingering arises at the 1 arXiv:2604.20372v1 [physics.flu-dyn] 22 Apr 2026 interface between two fluids of differing viscosities when a less viscous fluid displaces a more viscous fluid within a porous domain [7ā11]. This classical fluid mechanics phenomenon, primarily driven by the SaffmanāTaylor instability, has been exten- sively studied in both miscible and immiscible fluid systems [12ā14]. In miscible flow displacements, such as water displacing glycerin in Hele-Shaw cells [15,16] or CO 2 displacing crude oil or water in petroleum or geothermal reservoirs [10,17,18], the advancing interface evolves into intricate finger-like patterns that grow increasingly complex as the instability develops throughout the domain. Accurate prediction of such instabilities is crucial for numerous applications, including enhanced oil recovery, CO 2 sequestration, groundwater remediation, hydraulic fracturing, and the design of microfluidic systems in biomedical engineering [10,19ā22]. Uncontrolled finger propagation can significantly degrade operational performance by prolonging contam- inant removal, reducing reservoir sweep eļ¬iciency [23], and increasing the amount of CO 2 required for effective sequestration [24ā26]. On the other hand, enhancement of fingering has been proposed as a mechanism for faster mixing under laminar flow conditions [9,27]. Recently, the complexity and randomness of the fingers have been utilized to propose an anti-counterfeiting solution [28]. Predicting viscous fingering is challenging because small perturbations can rapidly evolve into complex flow patterns across multiple scales. Numerical simulation of vis- cous fingering presents notable challenges due to the nonlinear and multiscale nature of the phenomenon [27,29]. The system is governed by coupled partial differen- tial equations (PDEs) derived from the conservation of mass and momentum. These equations link physical parameters such as fluid viscosity and density to the evolving state variables like the fluid concentration field. Through non-dimensionalization, key dimensionless groups such as the viscosity ratioMand the Peclet number Pe emerge to characterize the dynamics of the instability. Solving these PDEs over large domains (equivalently, large Pe) and large viscosity ratios often relies on direct numerical simu- lation (DNS) techniques, commonly employing finite volume or finite element methods for spatial discretization [30], along with time-stepping schemes such as Euler integra- tion. However, due to the sensitivity of the flow to perturbations in medium properties and initial conditions and the complex feedback between advection and diffusion, many simulations diverge from experimental or field-scale observations, particularly as fingering patterns evolve into highly intricate structures. Recent progress in AI mod- eling of fluid flow, especially deep learning models, offers promise in addressing these challenges. This study shows that deep learning models may produce visually con- vincing predictions that nonetheless violate fundamental physical principles, a failure mode analogous to hallucinations in language models. Hallucination of AI models has been extensively documented and studied in large language modelsāprompting major efforts to understand its societal impact and develop mitigation strategies [31]. However, it has been largely assumed that physics- informed or data-driven AI models operating in scientific domains are immune to such failures. We report the first systematic evidence of hallucination in AI models of fluid dynamics. Using modern architectures including Vision Transformers, we demonstrate that AI models trained to predict viscous fingering dynamics can exhibit physically inconsistent behavior despite maintaining visual coherence. We further identify the 2 Fig. 1The DeepFingers architecture for modeling viscous fingering of the less viscous fluid (yellow) through a more viscous fluid (black). It applies FNO at the branch network and a fully connected (FC) layer at the trunk network. The outputs are merged before processing in a series of two U-FNO layers and a projection layer (Q) to generate the concentration map at the next time step. origin of hallucination in fluid-dynamics AI models as spectral bias, wherein learning architectures disproportionately favor certain length scales at the expense of others. Guided by this insight, we introduce a new deep learning framework, combining DeepONet [32] and Fourier Neural Operator (FNO) [33] frameworks, that enforces balanced learning across the full spectrum of spatial modes. We demonstrate how such a design is necessary to model unstable flows at the full-field scale or, equivalently, at large Peclet numbers representative of real-world systems. This raises awareness of the need for physical consistency when designing and evaluating deep learning models for scientific applications. Results Flows with hydrodynamic instabilities, and viscous fingering in particular, present a multi-fidelity modeling challenge. The bulk fluid region and the advancing fingers of the less viscous fluid evolve dynamically over time, dictated by partial differen- tial equations (PDEs) of mass and momentum balance, while intricate mechanisms emerge at the finger tips and along the interface between the two fluids. These regions exhibit sharp or smooth gradient transitions, depending on the interplay of nonlin- ear advection and diffusion. Overall spreading and mixing behavior is governed by globally defined flow metrics. Details of the governing PDEs, flow metrics, and their implications for instability are provided in the Supporting Information. On the AI modeling side, it is known that Fourier Neural Operator (FNO) layers, which implement integral kernel operators formulated from Greenās function, provide an expressive and eļ¬icient representation for solving PDEs whose solutions contain multiple length scales. The effectiveness of FNO has been demonstrated across a range of PDE benchmarks, where it outperforms other deep learning architectures in 3 both accuracy and computational eļ¬iciency [33ā36]. In the broader context of oper- ator learning, DeepONet has also shown strong performance across diverse physical systems [37ā42], further highlighting the advantages of learning solution operators directly from data. Inspired by these developments, we propose a new architecture for modeling flows with fingering instabilities: DeepFingers. The DeepFingers architecture (Fig.1) extends the DeepONet framework [32] by incorporating a FNO [33] in the branch network, a fully connected structure in the trunk network, and a series of U-FNO layers [43,44] that embed the well-known U-Net [45] structure to enhance feature extraction and multi-scale representation within the FNO architecture. The U-FNO layers enhance representational capacity by capturing high-frequency components that are not adequately resolved by the standard Fourier basis [43]. The framework consists of two separate components: a branch network and a trunk network. The branch network processes the input function at fixed sensors locations, while the trunk network receives the input parameters (often referred to as coordinates or locations in the original formulation). This design is well-suited to our objective of investigating the fingering dynamics, where the branch network represents the grid of concentration field and the trunk network process timetand viscosity ratioMas input parameters. The assessment of DeepFingers covers four aspects: computational eļ¬iciency, hal- lucination detection, comparison with baseline models using physics-based metrics and spectral analysis, and performance under uncertainty. Predictions are initialized from prescribed initial conditions and advanced auto-regressively in time. Because fingering becomes increasingly unstable as theMvalue increases, a consequence of the exponential viscosity function (see equations in the Supplementary Information), physically grounded diagnostics and spectral analysis are chosen over pixel-wise errors. To evaluate its effectiveness, we compare DeepFingers against two state-of-the- art AI modeling frameworks that employ a two-stage modeling approach, previously developed for cylindrical flow problems [46,47] and for Darcy flows with small Peclet numbers. In addition, we benchmark it against a Vision Transformer (ViT)-based model [48], a widely recognized architecture for image-related tasks, used here as a transformer-based baseline. We also note recent progress in applying transformer architectures to physics-based flow modeling, such as the CViT framework [49], which further illustrates the growing role of attention mechanisms in scientific machine learn- ing. DeepFingers is benchmarked against DAE-LSTM and ViT across a range ofM values and time horizons to provide a comprehensive measure of model robustness. Reconstruction of the concentration field from DNS and various deep learning models is presented in Fig.2. The comparison includes DeepFingers, ViT, and DAE- LSTM across a range ofMvalues and time. DNS provides a physical benchmark that begins with the emergence of narrow fingers, followed by repeated tip-splitting, merg- ing, and shielding interactions that drive the instability. Accurately reproducing these highly nonlinear and spatio-temporal dynamics directly from data, without solving the governing equations, is an exceptionally diļ¬icult task [11]. The proposed DeepFin- gers demonstrates a strong ability to recover these essential physical characteristics. At lowM, it correctly predicts short, slowly advancing fingers. AsMincreases, it 4 captures the accelerated finger growth, more frequent tip-splitting events, and the for- mation of intricate patterns. At the highest viscosity ratio (M= 55), the model even reproduces the onset of channeling [50], where high-mobility pathways dominate and significantly alter the fluid displacement process. Hallucinations in AI models of flows with fingers The Densenet Autoencoder model (DAE-LSTM) exhibits noticeable errors in the early stages of flow. Spurious patches of the less viscous fluid (yellow-colored) emerge within otherwise black regions of the more viscous fluid; see Fig.2row 4 column 1. At later times, the predicted fingers appear overly smooth, with diffused tips and insuļ¬icient splitting. In contrast, the ViT model struggles particularly at later time steps, producing unrealistic structures, such as black islands within yellow regions; see Fig.2row 7, column 5. These structures are inconsistent with the underlying phys- ical process. Although these artifacts represent errors in the modelās learning, they are qualitatively distinct from the inaccuracies typically observed in DNS simulations. Such unrealistic flow structures can be regarded as hallucinations of the DL model. These limitations underscore the inherent diļ¬iculty of achieving robust generalization across spatial and temporal scales using conventional neural network. In contrast, errors in DNS modeling stem primarily from the accuracy and stability of the spatial and temporal discretizations schemes. Originally introduced in Neural Machine Translation (NMT) [31], the term halluci- nation has since been widely adopted to describe a class of errors commonly observed in Large Language Models (LLMs) [51]. However, within the broader AI community, its definition remains vague and inconsistent. As highlighted in a recent review [52], no universally accepted characterization or criteria currently exist, and the interpre- tation of hallucination often varies across domains, sometimes even contradictorily. In general, the term refers to outputs that appear plausible and coherent yet are fun- damentally incorrect or unsupported by the input data [53]. AI-generated content, for instance, can produce highly convincing but fabricated or misleading information, raising growing concerns about misinformation and model reliability [54]. In this study, we extend the notion of hallucination to AI models of physics, in particular, physics of porous media flows, where such behavior manifests as predic- tions that appear visually coherent but violate fundamental physical laws or deviate from expected flow dynamics. Specifically, we use the term hallucination to describe distinct, physically unrealistic patterns produced by models, such as the DAE-LSTM and ViT, whose errors cannot be explained by known numerical or DNS-related arti- facts. This perspective bridges the conceptual gap between hallucinations in LLMs and spurious behaviors in physics-based DL model, offering a framework to identify, interpret, and mitigate nonphysical predictions in data-driven simulations of fluid dynamics. Overall, the results confirm that while existing architectures such as DAE- LSTM and ViT fail to capture the rich multiscale complexity of VF, DeepFingers succeeds in reproducing the physically realistic evolution of thecfield across a range ofMand time. 5 Fig. 2Hallucination detection in AI models of fingering. DeepFingers, the proposed architecture, produces stable, realistic results that closely follow the physical evolution, whereas ViT introduces nonphysical black islands (hallucinations) at later times, and DAE-LSTM yields diffused finger tips and spurious artifacts (hallucinations) at early times.Mis the viscosity ratio. 6 Spectral analysis Spectral-based mode analysis provides a quantitative framework for assessing the multiscale fidelity of the reconstructed, non-stationary concentration fields of flow. We conduct spectral analysis using wavelets, instead of harmonic functions, because wavelets offer simultaneous localization in time and frequency, making them ideal for analyzing non-stationary, transient signals. The resulting spectrum reveals how features of different scales (incipient fingers to finger merging to channeling) evolve dynamically at each time step. As illustrated in Fig.3, the seven wavelet modes exhibit a clear redistribution of spectral energy over time, where lower modes correspond to large-scale, low-frequency structures and higher modes capture small-scale, high- frequency features. For the ViT model, the first mode remains consistently below DNS, reflecting a slower propagation of the dominant large-scale yellow region shown in Fig.2, with this bias later evident in the domain-averaged concentration Ģc(t). In contrast, the higher modes (2ā7) of the ViT tend to exceed DNS levels, indicating an overestimation of small-scale fluctuations such as fingers, which manifests in mean dissipation rateε c . The DAEāLSTM model exhibits a more complex,M-dependent pattern: while mode 1 is lower than DNS only forM= 13, at higher viscosity ratios (M= 28and M= 55) the higher modes (3ā7) fall below DNS. This deviation arises because the model produces fewer fingers overall, leading to weaker fine-scale activity and overly diffused boundaries especially at early time steps. In contrast, DeepFingers achieves a balanced distribution across all modes. Although pixel-level finger patterns do not match DNS exactly, an expected conse- quence of the intrinsic instability of fingering, the modes distribution across scales aligns closely with DNS. This balance indicates that DeepFingers avoids the spec- tral biases observed in ViT and DAEāLSTM and is therefore better suited in the physics-based evaluations presented in the next section. To better quantify the overall biases, the mean spectral energy for each viscos- ity ratioMis calculated by aggregating across all time steps. Fig. 4summarizes these biases for the compared models. DeepFingers aligns closely with DNS across all spectral modes, indicating balanced multiscale fidelity. In contrast, the ViT model consistently produces higher spectral energy in modes 2ā6, resulting in a larger over- all area relative to DNS. The behavior of the DAEāLSTM model depends onM. For M= 13the model overestimates energy, but forM= 28andM= 55the higher modes (3-7) fall below DNS, showing the opposite tendency of ViT by underestimating fine-scale finger structures and producing overly diffused boundaries. Addressing hallucination by spectral debiasing The ViT model used in this study relies on the standard attention mechanism of the original architecture, which was primarily designed for static image tasks such as seismic interpretation [55]. Consequently, it lacks the capacity to capture temporal dependencies. Extending the architecture to incorporate temporal awareness therefore represents a promising direction for improving predictive performance. In contrast, the CViT architecture [ 49] was explicitly developed for spatio-temporal physics problems, 7 Fig. 3Spectral modes comparison among DNS, DeepFingers, ViT, and DAEāLSTM. DNS shows the expected redistribution of spectral modes from coarse to fine scales. DeepFingers closely follows this evolution, maintaining a balanced distribution across all modes. ViT underestimates the dominant large-scale mode while overestimating higher modes, leading to exaggerated small-scale fluctuations and nonphysical artifacts. DAEāLSTM exhibitsMdependent deviations, with higher modes sup- pressed due to the smaller number of fingers and overly diffused boundaries. Overall, DeepFingers avoids the spectral biases observed in ViT and DAEāLSTM, reproducing the multiscale dynamics more faithfully. Fig. 4Mean spectral energy comparison among DNS, DeepFingers, ViT, and DAEāLSTM across viscosity ratiosM, aggregated over all time steps. DeepFingers matches DNS closely across all modes, demonstrating balanced multiscale fidelity. ViT consistently overestimates spectral energy in modes 2ā6. In contrast, DAEāLSTM exhibitsMdependent behavior: it overestimates energy atM= 13, but underestimates higher modes atM= 28andM= 55, showing the opposite tendency of ViT due to fewer fingers and overly diffused boundaries. 8 Fig. 5Addressing hallucination in ViT by spectral debiasing. Improved ViT includes cross attention blocks, which removes the nonphysical black (more viscous fluid) islands observed in the original ViT implementation, demonstrating improved temporal prediction and spatial continuity. employing cross-attention mechanisms tailored to solve dynamic problem. However, its trainable grid-based coordinate embeddings incur substantial memory costs when evaluated over all spatial coordinates during training, necessitating random coordinate sampling. Because fingering dynamics are dominated by fine-scale behavior between two fluids, this sampling strategy prevents CViT from learning the interface evolution effectively, leading to severely blurry result. Motivated by these limitations, we incorporate cross-attention blocks inspired by CViT into the ViT framework. This modification substantially improves temporal coherence and spatial continuity, eliminating artifacts such as the nonphysical black islands previously observed within yellow regions (Fig.2row 7, column 5). Nonethe- less, challenges remain: atM= 13, the reconstructed fingers are still shorter and more diffused than the ground truth, as shown in Fig.5. These findings underscore both the strong potential of transformer-based architectures for modeling fingering and the need for continued architectural refinement. We further argue that hybrid strate- gies, such as two-stage frameworks exemplified by DAE-LSTM, represent particularly promising avenues for future research. Global metrics of fluid mixing from fingering We consider four metrics: domain-averaged concentration, breakthrough concentra- tion, degree of mixing, and rate of mixing. See SI for details. 9 Average concentration We compare the temporal evolution of the average concentration Ģc(t)from DNS against different AI modelsāDeepFingers, ViT, and DAE-LSTMāfor various vis- cosity ratioMvalues. This comparison plot in Fig. S3 highlights the proposed DeepFingers frameworkās ability to capture the dynamics from early to late times. The reason is that Ģc(t)is strongly influenced by the movement of the less viscous yel- low fluid and DeepFingers closely approximates this movement, as shown by Fig.2, rows 2, 6, and 10. In contrast, the ViT model deviates significantly and consistently underestimates the Ģc(t)values. Over time, while ViT captures the growth of the fin- gers, the less viscous fluid regions do not progress as expected. This, combined with the effect of hallucinations manifested as black colored (more viscous fluid) islands of the viscous fluid within the less viscous region causes Ģc(t)to remain consistently underestimated. See Fig.2, rows 3, 7, and 11, for ViT results. For the predictions of DAE-LSTM, the model generally follows the overall trend but exhibits noisy and oscillatory behavior, reflecting limited temporal smoothness and spatial continuity. This indicates that while DAE-LSTM can capture the overall growth of the concentration field, it struggles to maintain smooth spatial evolution over time. ForM= 13(row 4 of Fig.2), the slow movement of the less viscous fluid causes Ģc(t)to remain consistently below the DNS result. In contrast, forM= 28 (row 8), the less viscous fluid moves faster than in DNS, leading Ģc(t)to exceed the reference values in the later time steps. ForM= 55, DAE-LSTM approximates Ģc(t) more accurately, but noise and oscillations persist due to spurious patches of less viscous fluid appearing in the more viscous region at early times and diffused finger tips. Overall, DAE-LSTM struggles to produce consistent performance across different Mvalues. Table 1supports these observations with quantitative evaluation results; DeepFingers outperforms both ViT and DAEāLSTM. Breakthrough concentration at outlet boundary The concentration breakthrough curve is an important diagnostic of the transport process and is, therefore, often measured in field and lab experiments. The DeepFin- gers model successfully approximates the breakthrough curve across the range ofM values (Fig. S4). Other DL models, ViT and DAE-LSTM, hallucinate in predicting the breakthrough behavior. The ViT model predicts too fast breakthrough for the low viscosity contrast,M= 13, because the ViT fingers grow more rapidly and reach the outlet boundary earlier; see Fig.2row 3, column 4. TheM= 13breakthrough is faster than the breakthroughs at higherMvalues (see Fig. S4), which is counterin- tuitive because the breakthrough time should decrease as theMincreases. This is an evidence of hallucination. At later time steps forM= 55, hallucinations in the form of black colored (more viscous fluid) islands within the yellow colored (less viscous fluid) region emerge, while the fingers arriving at the outlet boundary remain thin and exhibit limited growth over time. Consequently, Ģc out (t)tends to flatten, showing a slower increase compared to the other models. This shows that the ViT model lacks spatial continuity and accurate temporal prediction. 10 Table 1Global mixing metrics aggregated over all values of the viscosity ratioMin the test dataset. RMSE values are relative to DNS. Bold entries indicate the best performance, highlighting DeepFingers as the leading model capable of addressing hallucinations observed in ViT and DAE-LSTM. ModelsRMSE Ģc Ģc out Ļ 2 c ε c DeepFingers0.004100.009660.003272.67Ć10 ā7 ViT0.11463 0.02389 0.005239.10Ć10 ā7 DAE-LSTM 0.05334 0.03254 0.003858.66Ć10 ā7 As for the DAE-LSTM model, a hallucination in the form of a yellow patch within the black region causes a small increase in Ģc out (t)for a short period at an early time step, as shown in Fig. S4 forM= 28. This suggests that DAE-LSTM also lacks spatial continuity. Similar to the ViT model, forM= 13the fingers grow faster and are too long (compared to DNS). For smallM, the finger growth should be slower due to the dominance of diffusive mechanisms. Consequently, the fingers predicted by DAE-LSTM reach the outlet boundary earlier; see Fig.2, row 4, column 4. AsM increases, the breakthrough time remains nearly constant, which is counterintuitive. increasingMshould result in faster breakthrough, as confirmed by DNS; see Fig. S4. This indicates that DAEāLSTM struggles to produce accurate temporal trajectory predictions. The primary reason is that the latent features generated by the DAE part exhibit a highly complex structure, making it diļ¬icult for the LSTM component to learn the corresponding temporal evolution effectively. Simpler latent structures are still captured effectively by DAE during early stages of fingering, when the patterns remain relatively wellāseparated and dominated by growth rather than merging or other complex interactions. Uncertainty propagation and quantification In geoscience applications, quantifying uncertainty is essential due to the uncertainty in our knowledge about rockās heterogeneity and fluidās initial condition. In unstable flows, spatial variability in rock properties can significantly influence the initial condi- tion and subsequent flow dynamics. To evaluate how uncertainty impacts fluid mixing and spreading in the domain, we analyze the temporal evolution of probability den- sity functions of the metrics introduced above, as illustrated in Fig.6. The violin plots demonstrate that DeepFingers provides a reliable approximation of the uncertainty induced by varying initial conditions. The first three metrics, Ģc(t), Ģc out (t), andĻ 2 c show that DeepFingers closely captures both the variance and median behavior observed. For theε c , atM= 28during early time steps, DeepFingers exhibits a slightly higher variance than DNS, but converges toward the DNS distribution as time progresses. AtM= 55, DeepFingers initially underestimates the distribution relative to DNS, yet aligns more closely at later time steps. 11 Fig. 6Uncertainty propagation in AI model: time evolution of probability density functions (PDFs) of four mixing metrics c,c out ,Ļ 2 c , andε c . Comparison of DeepFingers predictions (gray shading) with DNS (orange shading) illustrates that DeepFingers successfully propagates uncertainty. Source of uncertainty is initial condition variability. AsMincreases, the system exhibits more complex fingering behavior. Early time steps are dominated by advection, leading to dynamic and unstable finger growth, while later stages are increasingly governed by diffusion, resulting in smoother and more predictable patterns. Such fingering-driven uncertainty is visible even when the initial condition is certain, as shown by Fig. S2. The phenomenon exhibits dynamics that bear resemblance to chaotic and turbulent behavior. Discussion AI has been increasingly positioned as a transformative tool for physics, with ambi- tious claims ranging from accelerated discovery to the replacement of direct numerical simulation in complex flow systems. Deep Learning (DL) models offer a powerful alter- native to traditional numerical solvers for approximating complex flow physics directly from data. This approach addresses long-standing challenges in simulating multiscale, 12 nonlinear processes such as viscous fingering, where conventional methods can be com- putationally expensive or become numerically unstable at high values of the viscosity ratio or Peclet number. We show that state-of-the-art AI models can generate predic- tions that are visually plausible yet fundamentally nonphysical, violating governing principles of fluid motion. We formalize this phenomenon by extending the concept of hallucination from large language models (LLMs) to physics-based AI, thereby providing a unifying framework to interpret spurious yet convincing predictions in data-driven simulations. While hallucination has been widely discussed in the context of natural language and computer vision models, its appearance in scientific modeling remains unexplored. Our finding has significant implications for both the foundations of AI and the relia- bility of data-driven modeling in the physical sciences. This conceptual bridge clarifies that hallucination is not domain-specific, but instead reflects deeper inductive biases of learning architectures when applied to multiscale physical systems. By integrating the Fourier neural operator within a deep operator network to address spectral bias, we design a new AI architecture for fingering dynamics that can not only replicate the intricate spatio-temporal evolution of fingering but also adhere to key underlying physics that govern fluid mixing. Another major contribution is the explicit treatment of uncertainty, which lies at the core of all subsurface processes because natural heterogeneity lead to inherently stochastic flow behaviors. Conventional AI-based surrogate models often overlook this aspect, producing single-point estimates that fail to represent the range of plausible outcomes. We show how to leverage the intrinsic stochasticity to generate a sets of physically consistent solutions, thereby enabling the propagation of uncertainty in both space and time. This capability is especially significant for decision-making in reservoir management and other geoscientific applications, where understanding the variability of outcomes is crucial. Methods Fourier Neural Operator (FNO) The goal of the FNO is to learn the operatorG:A7!B. Here,Adenotes the input function andBthe corresponding output function: A:r7!A(r),B:r7!B(r),r2D(1) whererrepresents a point in the domainDR d and may correspond to a spatial coordinate and can also include time if the problem is spatio-temporal, e.g., in viscous fingering. Thus, bothA(r)andB(r)are functions defined over the same domainD. The original implementation [33] introduces the operator layer as an iterative framework to learn the mapping from an input functionAto an output functionB: B=G(A) = ( QL (L) L (1) P ) (A)(2) Here,denotes function composition, andLrepresents the number of layers. The operatorPlifts the input into the initial latent representationz (0) through a linear 13 layer, while eachL (ā) corresponds to theā-th non-linear operator layer. The operator Qprojects the final latent representationz (L) to obtain the output functionB. The non-linear operator defined in Equation (2) updates the representationz (ā) 7!z (ā+1) as follows: L (ā) ( z ( ā ) ) =Ļ ( W ( ā ) z ( ā ) +K ( ā ) ( z ( ā ) )) (3) whereĻis a non-linear activation function applied component-wise,W ( ā ) denotes a linear transformation, andK ( ā ) represents a kernel integral operator defined via the Fourier transform. The kernel operator is computed as: K ( ā ) ( z ( ā ) ) =F ā 1 ( R (ā) F ( z (ā) )) (4) whereFandF ā1 denote the Fourier and inverse Fourier transforms, respectively, and R (ā) is a complex-valued weight tensor applied to a truncated set of Fourier modes. The FNO was further extended to U-FNO [43], which improves the preservation of multi-scale features by concatenating fine and coarse level representations across layers. This is essential for modeling fingering, which is known to display spatial features over a range of length scales, from tiny incipient fingers at the pore-scale to long channels at the domain-scale. In particular, U-FNO is effective at retaining high- frequency Fourier modes, which are essential for capturing fine-scale details such as finger structures. Fig.1illustrates the key differences between FNO and U-FNO. Compared to the standard FNO, U-FNO incorporates an additional U-Net based convolutional operator, denoted byU (ā) . Leth (ā) is the representation feature that propagates through and is updated by the U-FNO layers. We define a new U-FNO operator layerU (ā) , which updates the representationh (ā) 7!h (ā+1) as follows: U (ā) ( h (ā) ) =Ļ ( W (ā) h (ā) +U (ā) ( h (ā) ) +K (ā) ( h (ā) )) (5) whereU (ā) introduces multi-scale convolutional features, whileK (ā) andW (ā) follow the definitions in Equation (3). By combining these components, the network con- structs a richer representation that is better suited to capturing underlying physics of multi-scale dynamics and complex interactions of VF. Deep Operator Networks (DeepONet) DeepONet approximates a nonlinear operatorG:A7!Bby mapping between two infinite-dimensional function spaces using two separate networks, namely the branch and trunk networks. The branch network processes the input functionA, while the trunk network encodes the query locationy2Y. Together, they approximate the output functionG(A)(y)orB. To be precise, let(r 1 ,...,r m )denote a collection ofmgrid points used to discretize the input functionAexpressed asA m = ( A(r 1 ),...,A(r m ) ) . The branch network transforms the discretized inputA m into a vector of branch outputsb= (b 1 ,...,b q )2 R q . Simultaneously, the trunk network evaluates the query locationy2Yto produce the associated trunk outputsĻ= ( Ļ 1 ,...,Ļ q ) 2R q . The output of DeepONet is 14 obtained by combining the branch outputs with the trunk outputs through a merging operator. B=G(A m )(y) = q ā i=1 b i (A m )Ļ i (y)(6) This formulation combines the two representations and allows DeepONet to generalize across different function spaces and spatial domains, with theoretical guarantees of universal approximation [32]. DeepFingers The proposed DeepFingers architecture, as shown in Fig.1, introduces a hybrid frame- work for modeling the fingering problem by combining the strengths of DeepONet and FNO. While retaining the general structure of DeepONet, the branch network is augmented with a single FNO layer (L= 1), as defined in Equation (2), to eļ¬iciently capture global, non-local dependencies of the input function. However, we remove the operatorQin Equation (2). The outputs of the branch and trunk networks are subsequently merged and refined through two consecutive U-FNO layers, which are essential for extracting multi-scale features and resolving the sharp interfacial patterns characteristic of fingering dynamics. A final projection layerSis applied to produce a single-channel output solution in Fig.1. DeepFingers takes the concentration fieldc t as the input function, denoted byA in Equations (2) and (6). In the trunk network, we define the parameter vectorξ, consisting of two scalar values: the viscosity ratioMand the future time step(t+ 1). Formally, ξ= [M,t+ 1] whereξcorresponds toyin Equation ( 6). Accordingly, the VF problem addressed by DeepFingers is formulated as learning the operator mapping: G: (c t ,ξ)7!c t+1 Both the input functionc t and the output functionc t+1 are discretized on a cartesian two-dimensional grid of shape(n ā² y ,n ā² x ) = (64,128)representing the spatial domain of the concentration field, channelC in = 1. Thus, thec t andc t+1 tensors have a shape of(1,64,128). Letb out andt out are the output of branch and trunk networks respectively. Both are lifted to produce output channel ofC out = 64thru operatorPin FNO layer of branch network and linear transformation in trunk network. The results are merged thru pointwise multiplication operation as indicated by crossed node in Fig.1. h (0) =b out t out (7) whereh (0) has shape(C out ,n ā² y ,n ā² x ) = (64,64,128). Representationh (0) is the initial representationhbefore updated in Equation (5). We employ a series of two U-FNO layers to update representationh (0) sequentially according to Equation (5): 15 h (1) =Ļ ( K (1) (h (0) ) +U (1) (h (0) ) +W (1) h (0) ) h (2) =Ļ ( K (2) (h (1) ) +U (2) (h (1) ) +W (2) h (1) ) Lastly, the operatorSprojectsh (2) back to a singleāchannel field, ensuring that the output dimensionality matchesC in = 1. c t+1 =S ( h (2) ) This projection completes the operator mapping by producing the predicted concen- tration field at the next time stepc t+1 . Funding Declaration R.W. acknowledges funding support from the USC Don Paul Fellowship. B.J. acknowledges funding support from the Aramco Research Project. Author Contribution B.J. designed the study and created the DNS model. R.W. created and trained the AI models and prepared the plots. R.W. and B.J. interpreted and analyzed the plots. R.W. and B.J. wrote the initial manuscript draft. B.J. wrote the final manuscript. References [1]Losey, M.W., Schmidt, M.A., Jensen, K.F.: Microfabricated multiphase packed- bed reactors: Characterization of mass transfer and reactions. Ind. Eng. Chem. Res.40, 2555 (2001) [2]Dunn, D.A., Feygin, I.: Challenges and solutions to ultra-high-throughput screen- ing assay miniaturization: submicroliter fluid handling. Drug Discov. Today5, 84 (2000) [3]Baldyga, J., Czarnocki, R., Shekunov, B.Y., Smith, K.B.: Particle formation in supercritical fluids: scale-up problem. Chem. Eng. Res. Des.88, 331 (2010) [4]Cullen, P.J.: Food Mixing: Principles and Applications. Wiley, ??? (2009) [5]Olson, P., Silver, P.G., Carlson, R.W.: The large scale structure of convection in the earthās mantle. Nature344, 209 (1990) [6]Dentz, M., Le Borgne, T., Englert, A., Bijeljic, B.: Mixing, spreading and reaction in heterogeneous media: A brief review. Journal of Contaminant Hydrology120, 1ā17 (2011) 16 [7]Homsy, G.M.: Viscous fingering in porous media. Annual Review of Fluid Mechanics19(1), 271ā311 (1987) [8]Araktingi, U.G., Orr, F.: Viscous fingering in heterogeneous porous media. SPE advanced technology series1(01), 71ā80 (1993) [9]Jha, B., Cueto-Felgueroso, L., Juanes, R.: Fluid mixing from viscous fingering. Physical review letters106(19), 194502 (2011) [10]Zheng, Z., Kim, H., Stone, H.A.: Controlling viscous fingering using time- dependent strategies. Physical review letters115(17), 174501 (2015) [11]Pinilla, A., Asuaje, M., Ratkovich, N.: Experimental and computational advances on the study of viscous fingering: An umbrella review. Heliyon7(7) (2021) [12]Yazdi, A.A., Norouzi, M.: Numerical study of saffmanātaylor instability in immiscible nonlinear viscoelastic flows. Rheologica Acta57(8), 575ā589 (2018) [13]Chui, J.Y., Anna, P., Juanes, R.: Interface evolution during radial miscible viscous fingering. Physical Review E92(4), 041003 (2015) [14]Moyles, I., Wetton, B.: Fingering phenomena in immiscible displacement in porous media flow. Journal of Engineering Mathematics90(1), 83ā104 (2015) [15]Bacri, J.C., Rakotomalala, N., Salin, D., Woumeni, R.: Miscible viscous fingering: Experiments versus continuum approach. Phys. Fluids4, 1611 (1992) [16]Petitjeans, P., Maxworthy, T.: Miscible displacements in capillary tubes. part 1. experiments. Journal of Fluid Mechanics326, 37ā56 (1996) [17]Ren, J., Xiao, W., Cheng, Q., Song, P., Bai, X., Xie, Q., Pu, W., Zheng, L.: Experimental study on water/co2 flow of tight oil using hthp microscopic visu- alization and nmr technology. Geoenergy Science and Engineering250, 213834 (2025) [18]Park, C.-W., Gorell, S., Homsy, G.: Two-phase displacement in hele-shaw cells: experiments on viscously driven instabilities. Journal of Fluid Mechanics141, 275ā287 (1984) [19]Gao, T., Mirzadeh, M., Bai, P., Conforti, K.M., Bazant, M.Z.: Active control of viscous fingering using electric fields. Nature communications10(1), 4002 (2019) [20]Escala, D.M., MuƱuzuri, A.P.: A bottom-up approach to construct or deconstruct a fluid instability. Scientific reports11(1), 24368 (2021) [21]Tran, M., Jha, B.: Coupling between transport and geomechanics affects spread- ing and mixing during viscous fingering in deformable aquifers. Advances in Water Resources136, 103485 (2020) 17 [22]Tu, T.-Y., Shen, Y.-P., Lim, S.-H., Wang, Y.-K.: A facile method for generat- ing a smooth and tubular vessel lumen using a viscous fingering pattern in a microfluidic device. Frontiers in Bioengineering and Biotechnology10, 877480 (2022) [23]Nicolaides, C., Jha, B., Cueto-Felgueroso, L., Juanes, R.: Impact of viscous fin- gering and permeability heterogeneity on fluid mixing in porous media. Water Resources Research51(4), 2634ā2647 (2015) [24]Wang, Y., Zhang, C., Wei, N., Oostrom, M., Wietsma, T.W., Li, X., Bonneville, A.: Experimental study of crossover from capillary to viscous fingering for super- critical co2āwater displacement in a homogeneous pore network. Environmental science & technology47(1), 212ā218 (2013) [25]Gooya, R., Silvestri, A., Moaddel, A., Andersson, M., Stipp, S., SĆørensen, H.: Unstable, super critical co2āwater displacement in fine grained porous media under geologic carbon sequestration conditions. Scientific reports9(1), 11272 (2019) [26]Berg, S., Ott, H.: Stability of co2ābrine immiscible displacement. International Journal of Greenhouse Gas Control11, 188ā203 (2012) [27]Jha, B., Cueto-Felgueroso, L., Juanes, R.: Synergetic fluid mixing from viscous fingering and alternating injection. Physical Review Letters111(14), 144501 (2013) [28]Zhao, W., Wang, S., Zhou, Y., Wang, M., Yang, L., Liu, J., Wang, L., Zhu, P.: Viscous fingering instability in one-end-lifted hele-shaw cells for producing three- dimensional hierarchical structures. Journal of Colloid and Interface Science, 138202 (2025) [29]Tan, C., Homsy, G.: Simulation of nonlinear viscous fingering in miscible displacement. The Physics of fluids31(6), 1330ā1338 (1988) [30]Casademunt, J.: Viscous fingering as a paradigm of interfacial pattern forma- tion: Recent results and new challenges. Chaos: An Interdisciplinary Journal of Nonlinear Science14(3), 809ā824 (2004) [31]Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., Car- lini, N.: Hallucinations in neural machine translation. In: NIPS 2018 Workshop on Interpretability and Robustness in Audio, Speech, and Language (IRASL) (2018). Neural Information Processing Systems [32]Lu, L., Jin, P., Karniadakis, G.E.: Deeponet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators. arXiv preprint arXiv:1910.03193 (2019) 18 [33]Li, Z., Kovachki, N., Azizzadenesheli, K., Liu, B., Bhattacharya, K., Stuart, A., Anandkumar, A.: Fourier neural operator for parametric partial differential equations. arXiv preprint arXiv:2010.08895 (2020) [34]Li, Z., Huang, D.Z., Liu, B., Anandkumar, A.: Fourier neural operator with learned deformations for pdes on general geometries. Journal of Machine Learning Research24(388), 1ā26 (2023) [35]Raonic, B., Molinaro, R., De Ryck, T., Rohner, T., Bartolucci, F., Alaifari, R., Mishra, S., BĆ©zenac, E.: Convolutional neural operators for robust and accurate learning of pdes. Advances in Neural Information Processing Systems36, 77187ā 77200 (2023) [36]Liu, C., Murari, D., Liu, L., Li, Y., Budd, C., Schƶnlieb, C.-B.: Enhancing fourier neural operators with local spatial features. arXiv preprint arXiv:2503.17797 (2025) [37]Sousa Almeida, J.L., Rocha, P.R.B., Carvalho, A.M., Nogueira Jr, A.C.: A coupled variational encoder-decoder-deeponet surrogate model for the rayleigh- bĆ©nard convection problem. In: When Machine Learning Meets Dynamical Systems: Theory and Applications (2023) [38]MichaÅowska, K., Goswami, S., Karniadakis, G.E., Riemer-SĆørensen, S.: Don- lstm: Multi-resolution learning with deeponets and long short-term memory neural networks. arXiv preprint arXiv:2310.02491 (2023) [39]Haghighat, E., Waheed, U., Karniadakis, G.: En-deeponet: An enrichment approach for enhancing the expressivity of neural operators with applications to seismology. Computer Methods in Applied Mechanics and Engineering420, 116681 (2024) [40]Moya, C., Lin, G.: Fed-deeponet: Stochastic gradient-based federated training of deep operator networks. Algorithms15(9), 325 (2022) [41]Kontolati, K., Goswami, S., Em Karniadakis, G., Shields, M.D.: Learning non- linear operators in latent spaces for real-time predictions of complex dynamics in physical systems. Nature Communications15(1), 5101 (2024) [42]Garg, S., Chakraborty, S.: Variational bayes deep operator network: a data- driven bayesian solver for parametric differential equations. arXiv preprint arXiv:2206.05655 (2022) [43]Wen, G., Li, Z., Azizzadenesheli, K., Anandkumar, A., Benson, S.M.: U-fnoāan enhanced fourier neural operator-based deep-learning model for multiphase flow. Advances in Water Resources163, 104180 (2022) 19 [44]Zhu, M., Feng, S., Lin, Y., Lu, L.: Fourier-deeponet: Fourier-enhanced deep opera- tor networks for full waveform inversion with improved accuracy, generalizability, and robustness (2023)https://doi.org/10.2139/ssrn.4461079 [45]Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical Image Computing and Computer-assisted Intervention, p. 234ā241 (2015). Springer [46]Geneva, N.D., Zabaras, N.: Transformers for modeling physical systems. Neural networks146, 272ā289 (2022) [47]Solera-Rico, A., Sanmiguel Vila, C., Gómez-López, M., Wang, Y., Almashjary, A., Dawson, S.T., Vinuesa, R.:β-variational autoencoders and transformers for reduced-order modelling of fluid flows. Nature Communications15(1), 1361 (2024) [48]Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020) [49]Wang, S., Seidman, J.H., Sankaran, S., Wang, H., Pappas, G.J., Perdikaris, P.: Cvit: Continuous vision transformer for operator learning. arXiv preprint arXiv:2405.13998 (2024) [50]Alasker, M., Liu, Z., Ghazvini, F., Barros, F., Jha, B.: Scalar transport during miscible viscous fingering experiments in consolidated porous media. Physics of Fluids37(11) (2025) [51]Maynez, J., Narayan, S., Bohnet, B., McDonald, R.: On faithfulness and factu- ality in abstractive summarization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), p. 1906ā1919 (2020). Association for Computational Linguistics.https://aclanthology.org/2020.acl- main.173/ [52]Maleki, N., Padmanabhan, B., Dutta, K.: Ai hallucinations: a misnomer worth clarifying. In: 2024 IEEE Conference on Artificial Intelligence (CAI), p. 133ā138 (2024). IEEE [53]Dillion, D., Tandon, N., Gu, Y., Gray, K.: Can ai language models replace human participants? Trends in Cognitive Sciences27(7), 597ā600 (2023) [54]Salah, M., Al Halbusi, H., Abdelfattah, F.: May the force of text data analysis be with you: Unleashing the power of generative ai for social psychology research. Computers in Human Behavior: Artificial Humans1(2), 100006 (2023) [55]Sheng, H., Wu, X., Si, X., Li, J., Zhang, S., Duan, X.: Seismic Foundation Model 20 (SFM): a new generation deep learning model in geophysics. arXiv preprint arXiv:2309.02791 (2023) 21