Paper deep dive
AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging
Eunji Ko, Patrick Ross, Corey Hart, Wolfgang Losert
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during reconstruction. Nevertheless, this study explores two cases in which optical scattering may serve a beneficial role in image reconstruction tasks. We compared the No Scattering MNIST dataset with three Scattering MNIST datasets, each generated under distinct scattering conditions. To assess the information content of the resulting speckle patterns, we employed a Variational Autoencoder (VAE) approach which achieves accuracy comparable to state-of-the-art deep learning approaches, but has an interpretable latent space. We find that scattering can enhance data robustness against spatial pixel loss by effectively distributing information. We also demonstrate that scattering can enable distinctions of focal depth information. We anticipate that these findings will contribute to more efficient imaging techniques, particularly in the presence of obstacles and three-dimensional signals.
Tags
Links
- Source: https://arxiv.org/abs/2607.22867v1
- Canonical: https://arxiv.org/abs/2607.22867v1
Trouble viewing inline? Open PDF directly β
Full Text
26,387 characters extracted from source content.
Expand or collapse full text
*Contact author: wlosert@umd.edu AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging Eunji Ko 1 , Patrick Ross 2 , Corey Hart 2 , Wolfgang Losert 1,* 1 University of Maryland β College Park 2 Lockheed Martin ABSTRACT. Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during reconstruction. Nevertheless, this study explores two cases in which optical scattering may serve a beneficial role in image reconstruction tasks. We compared the No Scattering MNIST dataset with three Scattering MNIST datasets, each generated under distinct scattering conditions. To assess the information content of the resulting speckle patterns, we employed a Variational Autoencoder (VAE) approach which achieves accuracy comparable to state-of-the-art deep learning approaches, but has an interpretable latent space. We find that scattering can enhance data robustness against spatial pixel loss by effectively distributing information. We also demonstrate that scattering can enable distinctions of focal depth information. We anticipate that these findings will contribute to more efficient imaging techniques, particularly in the presence of obstacles and three-dimensional signals. I. INTRODUCTION. Imaging through scattering media is a challenging problem across a broad range of applications, from fluorescence microscopy in turbid biological tissue to depth-resolved sensing in aerosols and colloidal suspensions. In these settings, multiple scattering events degrade spatial coherence and the signal captured by detectors. In many previous studies, scattering has therefore been treated as noise to be corrected, either by measuring the transmission matrix of the scattering process to characterize the full inputβ output field relationship [1, 2], or by training deep neural networks to invert the scattering transformation with U-Net or transformer-based architectures [3β6]. These deep learning models, trained on paired speckleβobject datasets, can learn effective end-to-end mappings without explicit physical modeling of the scattering process. While these approaches reconstruct the original information effectively, they offer no mechanism for scattering to assist reconstruction or classification tasks. Variational Autoencoders (VAEs) [7, 8] offer a generative alternative that addresses this limitation. By mapping observations to a structured probabilistic latent space, VAEs learn compact representations that capture the underlying generative factors of the data rather than merely fitting their parameters to a particular dataset. Unlike discriminative models, the VAE latent space affords a degree of interpretability: individual latent dimensions can be traversed and inspected to reveal which physical factors of variation they encode [9, 10], making VAEs particularly well- suited to problems where understanding the generative structure of the data is as important as predictive accuracy. Two fundamental questions arise when applying such models to speckle imaging. The first is practical: how much of the detector plane is needed? Scattering redistributes light across the plane, but it is not known whether object-relevant information is concentrated near the optical axis or spread across the entire speckle field and how this depends on scattering strength. The second question concerns whether the scattering process itself encodes physically meaningful information beyond the object. Specifically, whether the focal depth of the object within the scattering volume leaves a decodable signature in the speckle pattern. Prior work has established that depth information survives scattering and can be recovered from a single speckle pattern [11, 12], suggesting that the speckle field simultaneously carries both object content and depth-related information. However, no existing framework has attempted to jointly decode and explicitly separate these two contributions within a single learned latent representation. *Contact author: wlosert@umd.edu In this work, we address both questions with VAE frameworks. We first apply circular occlusion masks of increasing radius to the speckle input and track reconstruction performance as a function of mask radius. Performance degrades in different ways for No -Scattering MNIST and Scattering MNIST datasets with increasing mask size, indicating that object information is encoded across the speckle field when scattered. In the second part, we present a split-latent VAE architecture, inspired by Domain-invariant VAE [13], that partitions the latent space into two disjoint subspaces: a βdigitβ subspace capturing structural features, and a βfocalβ subspace encoding depth- discriminative information reinforced by a jointly trained classifier head. The model reconstructs clean and scattered image estimates and the focal depth information from the latent space without requiring separate networks for reconstruction and depth estimation. I. METHODS Figure 1. Optical scattering imaging setup. A linearly polarized, telescope-expanded laser beam illuminates an SLM, and the modulated output is imaged through a scattering medium onto a CMOS camera via a 4f relay with zero-order blocking at the Fourier plane between lens 3 and lens 4 A. Experimental setup Speckle patterns were generated by passing spatially modulated images, formed using a phase-only SLM (Holoeye LCR-2500, 1024 Γ 768 resolution, 19 ΞΌm pixel pitch), through a custom scattering medium, as shown in Fig 1. The input beam was provided by a collimated 473 nm laser (OBIS LX, 200 mW), expanded 20Γ to match the SLM aperture. Polarization was optimized using a polarizer and a half-wave plate to ensure efficient SLM modulation. The phase patterns were relayed through a 4f system. Unmodulated light was filtered at the Fourier plane, and the resulting speckle patterns were imaged onto a CMOS sensor (PointGrey, 1600 Γ 1200, 1/1.8"). A zero-order blocker is used in this setup with SLM to block the non-diffracted beam that passes straight through without modulation. This zero-order light often appears as a bright central spot and carries no useful information, therefore, it can obscure the modulated or diffracted patterns we want to analyze. The scattering medium was fabricated by pouring Polydimethylsiloxane (PDMS) into a 3D-printed mold. PDMS is commonly used as a scattering medium because it is optically transparent, biocompatible, and easy to mold or tune. We created three scattering media (low, moderate, high scattering) with varying scattering intensities by incorporating three distinct concentrations of zinc-oxide (ZnO) microparticles into PDMS. When the spatially structured light from the SLM passed through this medium, it generated complex speckle patterns that were employed for further analysis. This study examines four distinct datasets. The first dataset, referred to as the 'No Scattering' MNIST dataset, was created by transmitting MNIST digit pattern generated by a spatial light modulator (SLM) directly to a camera, without the presence of a scattering medium. And there are three datasets labeled as 'Low Scattering', 'Moderate Scattering', and 'High Scattering' MNIST datasets. These were produced by transmitting the MNIST digit pattern generated by the SLM through a medium with different scattering strength. B. Computational methods Figure 2. Illustration of variational autoencoder (VAE) *Contact author: wlosert@umd.edu 1. Variational autoencoder We adopted a variational autoencoder (VAE) framework to reconstruct original information from the input images to target images. VAE is one of generative models that learns to encode input data into a probabilistic latent space and decode from the space back into output images. a. Standard VAE As shown in Fig 2, given an observed image ν±β β !Γ# , the VAE learns an approximate posterior distribution ν $ ( ν³β£ν± ) over latent variables ν³, parameterized by an encoder network with parameters ν. A decoder network with parameters ν then models the likelihood ν % (ν±β£ν³). Training is performed by maximizing the Evidence Lower Bound (ELBO): β ELBO =νΌ & ! (ν³β£ν±) [ logν % (ν±β£ν³) ] βν· KL 7ν $ (ν³β£ν±) β₯ ν(ν³)9 where the first term is the reconstruction likelihood and the second term is the KullbackβLeibler (KL) divergence from the approximate posterior to the standard normal prior ν(ν³)=ν©(ν,ν). The encoder outputs the parameters of a diagonal Gaussian posterior: ν $ (ν³β£ν±)=ν©7ν³;ν $ (ν±), diag(ν $ , (ν±))9 Latent samples are drawn via the reparameterization trick to enable gradient backpropagation: ν³=ν $ ( ν± ) +ν $ ( ν± ) βν,νβΌν©(ν,ν) The KL divergence term for a diagonal Gaussian has a closed-form expression: ν· KL 7ν $ (ν³β£ν±) β₯ ν(ν³)9 =β 1 2 G71+logν - , βν - , βν - , 9 . -/0 where ν is the dimensionality of the latent space. b. Split-Latent VAE To disentangle the scattering-invariant image content from scattering-dependent focal depth information, we divide the latent space into two disjoint subspaces: ν³=[ν³ digit , ν³ focal ] where ν³ digit ββ . " captures digit features common across focal depths, and ν³ focal ββ . # encodes focal depth information that is blended with scattering information. The encoder produces a single concatenated posterior parameter vector (ν,logν , )β β . " 1. # , from which the subspace-specific parameters are obtained by index partitioning: ν 2 =ν [:. " ] ,ν 6 =ν [. " :] and analogously for logν 2 , and logν 6 , . Latent samples are drawn via the reparameterization trick and split accordingly into ν³ digit and ν³ focal . The decoder is split into two separate paths with distinct inputs. The clean image reconstruction ν±K clean is decoded exclusively from ν³ digit , enforcing that the digit subspace alone must carry sufficient information to recover the digit content. The scattered image reconstruction ν±K scat is decoded from the full latent code ν³=[ν³ digit ,ν³ focal ], allowing the focal subspace ν³ focal to encode also the scattering information needed to reproduce the observed speckle pattern, not only the coal depth information. The full training loss consists of reconstruction, KL regularization, and focal classification terms: β=β recon +ν½β β KL +νβ β focal The reconstruction loss is computed as the mean squared error between the model outputs and their respective targets: β recon =β₯ν± clean βν±K clean β₯ , , +β₯ν± scat βν±K scat β₯ , , The KL loss is summed over both latent subspaces: β KL =ν· KL 7ν $ (ν³ digit β£ν±) β₯ ν(ν³)9 +ν· KL 7ν $ (ν³ focal β£ν±) β₯ ν(ν³)9 *Contact author: wlosert@umd.edu The focal depth classifier is a lightweight network ν 7 that takes ν³ focal as input and predicts a probability distribution over νΆ discrete focal depth classes. The classification loss is the standard cross-entropy: β focal =βGν¦ 8 9 8/0 logν Μ 8 ,ν Μ 8 =softmax7ν 7 (ν³ focal )9 8 where ν¦ 8 β0,1is the one-hot ground-truth focal depth label and ν Μ 8 is the predicted probability for class ν. Overall, the hyperparameter ν½controls the regularization strength of the KL term, and ν weights the classification objective relative to reconstruction. 2. Computation details To assess the semantic fidelity of the reconstructed images, we used a Stochastic Optimization of Plain Convolutional Neural Networks (SOPCNN) to classify them [14]. The SOPCNN was trained on original MNIST datasets and then tested on reconstructed outputs. This classification accuracy was used to quantify the preservation of class- discriminative features. To mitigate dataset sampling bias, each experimental condition was run across five independently drawn random seeds, with all reported metrics reflecting test- set performance exclusively. All models were implemented using PyTorch and trained on a workstation equipped with an NVIDIA GeForce RTX 4090 GPU with CUDA 12.3 support. GPU monitoring and performance management were conducted via NVIDIA-SMI. I. RESULTS & DISCUSSION Figure 3. Visual examples of VAE reconstruction results with (a) the No-Scattering MNIST dataset and (b) the Moderate Scattering dataset. The first column shows the input, the second column displays the output, and the last column presents the ground truth (target). Each row represents the cases where a different radius of mask applied to the input images. *Contact author: wlosert@umd.edu Fig 4. The reconstruction outcomes of center-occluded MNIST images using VAE, assessed by the accuracy of SOPCNN image classification, over increasing radius. No Scattering data was fitted with Sigmoid function, while Low/Moderate/High Scattering datasets were fitted with a power law function. All the data was fitted on average accuracies over N = 5 independent training seeds Table I. Fitted power law parameter νΌ for Scattering datasets. νΌβ1reflects increasingly uniform speckle statistics. νΌ Low Scattering 0.692 Β± 0.030 Moderate Scattering 0.694 Β± 0.022 High Scattering 0.757 Β± 0.047 A. Circular occlusion masking In the initial segment of our results, we sought to determine whether the peripheral regions of images retain the original information when the central portion is obscured by a circular mask. Specifically, we applied a black circular mask, composed of zero- pixel intensities, to the center of the images. To assess the impact of increasing mask radius on image reconstruction, we varied the radius from 20 to 192 pixels, which is half the length of 384x384 images. Example of input, output, and target images for three different mask radii (60, 120, and 192) are presented in Fig. 3 (a)-(b). Notably, the reconstructed output images gradually lose detail and completely fail to reconstruct at a radius of 192 in the No Scattering MNIST dataset. Conversely, the output in the Moderate Scattering MNIST dataset retains its details as the mask radius increases, even when the images are nearly covered with a r=192 px. mask. The ability of the Scattered datasets to reconstruct the original images despite increasing radius is illustrated in greater detail in Fig. 4. In SOPCNN image classification task, the No Scattering MNIST dataset deteriorates rapidly as the mask radius increases, falling to approximately 10% in SOPCNN accuracy, indicative of a random guess. In contrast, the decline in the Scattering datasets is much less pronounced. To investigate how scattering redistributes spatial information analytically, we fitted each dataset to analytical functions that can describe each datasetβs behavior. For No Scattering MNIST, accuracy seems to follow a sigmoid decay, as shown in Fig 4: ν΄(ν)=ν΄ :;< + ν΄ :=> βν΄ :;< 1+ [ ν ν 8 \ ? where ν 8 is the radius at which accuracy falls to the midpoint between its maximum and minimum values, and νis the steepness coefficient describing how sharply accuracy collapses near ν 8 . The large (fitted) steepness coefficient ν=6.69 indicates that accuracy remains nearly unchanged for mask radii well below (fitted) ν 8 , then collapses rapidly over a narrow radius window centered at ν 8 =141.3px, which turns out to be near the radius the mask reaches the boundary of digit stroke region. For all Scattering datasets, a power law effectively characterizes their accuracy trends: ν΄(ν)=ν΄ :;< +(ν΄ :=> βν΄ :;< )c1β ν , ν , e @ where 1βν , /ν , is the unmasked area fraction and νΌgoverns how nonlinearly accuracy responds to visible area. Also, R=192 px. is the maximum mask radius applied, corresponding to the inscribed circle of the 384Γ384 size of image. When νΌ=1, accuracy degrades in direct proportion to masked area, which means spatially uniform information distribution. And νΌ increases monotonically with scattering density (0.692β0.757), approaching the uniform limit νΌ= 1 at the heaviest scattering condition, which is consistent with heavier scattering producing more spatially homogeneous speckle statistics. *Contact author: wlosert@umd.edu Fig 5. (a) The illustration of generating three different focal depth datasets (b) the schematic of the split- latent VAE we developed for the focal depth discrimination task Fig 6. The outcomes of the split-latent VAE (a) SOPCNN accuracy for reconstructed clean MNIST images under different scattering conditions, and (b) focal classification accuracy on ν³ focal ββ A across different scattering scenarios. Each datapoints is from five different independent training seeds. B. Focal depth discrimination In this section of our study, we explore whether optical scattering could effectively distinguish between the focal depths of three-dimensional scattering medium. Here, we projected the same MNIST pattern onto each of three focal depths of the scattering medium and collected speckle patterns for each depth separately as demonstrated in Fig 5(a). We then simultaneously input the three datasets into the encoder of split-latent VAE, as depicted in Fig 5(b). We designed the structure of this model such that the initial latent vector, ν³ digit is responsible solely for encoding MNIST digit data, while the second latent vector, ν³ focal is dedicated to encoding focal depth information. This is achieved by using only ν³ digit to decode into clear MNIST images and employing both ν³ digit and ν³ focal to decode into dispersed MNIST images. We deliberately chose not to include a third latent space for encoding scattering information, as we aimed to integrate this information specifically into ν³ focal to leverage the scattering effect. And ν³ focal is reinforced by a jointly trained classifier head. The SOPCNN accuracy of reconstructed images and focal classification accuracy on ν³ focal latent vector is shown in Fig 6(a) and (b), respectively. When scattering was low to moderate, the accuracy ranged from 96 to 97%, but with high scattering, the SOPCNN accuracy dropped to about 85%. Interestingly, even though high scattering had the lowest SOPCNN accuracy, it resulted in the highest accuracy for focal classification ~98%. Additionally, the accuracy of focal classification increased gradually as the scattering condition became more intense. Table I. Silhouette score (S) of ν³ focal across scattering densities, averaged over N = 5 independent training seeds (mean Β± std). S Low Scattering 0.4237 Β± 0.0230 Moderate Scattering 0.4420 Β± 0.0140 High Scattering 0.5825 Β± 0.0288 To assess focal depth discriminability in the latent space, we computed the silhouette score on ν³ focal ββ A . For each scattering condition and each of the five training seeds, 1,000 samples per focal position were encoded using the posterior mean ν, yielding 3,000 points in β A with known focal labels ν B , ν 0 , ν , . For each point ν³ C in cluster νΆ D , we compute ν ( ν³ C ) , the mean Euclidean distance to all other points in the same *Contact author: wlosert@umd.edu cluster, and ν(ν³ C ), the mean Euclidean distance to all points in the nearest other cluster: ν ( ν³ C ) = 1 β£νΆ D β£β1 Gβ₯ ν³ $ β9 % ,GHC ν³ C βν³ G β₯ , , ν(ν³ C )=min -HD 1 β£νΆ - β£ Gβ₯ ν³ $ β9 & ν³ C βν³ G β₯ , The silhouette coefficient is ν (ν³ C )=(νβν)/ max(ν,ν), and the silhouette score is the mean of ν across all points, ranging from β1 (point closer to another cluster than its own) to +1 (point well within its own cluster). As shown in Table I, the silhouette score remains lower in Low and Moderate Scattering then increases substantially in High Scattering, confirming that heavier scattering produces better- separated focal depth clusters in the latent space. The consistency of this effect across all five seeds indicates that this improvement reflects a property of the physical scattering regime rather than artifact from a particular training run. Table I. Pairwise inter-cluster distances between focal-depth centroids in the focal latent subspace ( ν³ focal ) across scattering densities, averaged over N = 5 independent training seeds (mean Β± std). Inter- cluster distance is the Euclidean distance between cluster centroids in ββΆ space. Inter-cluster distance fββfβ fββfβ fββfβ Low Scattering 2.44Β±0.15 4.43Β±0.35 2.30Β±0.26 Moderate Scattering 2.61Β±0.19 4.25Β±0.22 2.52Β±0.15 High Scattering 12.41Β±2.95 19.94Β±4.71 7.61Β±1.75 To further characterize the geometric structure underlying this improvement, we computed pairwise inter-cluster distances between the three focal depthsβ cluster centroids in ν³ focal . The centroid of each cluster νΆ D is the mean of all encoded samples belonging to that focal position, and the inter-cluster distance between two clusters is the Euclidean distance between their centroids: ν(νΆ D ,νΆ - )=β₯ν³ Μ ( D ) βν³ Μ ( - ) β₯ , , ν³ Μ ( D ) = 1 β£νΆ D β£ Gν³ C ν³ ' β9 % As shown in Table I, inter-cluster distances remain lower in Low and Moderate Scattering, then increase with huge steps in High Scattering, as expected in S results in Table I. Across all scattering conditions, ν B βν , consistently yields the largest inter-cluster distance, reflecting the physical expectation that the two extreme focal planes produce the most distinct speckle statistics. In Low and Moderate Scattering, the adjacent pairs ν B βν 0 and ν 0 βν , are nearly equal, indicating a geometrically symmetric representation. In High Scattering, this symmetry breaks and ν B β ν 0 (12.41) becomes notably larger than ν 0 βν , (7.61), suggesting that heavy scattering produces greater speckle divergence between ν B and ν 0 focal planes than between ν 0 and ν , planes, which the encoder captures as an asymmetric reorganization of ν³ focal . V. CONCLUSIONS In this work, we investigated the impact of circular occlusion masking with varying radius on reconstruction quality. We found that the original information still can be found in the peripheral region of images in Scattering datasets compared to No- Scattering dataset that failed drastically as the radius of mask increases. Additionally, we applied curve fitting (Sigmoid for No Scattering dataset and Power law for Scattering datasets), and the resulting parameters effectively illustrated how the Scattering dataset contains more evenly distributed information across the images compared to the No Scattering dataset, with heavier scattering leading to a more uniform distribution. We also examined whether optical scattering could enable identifying the focal depth of information by projecting the same MNIST pattern onto each of three focal depths of the scattering medium and engineering VAE architecture into split-latent VAE. We discovered an improvement in focal depth classification accuracy as the scattering intensity increased even when high scattering had the lowest SOPCNN image classification accuracy. Furthermore, we calculated silhouette score (S) and pairwise inter- cluster distances between the three focal latent vector centroids to assess how scattering intensity influences the learned focal latent space. Consequently, we *Contact author: wlosert@umd.edu observed that heavy scattering not only actively enhances focal depth discriminability in the learned latent space but also captures and amplifies the asymmetry between focal planes. The focal depth discrimination capability is consistent with the 3D nature of a scattering medium and reveals two competing optimization choices: More scattering elements along a 3D path enable better depth discrimination, but eventually lead to a drop-off in the transmitted information. Overall, this study presents a proof-of-concept that scattering distributes information uniformly across the imaging sensor, and enables AI-assisted focal depth discrimination of optical signals. Our work raises the question whether optical scattering could be used in conjunction with AI-interpretation to develop robust and 3D sensing capabilities in a broad range of science and engineering domains. ACKNOWLEDGMENTS This work was supported by Lockheed Martin University Engagement Funds. We also thank Aiden Heggs and Prof. John Fourkas from University of Maryland, College Park for their help in preparing the scattering media. REFERENCES [1] Vellekoop, I. M., & Mosk, A. P. (2007). Focusing coherent light through opaque strongly scattering media. Optics letters, 32(16), 2309-2311. [2] Popoff, S. M., Lerosey, G., Carminati, R., Fink, M., Boccara, A. C., & Gigan, S. (2010). Measuring the Transmission Matrix in Optics: An Approach to the Study and Control of Light Propagation in Disordered Media. Physical review letters, 104(10), 100601. [3] Li, Y., Xue, Y., & Tian, L. (2018). Deep speckle correlation: a deep learning approach toward scalable imaging through scattering media. Optica, 5(10), 1181-1190. [4] Zhu, S., Guo, E., Gu, J., Bai, L., & Han, J. (2021). Imaging through unknown scattering media based on physics-informed learning. Photonics Research, 9(5), B210-B219. [5] Chen, Y. C., Mi, S. X., Tian, Y. P., Hu, X. B., Yuan, Q. Y., Chew, K. H., & Chen, R. P. (2025). Adaptive vectorial restoration from dynamic speckle patterns through biological scattering media based on deep learning. Sensors, 25(6), 1803. [6] Xia, W., Li, X., He, G., Luo, Z., Wu, X., & Huang, B. (2025). Imaging through highly scattering media via global transformer mapping. Journal of Optics, 27(4), 045603. [7] Kingma, D. P., & Welling, M. (2013). Auto- encoding variational bayes. arXiv preprint arXiv:1312.6114. [8] Rezende, D. J., Mohamed, S., & Wierstra, D. (2014, June). Stochastic backpropagation and approximate inference in deep generative models. In International conference on machine learning (p. 1278-1286). PMLR. [9] Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., & Lerchner, A. (2018). Understanding disentangling in Ξ²-VAE. arXiv preprint arXiv:1804.03599. [10] Valleti, M., Ziatdinov, M., Liu, Y., & Kalinin, S. V. (2024). Physics and chemistry from parsimonious representations: image analysis via invariant variational autoencoders. npj Computational Materials, 10(1), 183. [11] Zhu, S., Guo, E., Cui, Q., Bai, L., Han, J., & Zheng, D. (2020). Locating and imaging through scattering medium in a large depth. Sensors, 21(1), 90. [12] Fan, W., Chen, T., Xu, X., Chen, Z., Hu, H., Zhang, D., ... & Zhu, S. Y. (2021). Recognizing three- dimensional phase images with deep learning. arXiv preprint arXiv:2107.10584. [13] Ilse, M., Tomczak, J. M., Louizos, C., & Welling, M. (2020, September). Diva: Domain invariant variational autoencoders. In Medical imaging with deep learning (p. 322-348). PMLR. [14] Assiri, Y. (2020). Stochastic optimization of plain convolutional neural networks with simple methods. arXiv preprint arXiv:2001.08856.