Paper deep dive
Cryo-SWAN: the Multi-Scale Wavelet-decomposition-inspired Autoencoder Network for molecular density representation of molecular volumes
Rui Li, Artsemi Yushkevich, Mikhail Kudryashev, Artur Yakimovich
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/21/2026, 1:18:48 AM
Summary
The paper introduces Cryo-SWAN, a multi-scale wavelet-decomposition-inspired variational autoencoder designed for learning robust 3D shape representations from voxelized molecular density volumes, particularly those from cryo-EM. The model employs conditional coarse-to-fine latent encoding and recursive residual quantization to capture both global geometry and high-frequency structural details. It is evaluated on standard benchmarks (ModelNet40, BuildingNet) and a newly curated dataset, ProteinNet3D, demonstrating superior reconstruction quality compared to state-of-the-art 3D autoencoders. The study also highlights downstream applications in denoising and conditional shape generation using diffusion models.
Entities (9)
Relation Signals (11)
Cryo-SWAN → evaluatedon → ProteinNet3D
confidence 95% · Evaluated on ModelNet40, BuildingNet, and a newly curated dataset of cryo-EM volumes, ProteinNet3D
Cryo-SWAN → isbasedon → Variational Autoencoder
confidence 95% · We present Cryo-SWAN, a voxel-based variational autoencoder inspired by multi-scale wavelet decomposition.
Cryo-SWAN → usestechnique → Recursive Residual Quantization
confidence 95% · The model performs conditional coarse-to-fine latent encoding and recursive residual quantization across perception scales
ProteinNet3D → derivedfrom → EMDB
confidence 90% · curated a 3D volumes dataset of diverse macromolecules obtained from the publicly available Electron Microscopy Data Bank (EMDB)
Cryo-SWAN → evaluatedon → BuildingNet
confidence 90% · Evaluated on ModelNet40, BuildingNet, and a newly curated dataset
Cryo-SWAN → evaluatedon → ModelNet40
confidence 90% · Evaluated on ModelNet40, BuildingNet, and a newly curated dataset
Cryo-SWAN → enables → Diffusion Models
confidence 85% · integration with diffusion models enables denoising and conditional shape generation
Cryo-SWAN → outperforms → HQ-VAE
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Learning robust representations of 3D shapes from voxelized data is essential for advancing AI methods in biomedical imaging. However, most contemporary 3D computer vision approaches operate on point clouds, meshes, or octrees, while volumetric density maps, the native format of structural biology and cryo-EM, remain comparatively underexplored. We present Cryo-SWAN, a voxel-based variational autoencoder inspired by multi-scale wavelet decomposition. The model performs conditional coarse-to-fine latent encoding and recursive residual quantization across perception scales, enabling accurate capture of both global geometry and high-frequency structural detail in molecular density volumes. Evaluated on ModelNet40, BuildingNet, and a newly curated dataset of cryo-EM volumes, ProteinNet3D, Cryo-SWAN consistently improves reconstruction quality over state-of-the-art 3D autoencoders. We demonstrate that the molecular densities organize in learned latent space according to shared geometric features, while integration with diffusion models enables denoising and conditional shape generation. Together, Cryo-SWAN is a practical framework for data-driven structural biology and volumetric imaging.
Tags
Links
- Source: https://arxiv.org/abs/2603.03342v1
- Canonical: https://arxiv.org/abs/2603.03342v1
Trouble viewing inline? Open PDF directly →
Full Text
54,321 characters extracted from source content.
Expand or collapse full text
Cryo-SWAN: the Multi-Scale Wavelet-decomposition-inspired Autoencoder Network for molecular density representation of molecular volumes Rui Li 1,2 , Artsemi Yushkevich 3,4 , Mikhail Kudryashev 3,5 , Artur Yakimovich 1,2,6,7 1 Center for Advanced Systems Understanding (CASUS), Görlitz, Germany. 2 Helmholtz-Zentrum Dresden-Rossendorf e. V. (HZDR), Dresden, Germany. 3 In situ Structural Biology, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany. 4 Department of Physics, Humboldt University of Berlin, Berlin, Germany. 5 Institute of Medical Physics and Biophysics, Charite-Universitätsmedizin, Berlin, Germany. 6 Institute of Computer Science, University of Wrocaw, Wrocaw, Poland. 7 Cluster of Excellence Physics of Life, TU Dresden, Dresden, Germany. Contributing authors: mikhail.kudryashev@mdc-berlin.de;a.yakimovich@hzdr.de; Abstract Learning robust representations of 3D shapes from voxelized data is essential for advancing AI methods in biomedical imaging. However, most contemporary 3D computer vision approaches operate on point clouds, meshes, or octrees, while volumetric density maps, the native format of structural biology and cryo-EM, remain comparatively underexplored. We present Cryo-SWAN, a voxel-based variational autoencoder inspired by multi-scale wavelet decomposition. The model performs conditional coarse-to-fine latent encoding and recursive residual quantization across per- ception scales, enabling accurate capture of both global geometry and high-frequency structural detail in molecular density volumes. Evaluated on ModelNet40, BuildingNet, and a newly curated dataset of cryo-EM volumes, ProteinNet3D, Cryo-SWAN consistently improves reconstruction quality over state-of-the-art 3D autoencoders. We demonstrate that the molecular densities organize in learned latent space according to shared geometric features, while integration with diffusion models enables denoising and conditional shape generation. Together, Cryo-SWAN is a practical framework for data-driven structural biology and volumetric imaging. Keywords:Cryo-EM, Deep learning, Molecular shape representation Introduction The shape of 3D objects is often directly connected to their function. Learning a meaningful repre- sentation of objects’ shapes from noisy input data is paramount to accomplishing advanced machine learning tasks downstream. Such tasks could facilitate advancements in scientific fields, including cellular and structural biology, aiming to connect the shape of organelles and molecules to their biolog- ical functions. Through this, structural biology facilitates the understanding of important biological mechanisms and ultimately therapies. In biomedical imaging, 3D shapes are typically stored in voxel space. However, in 3D computer vision, most methods operate on processed geometric representations such as octrees [1] or VDB 1 arXiv:2603.03342v1 [eess.IV] 18 Feb 2026 grids [2]. Approaches such as Craftsman [3], XCube [4], and Dora-VAE [5] exploit point clouds, sparse voxels, or meshes. In parallel to the biomedical imaging, the task of 3D shape generation [3–9] has gained increasing attention due to rapid progress in AR/VR, animation, and autonomous driving. This led to the voxel-based representation learning being largely underexplored. Meaningful representations are also crucial for the tasks involving modern generative models guided by learned priors [10–15]. Once a desired shape can be generated it opens versatile opportunities in drug design andin silicoscreening. Most state-of-the-art (SOTA) 3D generative frameworks adopt a two-stage paradigm. First, they learn expressive latent representations using an autoencoder (AE) [16]. Second, they train generative models, such as diffusion models [17], generative adversarial networks (GANs) [18], or transformers [19] on the learned latent space. A typical AE [16,20] consists of an encoder, a latent layer, and a decoder. The encoder maps a high-dimensional inputx i ∈ Xto a latent vectorz i ∈ Zin a process known as representation learning [21], and a decoder reconstructs it asx ′ i ∈X ′ . The quality of the learned representation determines downstream generative performance, especially for high- dimensional and geometry-rich data such as 3D shapes. Most AE-based approaches in structural biology rely on either 2D projections or atomic models. Yet, cryo-EM data, stored as 3D density volumes, rather than atomic structures, leaves representation learning on 3D cryo-EM densities underexplored. The AE-based reconstruction process is governed by an encoding distributionq(Z | X)and a decoding distributionp(X ′ | Z). The conventional AEs optimize only the reconstruction loss |x ′ i −x i | 2 and do not constrain the latent structure. Variational autoencoders (VAEs), on the other hand, introduced a Gaussian prior over the latent space [ 16], enabling sampling and regularization via the KL-divergence [22]. The latent variables are constrained as (z i ∼ N(0,1)), and training includes a regularization term (βD KL (q φ (z|x)||p(z))). Although VAEs improve latent organization and training stability, they often produce blurry reconstructions due to overly coarse embeddings [23,24]. As a result, recent research has focused on improving latent expressiveness. VQ-VAE [25] replaced continuous latent variables with a discrete codebooke∈R K×D . Given an encoder output z e (x), discretization is performed via nearest-neighbor lookup: q(z=k|x) = ( 1ifk= arg min j |z e (x)−e j | 2 , 0otherwise. Training includes a reconstruction loss and a commitment loss, |sg[z e (x)]−e| 2 2 +β|z e (x)−sg[e]| 2 2 , where sg[·]denotes the stop-gradient operator. The authors of the RQ-VAE introduced recursive residual quantization [ 11]. Starting with the encoder outputr (0) =z e , the remaining residual errors are quantized acrossLrecursive levels using codebooksC (l) =e (l) k . At each residual level, c (l) = arg min k |r (l−1) −e (l) k | 2 , r (l) =r (l−1) −e (l) c (l) , and the final latent approximation isˆz L = P L l=1 e (l) c (l) . HQ-VAE [14] introduced hierarchical codebooks across multiple spatial scales. The encoder fea- ture map is discretized into latent variablesZ 1:L =Z (l) L l=1 , where each spatial scale level has a dedicated codebookC (l) =e (l) k . Commitment losses are computed hierarchically: L code = L X l=1 d l X i=1 sg[ˆz (l) i ]−e (l) k 2 . VAR-VAE [15] adopted a multi-scale tokenization strategy inspired by transformer-based vision models [26–28]. Latent mapsZ (l) are embedded at progressively coarser resolutions and modeled 2 autoregressively: p(Z (1:L) ) = L Y l=1 p Z (l) |Z (<l) . In this work we introduce cryo-SWAN, a multi-Scale quantized Wavelet-inspired [29,30] AutoeN- coder, capable of effectively learning meaningful shape representations from the voxel space. Cryo-SWAN performs multi-level latent quantization combined with recursive residual optimization at each scale, enabling high-fidelity representation learning directly from raw density volumes without geometric or atomic priors. Multi-scale representations are widely used in image restoration and sig- nal approximation [29,31,32]. For example, MPRNet [31] and BCR-Net [29,30] demonstrated that recursive, multi-level decomposition effectively preserves high-frequency information, which is criti- cal for accurate reconstruction. Here, we demonstrate the applicability of multi-scale representations to 3D shapes. We evaluate cryo-SWAN on two standard computer vision benchmarks: ModelNet [ 33] and BuildingNet [34]. Furthermore, we introduce ProteinNet3D, a new dataset comprising over 24k experimental cryo-EM density maps from EMDB [35], manually curated. Cryo-SWAN is compared to the state-of-the-art voxel-based AE models, including HQ-VAE [14], VAR-VAE [15], RQ-VAE [11], and VQ-VAE [25]. We report IoU, F1-score, MSE, and PSNR across all datasets, and addi- tionally evaluate reconstruction resolution on ProteinNet3D using Fourier Shell Correlation (FSC) [36–38]. Cryo-SWAN consistently outperforms competing methods, particularly in capturing high- frequency structural details critical for accurate 3D shape representation. Finally, we demonstrate two downstream applications: geometry-based hub detection using latent molecular embeddings and conditional latent diffusion for high-fidelity molecular volume generation, highlighting cryo-SWANs potential as a general backbone for data-driven structural biology. Methods Benchmark datasets: ModelNet40 and BuildingNet The ModelNet40 dataset includes 40 categories of common 3D objects in.offformat, spanning low- frequency shapes (e.g., chairs) to high-frequency structures (e.g., plants with intricate foliage). The BuildingNet dataset contains.objfiles of architectural structures. Unlike ModelNet, architectures are hollow, structurally detailed, and rich in high-frequency components. Thus, it poses more chal- lenges. To match the volumetric nature of cryo-EM data, we voxelized the.offand.objfiles as densities. ProteinNet3D dataset To evaluate our model, we curated a 3D volumes dataset of diverse macromolecules obtained from the publicly available Electron Microscopy Data Bank (EMDB). EMDB is a comprehensive, annotated repository of volumetric data derived from experimental and computational cryo-EM. It covers a wide range of macromolecules, molecular complexes, and subcellular structures. We retrieved the complete EMDB metadata as of late November 2022 via the EMDB API, comprising approximately 28,700 molecular density volumes. To ensure consistency and enable mean- ingful evaluation, we restricted our selection to entries derived from cryo-EM single-particle analysis (SPA) or cryo-electron tomography (cryo-ET) subtomogram averaging (STA). Focusing on individ- ual macromolecules, we further filtered the dataset by molecular weight, retaining entries within the 100–1500 kDa range. This criterion excluded very small structures (e.g., individual domains) as well as excessively large complexes or subcellular assemblies. To ensure uniform scaling of volumetric fea- tures, we normalized voxel spacing by resampling all volumes to a common isotropic pixel. Serving for the deep learning training, we enriched the data based on reported volume size, voxel spac- ing, and resolution from EMDataResource (http://w.emdataresource.org). Since resampling can affect resolution due to the Nyquist limit, we carefully considered potential resolution degradation introduced by the new sampling rate. Following the filtering procedure, 5,222 qualified EMDB entries were retained. The associated volumetric data were retrieved through the EMDB API using their accession identifiers. We masked the target volumes to suppress background noise and extraneous regions based on the provided contour level annotation on each entry. All volumes were uniformly resampled to an isotropic voxel spacing of 4 Å, allowing for the optimal volume size (Supplementary Fig.S1, top) and the optimal 3 Nyquist limit relative to the entry resolution and pixel size originally reported (Supplementary Fig.S1, bottom). To mitigate aliasing artifacts introduced during resampling, a low-pass filter at the Nyquist frequency was applied. The processed volumes were then adjusted to a fixed spatial resolution of64 3 voxels via cropping or zero-padding. We then normalized all entries as zero mean and unit variance. To improve data diversity, we augmented each volume with five randomly generated 3D rotations. This yields a total of 26,110 samples, while each accompanied by its corresponding metadata. The breadth of the ProteinNet3D dataset is evidenced by its molecular weight distribution (Sup- plementary Fig.S2, top), which ranges from 100 to 1500 kDa (Supplementary Fig.S2, bottom). Additionally, since the dataset originates from experimentally reconstructed cryo-EM volumes, it contains both canonical macromolecular structures (Supplementary Fig.S3, top) and reconstruction- related artifacts (Supplementary Fig.S3, middle). In addition, because numerous proteins operate within membrane-associated environments, membrane densities are frequently retained in the reconstructed volumes (Supplementary Fig.S3, bottom). Model training hyperparameters All datasets were split into training, validation, and test sets with a ratio of [0.8, 0.1, 0.1]. In our experiments, we setL= 10for the positional encoding step. The model is trained using the Adam optimizer with a learning rate of1×10 −3 and a weight decay of1×10 −5 . Perception scales for RQ1 are set to [1, 4, 6, 8, 10, 12, 14, 16], and for RQ2 to [1, 2, 4, 8]. The codebook size at each level is fixed at 4096. Experiments are conducted on two Nvidia A-100 GPUs. Voxel positional encoding Prior work demonstrates that models trained directly on raw pixels or spatial coordinates tend to focus on low-frequency content [39] and ignore the high-frequency structures. Positional encoding (PE) [39,40] contributes to the recovery of high-frequency details in 3D CV tasks. To address this, we adopt a sine/cosine PE strategy to represent the high-frequency information on the voxel. The encoding function is defined asf(v i ), wherev i represents a voxel unit in the density volume, andL denotes the number of encoding levels: f(v i ) = sin(2 0 πv i ),cos(2 0 πv i ),·,sin(2 L−1 πv i ),cos(2 L−1 πv i ) .(1) Multi-scale conditional approximations The key to VAEs’ performance is the quality of latent space approximation via codebook embeddings. Accurately representing the latent featurezthrough a well-structured embedding function is central to addressing this challenge. In the VAE quantization process, letˆzdenote the optimal quantized latent vector corresponding to the encoder outputz. The goal is to learn a mappingz→ˆzthrough an embedding functionf θ with parametersθ. This process is expressed asˆz=f θ ◦z. Inspired by wavelet approximation theory [29,30], we can approximatef θ using an operatorA θ . This reformulates the mapping as: ˆz≈A θ ◦z(2) In wavelet theory, the operatorA θ can be decomposed into a set of multi-level components h A (0) θ ,·,A (l) θ i so thatA θ = P l i=0 A (i) θ . Under standard linear assumptions, the operator at a given levellcan be expressed in the form: A (l) θ =W (l) " D (l) 1 D (l) 2 D (l) 3 A (l−1) # W (l) ⊤ , (3) W (l) denotes the wavelet transformation matrix at levell.D (l) 1 , D (l) 2 , D (l) 3 are decomposition components. For a detailed theoretical foundation, refer to [ 30,41]. This hierarchical formula- tion reveals that the operator at levellis conditioned on the previous levell−1. This recursive dependency is captured asA (l+1) θ =A (l+1) θ | A (l) θ . Another wavelet work [29] demonstrates that incorporating residual structures to fuse conditions enhances signal approximation in nonlinear set- tings. Inspired by this, we denote the residual conditioning operator asM. Updating the recursive 4 formulation we haveA (l+1) θ =M(A l θ ).Mrepresents the residual operation at levell. This allows the approximation operator to capture recursive dependencies across levels, which can be general- ized asA (l+1) θ =M A l θ |A 0 θ , . . . ,A l−1 θ . By incorporating this relationship into Equation2, the approximation process is formulated as: ˆz≈ L X i=0 M A (i) θ |A (0) θ , . . . ,A (i−1) θ ◦z(4) This formulation highlights how recursive conditioning improves approximation quality, offering insight into the improvements of HQ-VAE. Inspired by this, our approach cryo-SWAN explicitly conditions each hierarchical level on the outputs of preceding levels. It differs from HQ-VAE in a key way: each level explicitly conditions on prior outputs through a conditional fusion structure. Recursive residual quantization Besides the multi-level global structure, cryo-SWAN embeds features into the latent space through codebook-based quantization at levell. This operation is defined asA (l) θ ◦z= ˆz. Inspired by the recursive perception strategy in VAR-VAE, we decompose the encoder feature mapzinto a multi- scale representation:Z=z 0 , z 1 ,·, z n . Each perception scale corresponds to a distinct spatial dimensionL i ×W i ×z i . In contrast to RQ-VAE, each perception level in cryo-SWAN operates with a unique window size and is explicitly conditioned on the preceding level, denoted asz i =z i |z i−1 . At scalel, the local codebook is defined asC (l) =e (l) 1 ,·, e (l) k .e (l) i denotes thei-th embedding vector. Quantization operations at levellshareC (l) . The optimization objective at levellis given by C (l) = arg min j Z (l) |Z (l−1) −e (l) k 2 . The corresponding loss at levellis defined as in Equation 5. L l (A (l) θ ) = k X i=1 sg h z (l) i |z (l−1) i i −e (l) k 2 (5) At the global multi-scale level, the total loss is aggregated asL global = P L i=1 L i (A i |A i−1 ). This formulation captures the global commitment loss across all levels in the hierarchy. Cryo-SWAN Fig.1illustrates the architectures of cryo-SWAN. We set the decomposition depth toL= 2. The architecture includes two embedding stages (Fig.1a). The feature maps at each level are quantized in a recursive residual manner (Fig.1b). Between levels, features are integrated via conditional fusion and propagated to the next stage. Pseudocode in Algorithms1and2details the procedures corresponding to Fig.1a and Fig.1b. Globally, the total loss consists of the reconstruction loss logp(x|ˆz(x))and the sum of commitment losses across all scales, as defined in previous section. 5 Algorithm 1:cryo-SWAN Input:x∈R B×1×D×H×W 2:Params:ScalesK, cond. Output:ˆx,L k K 1 ,I k K 1 4:procedureForward(x) Phase 1: Encoding 6:Enc.←[ ], z←x fork= 1toKdo 8:Enc..push(encoder k (z)), z←Enc.[k−1] end for 10:Phase 2: Quant and Decoding L k ,I k ←[ ],[ ] 12:fork=K to 1do z c ← ( Enc.[k−1]ifk=K concat(r k+1 ,Enc.[k−1])else 14:ifcond.∧k̸=K:z c ←combi_conv k (z c ) ˆz k , I k ,L k ←quantizer k (quant_conv k (z c )) 16:r k ←decoder k (post_quant_conv k (ˆz k )) L k .append(L k ),I k .append(I k ) 18:end for returnr 1 ,L k ,I k 20:end procedure Algorithm 2:Residual Quantizer 1:Input:encoder mapf leveli ∈R B×D×H×W 2:Parameters:Perception scalesM 3:Output:f quant , TokensR, LossL i 4:Initialize:R i ←[ ],L i ←0, CodebookZ∈R V×C 5:foreach scalem∈1, . . . , Mdo 6:r m ←arg min j ∥f m −Z j ∥ 2 7:R←R∪r m 8:z m ←lookup(Z, r m ) 9:ˆz m ←Conv3D(z m ) 10:L←L+β∥f quant −f leveli ∥ 2 +∥ˆz m −f m ∥ 2 11:f quant ←f quant + ˆz m 12:end for 13:L i ←L i /M Determining resolution with Fourier Shell Correlation (FSC) To assess the restoration quality of EMDB density maps, we compute the 3D Fourier Shell Corre- lation (FSC) [42,43]. In Equation6,k i denotes a voxel in the Fourier spaceFVcorresponding to spatial frequencyk(3D Fourier shell) . Computing FSC across the frequency spectrum between reconstructedV pred andV GT yields the FSC curve, which reflects the similarity between volumes at different frequencies. This curve captures the signal-to-noise ratio as a function of frequency. The resolution is defined as the highest frequencyk i where the FSC drops below the 0.143 threshold, indicating the last statistically reliable match between the predicted and ground truth volumes [44]. FSC(k) = P k i ∈k FV pred (k i )·F ∗ V GT (k i ) q P k i ∈k ∥FV pred (k i )∥ 2 · P k i ∈k ∥F ∗ V GT (k i )∥ 2 (6) 6 Results Architecture design for voxel-based 3D shape representation To design an architecture capable of learning meaningful representations of shape from voxel-based 3D data, we started with an AE-style structure (Fig.1a). This was primarily motivated by AEs’ abilities to be trained in a self-supervised fashion, allowing them to take advantage of the existing voxel-level datasets. Yet, traditional autoencoders struggle to capture complex geometric properties, such as protein structures, when relying solely on raw density values. Motivated by prior work on multi-level decomposition[29], we introduced residual quantization strategies into the latent space and propose a multi-Scale Wavelet-inspired AutoeNcoder cryo-SWAN. Cryo-SWAN iteratively mod- els the latent representation in a multi-scale manner. This enables it to effectively capture both low- and high-frequency geometric information in protein densities. In this work, specifically, we showcase that the decomposition depth of 2 can provide an efficient density representation up to the used sampling limit of 8 Å. We therefore set the decomposition depth toL= 2. The architecture includes two embedding stages. The feature maps at each level are quantized in a recursive residual manner (Fig.1b). Between levels, features are integrated via conditional fusion and propagated to the next stage (see Methods). Once trained, cryo-SWAN model can be used to perform downstream tasks, such as denoising noisy input volumes or conditional generation of new shapes (Fig.1c). Fig. 1:Principle architecture of Multi-Scale Wavelet-decomposition-inspired Autoen- coder Network. (a) Multi-scale decomposition of the input signal. (b) The residual quantization process. A residual loss is computed during codebook lookup at each scale. (c) Downstream applica- tions of the voxel-based 3D shape representations. Representation learning performance of SWAN on benchmarks To evaluate the design of the cryo-SWAN architecture, we conducted representation learning experi- ments on two benchmarks: ModelNet and BuildingNet. In the first, we selected ModelNet40, which 7 contains 3D objects from 40 categories, ranging from low- to high-frequency shapes. In the sec- ond, we used the entire BuildingNet dataset containing structurally complex architectural shapes with rich high-frequency details. To make them comparable with cryo-EM data, we voxelized both datasets into volumetric densities (see Methods). For comparison, we included the SOTA models from computer vision: VAR-VAE, HQ-VAE, RQ-VAE, and the foundational VQ-VAE. Importantly, we focused solely on the VAE components of these models without the generative models. In addi- tion, we evaluated methods tailored to cryo-EM applications, cryo-DRGN[45] and cryo-Target[46], which are parts of comprehensive toolkits designed for specific cryo-EM processing tasks. There- fore, to ensure a fair comparison, we evaluate only their autoencoder components, which we call DRGN-VAR and Target-VAE, with respect to representation quality (Fig.2). To facilitate the qualitative performance comparison on the ModelNet benchmark, we evalu- ated all models under varying levels of geometric detail. We did this by splitting the objects into low, mix, and high categories, which correspond to predominantly low-frequency, mixed-frequency, and high-frequency geometric information, respectively (Fig.2, ModelNet). Results suggest that as the proportion of high-frequency content increases, the representation task becomes progressively more challenging. Specifically, in low-frequencydominated setting, most methods achieve satisfac- tory performance. Albeit DRGN-VAE and Target-VAE did not adequately represent the densities in this setting. This limitation likely stems from their reliance on classical VAE backbones, which offer limited representational capacity for complex 3D geometry. As higher-frequency information increased, both RQ-VAE and VQ-VAE lost fine-grained details at the mixed-frequency level. When high-frequency components dominated, cryo-SWAN was the only model that consistently preserved the intricate geometric structures. Results on BuildingNet, which isdominated by high-frequency geometric details, suggested the same trends (Fig. 2, BuildingNet). Cryo-SWAN effectively captured such frequency information and demonstratedsuperior performance over all competing methods. Next, to quantitatively evaluate the performance, we used four standard metrics: MSE (0.007), PSNR (21.726), IoU (0.993), and F1 score (0.996), shown in Tab.1. Results suggest that across gen- eral computer vision (CV) benchmarks, our model consistently outperformed all competing methods on all metrics, demonstrating a clear performance advantage. ModelModelNetBuildingNetProteinNet3D MSE↓PSNR↑IoU↑F1↑MSE↓PSNR↑IoU↑F1↑MSE↓PSNR↑IoU↑F1↑FSC (0.5)↓ SWAN0.007 21.726 0.993 0.996 0.003 25.585 0.839 0.912 0.003 27.063 0.672 0.7999.10 VAR0.025 17.508 0.836 0.904 0.006 22.736 0.616 0.761 0.006 22.790 0.4000.55814.01 HQ0.01319.6120.9240.9580.00523.5880.6720.8020.007 22.558 0.317 0.46618.20 RQ0.027 16.326 0.790 0.875 0.008 21.886 0.476 0.642 0.008 21.990 0.267 0.40924.08 VQ0.044 13.931 0.726 0.834 0.025 16.374 0.327 0.489 0.00524.1070.248 0.38823.49 DRGN 0.055 13.201 0.518 0.660 0.025 16.983 0.207 0.341 0.012 20.210 0.182 0.30055.75 Target 0.119 9.740 0.312 0.458 0.049 13.480 0.147 0.253 0.011 20.394 0.140 0.23877.82 Table 1: Evaluation of candidates on three datasets with PSNR, MSE, IoU, F1-score, and FSC resolution (only for Protein3DNet). Bold items indicate the best performance, underlined values indicate the second best. SWAN performance on 3D molecular shapes To assess cryo-SWAN on experimentally derived molecular data, we curated ProteinNet3D, com- prising over 24,000 cryo-EM density volumes retrieved from EMDB. All maps were resampled to a uniform isotropic voxel spacing of 4 Å, resulting in a Nyquist limit of 8 Å [47]. At this sampling, protein densities retain substantial structural complexity while exhibiting heterogeneous noise levels and diverse geometric features across multiple spatial scales, which presents a great challenge for representation learning algorithms due to highly complex geometries [48]. This makes ProteinNet3D a stringent benchmark for volumetric representation learning. Representative reconstructions produced by cryo-SWAN and competing methods are shown in Fig.3a. Remarkably, qualitative comparisons indicate that DRGN-VAE and Target-VAE predom- inantly reconstruct low-frequency structural components, with limited recovery of finer geometric features. VAR-VAE and HQ-VAE better capture intermediate detail, yet visibly attenuate high- frequency density variations. In contrast, cryo-SWAN preserves fine-grained structural elements while maintaining coherent global shape. We next evaluated the performance of all models quantitatively (Tab.1). Results suggest that the cryo-SWAN outperforms other models on all the metrics: MSE (0.003), PSNR (27.063), IoU (0.672), 8 Fig. 2:cryo-SWAN performance on common 3D shape benchmarks. (a) Performance on three representative objects from BuildingNet dataset. (b) Performance on three representative objects from ModelNet dataset. Here, low, mix, and high denote the frequency components, where mix captures a combination of low- and high-frequency details. and F1 (0.799). To further interrogate the performance of the models with respect to a commonly used cryo-EM measure, we evaluated the Fourier Shell Correlation (FSC) between the input volumes and reconstructions (Fig.3). Results suggested that the cryo-SWAN achieved a formidable resolution of 9.10 Å (with a cut-off of 0.5). Altogether, these results suggest that our architecture (cryo-SWAN) consistently outperforms other models in reconstructing complex 3D densities as protein structures across all the metrics, which can be attributed to the architecture’s ability to learn higher-quality representations. To further interrogate this point, we performed a dimensionality reduction analysis of the SWAN representations using UMAP [49] (Fig.3c). Specifically, we compared UMAP visualizations of the original EMDB data points with their latent representations from the cryo-SWAN model. Evidently, the cryo-SWAN embeddings exhibit richer and clearer structural organization of 3D protein densities. Visualisation suggests that certain data points are positioned more closely, likely reflecting their higher geometric similarity. Further analysis of such points in Fig.3d reveals that certain proteins exhibit strong geometric similarity, which we refer to as hubs. By zooming into a representative hub, we identify an anchor protein density and retrieve other geometrically similar densities within a defined neighborhood. For example, EMD-4960 (human p97/VCP), EMD-13186 (A. baumannii F 1 F o -ATP synthase), and EMD-27864 (E. coliRho factor) occupy a neighboring region of the latent manifold. At the 8 Å resolution considered here, these densities share common geometric features observable at the envelope level, including ring-like features with a rotational symmetry around the central axis, a resolvable internal cavity, and similar size and aspect ratio. We present more examples at the supplementary Fig.S4. These structurally similar densities highlight shared geometric characteristics and may provide insights into their underlying biological functions. Importantly, the observed latent space neighborhoods reflect similarity in volumetric geome- try rather than sequence identity or biochemical function. These results suggest that cryo-SWAN embeddings capture multi-level structural organization at the level of resolvable density features, indicating that the latent space encodes structural similarities within cryo-EM data. Together, these results demonstrate that cryo-SWAN not only improves reconstruction fidelity on complex molecu- lar volumes but also learns latent representations that encode meaningful geometric structure within cryo-EM data. Ablation study To verify that the performance gain of cryo-SWAN results from our design choices, we performed an ablation study. A key challenge in VQ-VAE-based models involves codebook collapse [14,23], where only a small subset of the codebook indices is used during embedding. Such a collapse often leads to inefficient codebook utilization and degraded embedding quality. An effective embedding strategy should promote both diverse index usage and semantically meaningful patterns. Therefore, to assess the codebook usage, we visualized the lateral view of the latent embedding indices from residual quantization level (RQ1; see Fig.1a,b) on ProteinNet3D (Fig.4a). We examined quantization 9 Fig. 3:Representation learning on the ProteinNet3D and explorations in the latent space.(a) The representation learning on the ProteinNet3D. (b) The FSC evaluation of all the representations. (c) Dimensionality reduction comparison: Cryo-SWAN latent vectors vs. original EMDB volumes. (d) UMAP projection of latent vectors reveals distinct hubs reflecting structural similarity. Molecules with similar 3D structures at 8 Å resolution (siblings). behavior across perception scales (4, 8, 16). Results suggest that the codebook index is well-utilized throughout with no obvious signs of collapse. We further zoomed in on specific patches (Fig.4b) across the perception scales RQ1 and RQ2 (see Fig.1a,b for architecture reference). The results reveal a clear coarse-to-fine quantization pattern: lower scales use broader index assignments, while higher scales introduce localized details. Transitions across scales remain smooth and coherent, suggesting that the model recursively builds hierarchical features in a structured and consistent manner. Finally, we examined the molecular shape reconstructions performance from RQ1 in com- parison to the full representation (Fig.4c) visually and quantitatively (Fig.4d). The single-scale 10 variant exhibits noticeably degraded reconstruction quality, particularly in high-frequency regions, and performs significantly worse across evaluation metrics. These findings suggest that the single- scale performance (RQ1) is significantly worse than the multi-scale (RQ1+RQ2), justifying our architectural design choices. Fig. 4:Ablation study for codebook collapsing and multi-scale representations.(a) The calls for all the codebook entries. (b) The latent code representation for both scales (coarse and fine). (c) The molecular density volume: original (GT), at a certain representation scale (RQ1) and at a full representation (RQ1+RQ2). (d) Evaluation of representation performance at a certain representation scale (RQ1) compared to the full representation (RQ1+RQ2). Downstream tasks of cryo-SWAN Finally, to demonstrate how cryo-SWAN high-quality representations can be used for down- stream applications, we explore two tasks: conditional generation of 3D molecular structures and unsupervised 3D denoising. In the first task, the high-quality representations cryo-SWAN learned can be employed to gen- erate similar shapes in high resolution (pixel size as 4 Å in this work), given an input shape as a condition. Albeit indirectly, this can serve a number of applications fromde novomolecular design to synthetic data generation. To accomplish this, we used the denoising diffusion probabilistic model (DDPM) [17] to learn the distribution of latent vectors from the cryo-SWAN quantization (Fig.5). During generation, we condition the molecules’ latent vectors on DDPM, enabling the synthesis of structurally similar variants (Fig.5a). These changed latent vectors are decoded by the cryo-SWAN decoder to produce synthetic 3D molecular density maps (Fig.5b). Results suggest that the gen- erated densities preserve key structural features of the anchor. For more examples, please refer to supplementary Fig.S5. Next, we explored the downstream task of unsupervised 3D denoising (Fig.5c). To simulate the reconstruction of the noisy input we first added artificial Gaussian noise (SNR=0.5) to the shapes from the ModelNet40 benchmark and then restored it using a band-pass filter (low-0.25, high-0.01) or cryo-SWAN (see Methods). By projecting latent codes of noisy inputs onto the latent space of clean densities, our approach achieves effective denoising while preserving high-frequency information. Results suggest that cryo-SWAN was able to better preserve the high-frequency (fine-grained) 3D structural details of the 3D shape compared to band-pass (Fig.5c, zoomed-in boxes). Furthermore, our approach outperformed the baseline with respect to PSNR (14.98), SSIM (0.26), and RMSE 11 Fig. 5:Downstream applications of cryo-SWAN including shape denoising and condi- tional molecular shape generation.(a) The conditional generation from diffusion models based on cryo-SWAN representations. (b) Examples of conditional generations. The variants are similar in the geometric perspective to the real protein densities in EMDB (anchor point). (c) The unsuper- vised denoising based on the representation learning for 3D densities from ModelNet. (0.19) metrics on the test dataset. These results indicate that the learned latent representations can 12 act as a structured prior over valid 3D shapes, enabling restoration beyond simple frequency-domain filtering. Discussion Learning representations directly from the voxelized shapes is an under-researched avenue, yet it bears promise for creating advanced algorithms in the realm of biomedical imaging. It opens an avenue for handling data from modalities like cryo-ET directly on the averaged densities, allowing us to learn high-quality representations of molecular shapes. In this paper, we propose addressing this challenge by introducing cryo-SWAN a compute- efficient architecture inspired by wavelet decomposition. Cryo-SWAN performs conditional, coarse-to- fine residual quantization at multiple scales, allowing to capture both global shape and fine structural details. We evaluated cryo-SWAN on three datasets: ModelNet and BuildingNet (voxelized), and ProteinNet3D a large-scale cryo-EM dataset we curated using EMDB[35] as a source. We demon- strate that cryo-SWAN outperformed SOTA pixel/voxel-based VAE baselines (HQ-VAE, VAR-VAE, RQ-VAE, and VQ-VAE) on metrics: MSE, PSNR, IoU, and F1-score across all datasets. These improvements arise from two core design principles. First, the coarse-to-fine conditional architecture explicitly distributes representational capacity: the coarse level captures overall geom- etry, while the fine level focuses on localized high-resolution features. This hierarchical separation mitigates the blurring and detail loss common to single-scale latent models. Second, recursive residual quantization enforces structured latent approximations that promote diverse and semantically mean- ingful codebook utilization, reducing collapse phenomena often observed in vector-quantized models. Ablation experiments confirm that this multi-scale design is essential to the observed performance gains. Furthermore, we explored the quality of the learned representations through UMAP dimension- ality reduction. Our analysis revealed clear data points aggregates (hubs). Molecules in these hubs shared similar structural characteristics. It is tempting to speculate that cryo-SWAN and future approaches could be used for structural mining, linking molecular geometry with putative links to biological function. Finally, we explored conditional shape generation a downstream task enabled by the high-quality representations learned by cryo-SWAN. We implemented a latent diffusion pipeline based on cryo-SWAN and DDPM[17]. By conditioning the DDPM on an anchor protein, we gen- erated synthetic density volumes with similar structural characteristics, enabling controllable and realistic molecular variation. Notably, some limitations persist due to the scope of this work. Firstly, we explored only a two- scale decomposition. However, further scales could be possible, albeit likely with diminishing returns. Secondly, cross-scale codebook sharing remains an open question, as does using different codebook sizes per scale. We aim to explore these questions in future. Practically, cryo-SWAN may enable the construction of data-driven molecular shape priors directly from experimental density repositories. Such priors may improve reconstructions, support realistic data augmentation, and reduce reliance on models when they are incomplete or unavailable. Potential extensions include molecular identification through latent-space, restoration of tomographic "missing-wedge" artifacts [50], and orientation estimation of molecules for subtomogram averaging [51]. Taken together, cryo-SWAN provides a general-purpose framework for learning 3D representa- tions of volumetric densities. Our work opens several promising avenues for representation learning in structural biology and possibly other fields relying on 3D biomedical imaging. Looking forward, Cryo-SWANs architecture could be extended with multi-modal data by incorporating a sequence encoder [52,53]. Coupled with its high-resolution representation capability, this could further facilitate the rational design of molecular sequences with defined structures. Acknowledgments This work was partially funded by the Center for Advanced Systems Understanding (CASUS), which is financed by Germanys Federal Ministry of Research, Technology and Space (BMFTR) and by the Saxon Ministry for Science, Culture, and Tourism (SMWK) with tax funds based on the budget approved by the Saxon State Parliament. The work was supported by the Helmholtz Imaging IVF grant CryoFocal. MK is supported by the Heisenberg award from the DFG (KU 3222/2-1), and funding from the Helmholtz Association. AY is supported by the Helmholtz Association Initiative 13 and Networking Fund in the frame of Helmholtz AI as well as by the Helmholtz Foundation Model Initiative within the project PROFOUND. The authors thank HelmholtzAI (grant tomoCAT). The authors gratefully acknowledge the Gauss Centre for Supercomputing e.V. (w.gauss-centre.eu) for funding this project by providing computing time through the John von Neumann Institute for Computing (NIC) on the GCS Supercomputer JUWELS at Jülich Supercomputing Centre (JSC). References [1]Xiong, B., Wei, S.-T., Zheng, X.-Y., Cao, Y.-P., Lian, Z., Wang, P.-S.: OctFusion: Octree-based Diffusion Models for 3D Shape Generation. arXiv (2024).https://doi.org/10.48550/arXiv.2408. 14732 [2]Museth, K.: VDB: High-resolution sparse volumes with dynamic topology. ACM Transactions on Graphics (TOG)32(3), 1–22 (2013)https://doi.org/10.1145/2487228.2487235 [3]Li, W., Liu, J., Chen, R., Liang, Y., Chen, X., Tan, P., Long, X.: CraftsMan: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner. arXiv (2024). https://doi.org/10.48550/arXiv.2405.14979 [4]Ren, X., Huang, J., Zeng, X., Museth, K., Fidler, S., Williams, F.: XCube: Large-scale 3D gen- erative modeling using sparse voxel hierarchies. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 4209–4219 (2024).https://doi.org/10.1109/ CVPR52733.2024.00403 [5]Chen, R., Zhang, J., Liang, Y., Luo, G., Li, W., Liu, J., Li, X., Long, X., Feng, J., Tan, P.: Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders. arXiv (2025). https://doi.org/10.48550/arXiv.2412.17808 [6]Cheng, Y.-C., Lee, H.-Y., Tulyakov, S., Schwing, A., Gui, L.: SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation. arXiv (2023).https://doi.org/10.48550/arXiv. 2212.04493 [7]Mittal, P., Cheng, Y.-C., Singh, M., Tulsiani, S.: AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation. arXiv (2023).https://doi.org/10.48550/arXiv.2203.09516 [8]Zhao, Z., Liu, W., Chen, X., Zeng, X., Wang, R., Cheng, P., Fu, B., Chen, T., Yu, G., Gao, S.: Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation. arXiv (2023).https://doi.org/10.48550/arXiv.2306.17115 [9]Li, X., Zhang, Q., Kang, D., Cheng, W., Gao, Y., Zhang, J., Liang, Z., Liao, J., Cao, Y.-P., Shan, Y.: Advances in 3D generation: A survey. arXiv preprint (2024)https://doi.org/10.48550/ arXiv.2401.17807 [10]Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S.W., Fidler, S., Kreis, K.: Align your latents: High-resolution video synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 22563–22575 (2023).https://doi.org/10.1109/CVPR52729.2023.02161 [11]Lee, D., Kim, C., Kim, S., Cho, M., Han, W.-S.: Autoregressive image generation using residual quantization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11523–11532 (2022).https://doi.org/10.48550/arXiv.2203.01941 [12]Razavi, A., Oord, A., Vinyals, O.: Generating diverse high-fidelity images with VQ-VAE-2. In: Advances in Neural Information Processing Systems (NeurIPS), vol. 32 (2019).https://dl.acm. org/doi/10.5555/3454287.3455618 [13]Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthe- sis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 10674–10685 (2022).https://doi.org/10.1109/ CVPR52688.2022.01042 14 [14]Takida, Y., Ikemiya, Y., Shibuya, T., Shimada, K., Choi, W., Lai, C.-H., Murata, N., Uesaka, T., Uchida, K., Liao, W.-H., Mitsufuji, Y.: HQ-VAE: Hierarchical discrete representation learning with variational bayes. Transactions on Machine Learning Research (2024)https://doi.org/10. 48550/arXiv.2401.00365 [15]Tian, K., Jiang, Y., Yuan, Z., Peng, B., Wang, L.: Visual autoregressive modeling: Scalable image generation via next-scale prediction. In: Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS) (2024).https://dl.acm.org/doi/10.5555/ 3737916.3740610 [16]Kingma, D.P., Welling, M.: Auto-encoding variational bayes. In: 2nd International Conference on Learning Representations (ICLR) (2014).https://doi.org/10.48550/arXiv.1312.6114 [17]Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural infor- mation processing systems (NeurIPS)33, 6840–6851 (2020)https://doi.org/10.48550/arXiv. 2006.11239 [18]Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial networks. Communications of the ACM63(11), 139–144 (2020)https://doi.org/10.1145/3422622 [19]Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polo- sukhin, I.: Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS)30(2017) https://doi.org/10.48550/arXiv.1706.03762 [20]Bank, D., Koenigstein, N., Giryes, R.: Autoencoders. In: Rokach, L., Maimon, O., Shmueli, E. (eds.) Machine Learning for Data Science Handbook, p. 353–374. Springer.https://doi.org/10. 1007/978-3-031-24628-9_16 [21]Bengio, Y., Courville, A., Vincent, P.: Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence35(8), 1798–1828 (2013)https: //doi.org/10.1109/TPAMI.2013.50 [22]Odaibo, S.: Tutorial: Deriving the Standard Variational Autoencoder (VAE) Loss Function. arXiv. arXiv:1907.08956 [cs.LG] (2019).https://doi.org/10.48550/arXiv.1907.08956.https:// doi.org/10.48550/arXiv.1907.08956 [23]Kossale, Y., Airaj, M., Darouichi, A.: Mode Collapse in Generative Adversarial Networks: An Overview. In: 2022 8th International Conference on Optimization and Applications (ICOA), p. 1–6 (2022).https://doi.org/10.1109/ICOA55659.2022.9934291 [24]Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., Tagliasacchi, M.: SoundStream: An end-to- end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing 30, 495–507 (2022)https://doi.org/10.1109/TASLP.2021.3129994 [25]Van Den Oord, A., Vinyals, O.,et al.: Neural discrete representation learning. In: Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), p. 6309–6318 (2017).https://dl.acm.org/doi/10.5555/3295222.3295378 [26]Esser, P., Rombach, R., Ommer, B.: Taming transformers for high-resolution image synthesis. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 12868–12878 (2021).https://doi.org/10.1109/CVPR46437.2021.01268 [27]Prabhakar, C., Li, H., Yang, J., Shit, S., Wiestler, B., Menze, B.: ViT-AE++: improving vision transformer autoencoder for self-supervised medical image representations. In: Medical Imaging with Deep Learning, p. 666–679 (2024).https://doi.org/10.48550/arXiv.2301.07382 [28]Chefer, H., Gur, S., Wolf, L.: Transformer interpretability beyond attention visualization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 782–791 (2021).https://doi.org/10.1109/CVPR46437.2021.00084 15 [29]Li, R., Kudryashev, M., Yakimovich, A.: Solving the inverse problem of microscopy deconvo- lution with a residual Beylkin-Coifman-Rokhlin neural network. In: European Conference on Computer Vision, p. 378–395 (2024).https://doi.org/10.1007/978-3-031-73226-3_22 [30]Fan, Y., Bohorquez, C.O., Ying, L.: BCR-Net: A neural network based on the nonstandard wavelet form. Journal of Computational Physics384, 1–15 (2019)https://doi.org/10.1016/j.jcp. 2019.02.002 [31]Zamir, S.W., Arora, A., Khan, S., Hayat, M., Khan, F.S., Yang, M.-H., Shao, L.: Multi-stage pro- gressive image restoration. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14816–14826 (2021).https://doi.org/10.1109/CVPR46437.2021. 01458 [32]Zhou, Z., Wang, B., Li, S., Dong, M.: Perceptual fusion of infrared and visible images through a hybrid multi-scale decomposition with Gaussian and bilateral filters. Information Fusion30, 15–26 (2016)https://doi.org/10.1016/j.inffus.2015.11.003 [33]Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., Xiao, J.: 3D ShapeNets: A deep representation for volumetric shapes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 1912–1920 (2015).https://doi.org/10.1109/CVPR. 2015.7298801 [34]Selvaraju, P., Nabail, M., Loizou, M., Maslioukova, M., Averkiou, M., Andreou, A., Chaud- huri, S., Kalogerakis, E.: BuildingNet: Learning to label 3D buildings. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 10377–10387 (2021). https://doi.org/10.1109/ICCV48922.2021.01023 [35]The wwPDB Consortium, Turner, J., Abbott, S., Fonseca, N., Pye, R., Carrijo, L.,et al.: EMDBthe Electron Microscopy Data Bank. Nucleic Acids Research52(D1), 456–465 (2024) https://doi.org/10.1093/nar/gkad1019 [36]Van Heel, M., Schatz, M.: Fourier shell correlation threshold criteria. Journal of Structural Biology151(3), 250–262 (2005)https://doi.org/10.1016/j.jsb.2005.05.009 [37]Aiyer, S., Zhang, C., Baldwin, P.R., Lyumkis, D.: Evaluating local and directional resolution of cryo-EM density maps. In: Cryo-EM: Methods and Protocols, p. 161–187 (2020).https: //doi.org/10.1007/978-1-0716-0966-8_8 [38]Agard, D., Cheng, Y., Glaeser, R.M., Subramaniam, S.: Single-particle cryo-electron microscopy (cryo-EM): Progress, challenges, and perspectives for further improvement. Advances in Imag- ing and Electron Physics 185 , 113–137 (2014) https://doi.org/10.1016/B978-0-12-800144-8. 00002-1 [39]Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., Courville, A.: On the spectral bias of neural networks. In: Proceedings of the 36th International Confer- ence on Machine Learning (ICML), p. 5301–5310 (2019).https://doi.org/10.48550/arXiv.1806. 08734 [40]Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65(1), 99–106 (2021)https://doi.org/10.1145/3503250 [41]Beylkin, G., Coifman, R., Rokhlin, V.: Fast wavelet transforms and numerical algorithms I. Communications on Pure and Applied Mathematics44(2), 141–183 (1991)https://doi.org/10. 1002/cpa.3160440202 [42]Williams, D.B., Carter, C.B.: The transmission electron microscope. In: Transmission Electron Microscopy: A Textbook for Materials Science, p. 3–22. Springer, Boston, MA (2009).https: //doi.org/10.1007/978-0-387-76501-3_1 16 [43]Harauz, G., Heel, M.: Exact filters for general geometry three dimensional reconstruction. Optik 73(4), 146–156 (1986) [44]Rosenthal, P.B., Henderson, R.: Optimal determination of particle orientation, absolute hand, and contrast loss in single-particle electron cryomicroscopy. Journal of Molecular Biology333(4), 721–745 (2003)https://doi.org/10.1016/j.jmb.2003.07.013 [45]Zhong, E.D., Bepler, T., Berger, B., Davis, J.H.: CryoDRGN: reconstruction of heterogeneous cryo-EM structures using neural networks. Nature Methods18(2), 176–185 (2021)https://doi. org/10.1038/s41592-020-01049-4 [46]Nasiri, A., Bepler, T.: Unsupervised object representation learning using translation and rota- tion group equivariant vae. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS) (2022).https://dl.acm.org/doi/10.5555/3600270. 3601380 [47]Shannon, C.E.: Communication in the presence of noise. Proceedings of the IRE37(1), 10–21 (1949)https://doi.org/10.1109/JRPROC.1949.232969 [48]Harris, J.: A bound on the geometric genus of projective varieties. Annali della Scuola Normale Superiore di Pisa-Classe di Scienze8(1), 35–68 (1981) [49]McInnes, L., Healy, J., Saul, N., GroSSberger, L.: UMAP: Uniform manifold approximation and projection. Journal of Open Source Software3(29), 861 (2018)https://doi.org/10.21105/joss. 00861 [50]Wiedemann, S., Heckel, R.: A deep learning method for simultaneous denoising and missing wedge reconstruction in cryogenic electron tomography. Nature Communications15(1), 8255 (2024)https://doi.org/10.1038/s41467-024-51438-y [51]Leigh, K.E., Navarro, P.P., Scaramuzza, S., Chen, W., Zhang, Y., Castaño-Díez, D., Kudryashev, M.: Chapter 11 - subtomogram averaging from cryo-electron tomograms. In: Three-Dimensional Electron Microscopy. Methods in Cell Biology, vol. 152, p. 217–259. Academic Press, ??? (2019). https://doi.org/10.1016/bs.mcb.2019.04.003 [52]Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z.,et al.: Evolutionary-scale prediction of atomic-level protein structure with a language model. Science379(6637), 1123–1130 (2023)https://doi.org/ 10.1126/science.ade2574 [53]Brandes, N., Ofer, D., Peleg, Y., Rappoport, N., Linial, M.: ProteinBERT: a universal deep- learning model of protein sequence and function. Bioinformatics 38 (8), 2102–2110 (2022)https: //doi.org/10.1093/bioinformatics/btac020 Supplementary Information ProteinNet3D Here we offer additional details on the preparation and contents of the ProteinNet3D dataset, includ- ing the optimal volume and pixel size choice (Supplementary Fig.S1), the diversity of the molecular weights (Suppl. Fig.S2), and the variability of background scenarios (Supplementary Fig.S3). 17 Fig. S1:Optimal choice of the pixel size and the volume size.(top) Volume size distribution at various resampled pixel size. (bottom) Volumes distribution respective to the originally reported resolution and the pixel size: ProteinNet3D volumes, preserving original resolution after resampling (blue) and degraded to8Å after resampling due to the Nyquist limit (green), as well as filtered out volumes due to the cutoffs applied (white). 18 Fig. S2:ProteinNet3D diversity: molecular weights range.(a) Molecular weights distribution. (b) Example molecular volume views Fig. S3:ProteinNet3D complexity: membrane-related signal and experimental noise presence.(top) Normal protein densities in ProteinNet. (middle) ProteinNet example volumes containing membrane-related signal. (bottom) ProteinNet example volumes containing experimental noise 19 Applications We present additional examples of geometry-similarity detection in Supplementary Fig. S4. Lever- aging Cryo-SWAN, we identify proteins with similar structural geometry at 8 Å resolution. These findings may offer insights into the relationship between molecular shape and biological function. Fig. S4:Molecular shapes similarity within hubs identified with Cryo-SWAN. Additional generation examples are shown in Supplementary Fig.S5. Using the latent represen- tation of a template protein (anchor), Cryo-SWAN enables the generation of structurally similar 3D densities. Note that these results were produced using a basic 3D U-Net DDPM, leading to some variation and scatter in the outputs. Higher-quality generation could be achieved by incorporating more advanced generative models, such as GANs, transformers, and Stable Diffusion. Fig. S5:The latent diffusion generation based on cryo-SWAN representation. 20