Paper deep dive
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/2/2026, 1:07:17 PM
Summary
The paper introduces dRAE (discrete Representation Autoencoder) and Hyper-Spherical Quantization (HSQ) to address codebook collapse in high-dimensional visual representation autoencoders. By decoupling semantic orientation from feature magnitude via angular routing, HSQ enables scalable codebook usage (up to 131,072 codes) with 100% utilization, improving both image reconstruction fidelity and multimodal understanding performance compared to standard Euclidean Vector Quantization (VQ).
Entities (9)
Relation Signals (6)
dRAE → uses → HSQ
confidence 95% · The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction... powered by Hyper-Spherical Quantization (HSQ).
HSQ → solves → codebook collapse
confidence 92% · HSQ... prevents code assignment from being dominated by scale rather than meaning... significantly improves tokenization quality... eliminating the high-variance outliers that trigger codebook collapse
HSQ → enables → Scalable Codebook
confidence 90% · Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131,072, along with 100% codebook utilization
HSQ → replaces → Euclidean Distance
confidence 90% · HSQ replaces the standard Euclidean assignment with angular routing
VQRAE → suffersfrom → codebook collapse
confidence 88% · VQRAE... is prone to codebook collapse and of limited scaling potential.
dRAE → outperforms → VQRAE
confidence 85% · dRAE outperforms VQRAE with significant margins.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.
Tags
Links
- Source: https://arxiv.org/abs/2607.22148v1
- Canonical: https://arxiv.org/abs/2607.22148v1
Trouble viewing inline? Open PDF directly →
Full Text
59,349 characters extracted from source content.
Expand or collapse full text
dRAE: Representation Autoencoder with Hyper-Spherical Codes Tianren Ma1,4 Lin Long2,4 Chuyan Chen3,4 Mu Zhang1 Junbo Zhao2,4∗ Tong Zhang1 Qixiang Ye1 Corresponding authors Abstract In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models — a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131,072, along with 100% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks. 1 University of Chinese Academy of Sciences 2 Zhejiang University 3 Peking University 4 Ant Group Project Page: https://drae-hsq.github.io 1 Introduction The trajectories of visual understanding and generation have conventionally been separated by their representation designs. Visual encoders optimized for understanding [42, 37, 55], excel at mapping images to high-dimensional, semantically rich spaces via vision transformers (ViTs) [13, 20], whereas generative frameworks typically rely on variational autoencoders (VAEs) [24, 43, 15, 38] to compress images into low-dimensional manifolds. While modern VAEs are highly optimized for localized pixel reconstruction, scaling their latent dimensions is constrained by the KL-divergence prior and practical difficulties [65]. Consequently, most unified multimodal models resort to disjointed pipelines—employing a ViT for comprehension and an independent VAE for generation. Recently, a promising trend has emerged to bridge this architectural gap. Frameworks [61, 70] such as the representation autoencoder (RAE) demonstrate that high-dimensional features from contrastive or self-supervised models can be directly used to reconstruct images via a simple ViT. As these pretrained representations inherently capture global semantic alignment, visual tokens derived from this space carry significantly higher generalizable information than that from a pure reconstruction VAE. This high semantic bandwidth is desirable for multimodal models, as it minimizes the need for complex cross-modal alignment [54, 53], offering a more unified foundation for integrated understanding and generation. Figure 1: Illustration of the code space defined by VQ (a) and HSQ (b), and distribution of their learned codebook embeddings (c). Sim(⋅)Sim(·) for cosine similarity, and sg(⋅)sg(·) for stop-gradient operation. For the angular analysis in c, we sample features from the central region with similar magnitude ranges and visualize them with PCA dimension reduction [58]. However, transporting such signals via discrete indices presents a significant optimization challenge. Existing vector quantization (VQ) methods [56], Fig. 1 a, which are designed for low-dimensional autoencoder bottlenecks, typically result in severe codebook collapse when applied on high-dimensional features (see Fig. 2). Empirical observations in recent baselines [14] reveal that the tokenizer plateaus when scaled to a vocabulary larger than 16K codes. This phenomenon neutralizes the primary advantage of representation-based designs, as the constrained vocabulary cannot adequately capture the diverse semantic capacity of the latent space. In this study, we argue that the observed codebook collapse is largely attributable to the fundamental metric mismatch. While semantics in vision foundation models are predominantly encoded in the orientation of latent vectors [58], VQ-based frameworks rely on Euclidean distance for codebook assignment. This inconsistency makes the objective sensitive to the anisotropic scale of pretrained spaces, resulting in codebook embeddings with high-variance magnitude and clustered angular distributions, as evidenced in Fig. 1 c. Such magnitude-induced bias allows latent vectors with disparate norms to disproportionately dominate the tessellation, enabling certain codebook entries to "hijack" assignments based on numerical magnitude alone, regardless of their semantic alignment with the input. At the same time, simply imposing a unit-sphere prior on inputs, which means discarding the magnitude, would deteriorate reconstruction, as this operation is non-uniform across tokens, leading to structural and textural nuances lost. To resolve this dilemma, we propose the discrete Representation Autoencoder (dRAE) powered by Hyper-Spherical Quantization (HSQ). The core innovation lies in explicitly decoupling the metric used for codebook routing from the quantization objective. By shifting from Euclidean distance to angular similarity, HSQ assigns codes based on semantic orientation while preserving essential magnitude information for the decoder. This approach significantly improves tokenization quality through two primary mechanisms: magnitude homogenization and angular uniformity. As illustrated in Fig. 1 c, HSQ constrains the codebook to a significantly more stable magnitude scale, effectively eliminating the high-variance outliers that trigger codebook collapse in vanilla VQ or SimVQ [71]. By reducing the magnitude-induced noise, the latent space attains a more isotropic, uniform distribution on the hyper-sphere. This ensures that the codebook utilizes its full capacity to represent distinct semantic directions rather than being dominated by arbitrary feature scales. Consequently, dRAE preserves the fine structural details necessary for high-fidelity synthesis while maintaining the semantic integrity required for large-scale multimodal understanding. Through extensive experiments, we demonstrate that the proposed HSQ not only shows continuous gain on reconstruction by scaling vocabulary up to 131,072131,072, but also keeps high semantic fidelity for image understanding tasks. We also validate the practical efficacy of dRAE by integrating it with a generative pipeline for text-to-image generation, confirming its viability as a high-throughput tokenizer for both visual understanding and generation tasks. 2 Related Work Visual Semantic and Generative Latent. The evolution of self-supervised and contrastive learning has established a robust foundation for multimodal models. By utilizing objectives such as masked modeling [20], self-distillation [7, 37] and image-text alignment [42, 55], visual models embed image patches into high-dimensional latents (e.g. d∈[768,1536]d∈[768,1536]), and learn semantically structured representation that generalizes across visual understanding tasks. On the other hand, state-of-the-art generative models still operate on heavily compressed manifolds (d∈[8,64]d∈[8,64]), i.e., diffusion is typically built in reconstruction-trained VAE space [43, 65]. This creates a dimensional incompatibility between the semantic-rich embeddings and the compact generative latents. Unifying Comprehension and Generation. One direct solution of aligning these disparate paths is to explore an ensemble of semantic and reconstruction objectives. VTP [64] explores end-to-end pretraining to learn representations that can unify comprehension with generation from scratch. However, it imposes extremely high demand on data collection and training budgets. A more promising trend is to utilize a pretrained encoder for more downstream tasks. VILA-U [61] uses residual quantization to align image latents with text embeddings, while TokLIP [30] bakes CLIP features into the pretrained VQGAN tokenizer through distillation and contrastive training. Recently, RAE [70] proposes that the high-dimensional features, optimized for semantics, can be directly decoded for image reconstruction. It also shows that high-dimensional vectors, which jointly capture high semantic density and fine-rained structural, are of great potential in visual world modeling [53]. Its discrete counterpart, VQRAE [14], uses SimVQ [71] to quantize the high-dimensional tokens, but is prone to codebook collapse and of limited scaling potential. In this study, we attempt to address these limitations by introducing a scalable quantization method that stabilizes training while preserving semantic depth. 3 Preliminary Image Reconstruction. This requires the model to learn a mapping from the input image x∈[0,1]h×w×3x∈[0,1]^h× w× 3 to latents and reconstruct them back to pixel space. Formally, an encoder E maps x to latent features Z:=E(x)=znn=1N∈ℝN×dZ:=E(x)=\z_n\_n=1^N ^N× d, where N=hw/p2N=hw/p^2 denotes the number of visual tokens. We will also use z∈ℝdz ^d to denote an un-quantized embedding. A decoder D reconstructs the image x^=D(Z) x=D(Z). The model is trained to minimize a objective that encourages both pixel-level fidelity (L1L_1 distance) and perceptual quality (perceptual ℒpercL_perc and discriminator loss ℒdiscL_disc): ℒrec=L1(x^,x)+ωpℒperc(x^,x)+ωdℒdisc(x^,x).L_rec=L_1( x,x)+ _pL_perc( x,x)+ _dL_disc( x,x). (1) Representation Autoencoder. Conventional approaches such as VAEs compress images into low-dimensional latent spaces, i.e.i.e., the bottleneck, which facilitates convergence but limits the capacity to encode semantic information. In contrast, RAE removes the bottleneck design, and instead leverages high-dimensional, semantically structured embeddings directly from pretrained Vision Foundation Models (VFMs) as its latents. Specifically, RAE employs a frozen VFM as E and a trainable ViT as D. By avoiding aggressive compression, RAE maintains features that are beneficial for downstream visual understanding and generation tasks. Vector-Quantizing Semantics. VFM features can be quantized for discrete modeling. Let Q(⋅)Q(·) denote the quantization operation, the quantized features as zq:=Q(z)z_q:=Q(z) and Zq:=Q(Z)Z_q:=Q(Z). VQRAE projected the image features to 1536 dimensions, and each vector znz_n is mapped to its nearest entry from a learnable codebook =cii=1KC=\c_i\_i=1^K, where K is the codebook capacity: zq=argminci∈‖z−ci‖2.z_q= _c_i \|z-c_i\|_2. (2) The quantization objective consists of a codebook loss and a commitment loss: ℒVQ=‖sg[z]−zq‖22⏟ℒcodebook+β‖z−sg[zq]‖22⏟ℒcommit,L_VQ= \|sg[z]-z_q\|_2^2_L_codebook+β \|z-sg[z_q]\|_2^2_L_commit, (3) where sg[⋅]sg[·] denotes the stop-gradient operator and β controls the commitment strength. To preserve the semantic structure while improving reconstruction fidelity, VQRAE also adopts a self-distillation strategy. A teacher model T (initialized from the original VFM) provides supervision to the unfrozen D, encouraging the learned features ZIZ_I to remain close to the teacher representations: ℒtotal=ℒrec+ℒVQ+λ‖Z−T(x)‖22.L_total=L_rec+L_VQ+λ\|Z-T(x)\|_2^2. (4) Additionally, to improve codebook utilization, a learnable projection matrix W∈ℝd×dW ^d× d is introduced by SimVQ [71] to transform codebook entries as WciWc_i, allowing for a more flexible alignment. 4 Methodology 4.1 Anisotropy in High-dimensional Space In high-dimensional Euclidean space, a standard Gaussian distribution exhibits the thin-shell concentration property [57, 25]: if z∼(0,Id)z (0,I_d), its norm ‖z‖2\|z\|_2 concentrates sharply around d d as d→∞d→∞. Vision encoders such as CLIP [42] are trained with contrastive objectives that encourage features to spread uniformly across directions-a pressure analogous to the isotropic structure of a Gaussian [5]. Related self-supervised approaches like DINO [7, 37] series and JEPA [1] also yield highly regular representations [4], which have been shown to encode density-related structure that can be exploited with Gaussian models. Their patch embeddings, as suggested by recent studies [26, 58], similarly concentrate on a thin spherical shell, with meaningful information encoded in both the direction (angular) and the radius (magnitude) of each feature vector. We conduct diagnostic experiments to empirically investigate their function, with the aim of validating two hypotheses: H1. Semantic information roughly lies on a hyper-sphere for understanding. We train two variants of an MLLM following LLaVA’s ViT-MLP-LLM architecture [32]. The original version receives raw image features Z as MLP’s inputs. In the normalized version, we apply ℓ2 _2 normalization such that all visual features Z′=Z/‖Z‖2Z =Z/\|Z\|_2 reside on a unit hyper-sphere. As shown in Tab. 1, the performance gap across diverse benchmarks—ranging from visual perception (MMBench [33]) to hallucination (POPE [29]) —is negligible (∼1% 1\%). This confirms that the semantic signal is almost entirely preserved in the feature’s directional orientation for high-level understanding. H2. Magnitude is structurally essential for image reconstruction. Following RAE, we train a ViT decoder to reconstruct pixels from either Z or Z′Z . As shown in Tab. 1, models trained on normalized features struggle to recover global structure, resulting in worse performance on reconstruction metrics. Table 1: Diagnostic Analysis. Comparison of understanding (a) and reconstruction (b) performance between raw and normalized image features as inputs. (a) Multimodal Understanding Variant MMBench ↑ TextVQA ↑ POPE ↑ Raw Z 82.2 61.3 85.2 Norm Z′Z 81.1 60.7 84.3 (b) Image Reconstruction Fidelity Variant PSNR ↑ SSIM ↑ rFID ↓ Raw Z 22.5 0.62 4.62 Norm Z′Z 20.6 0.51 9.57 The diagnostic results suggest that, to build a tokenizer that captures both high semantic density and fine-grained structural fidelity, we need a scalable quantization algorithm that preserves both the angular and magnitude information of high-dimensional features. 4.2 Hyper-Spherical Quantization The Unreliable Euclidean. As discussed above, representation features approximately concentrate on a hyper-spherical shell rather than a full Euclidean space. However, VQ optimizes codebooks using the ℓ2 _2 distance, as in Eq. 3. Such an objective is linearly sensitive to variations in feature norms, which biases the partitioning toward radial differences rather than angular structure. Consequently, samples with larger or varying norms can disproportionately influence the tessellation, irrespective of their directional alignment, as shown in Fig. 1. We posit that this mismatch is a primary cause of the codebook collapse phenomenon observed in Fig. 2. Figure 2: Left: Training curve under identical settings. Right: Gradient norm of codebook during training, where VQ’s collapse mode can be observed. Algorithm 1 Classic Vector Quantization 1: Input: Z∈ℝN×d,C,βZ ^N× d,C,β 2: Di,j←‖Zi−Cj‖22D_i,j←\|Z_i-C_j\|^2_2 3: Ii←argminjDi,jI_i← _jD_i,j 4: Zq←C[I]Z_q← C[I] 5: ℒcodebook←‖sg[Z]−Zq‖22L_codebook←\|sg[Z]-Z_q\|^2_2 6: ℒcommit←‖Z−sg[Zq]‖22L_commit←\|Z-sg[Z_q]\|^2_2 7: ℒVQ←ℒcodebook+βℒcommitL_VQ _codebook+ _commit 8: Return: ℒVQL_VQ Hyper-Spherical Quantization # Compute cosine similarity 3: Si,j←Zi‖Zi‖2⋅Cj‖Cj‖2S_i,j← Z_i\|Z_i\|_2· C_j\|C_j\|_2 # Angular routing 4: Ii←argmaxjSi,jI_i← _jS_i,j 5: Zq←C[I]Z_q←C[I] # Spherical codebook loss 6: ℒcodebook←1−sg[Z‖Z‖2]⋅Zq‖Zq‖2L_codebook← 1-sg[ Z\|Z\|_2]· Z_q\|Z_q\|_2 Objective Decoupling. To mitigate these challenges, we propose Hyper-Spherical Quantization (HSQ). Our core insight is to decouple the semantic orientation of a feature from its magnitude during the quantization step. Following VQRAE, we add a learnable linear projection on the raw image features, then, HSQ replaces the standard Euclidean assignment with angular routing: Ii=argmaxjZi⋅Cj‖Zi‖2‖Cj‖2.I_i= _j Z_i·C_j\|Z_i\|_2\|C_j\|_2. (5) By transitioning to an angular-based metric, we assign the codebook based on purely the directional relationships learned during pretraining. To maintain this constraint, we also rewrite the codebook loss ℒcodebookL_codebook as a spherical objective, which constrains codebook updates to the tangent space. ℒcodebook=1−sg[Z‖Z‖2]⋅Zq‖Zq‖2.L_codebook=1-sg [ Z\|Z\|_2 ]· Z_q\|Z_q\|_2. (6) How is the magnitude optimized? The magnitude scaling of code vectors is not arbitrary. In fact, replacing all objectives with angular-based methods leads to unstable and slow convergence, since the magnitude is completely unsupervised (see Sec. 6 for details). Therefore, we retain the Euclidean objective for commitment loss deliberately, as it provides magnitude-aware guidance to the projected features. The core comparison is shown above. Please refer to Append. A for the complete training algorithm, where full implementation like the straight-through estimator (STE) and codebook projection are included. 4.3 Distinction from Existing Quantization Schemes To fully appreciate our design, it is instructive to compare HSQ with two popularized quantization paradigms: Finite Scalar Quantization (FSQ) [36] and Spherical Vector Quantization (SVQ) [68]. FSQ dispenses with a learnable codebook entirely, instead bounding the continuous latents via a squashing function (e.g., tanh ) and rounding them to a predefined discrete grid. While FSQ effectively eliminates collapse by removing competitive codebook learning, it is confined to low dimensions. The implicit vocabulary size in FSQ scales exponentially with the dimension d as K=LdK=L^d, where L is the number of bins per dimension. This will yield intractable vocabulary size for high-dimensional features. In contrast, HSQ operates directly in the native high-dimensional space while maintaining a highly utilized codebook. SVQ projects latents onto a unit sphere, but it traditionally do so to impose a uniform yet strict prior on the latent space for training generative models from scratch. Consequently,it passes the normalized vectors to the decoder. Crucially, HSQ differs by explicitly decoupling the angular routing from the magnitude-aware reconstruction, which does not enforce SVQ’s unit sphere prior on the inputs. As detailed in Sec. 4.2, we only use the spherical geometry during the nearest-neighbor assignment. For the actual representation passed to the decoder (and used in the commitment loss), we retrieve the unnormalized code vector. 5 Experiment 5.1 Setups We provide a brief overview of the implementation in this section. Full experimental details and hyperparameters are included in Sec. A. Throughout the experimental section, VQ denotes vector quantization trained with the conventional ℓ2 _2 objective. Unless specified, all VQ implementations follow the SimVQ design, where the codes are projected with an linear projection layer. Implementation. Following VQRAE, we utilize a pretrained SigLIP2 ViT-So400M [55] as the image encoder, and use a ViT of symmetric design for the decoder. Unlike previous methods [14, 30] that may use two-stage training for stability and performance trade-off, dRAE is trained end-to-end with distillation loss in one stage. With no specific embedding initialization, anti-collapse trick nor curriculum learning, the training pipeline is much simplified. For image understanding, we employ Qwen2.5-7B [3] as the LLM backbone, and adapt the commonly used visual instruction tuning pipeline as in LLaVA-1.5 [32]. For text-to-image (T2I) generation, we use Flan-T5-XXL [10] for text conditioning, and a diffusion transformer with discrete prediction head [35, 44]. The generative training follows the discrete diffusion paradigm, where images are tokenized into discrete indices and masked randomly during training. The training process is optimized via a time-weighted masked cross-entropy loss. We also train several class-to-image (C2I) generative variants on ImageNet for ablative study. See A.4 for details. Evaluation Metrics. We evaluate reconstruction quality using reconstruction FID (rFID) [21], Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index Measure (SSIM) on the ImageNet-1K [11] validation set. For multimodal understanding, we evaluate on GQA [23], TextVQA (TQA) [47], MMBench-en (MMB) [33], MME-Perception (MME-P) [16], and SEEDBench-Img (SEED) [28]. For C2I generation, we evaluate using generation FID (gFID) and Inception Score (IS) [45]. For T2I generation, we evaluate on GenEval [18] and DPG-Bench [22]. 5.2 Image Reconstruction 2142^142152^152162^162172^175510101515Active Codes (×103× 10^3)VQHSQ 2142^142152^152162^162172^17202021212222PSNR ↑ 2142^142152^152162^162172^17111.51.5222.52.5rFID ↓ Figure 3: Scaling behavior with K=214∼217K=2^14 2^17 vocabulary size. From left to right: active code per batch, reconstruction PSNR, and rFID. Unlike prior observations that suggest limited gains from increasing vocabulary size [14], the proposed approach demonstrates consistent scaling behavior. As shown in Fig. 3, increasing the vocabulary size leads to higher code utilization, which correlates with improved reconstruction fidelity. As a contrast, the VQRAE (VQ) baseline demonstrates marginal improvement of active code number and PSNR when scaling up the vocabulary size. Note that the active code number is calculated by summing up the unique code usage per batch during training. Our hyper-spherical code design exhibits strong resistance to codebook collapse. Under a vocabulary size of K=65,536K=65,536 with batch size 64 per GPU, the VQRAE baseline only activates ∼ 6K codes per forward pass, whereas our method consistently more than ∼ 12K codes, while keeping a high global code utilization ratio (>90%>90\%) without stochastic sampling tricks. For HSQ, larger vocabularies also yield consistently better reconstruction performance, while VQ method’s improvement (before collapse) is relatively marginal. As can be seen in Tab 2, under the same codebook budget (16,38416,384), dRAE outperforms VQRAE with significant margins. With a 131,072131,072 codebook, dRAE achieves 0.42 rFID, and its PSNR and SSIM are also the best among the compared discrete tokenizers. Table 2: Evaluation on reconstruction quality. NA refers to direct utilization of pretrained features. Dim. indicates latent dimension (bottleneck) for autoencoders. *: See A.6 for evaluation details. Autoencoders Method Codebook Dim. rFID ↓ PSNR ↑ SSIM ↑ Low-dimensional SD-VAE [43] VAE – 32 0.49 26.10 0.72 UniLIP* [51] VAE – 32 0.79 22.99 0.67 MUSE-VL [63] VQ 3276832768 8 2.26 20.14 0.65 QLIP [69] BSQ 2282^28 28 3.21 23.16 0.63 Representational RAE [70] NA – 768 0.62 19.20 0.44 VILA-U [61] RQ 1638416384 1152 1.25 – – VQRAE* [14] VQ 1638416384 1152 1.31 22.33 0.67 dRAE HSQ 1638416384 1152 0.69 24.03 0.70 dRAE HSQ 131072131072 1152 0.42 24.52 0.72 5.3 Multimodal Tasks Image Understanding. dRAE demonstrates competitive performance on multimodal understanding benchmarks. We follow VQRAE to evaluate the tokenizer’s performance using its tuned image encoder. Note that VQRAE is trained with original encoder but tested with the tuned one, while our method uses the encoder optimized for reconstruction throughout MLLLM training and testing, consistent with other baselines. As shown in Tab. 3, while the tokenizer training is end-to-end, dRAE maximally preserved the semantic knowledge encoded in the pretrained encoder. Table 3: Evaluation on multimodal understanding. Res. for image resolution. All the models listed use ∼7 7B LLM as the backbone. Top-two best results are highlighted. Method Vision Encoder Res. GQA TQA MMB MME-P SEED w/ Understanding Only LLaVA-1.5 [32] CLIP ViT-L 336 62.0 46.1 64.3 1510.7 58.6 LLaVA-OV [27] SigLIP ViT-So 384 – – 80.8 1580.0 75.4 w/ Unified Tokenizer VILA-U [61] SigLIP ViT-So 384 60.8 60.8 - 1401.8 59.0 UniTok [34] Vitamin-L 256 - - - 1448.0 - QLIP [69] CLIP ViT-L 392 61.8 55.2 - 1498.3 - TokLIP [30] SigLIP2 ViT-So 384 59.5 - 67.6 1488.4 70.4 Tar [19] SigLIP2 ViT-So 384 61.3 - 74.4 1571.0 72.1 VQRAE [14] SigLIP2 ViT-So 512 63.6 58.8 67.6 1494.2 62.8 dRAE SigLIP2 ViT-So 384 62.1 63.8 81.4 1468.8 71.1 dRAE SigLIP2 ViT-So 512 62.4 67.7 81.5 1535.1 72.7 Visual Generation. We evaluate the T2I generation performance of dRAE in Tab. 4. Despite its lightweight architecture (∼ 700M trainable parameters) and relatively modest training scale, our generative model demonstrates competitive capabilities against baselines trained on significantly larger datasets. This underscores that semantic latents constructed from VFMs serve as a powerful and highly efficient foundation for generative downstream tasks. For C2I generation, models utilizing HSQ-derived tokenizers consistently outperform their VQ-derived counterparts across all codebook sizes, yielding substantial improvements in both gFID and Inception Score (IS), as detailed in Tab. 8. Furthermore, we integrate iREPA [48] to enhance internal representation alignment. As shown in Tab. 8, the HSQ-based approach benefits more pronouncedly from this alignment strategy. Specifically, while iREPA accelerates early convergence for both tokenizers, the VQ-based baseline suffers from an early performance plateau and noticeable training oscillations. Table 4: Evaluation on text-to-image generation. Data refers to the image-text pairs traversed during training. We leave the exhaustive data and model curation for future study. Model GenEval DPG-Bench Data Single Two Count Color Pos. Attr. Overall Overall Continuous Generation SDv1.5 [43] 0.97 0.38 0.35 0.76 0.04 0.06 0.43 63.18 >1>1B DALL-E 3 [31] 0.96 0.87 0.47 0.83 0.43 0.45 0.67 83.50 >1>1B Discrete Generation LlamaGen [50] 0.71 0.34 0.21 0.58 0.07 0.04 0.32 64.84 60M TokenFlow [40] – – – – – – 0.55 73.38 60M Show-o [62] 0.95 0.52 0.49 0.82 0.11 0.28 0.53 – 35M Meissonic [2] 0.99 0.66 0.42 0.86 0.10 0.22 0.54 – >>200M Janus [60] 0.97 0.68 0.30 0.84 0.46 0.42 0.61 79.68 >>100M dRAE 0.98 0.69 0.45 0.92 0.30 0.39 0.63 80.58 12M 6 Ablative Study Table 5: Ablation on C2I generation. Using simplest sampler without hyperparameter searching. Method Codebook rFID gFID↓ IS↑ VQ 32768 2.25 5.37 251.3 VQ 65536 2.21 5.51 264.3 HSQ 32768 2.20 4.83 268.4 HSQ 65536 2.14 4.45 287.3 Table 6: Ablation on alignment. iREPA for improved representation alignment [48]. Method Configuration gFID↓ IS↑ VQ w/o iREPA 6.89 215.3 VQ w/ iREPA 7.16 226.1 HSQ w/o iREPA 6.65 251.4 HSQ w/ iREPA 6.11 256.3 Table 7: Ablation of objective metric on reconstruction tasks. ℓ2 _2 for standard Eucildean metric, and θ for cosine-similarity based method. Assign ℒcodebookL_codebook ℒcommitL_commit rFID SSIM ℓ2 _2 ℓ2 _2 ℓ2 _2 3.59 0.62 θ ℓ2 _2 ℓ2 _2 3.31 0.62 θ θ ℓ2 _2 3.02 0.65 θ θ θ 12.7 0.33 Table 8: Ablation of tokenizer design on multimodal understanding tasks under identical visual instruction tuning setting. Method Codebook GQA MMB VQ 16384 34.1 43.8 VQ 65536 33.8 42.0 HSQ 16384 34.3 44.1 HSQ 65536 35.8 45.6 Loss Design. We evaluate the contribution of each component in our design by systematically ablating on the routing mechanism and loss formulations. Table 8 summarizes the evaluated variants. The baseline (line 1) uses Euclidean routing with ℓ2 _2-based codebook and commitment loss. Replacing Euclidean distance with cosine similarity (line 2) already improves performance, indicating the benefit of directional matching in the latent space. Building on this, our method (line 3) introduces a cosine-based ℒcodebookL_codebook while retaining an ℓ2 _2 commitment loss. The full-spherical variant (line 4) applies cosine metrics to all components. The results show that the best performance is achieved by combining a cosine-based ℒcodebookL_codebook with an ℓ2 _2 commitment loss. This allows the codebook to align with the directional structure of the latent space, while the Euclidean commitment term preserves feature magnitude and guides the projection to meet with codebook’s distribution. When both losses are defined in cosine space, performance degrades significantly, suggesting that magnitude information remains important for convergence. We also provide additional experiment results with DINOv2 as the encoder in A.3. Impact on Semantics. We further explore the use of quantized features for training multimodal models and evaluate the impact of different codebook designs on downstream understanding performance. As shown in Tab 8, performance with HSQ consistently improves as the codebook size increases, whereas VQ-based methods exhibit relatively stagnant behavior. This observation further supports our claim that semantically aligned encoding leads to superior performance.In practice, we also observe that VQ-based approaches suffer from notable instability during training: the reconstruction loss exhibits persistent oscillations, while the distillation loss tends to undergo an irreversible increase in the later stages of training, indicating a growing deviation from the pretrained teacher model. Such issues are not observed with HSQ. 7 Beyond Image Reconstruction The main experiments in this paper follow the RAE paradigm, where image reconstruction serves as a proxy task for training the tokenizer. Nevertheless, the core idea of HSQ is to learn an effective mechanism for quantizing high-dimensional representations, which should be applicable regardless of the choice of proxy task. Several recent works [17, 12, 52, 59, 39] have discussed related directions. We explore alternative proxy and report two preliminary experiments: Feature Reconstruction. With the encoder E frozen, we get the encoded image features z=E(x)z=E(x) and train a codebook together with a ViT decoder to reconstruct the features z z. The reconstruction objective is defined as ℒrec=L2(z^,z)L_rec=L_2( z,z). We train the quantizer and decoder for 100K steps and evaluate the PSNR and cosine similarity between z and z z. As shown in Tab. 9, HSQ significantly outperforms VQ. Nevertheless, reconstructing high-dimensional patch embeddings is more challenging than reconstructing pixel values, which we attribute to the intrinsic noise in the encoder representations and the absence of perceptual supervision [67], from which VAEs benefit. Table 9: Comparison of VQ and HSQ on feature reconstruction. Method PSNR ↑ Sim. ↑ VQ 4.91 0.81 HSQ 8.24 0.92 Joint Image-Semantic Modeling. We insert the quantizer between the frozen visual encoder E and the trainable language backbone DLLMD_LLM of Qwen-3.5 [41], while attaching a trainable ViT decoder DpixelD_pixel in parallel. The codebook is then trained end-to-end using both image reconstruction and image understanding objectives, allowing supervision to flow through a unified training objective. Let zqz_q denote the quantized image features, the reconstructed image x^=Dpixel(z) x=D_pixel(z), and y as the language response for query q about image x. We use 4M image-query-response pairs x,q,y\x,q,y\ from LLaVA-OneVision [27] as the training dataset. The overall objective combines the language modeling loss and the image reconstruction loss as ℒrec=−∑tlogpDLLM(yt∣y<t,q,zq)+L1Dpixel(x^,x)L_rec=- _t p_D_LLM\! (y_t y_<t,q,z_q )+L1_D_pixel( x,x). Through experiments, we found that an entropy regularization loss [8, 66] is essential for effective codebook utilization when an LLM is used as the decoder. Since this loss is already included in the IBQ family, we plug it to HSQ and VQ for a fair comparison. We compare HSQ with VQ, IBQ [46], and the ℓ2 _2-normalized variant of IBQ [17], with a vocabulary size of K=131,072K=131,072. The entropy regularization uses a temperature-scaled softmax with temperature τ=0.01τ=0.01 for all methods. As shown in Fig. 4, VQ performs the worst among all compared methods, exhibiting slow convergence and poor codebook usage. The IBQ variants benefit from stochastic sampling and consequently achieve higher code utilization. HSQ, despite not relying on randomized sampling, achieves the highest codebook utilization and the fastest convergence in both image reconstruction and understanding tasks by simply adopting angular-based routing and codebook updates. These results further demonstrate the effectiveness of HSQ for quantizing high-dimensional representations. Figure 4: Comparison of quantization methods under the joint image-semantic optimization. 8 Conclusion We presented dRAE, a Representation Autoencoder with Hyper-Spherical Codes, directly tackling the pervasive codebook collapse issue in high-dimensional discrete visual tokenizers. By aligning the codebook update dynamics more closely with the underlying semantic latent distribution, our approach achieves exceptionally high codebook utilization, fast convergence and robust reconstruction quality with minimal anti-collasping tricks. This study provides a fresh insight for training visual tokenizers, as well as laying a scalable foundation for training next-generation MLLMs. 9 Acknowledgments This work was supported by Ant Group Research Intern Program. References [1] M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, and N. Ballas (2023) Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. arXiv preprint arXiv:2301.08243. Cited by: §4.1. [2] J. Bai, T. Ye, W. Chow, E. Song, Q. Chen, X. Li, Z. Dong, L. Zhu, and S. Yan (2025) Meissonic: revitalizing masked generative transformers for efficient high-resolution text-to-image synthesis. In ICLR, Cited by: Table 4. [3] S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin (2025-02) Qwen2.5-VL Technical Report. arXiv. External Links: 2502.13923, Document Cited by: §5.1. [4] R. Balestriero, N. Ballas, M. Rabbat, and Y. LeCun (2025) Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density. arXiv preprint arXiv:2510.05949. Cited by: §4.1. [5] R. Betser, E. Gofer, M. Y. Levi, and G. Gilboa (2026) InfoNCE Induces Gaussian Distribution. In ICLR, Cited by: §4.1. [6] O. Boer Bohan (2024-06) Megalith-10m. dataset. Note: https://huggingface.co/datasets/madebyollin/megalith-10m Cited by: §A.1. [7] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin (2021) Emerging properties in self-supervised vision transformers. In ICCV, p. 9650–9660. Cited by: §2, §4.1. [8] H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M. Yang, K. P. Murphy, W. T. Freeman, M. Rubinstein, Y. Li, and D. Krishnan (2023) Muse: Text-To-Image Generation via Masked Generative Transformers. In ICML, Vol. 202, p. 4055–4075. Cited by: §7. [9] J. Chen, Z. Xu, X. Pan, Y. Hu, C. Qin, T. Goldstein, L. Huang, T. Zhou, S. Xie, S. Savarese, et al. (2025) Blip3-o: a family of fully open unified multimodal models-architecture, training and dataset. arXiv preprint arXiv:2505.09568. Cited by: §A.1. [10] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al. (2024) Scaling instruction-finetuned language models. J. Mach. Learn. Res. 25 (70), p. 1–53. Cited by: §5.1. [11] J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei (2009) ImageNet: a large-scale hierarchical image database. In CVPR, p. 248–255. External Links: Document Cited by: §5.1. [12] B. Ding, C. Chu, D. Zang, H. Li, J. Cao, K. Gai, M. Wei, R. Tang, S. Wang, S. Mao, X. Luo, Y. Liu, Z. Ling, Z. Yang, Z. Li, C. Song, G. Zhou, G. Zhang, H. Peng, H. Wang, J. Deng, J. Ouyang, J. Zhang, L. Ren, Q. Wang, Q. Hu, T. Wang, X. Wang, Y. Yang, Z. Zhang, and Z. Wang (2026) Kelix Technical Report. arXiv preprint arXiv:2602.09843. Cited by: §7. [13] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In ICLR, Cited by: §1. [14] S. Du, J. Guo, B. Li, S. Cui, Z. Xu, Y. Luo, Y. Wei, K. Gai, X. Wang, K. Wu, et al. (2025) VQRAE: representation quantization autoencoders for multimodal understanding, generation and reconstruction. arXiv preprint arXiv:2511.23386. Cited by: §A.1, §1, §2, §5.1, §5.2, Table 2, Table 3. [15] P. Esser, R. Rombach, and B. Ommer (2021) Taming Transformers for High-Resolution Image Synthesis. In CVPR, p. 12873–12883. External Links: Document Cited by: §1. [16] C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, et al. (2023) MME: a comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394. Cited by: §5.1. [17] Z. Geng, Y. Wang, Y. Ma, C. Li, Y. Rao, S. Gu, Z. Zhong, Q. Lu, H. Hu, X. Zhang, Linus, D. Wang, and J. Jiang (2025) X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again. arXiv preprint arXiv 2507.22058. Cited by: §7, §7. [18] D. Ghosh, H. Hajishirzi, and L. Schmidt (2023) Geneval: an object-focused framework for evaluating text-to-image alignment. NeurIPS 36, p. 52132–52152. Cited by: §5.1. [19] J. Han, H. Chen, Y. Zhao, H. Wang, Q. Zhao, Z. Yang, H. He, X. Yue, and L. Jiang (2025) Vision as a dialect: unifying visual understanding and generation via text-aligned representations. arXiv preprint arXiv:2506.18898. Cited by: Table 3. [20] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. B. Girshick (2022) Masked autoencoders are scalable vision learners. In CVPR, p. 15979–15988. Cited by: §1, §2. [21] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, p. 6626–6637. Cited by: §5.1. [22] X. Hu, R. Wang, Y. Fang, B. Fu, P. Cheng, and G. Yu (2024) Ella: equip diffusion models with llm for enhanced semantic alignment. arXiv preprint arXiv:2403.05135. Cited by: §5.1. [23] D. A. Hudson and C. D. Manning (2019) Gqa: a new dataset for real-world visual reasoning and compositional question answering. In CVPR, p. 6700–6709. Cited by: §5.1. [24] D. P. Kingma and M. Welling (2014) Auto-encoding variational bayes. In ICLR, Cited by: §1. [25] M. Ledoux (2001) The concentration of measure phenomenon. American Mathematical Soc.. Cited by: §4.1. [26] M. Y. Levi and G. Gilboa (2025) The double-ellipsoid geometry of CLIP. In ICML, Vol. 267, p. 33999–34019. Cited by: §4.1. [27] B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al. (2024) Llava-onevision: easy visual task transfer. arXiv preprint arXiv:2408.03326. Cited by: §A.1, Table 3, §7. [28] B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan (2023) Seed-bench: benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125. Cited by: §5.1. [29] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. Wen (2023) Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355. Cited by: §4.1. [30] H. Lin, T. Wang, Y. Ge, Y. Ge, Z. Lu, Y. Wei, Q. Zhang, Z. Sun, and Y. Shan (2025) TokLIP: marry visual tokens to CLIP for multimodal comprehension and generation. CoRR abs/2505.05422. Cited by: Table 14, §2, §5.1, Table 3. [31] Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan (2024) Evaluating text-to-visual generation with image-to-text generation. In ECCV, p. 366–384. Cited by: Table 4. [32] H. Liu, C. Li, Y. Li, and Y. J. Lee (2024) Improved baselines with visual instruction tuning. In CVPR, p. 26296–26306. Cited by: §A.1, Table 14, §4.1, §5.1, Table 3. [33] Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al. (2024) Mmbench: is your multi-modal model an all-around player?. In ECCV, p. 216–233. Cited by: §4.1, §5.1. [34] C. Ma, Y. Jiang, J. Wu, J. Yang, X. Yu, Z. Yuan, B. Peng, and X. Qi (2025) Unitok: a unified tokenizer for visual generation and understanding. arXiv preprint arXiv:2502.20321. Cited by: Table 14, Table 3. [35] T. Ma, X. Zhang, B. Yang, J. Feng, and Q. Ye (2026) ReDDiT: Rehashing Noise for Discrete Visual Generation. In ICLR, Cited by: §5.1. [36] F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen (2024) Finite scalar quantization: VQ-VAE made simple. In ICLR, Cited by: §4.3. [37] M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P. Huang, S. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jégou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski (2024) DINOv2: learning robust visual features without supervision. Trans. Mach. Learn. Res. 2024. Cited by: §A.3, §1, §2, §4.1. [38] W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In ICCV, p. 4172–4182. Cited by: §1. [39] W. Peng, L. Meng, Y. Cai, X. Zhuang, Y. Yang, R. Fang, C. Wu, J. Lin, Z. Wu, and S. Bai (2026) Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification. arXiv preprint arXiv:2606.18249. Cited by: §7. [40] L. Qu, H. Zhang, Y. Liu, X. Wang, Y. Jiang, Y. Gao, H. Ye, D. K. Du, Z. Yuan, and X. Wu (2025) Tokenflow: unified image tokenizer for multimodal understanding and generation. In CVPR, p. 2545–2555. Cited by: Table 4. [41] Qwen Team (2026-02) Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §7. [42] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In ICML, p. 8748–8763. Cited by: §1, §2, §4.1. [43] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In CVPR, p. 10674–10685. Cited by: §1, §2, Table 2, Table 4. [44] S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V. Kuleshov (2024) Simple and Effective Masked Diffusion Language Models. In NeurIPS, Cited by: §A.5, §5.1. [45] T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen (2017) Improved techniques for training gans. In NeurIPS, p. 2226–2234. Cited by: §5.1. [46] F. Shi, Z. Luo, Y. Ge, Y. Yang, Y. Shan, and L. Wang (2025) Scalable Image Tokenization with Index Backpropagation Quantization. arXiv preprint arXiv:2412.02692. Cited by: §7. [47] A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach (2019) Towards vqa models that can read. In CVPR, p. 8317–8326. Cited by: §5.1. [48] J. Singh, X. Leng, Z. Wu, L. Zheng, R. Zhang, E. Shechtman, and S. Xie (2025) What matters for Representation Alignment: Global Information or Spatial Structure?. arXiv preprint arXiv:2512.10794. Cited by: §5.3, Table 8, Table 8. [49] J. Singh, B. Zheng, Z. Wu, R. Zhang, E. Shechtman, and S. Xie (2026) Improved baselines with representation autoencoders. arXiv preprint arXiv:2605.18324. Cited by: §A.3. [50] P. Sun, Y. Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan (2024) Autoregressive model beats diffusion: llama for scalable image generation. arXiv preprint arXiv:2406.06525. Cited by: Table 4. [51] H. Tang, C. Xie, X. Bao, T. Weng, P. Li, Y. Zheng, and L. Wang (2025) Unilip: adapting clip for unified multimodal understanding, generation and editing. arXiv preprint arXiv:2507.23278. Cited by: §A.1, Table 2. [52] M. L. Team (2026) LongCat-Next: Lexicalizing Modalities as Discrete Tokens. arXiv preprint arXiv:2603.27538. Cited by: §7. [53] S. Tong, D. Fan, J. Nguyen, E. Brown, G. Zhou, S. Qian, B. Zheng, T. Vallaeys, J. Han, R. Fergus, N. Murray, M. Ghazvininejad, M. Lewis, N. Ballas, A. Bar, M. Rabbat, J. Verbeek, L. S. Zettlemoyer, K. Sinha, Y. LeCun, and S. Xie (2026) Beyond language modeling: an exploration of multimodal pretraining. arXiv preprint arXiv:2603.03276. Cited by: §1, §2. [54] S. Tong, B. Zheng, Z. Wang, B. Tang, N. Ma, E. Brown, J. Yang, R. Fergus, Y. LeCun, and S. Xie (2026) Scaling text-to-image diffusion transformers with representation autoencoders. arXiv preprint arXiv:2601.16208. Cited by: §1. [55] M. Tschannen, A. A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, O. J. Hénaff, J. Harmsen, A. Steiner, and X. Zhai (2025) SigLIP 2: multilingual vision-language encoders with improved semantic understanding, localization, and dense features. CoRR abs/2502.14786. Cited by: §1, §2, §5.1. [56] A. Van Den Oord, O. Vinyals, et al. (2017) Neural discrete representation learning. NeurIPS 30. Cited by: Figure 5, Figure 5, §1. [57] R. Vershynin (2018) High-dimensional probability: an introduction with applications in data science. Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press. External Links: ISBN 978-1-108-41519-4, Document Cited by: §4.1. [58] T. Wang and P. Isola (2020) Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In ICML, Vol. 119, p. 9929–9939. Cited by: Figure 1, Figure 1, §1, §4.1. [59] Y. Wang, Z. Lin, C. Yang, Y. Zhao, F. Xiao, H. He, Q. Zhao, Z. Ding, F. Wang, S. Wang, Y. Zhang, H. Fan, and X. Liu (2026) Representation Forcing for Bottleneck-Free Unified Multimodal Models. arXiv preprint arXiv:2605.31604. Cited by: §7. [60] C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan, et al. (2024) Janus: decoupling visual encoding for unified multimodal understanding and generation. arXiv preprint arXiv:2410.13848. Cited by: Table 4. [61] Y. Wu, Z. Zhang, J. Chen, H. Tang, D. Li, Y. Fang, L. Zhu, E. Xie, H. Yin, L. Yi, S. Han, and Y. Lu (2025) VILA-U: a unified foundation model integrating visual understanding and generation. In ICLR, Cited by: Table 14, §1, §2, Table 2, Table 3. [62] J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou (2024) Show-o: one single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528. Cited by: Table 4. [63] R. Xie, C. Du, P. Song, and C. Liu (2025) Muse-vl: modeling unified vlm through semantic discrete encoding. In ICCV, p. 24135–24146. Cited by: Table 2. [64] J. Yao, Y. Song, Y. Zhou, and X. Wang (2025) Towards scalable pre-training of visual tokenizers for generation. arXiv preprint arXiv:2512.13687. Cited by: §2. [65] J. Yao, B. Yang, and X. Wang (2025) Reconstruction vs. generation: taming optimization dilemma in latent diffusion models. In CVPR, p. 15703–15712. Cited by: §1, §2. [66] L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y. Cheng, V. Birodkar, A. Gupta, X. Gu, A. G. Hauptmann, B. Gong, M. Yang, I. Essa, D. A. Ross, and L. Jiang (2024) Language Model Beats Diffusion – Tokenizer is Key to Visual Generation. arXiv preprint arXiv:2310.05737. Cited by: §7. [67] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. arXiv preprint arXiv:1801.03924. Cited by: §7. [68] Y. Zhao, H. Jiang, Z. Xu, C. Yang, E. Adeli, and P. Krähenbühl (2025) Spherical leech quantization for visual tokenization and generation. CoRR abs/2512.14697. Cited by: §4.3. [69] Y. Zhao, F. Xue, S. Reed, L. Fan, Y. Zhu, J. Kautz, Z. Yu, P. Krähenbühl, and D. Huang (2025) Qlip: text-aligned visual tokenization unifies auto-regressive multimodal understanding and generation. arXiv preprint arXiv:2502.05178. Cited by: Table 14, Table 2, Table 3. [70] B. Zheng, N. Ma, S. Tong, and S. Xie (2025) Diffusion transformers with representation autoencoders. CoRR abs/2510.11690. Cited by: §1, §2, Table 2. [71] Y. Zhu, B. Li, Y. Xin, Z. Xia, and L. Xu (2025) Addressing representation collapse in vector quantized models with one linear layer. In ICCV, p. 22968–22977. Cited by: §1, §2, §3. [72] K. Zou, Z. Zhao, B. Liu, and N. Yu (2026) Advancing aesthetic image generation via composition transfer. Int. J. Comput. Vis. 134, p. 252. External Links: Document Cited by: §A.1. Appendix A Implementation Details A.1 Datasets dRAE is trained to reconstruct the the open-sourced 36M images of BLIP3-o [9], following UniLIP [51] and VQRAE [14]. For image understanding, we follow the alignment and pretraining setting of LLaVA-1.5 [32], and conduct instruction tuning [27]. For T2I generation, we use 12M image-text pairs, including Megalith-10M [6] and T2I-2M [72] for generative pretraining, and BLIP3-o-60K [9] for instruction fine-tuning. A.2 Tokenizer Training Algorithm The complete training algorithm for HSQ is demonstrated in Alg. 2. We also provide an illustration in Fig 5. The training is conducted on 8×80G8× 80G NVIDIA GPUs for 3 days. Algorithm 2 Hyper-Spherical Quantization 1: Input: Feature tensor Z∈ℝN×dZ ^N× d. Base codebook embeddings C, learnable projection W, and commitment weight β 2: C^←W C← WC # Project codebook following SimVQ 3: Si,j←Zi‖Zi‖2⋅C^j‖C^j‖2S_i,j← Z_i\|Z_i\|_2· C_j\| C_j\|_2 # Compute cosine similarity 4: Ii←argmaxjSi,jI_i← _jS_i,j # Angular routing 5: Zq←C^[I]Z_q← C[I] # Lookup on projected codebook 6: ℒcodebook←1−sg[Z‖Z‖2]⋅Zq‖Zq‖2L_codebook← 1-sg[ Z\|Z\|_2]· Z_q\|Z_q\|_2 # Angular codebook loss 7: ℒcommit←‖Z−sg[Zq]‖22L_commit←\|Z-sg[Z_q]\|^2_2 # Commitment loss 8: ℒVQ←ℒcodebook+βℒcommitL_VQ _codebook+ _commit 9: Zout←Z+sg[Zq−Z]Z_out← Z+sg[Z_q-Z] # Straight-through estimator 10: Return: Zout,ℒVQ,IZ_out,L_VQ,I Figure 5: A demonstration of optimizing the codebook embeddings, following VQ-VAE [56]. Hyperparameter The training details of dRAE are provided in Tab. 13. The perceptual loss design follows the implementation of VQRAE, while the discriminator design follows RAE. Table 10: Ablation on reconstruction tasks with DINOv2-B Encoder. Assign ℒcodebookL_codebook ℒcommitL_commit rFID SSIM L2L_2 L2L_2 L2L_2 3.67 0.56 θ L2L_2 L2L_2 4.23 0.57 θ θ L2L_2 2.87 0.58 Table 11: Ablation on reconstruction tasks with RAEv2-style DINOv2-B Encoder. Assign ℒcodebookL_codebook ℒcommitL_commit rFID SSIM L2L_2 L2L_2 L2L_2 4.24 0.56 θ L2L_2 L2L_2 4.11 0.58 θ θ L2L_2 3.29 0.60 Table 12: Details for tokenizer training. Category Hyperparameter Value Architecture Image Encoder SigLIP2-So400M Encoder Input Size 512 Decoder Type ViT-XL Training Global Steps 150000 Global Batchsize 512 Optimizer AdamW (0.9,0.950.9,0.95) Learning Rate 2.0×10−42.0× 10^-4 Encoder Learning Rate 2.0×10−52.0× 10^-5 Weight Decay 0.0 EMA Decay 0.9978 LR Scheduler Type Cosine Warmup Epochs 1 Final Learning Rate 2.0×10−52.0× 10^-5 Discriminator Disc. Learning Rate 5.0×10−55.0× 10^-5 Disc. Loss Type Hinge Gen. Loss Type Vanilla Loss Weights Perceptual Weight ωp _p 1.0 Discriminator Weight ωd _d 0.1 Distillation Weight λ 1.0 Table 13: Details for generative models. Category Hyperparameter Value Architecture Text Encoder T5-XXL Max Text Length 64 Image Resolution 384 Decoder Type DiT-H Codebook 65536 Training Global Steps 125000 Global Batchsize 512 t Distribution Logit Norm Optimizer AdamW (0.9,0.950.9,0.95) Learning Rate 2.0×10−42.0× 10^-4 Weight Decay 0.0 EMA Decay 0.9999 LR Scheduler Type Constant Warmup Epochs 1 Inference Steps 5050 Schedule Cosine CFG 3.53.5 A.3 Additional Results We set a variant of dRAE at 256×256256× 256 resolution, using DINOv2-B [37] as the encoder, a 12 layer ViT as the decoder, and train it for 2000020000 steps with 512 global batch size. The ablation results are shown in Tab. 13. We also follow RAEv2 [49] and test a variant that adds multiple layer features of the encoder for better reconstruction quality. Specifically, we use the sum of hidden states from DINOv2’s 3,6,9,12 layers. The corresponding results are shown in Tab. 13. These experiments indicate that HSQ can also effectively capture the distribution of self-supervised visual encoders. A.4 MLLM Training Data usage and fair comparison For the main results table, the MLLM is trained using a combination of the publicly available datasets from LLaVA-1.5 and 1M randomly sampled image-text pairs from LLaVA-OneVision. We additionally train a controlled variant using a 384-resolution encoder, Vicuna-7B as the LLM backbone, and only the LLaVA-1.5 data. The comparison with other academic works is shown in Tab. 14. Under this more standardized training setup, our model—built upon an encoder fine-tuned for reconstruction—still demonstrates advantages across multiple comprehension metrics, further validating the potential of dRAE as a unified tokenizer. Table 14: Evaluation on multimodal understanding. All models use Vicuna-7B as the backbone and only LLaVA-1.5 data. Method Vision Encoder Res. POPE GQA TQA MMB SEED LLaVA-1.5 [32] CLIP ViT-L 336 85.9 62.0 46.1 64.3 58.6 VILA-U [61] SigLIP ViT-So 384 85.8 60.8 60.8 - 59.0 UniTok [34] Vitamin-L 256 81.7 - - - - QLIP [69] CLIP ViT-L 392 86.1 61.8 55.2 - - TokLIP [30] SigLIP2 ViT-So 384 84.1 59.5 - 67.6 70.4 dRAE SigLIP2 ViT-So 384 85.5 61.0 60.1 71.7 65.0 A.5 Generative Model Training Discrete Diffusion Model (DDM) DDM defines a forward process on discrete variables by gradually corrupting tokens to absorbing state m through a continuous-time Markov process [44]. We denote clean data as xt=0x_t=0 (x0x_0 for short), and noise it gradually as t→1t→ 1. Let αt _t be the noise scheduler (a monotonically decreasing survival function that satisfies α0=1,α1=0 _0=1, _1=0 ), the corrupted data distribution at time t is determined as xt∼q(xt|x0,t),q(xt|x0,t)=Cat(xt;αtx0+(1−αt)).x_t q(x_t|x_0,t),q(x_t|x_0,t)=Cat(x_t; _tx_0+(1- _t)m). (7) Let δ(x(t,i),m)δ(x_(t,i),m) be the indicator function that is activated only if the i-th position of xtx_t is m. For a linear scheduler, the objective is derived as the evidence lower bound (ELBO) of logπθ(x0|xt) _θ(x_0|x_t): ℒDDM=−t,x0,xt[1t∑i=1Lδ(x(t,i),)logπθ(x(0,i)|xt)]=−t,x0,xt[ℓπθ(xt,x0)].L_DDM=-E_t,\ x_0,\ x_t[ 1t _i=1^Lδ(x_(t,i),m) _θ(x_(0,i)|x_t)]=-E_t,\ x_0,\ x_t[ _ _θ(x_t,x_0)]. (8) For conditional generation where a prompt c is given, we write ℓπθ(xt,x0|) _ _θ(x_t,x_0|c) for simplicity. Following MDLM’s deduction, assume that the network can reconstruct x0x_0 perfectly, we use πθ(xt) _θ(x_t) to approximate this denoising process, and get the sampling rule as pθ(xs|xt)=1,if xs=xt,xt≠,1−αs1−αt,if xs=,xt=,αs−αt1−αtπθ(xt),if xs≠,xt=,0,otherwise.p_θ(x_s|x_t)= cases1,&if x_s=x_t,\ x_t ,\\ 1- _s1- _t,&if x_s=m,\ x_t=m,\\ _s- _t1- _t _θ(x_t),&if x_s ,\ x_t=m,\\ 0,&otherwise.\\ cases (9) Hyperparameter The details of T2I generative training are provided in Tab. 13. Since a resolution of 512 requires 1024 tokens per image, the batch size becomes constrained on each machine. Therefore, we additionally train a dRAE with SigLIP2 ViT-large-patch16-384, which compresses each image into 576 tokens, enabling more efficient generative training. In addition, since a large vocabulary also significantly increases the memory footprint of the output layer, we set the tokenizer vocabulary size to 65,536. The training is conducted on 16 GPUs for 3 days. For C2I generation, we use SigLIP2 ViT-base-patch16-256 as the visual encoder, and train the tokenizer at 256 resolution, which means compressing each image into 256 tokens. The generative model is a 12-layer transformer. The results in Tab. 8 are gained after training 300 epochs on ImageNet. The ablation in Tab. 8 are gained with 65536-vocab tokenizers after 100 epochs, using the default iREPA design for B-size DiT models. A.6 Evaluation Reconstruction We follow the evaluation protocol of RAE and assess reconstruction performance on the ImageNet validation set. During the reproduction of prior work, we observed that the SSIM evaluation in UniLIP contains a bug: the data-range passed to the evaluator is incorrectly set to 2.02.0 instead of the correct value of 1.01.0 corresponding to the actual image dynamic range, leading to an overestimation of SSIM by ∼10% 10\%. The VQRAE implementation, based on the same codebase, suffers from the same issue. Accordingly, we correct the reported SSIM values for both methods based on our experiments. Understanding We use the official script in the repository of LLaVA-1.5 for evaluation. Generation We follow the standard protocol of GenEval and DPG-Bench, which samples 4 images for each prompt and use their evaluation model to score. The images are sampled with 50 steps using a cosine schedule and a constant classifier free guidance (CFG) of 3.5.