Paper deep dive
StyleStream: Real-Time Zero-Shot Voice Style Conversion
Yisi Liu, Nicholas Lee, Gopala Anumanchipalli
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 3:08:37 PM
Summary
The paper introduces StyleStream, a real-time, zero-shot voice style conversion system that disentangles linguistic content from style attributes (timbre, accent, emotion). It utilizes a Destylizer with text supervision and a compact information bottleneck to extract content features, and a Stylizer based on a diffusion transformer (DiT) to generate stylized speech. The system achieves state-of-the-art performance with an end-to-end latency of 1 second.
Entities (10)
Relation Signals (8)
StyleStream → hascomponent → Destylizer
confidence 95% · StyleStream consists of two components: a Destylizer... and a Stylizer
StyleStream → hascomponent → Stylizer
confidence 95% · StyleStream consists of two components: ... and a Stylizer
Stylizer → usesarchitecture → Diffusion Transformer
confidence 92% · Stylizer, a diffusion transformer (DiT) that reintroduces target style
Destylizer → usesmodel → HuBERT-Large
confidence 90% · It consists of a frozen HuBERT-Large encoder
Destylizer → usestechnique → Finite Scalar Quantization
confidence 90% · apply finite scalar quantization (FSQ) as an information bottleneck
StyleStream → outperforms → CosyVoice 2
confidence 88% · StyleStream delivers state-of-the-art conversion quality, substantially improving accent and emotion similarity over prior work.
StyleStream → outperforms → Vevo
confidence 88% · StyleStream delivers state-of-the-art conversion quality... over prior work.
StyleStream → trainedon → Emilia
confidence 85% · the Stylizer is trained using the English portion of Emilia
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work has explored this problem, conversion quality remains limited, and real-time voice style conversion has not been addressed. We propose StyleStream, the first streamable zero-shot voice style conversion system that achieves state-of-the-art performance. StyleStream consists of two components: a Destylizer, which removes style attributes while preserving linguistic content, and a Stylizer, a diffusion transformer (DiT) that reintroduces target style conditioned on reference speech. Robust content-style disentanglement is enforced through text supervision and a highly constrained information bottleneck. This design enables a fully non-autoregressive architecture, achieving real-time voice style conversion with an end-to-end latency of 1 second. Samples and real-time demo: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2602.20113v1
- Canonical: https://arxiv.org/abs/2602.20113v1
Trouble viewing inline? Open PDF directly →
Full Text
56,075 characters extracted from source content.
Expand or collapse full text
StyleStream: Real-Time Zero-Shot Voice Style Conversion Yisi Liu 1 , Nicholas Lee 1 , Gopala Anumanchipalli 1 1 UC Berkeley, USA louisliu@berkeley.edu Abstract Voice style conversion aims to transform an input utterance to match a target speaker’s timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic con- tent from style. While prior work has explored this problem, conversion quality remains limited, and real-time voice style conversion has not been addressed. We propose StyleStream, the first streamable zero-shot voice style conversion system that achieves state-of-the-art performance. StyleStream con- sists of two components: a Destylizer, which removes style attributes while preserving linguistic content, and a Stylizer, a diffusion transformer (DiT) that reintroduces target style condi- tioned on reference speech. Robust content-style disentangle- ment is enforced through text supervision and a highly con- strained information bottleneck. This design enables a fully non-autoregressive architecture, achieving real-time voice style conversion with an end-to-end latency of 1 second. Samples and real-time demo: https://berkeley-speech-group. github.io/StyleStream/. Index Terms: voice style conversion, real-time, zero-shot 1. Introduction Zero-shot voice style conversion seeks to modify an input ut- terance so that it reflects the timbre, accent, and emotion of an unseen target speaker (collectively defined in this work as voice style) while preserving the original linguistic content, us- ing only a few seconds of reference speech. Although zero-shot voice style cloning has been extensively studied in the text-to- speech (TTS) domain [1, 2, 3, 4, 5, 6, 7], progress on zero-shot voice style conversion remains limited. In TTS, the linguistic content is taken directly from text, which is fully disentangled from the target speech that provides voice style. In contrast, zero-shot voice style conversion requires the model to first ex- tract the content features from the source utterance, and then generate speech in the target style. This content extraction re- lies on clean content-style disentanglement, which is the core challenge in speech-to-speech conversion. Existing strategies for disentanglement include information bottleneck [8, 9, 10, 6, 7, 11], signal perturbation [12, 13, 14, 15], mutual information minimization [16, 17], etc. While these approaches achieve reasonable content-timbre disentanglement, the extracted features often remain entangled with accent and emotion. For example, CosyVoice 2 [7] combines information bottlenecks with ASR for supervised disentanglement, yet its “semantic tokens” still encode considerable accent and emotion information due to the large codebook size (6561), as demon- strated in Section 6.3. The current state-of-the-art framework, Vevo [11], employs VQ-VAE [18] to quantize HuBERT [19] features with a compact codebook of size 32, producing dis- crete content tokens. Although Vevo demonstrates good con- version performance, its purely self-supervised training pro- vides no guarantee on what information is discarded by the bot- tleneck, and linguistic content is prone to degradation during quantization, leading to suboptimal intelligibility and conver- sion quality (see Section 6). Beyond disentanglement, another open challenge is real-time application: while real-time voice conversion (timbre only) has been widely studied [20, 21, 22], there is no existing work that addresses real-time voice style conversion (timbre, accent and emotion), leaving an important gap. Motivated by these challenges, we present StyleStream, a novel framework that, for the first time, supports zero-shot real- time voice style conversion across timbre, accent, and emotion, achieving state-of-the-art performance. As shown in Figure 1, the Destylizer first extracts content features that are disentan- gled from style information. The Stylizer then takes these con- tent features and, conditioned on the target speech, generates the stylized mel-spectrogram, which is subsequently converted into waveform by the vocoder. To achieve better content-style disentanglement, inspired by CosyVoice 2 [7], we train the Destylizer with an ASR loss and apply finite scalar quantization (FSQ) [23] as an informa- tion bottleneck, but unlike CosyVoice 2, where a large codebook size (6561) is used, we constrain the codebook size to 45. This combination of text supervision and a compact codebook en- ables cleaner disentanglement. The choice of content features is also crucial: rather than using the discrete codes directly, we adopt the continuous representations immediately preceding the FSQ layer as content features, following the intuition of SoftVC [24]. As shown in Section 6.4, this design is another critical fac- tor for effective disentanglement. For the Stylizer, we train a diffusion transformer (DiT) [25] with a spectrogram inpainting objective, similar to [1, 2, 3]. Since no autoregressive modules are involved, the input and out- put lengths remain identical, which makes streaming straight- forward and avoids lag or overlap caused by input-output length mismatches. In summary, our contributions include: • We introduce StyleStream, the first real-time voice style con- version system with an end-to-end latency of 1 second. • By combining ASR loss and a compact quantization code- book, the Destylizer achieves cleaner content-style disentan- glement compared to existing methods. • Trained on 50k hours of English data, StyleStream delivers state-of-the-art conversion quality, substantially improving accent and emotion similarity over prior work. arXiv:2602.20113v1 [cs.SD] 23 Feb 2026 DestylizerStylizerVocoder SourceTargetContentConverted “Kids are talking by the door.” “Kids are talking by the door!” Figure 1: System overview of StyleStream. The Destylizer extracts content features disentangled from style, and the Stylizer generates speech that preserves the source linguistic content while adopting the target timbre, accent, and emotion. Round Destylizer Content FSQ Down FSQ Up HuBERT Large Conformer Blocks ASR Decoder Kids are talking by the door. Figure 2: Destylizer architecture. The Destylizer, as part of the ASR encoder, is trained with a sequence-to-sequence ASR loss. The continuous representations immediately before the FSQ module are taken as content features. 2. Related Work 2.1. Voice Style Cloning Voice style cloning aims to reproduce a target speaker’s tim- bre, accent, and emotion from reference speech. A large body of prior work has explored this problem in the text-to-speech (TTS) domain, where textual input provides a clean and natu- rally disentangled representation of linguistic content. Modern TTS-based voice style cloning approaches can be broadly cat- egorized into two architectural classes: (1) non-autoregressive models, which primarily rely on diffusion transformers trained with feature inpainting objectives [1, 2, 3, 4]; and (2) autore- gressive models, which condition on text as a prefix and gener- ate speech tokens using next-token prediction [26, 5, 6, 7, 27]. These methods have demonstrated impressive synthesis quality and style fidelity, but their reliance on text limits their applica- bility to speech-to-speech conversion scenarios. In contrast, speech-to-speech voice style conversion op- erates directly on acoustic signals and thus requires explicit content–style disentanglement. Most prior work in this area focuses on transferring individual style attributes in isolation, such as timbre [8, 28, 29, 30], accent [31, 32, 33], or emotion [34, 35, 36]. While effective for targeted attribute transfer, these methods do not enable holistic voice style cloning that jointly matches the target speaker’s timbre, accent, and emotion. Only a limited number of works attempt to jointly clone timbre, accent, and emotion in a unified speech-to-speech set- ting. Among existing approaches, Vevo [11] is, to our knowl- edge, the only system that explicitly targets holistic voice style cloning. However, Vevo operates in an offline setting, relies on non-streamable architectures, and exhibits a clear tradeoff between content preservation and style fidelity. These limita- tions highlight the difficulty of achieving clean content-style disentanglement while maintaining high-quality and streamable speech generation. In this work, we address these challenges by proposing a streamable speech-to-speech framework that performs holistic voice style cloning, jointly transferring timbre, accent, and emo- tion from reference speech while preserving high intelligibility. 2.2. Speech Content Disentanglement A key difficulty in speech-to-speech conversion lies in isolating linguistic content from the input speech. Research on speech content disentanglement generally falls into a few main cate- gories: (1) Information bottleneck: one line of work introduces a quantization layer [24, 37, 31] or a low-dimensional hidden representation [38, 39, 40], often combined with proxy tasks such as ASR [34, 6, 7, 41] or autoencoding [8, 9, 11], to en- courage the separation of content from style; (2) Signal per- turbation: another family of methods [12, 13, 14, 29] modifies style-related cues in the signal (e.g., pitch randomization, for- mant shifting) while preserving linguistic content, and trains the model to produce similar representations for the original and perturbed inputs; (3) Training loss design: alternative ap- proaches incorporate objectives such as GAN loss [28, 42, 43], mutual information loss [16, 17, 44], or gradient reversal loss [10, 45] to explicitly promote disentanglement. However, con- tent features extracted by these methods often still leak accent and emotion information, which degrades the quality of voice style conversion. 2.3. Real-Time Voice Conversion The field of real-time voice conversion has seen rapid progress in both quality and latency [22, 21, 20, 46]. For example, RT- VC [20] achieves state-of-the-art real-time zero-shot voice con- version quality with a CPU latency of only 61.4ms. However, no existing work has addressed the problem of real-time zero- shot voice style conversion, where the goal is to modify not only timbre but also accent and emotion. This gap motivates our work, where we present StyleStream, the first system for real-time zero-shot voice style conversion. 3. Method 3.1. Destylizer: Content-Style Disentanglement To achieve clean content-style disentanglement, we introduce the Destylizer (Figure 2). It consists of a frozen HuBERT-Large encoder [19] followed by several conformer blocks [47]. Simi- lar to CosyVoice 2 [7], the Destylizer and the FSQ module to- gether form the ASR encoder. The continuous output features of the Destylizer, denoted as content features f c , are first pro- jected down to a D-dimensional space, and each dimension is quantized to the nearest integer in the range [−K i ,K i ], where V i = 2K i + 1, K i ∈ Z >0 (1) denotes the number of codes for the i-th dimension. The quan- tized low-dimensional features are then projected back to the original dimension and passed to the ASR decoder to predict text character tokens. The whole pipeline is trained end-to-end with a sequence-to-sequence ASR loss. At inference time, we use only the Destylizer to extract disentangled content features, operating at a sampling rate of 50 Hz. There are three keys to successful content-style disentan- glement: Text Supervision. Unlike Vevo [11], whose content tok- enizer is trained in a purely self-supervised manner, we fol- low CosyVoice 2 and impose text supervision for training the Destylizer. This directs linguistic content through the FSQ bot- tleneck, while style is suppressed since it does not aid ASR pre- diction and the bottleneck provides limited capacity. In contrast, Vevo relies on a reconstruction objective with no explicit cue for separating content from style, causing style leakage and partial loss of linguistic content, which degrades disentanglement qual- ity, as shown in Section 6. Compact Codebook. Unlike CosyVoice 2, where a large codebook size (6561) is used, we use a much smaller codebook to enforce a narrower information bottleneck, following the in- tuition of AutoVC [8]. Specifically, the vocabulary size V for FSQ levels [V 1 ,V 2 ,...,V D ] is given by: V = D Y i=1 V i (2) Adjusting D and V i controls the bottleneck width. As demon- strated in Section 6.3, CosyVoice 2’s “semantic tokens” remain largely entangled with style information, whereas our Destyl- izer provides much cleaner content features due to a compact codebook. Continuous Pre-quantization Features. Unlike Vevo [11] or CosyVoice 2 [7], which use discrete tokens as content, we in- stead adopt the continuous representations immediately before FSQ, inspired by SoftVC [24]. SoftVC demonstrated that such “soft units” strike an effective balance between disentanglement and content preservation. Moreover, as shown in Section 6.3, our continuous features contain even less style information than the discrete tokens of Vevo and CosyVoice 2, while using FSQ indices directly as content leads to unintelligible speech. 3.2. Stylizer: Stylized Acoustic Modeling After obtaining disentangled content features, the Stylizer gen- erates mel-spectrograms conditioned on the target style. As shown in Figure 3, it consists of two components: Diffusion Transformer. We use a diffusion transformer (DiT) [25] as the backbone of the Stylizer, as it has demon- strated strong performance in in-context voice style cloning [1, 2, 3]. The Stylizer is trained with a spectrogram in-painting objective: given a temporal binary mask m, the model recon- structs the masked segment m⊙x 1 of a mel-spectrogram x 1 ∈ R F×T , conditioned on the unmasked context (1− m)⊙ x 1 , the content features f c ∈ R D c ×T , and a style embedding e (de- scribed below), where⊙ denotes elementwise multiplication. Concretely, the noisy mel-spectrogram x t , context (1 − m) ⊙ x 1 and content features f c are concatenated along the channel dimension to form the DiT input. To ensure compati- bility, the mel-spectrogram is calculated at the same rate as f c (50 Hz). The flow time step t∼U[0, 1] is embedded with a si- nusoidal positional encoding, and added to the style embedding, which is then integrated into DiT via adaLN-Zero [25]. During training, we adopt the conditional flow matching (CFM) loss with an Optimal Transport (OT) path formulation. Let x 1 ∼ q(x) denote the ground-truth mel-spectrogram drawn from the data distribution, and x 0 ∼ p(x) = N(0,I) be the standard normal prior. We sample a time step t ∼ U[0, 1] and define the OT flow path as ψ t (x) = (1− t)x + tx 1 . Conse- quently, the noisy mel-spectrogram state at time t is generated as x t = ψ t (x 0 ). The predicted vector field is defined as: ˆv = v θ ψ t (x 0 ),t; f c , (1− m)⊙ x 1 , e (3) where m is the binary mask indicating the generation target, f c represents the content features, and e denotes the style embed- ding. The training objective is to minimize the difference be- tween ˆv and the target velocity, computed strictly on the masked regions: L CFM (θ) = E t,x 1 ,x 0 ,m m⊙ ˆv− d dt ψ t (x 0 ) 2 (4) During inference, we concatenate the content features of the target and source utterances along the time axis. The target utterance provides the mel-spectrogram context and the style embedding, while the source region is masked and generated by the Stylizer. To balance diversity with fidelity, we employ Classifier-Free Guidance (CFG): v θ,CFG = v θ (x t ,t;c) + α(v θ (x t ,t;c)− v θ (x t ,t; ∅))(5) where ∅ denotes null condition, c denotes all conditions, in- cluding f c , (1− m)⊙ x 1 ,e, and α is the CFG strength. Style Encoder. For style conditioning, we introduce a style encoder to capture global style attributes. Its architecture follows WavLM-TDNN 1 : representations from frozen WavLM layers [48] are aggregated using learnable coefficients, and the aggregated features are then passed to a Time Delay Neural Net- work (TDNN) [49]. Attentive statistics pooling [50] is applied to obtain the final style embedding, which is fused into DiT through adaLN-zero. The style encoder is trained jointly with the DiT in an end-to-end manner. 3.3. Vocoder We trained a causal vocoder to synthesize 16 kHz speech from 50 Hz mel-spectrograms for real-time inference. The design follows Vocos [51], with all convolution layers replaced by causal convolution. The training objective combines GAN loss, 1 https://huggingface.co/microsoft/ wavlm-base-plus-sv Source ContentTarget Content t Style Encoder Destylizer Diffusion Transformer Discarded Discarded To Generate Target Style TargetSource Figure 3: Stylizer architecture. The Stylizer contains a style encoder and a diffusion transformer. reconstruction loss, and feature matching loss, using the same configurations as Vocos. The original Vocos checkpoint 2 is used as a warm start. 3.4. Real-Time Design To enable streaming, we apply a chunked-causal attention mask to both the Destylizer and Stylizer, where each chunk attends to its own features and all preceding chunks, but not to future chunks. In the Destylizer, the HuBERT layers are unfrozen and made chunked-causal, and all convolution layers are con- verted to causal. The streaming Destylizer is initialized from its non-streaming checkpoint and trained with a mean squared er- ror (MSE) distillation loss against the non-streaming teacher’s content features. Once trained, the streaming Stylizer, warm- started from its non-streaming counterpart, is trained on top of the streaming Destylizer’s outputs. For real-time inference, we maintain a fixed-length target utterance and a ring buffer of input chunks of content features. This provides sufficient past context for the Stylizer while keep- ing latency manageable. The end-to-end latency L is calculated as: L = t chunksize + t proc (6) where t chunksize is the input speech chunk size and t proc is the processing time per chunk. As long as t proc < t chunksize , the model is streamable. 4. Experimental Setup 4.1. Dataset For the training data, the Stylizer is trained using the English portion of Emilia [53], which contains 50k hours of diverse in- the-wild English speech. The Destylizer is trained on a com- bined training set from LibriTTS [54], MSP-Podcast [55], and 2 https://github.com/gemelo-ai/vocos GLOBE [56], totaling approximately 1300 hours of speech that covers a wide range of speakers, emotions and accents. We refer to this combined dataset as LMG. LMG is also used for training in the analysis and ablation studies (Sections 6.3, 6.4, 6.5). For evaluation, we randomly sample 300 source utterances from a combination of the Emotion Speech Dataset (ESD) [57], GLOBE-test, and LibriTTS-test-clean. We also select 10 target utterances from ESD, RAVDESS [58], GLOBE-test, and L2- ARCTIC [59], covering 5 emotions (happy, angry, sad, fear- ful, calm) and 5 accents (British, American, Indian, Arabic, Mandarin). This results in 300 × 10 = 3000 source-target pairs, which we denote as StyleStream-Test. All the objec- tive and subjective evaluations in Section 6.1 are conducted on StyleStream-Test. Moreover, to evaluate content-style disen- tanglement (Section 6.3), we train style classifiers on top of the content features extracted from different models. The speaker classifier is trained on the VoxCeleb training set [60] and eval- uated on its test set, which contains 1,251 speakers. The accent classifier is trained on L2-ARCTIC, which includes 6 accents; 6 speakers (one per accent) are held out for testing, and the re- maining 18 speakers are used for training. The emotion classi- fier is trained on EmoV-DB [61], which covers 5 emotions; 2 speakers (each with all 5 emotions) are held out for testing, and the remaining 18 speakers are used for training. 4.2. Training Destylizer. We take the 18th layer of HuBERT-Large-ASR 3 as input to six conformer blocks, followed by an FSQ module and four transformer decoder layers (hidden size 768, FFN size 3072). ALiBi positional encoding [62] is adopted to reduce long-context reliance, making the model suitable for length extrapolation and real-time inference. FSQ levels are set to [5,3,3], yielding a compact codebook of 45 codes. The model is 3 https://huggingface.co/facebook/ hubert-large-ls960-ft Table 1: Results for zero-shot voice style conversion. The best results are shown in bold, and the second best results are underlined. ModelWER (%)↓ S-SIM↑ A-SIM↑ E-SIM↑ NMOS↑ A-SMOS↑ E-SMOS↑ S-SMOS↑ Ground Truth3.8—3.71±.11— FACodec [10]15.50.7630.4080.6682.55±.132.02±.132.76±.142.98±.14 CosyVoice 2.0 [7]9.50.7940.4500.6553.47±.112.26±.132.58±.153.25±.14 SeedVC v2 6 21.70.7660.5490.6882.65±.123.34±.123.21±.133.11±.13 Vevo [11]17.50.8180.5960.7123.38±.123.93±.113.49±.133.76±.13 Vevo 1.5 [52]20.50.7980.5540.6833.10±.123.15±.143.15±.143.34±.14 StyleStream (streaming)15.30.8550.6350.8033.34±.134.28±.094.37±.094.29±.10 StyleStream (offline)9.20.8520.6400.8273.42±.124.32±.094.42±.094.36±.10 Table 2: Chunk size vs. quality tradeoff Chunk size (ms)WER (%)↓S-SIM↑A-SIM↑E-SIM↑UTMOS↑ 20019.80.7970.5860.7042.33±.02 40015.00.8060.5970.7202.69±.03 60013.60.8110.6020.7342.85±.03 80012.70.8160.6070.7452.94±.03 100012.60.8200.6130.7513.00±.03 trained for 100k steps on 8 NVIDIA RTX A6000 GPUs (batch size 32) using AdamW [63] with a peak learning rate of 1e-4, 4k warm-up steps, and cosine annealing. The streaming vari- ant uses a chunk size of 600 ms but otherwise follows the same training configuration. Stylizer. The Stylizer consists of 16 transformer layers (hidden size 768, FFN size 3072). Speech is sampled at 16 kHz and converted to 100-bin mel-spectrograms with a hop size of 320, producing a 50 Hz frame rate. During training, 70−100% of mel-spectrogram frames are randomly masked for inpaint- ing. To support CFG inference, content features are dropped with probability 0.2, while context spectrograms and style em- beddings are dropped with probability 0.3. Training uses 6s segments for 400k steps on 8 A6000 GPUs (batch size 64), with AdamW at a peak learning rate of 1e-4, 2k warm-up steps, and cosine annealing. The streaming Stylizer also uses a 600 ms chunk size under the same configuration. Throughout the ex- periments, the CFG strength is fixed at 2 and the Number of Function Evaluations (NFE) is set to 16, utilizing the standard Euler sampling method. Vocoder. We follow the design choice of Vocos [51], but modify the convolution layers in ConvNext [64] blocks into causal convolutions. Training is initialized from the official checkpoint and performed on the LibriTTS training set for 100k steps, using 2 A6000 GPUs with a batch size of 64 and 2s input segments. The mel-spectrogram setup matches that of the Styl- izer, and all remaining training configurations follow Vocos. 5. Baselines We compare StyleStream against several representative voice conversion and voice style conversion systems. FACodec 4 [10] disentangles speech into phoneme content, normalized pitch, and speaker residuals using an information bottleneck with a gradient reversal layer, and performs conversion by swapping speaker embeddings. CosyVoice 2.0 5 [7] employs a supervised 4 https://github.com/open-mmlab/Amphion/tree/ main/models/codec/ns3_codec 5 https://github.com/FunAudioLLM/CosyVoice Table 3: Chunk size vs. processing time on RTX 4060 and RTX A6000. Entries show mean± standard deviation over runs. Chunk size (ms) RTX 4060 (s) RTX A6000 (s) 1000.537±.0530.427±.013 2000.574±.0590.429±.009 3000.615±.0570.432±.010 4000.647±.0900.441±.008 5000.653±.0510.431±.009 6000.668±.0480.429±.006 7000.673±.0410.429±.002 8000.673±.0400.431±.004 semantic speech tokenizer trained with an ASR loss and fi- nite scalar quantization (FSQ) with a 6561-entry codebook, fol- lowed by a flow-matching transformer trained with a spectro- gram in-painting objective conditioned on semantic tokens and speaker embeddings. SeedVC v2 6 is an upgraded version of SeedVC [29] that supports voice style conversion, for which no accompanying paper has been released. Vevo 7 [11], the previous state-of-the-art voice style conversion system, lever- ages self-supervised content tokens with a 32-code codebook, autoregressive content–style modeling, and Voicebox-style [1] acoustic modeling. Finally, Vevo 1.5 8 [52] extends Vevo to both voice and singing voice conversion by introducing a prosody to- kenizer that captures coarse-grained melody contours. 5.1. Metrics For objective evaluation, we report word error rate (WER) as a measure of intelligibility, computed with Whisper-large- v3 [65]. To evaluate style similarity, we extract embeddings 6 https://github.com/Plachtaa/seed-vc 7 https://github.com/open-mmlab/Amphion/tree/ main/models/vc/vevo 8 https://github.com/open-mmlab/Amphion/tree/ main/models/svc/vevosing for speaker 9 , accent 10 [66], and emotion 11 [67], and compute cosine similarity between generated and target speech, yield- ing speaker similarity (S-SIM), accent similarity (A-SIM), and emotion similarity (E-SIM). For subjective evaluation, we con- duct Mean Opinion Score (MOS) tests on a 5-point scale, in- cluding naturalness MOS (N-MOS) for converted speech and similarity MOS (SMOS) for speaker (S-SMOS), accent (A- SMOS), and emotion (E-SMOS) relative to the target. All sub- jective evaluations are carried out with crowd-sourced listen- ers recruited via Prolific 12 , where all participants are based in the US or UK, are native English speakers, and are familiar with common English accents. For each test condition, we col- lect a total of 400 ratings per model. In N-MOS, listeners rate the overall naturalness of each generated sample from 1 (com- pletely unnatural) to 5 (completely natural). In SMOS, listeners rate the similarity between a converted sample and the corre- sponding target on the same 5-point scale; for accent similarity, listeners are instructed to ignore speaker timbre, emotion, and recording quality, and focus solely on accent similarity. 6. Results 6.1. Zero-Shot Voice Style Conversion The main results are shown in Table 1. Note that StyleStream does not provide independent control over individual style factors; all evaluations reflect holistic target-style cloning, where speaker identity, accent, and emotion are transferred jointly. Overall, StyleStream consistently outperforms prior methods across intelligibility and multiple style similarity met- rics, demonstrating strong zero-shot generalization in both of- fline and streaming settings. In terms of intelligibility, the offline StyleStream achieves the lowest WER (9.2%), substantially outperforming the previ- ous state-of-the-art Vevo. Despite operating in a chunk-causal manner, the streaming variant remains competitive with Vevo, highlighting the effectiveness of the Destylizer in preserving linguistic content under streaming constraints.In contrast, methods such as SeedVC and FACodec exhibit higher WER, indicating weaker content preservation. For overall style fidelity, StyleStream achieves the high- est speaker, accent, and emotion similarity scores, as measured by both objective and subjective metrics. Specifically, the of- fline variant attains the best S-SIM, A-SIM, and E-SIM, while both offline and streaming models substantially outperform all baselines in S-SMOS, A-SMOS, and E-SMOS. These improve- ments are consistently reflected in both objective and subjective evaluations, indicating higher-fidelity and more perceptually ac- curate reproduction of the target speaking style. Taken together, these results show that StyleStream achieves a more favorable balance between content preserva- tion and holistic style fidelity than existing zero-shot voice style conversion systems, while maintaining robust performance in real-time streaming scenarios. 9 https://github.com/resemble-ai/Resemblyzer 10 https://huggingface.co/Jzuluaga/ accent-id-commonaccentecapa 11 https://github.com/ddlBoJack/emotion2vec 12 https://w.prolific.com/ 6.2. Streaming Analysis 6.2.1. Latency We measure the end-to-end streaming latency on a single NVIDIA RTX A6000 GPU. Using NFE=16, a chunk size of t chunksize = 600ms, a 5s target segment, and a 5s content ring buffer, with the target style embedding pre-extracted for infer- ence, the average processing time per chunk is t proc = 412.7ms. Since t proc < t chunksize , streaming is feasible. By Equation 6, the resulting end-to-end latency is L = 600 + 412.7 = 1012.7ms. 6.2.2. Chunksize-Latency Analysis We benchmark the processing time (t proc ) across varying chunk sizes on both a server-grade NVIDIA A6000 and a consumer- grade RTX 4060 Laptop GPU (NFE fixed at 16). As shown in Table 3, t proc on the A6000 remains effectively constant (≈ 0.43s) when varying the chunk size from 100ms to 800ms. This indicates that the DiT inference is dominated by fixed kernel launch overheads rather than computational saturation on high-end hardware. In contrast, the RTX 4060 exhibits a linear increase in processing time with chunk size, reflecting a compute-bound regime characteristic of consumer-grade de- ployment. 6.2.3. Chunksize-Quality Trade-off To decouple the effects of context length from training- inference mismatch, we analyze the chunksize-quality trade- off by simulating streaming inference using the offline check- point across chunk sizes ranging from 200ms to 1000ms (Table 2). We observe a monotonic improvement in both intelligibil- ity (WER) and style similarity as chunk size increases. This confirms that extended temporal context is essential for captur- ing accent and emotion and minimizing discontinuities at chunk boundaries. 6.3. Content-Style Disentanglement Analysis In this section, we analyze how different speech content rep- resentations affect downstream voice style conversion perfor- mance, with a particular focus on their degree of content-style disentanglement. Since the Destylizer is explicitly designed to remove timbre, accent, and emotion information while pre- serving linguistic content, we evaluate disentanglement both in- directly (via Stylizer performance) and directly (via auxiliary style classification probes). 6.3.1. Effect of Content Representations on Stylizer We first assess how different content features influence down- stream Stylizer performance. To this end, we train the Styl- izer from scratch on top of the Destylizer and several exist- ing speech content representations, and evaluate WER, S-SIM, A-SIM, E-SIM, and UTMOS [68]. Specifically, we compare against the continuous pre-quantization features and discrete representations from Vevo, as well as the raw 18th-layer fea- tures of HuBERT-Large-ASR, which are also incorporated as the first stage of the Destylizer. All Stylizers are trained on the LMG dataset (Section 4.1) using identical training configurations (Section 4.2), with train- ing limited to 100k steps for fair comparison. Results are re- ported in Table 4. Stylizers trained on Destylizer features achieve the best overall performance in WER, S-SIM, A-SIM, and E-SIM, while Table 4: Stylizer performance trained on different content features. Content FeaturesWER (%)↓S-SIM↑A-SIM↑E-SIM↑UTMOS↑ HuBERT-Large-ASR 18th12.50.7310.3890.6273.28±.020 Vevo Continuous12.60.7670.4200.6602.92±.029 Vevo Indices25.60.7570.5640.6572.65±.036 Destylizer (offline)10.70.8370.6260.7333.22±.029 Table 5: Style classification accuracy across different content features. “Acc.” stands for accuracy. Content FeaturesAccent Acc. ↓Emotion Acc. ↓Speaker Acc. ↓ HuBERT-Large-ASR 18th78.00%85.60%86.60% Vevo Continuous68.93%78.73%64.76% Vevo Indices55.60%53.40%23.00% CosyVoice 2.0 Indices50.60%56.20%25.00% Destylizer (offline)43.50%47.60%3.50% Destylizer (streaming)33.80%37.90%3.01% remaining competitive in UTMOS. In contrast, the HuBERT- based Stylizer yields substantially lower A-SIM and E-SIM, in- dicating limited style transfer capability. For Vevo, using dis- crete indices significantly improves A-SIM compared to con- tinuous features, but this comes at the cost of degraded content preservation, as reflected by the highest WER among all meth- ods. 6.3.2. Disentanglement Hypothesis We hypothesize that the observed performance differences are closely linked to the degree of content-style disentanglement in the underlying representations. When disentanglement is poor, residual style information may leak into the content features and be exploited by the Stylizer during training, causing it to recon- struct speech with source-style attributes. As a result, inference fails to faithfully reflect the target style, leading to degraded style fidelity despite strong training performance. To validate this hypothesis, we directly measure how much timbre, accent, and emotion information remain in each set of content features. 6.3.3. Style Classification Probing We train an ECAPA-TDNN [49] classifier on top of each con- tent representation to predict speaker identity, accent, and emo- tion. Since these attributes are precisely what the Destylizer is designed to remove, lower classification accuracy indicates stronger disentanglement, with random-guess performance as the ideal target. Results are shown in Table 5. Both variants of the Destyl- izer consistently achieve the lowest classification accuracy across all three tasks. In particular, Destylizer features yield approximately ∼3% accuracy on speaker classification, com- pared to 20% or higher for HuBERT, Vevo, and CosyVoice 2.0, indicating that most speaker information has been filtered out. The slightly above-random accuracy observed for the Destylizer may stem from residual correlations between prosody or duration and the extracted content features. This is expected, as the non-autoregressive, streaming-oriented de- sign preserves the source utterance duration. Nevertheless, the Destylizer yields the cleanest content representations among all compared methods, consistent with its superior downstream Stylizer performance. 6.4. Ablations on Architecture In this section,we present architectural ablations of StyleStream, all trained on the LMG dataset described in Sec- tion 4.1. Specifically, we first evaluate a variant without style embeddings and another that uses FSQ quantized indices in- stead of pre-quantization continuous features. Results are re- ported in Table 6.Removing the style embedding signifi- cantly reduces S-SIM, A-SIM and E-SIM, showing that rely- ing solely on the unmasked context mel-spectrogram is insuffi- cient for style modeling. The style encoder is therefore essential for faithful style conversion. Moreover, replacing continuous features with FSQ indices severely degrades intelligibility, as shown by the high WER. At first glance, this may seem sur- prising, since Vevo achieves reasonable results using discrete indices. However, as shown in Table 5, our Destylizer’s con- tinuous features already carry even less accent, emotion, and speaker information than Vevo’s discrete indices. Further quan- tization strips away critical linguistic content, which explains the poor downstream performance. We also ablate the effect of the FSQ bottleneck size by varying the FSQ levels during Destylizer training and retrain- ing the Stylizer on top (Table 7). Enlarging the codebook to 7× 5× 5× 5× 5 = 4375 improves UTMOS, since richer information is kept in the Destylizer features, making it easier for the Stylizer to reconstruct speech during training. However, this comes at the cost of a sharp drop in A-SIM and E-SIM, indicating substantial style leakage. Conversely, reducing the bottleneck size to 3× 3× 3 = 27 degrades the intelligibility, and a further reduction to 5× 3 = 15 severely damages content preservation, indicating an overly restrictive bottleneck. These results suggest that our chosen setting of [5, 3, 3] provides a fa- vorable trade-off: compact enough to filter out most of the style, yet wide enough to preserve linguistic content. 6.5. Ablations on Destylizer Training Data We study the impact of training data scale for the Destylizer by replacing the original 1.3k-hour LMG training set (Section 4.1) with the larger Emilia-EN dataset (50k hours), while retraining the Stylizer on top of the resulting Destylizer. The comparative Table 6: Ablation studies on the use of style encoder and the choice of Destylizer content features. ModelWER (%)↓S-SIM↑A-SIM↑E-SIM↑UTMOS↑ StyleStream (offline)10.70.8370.6260.7333.22±.029 w/o Style Emb15.30.7750.5090.6533.47±.024 w/ FSQ Indices123.50.8290.5730.7172.41±.024 Table 7: Ablation on Destylizer FSQ bottleneck size. “*” denotes the architecture used in all main experiments. FSQ LevelsWER (%)↓S-SIM↑A-SIM↑E-SIM↑UTMOS↑ [7,5,5,5,5]13.50.7450.4390.6383.58±.022 [5,5,3,3]14.90.8220.6000.7253.17±.031 [5,3,3]*10.70.8370.6260.7333.22±.029 [3,3,3]19.00.8250.6180.7323.13±.032 [5,3]101.40.8340.5820.7142.70±.029 Table 8: Ablation on the Destylizer training data configura- tions. Training data WER (%)↓ S-SIM↑ A-SIM↑ E-SIM↑ LMG (1.3k h)10.70.8370.6260.733 Emilia (50k h)10.30.8160.5940.687 results are reported in Table 8. Scaling the Destylizer training data to 50k hours does not yield a significant overall improve- ment. Although content preservation improves marginally, with WER decreasing from 10.7% to 10.3%, the style disentangle- ment metrics (S-SIM and A-SIM) exhibit slight degradation. We attribute this saturation effect to two main factors. First, the Destylizer performs a discriminative extraction task, com- pressing speech into content representations under ASR super- vision, which fundamentally differs from the Stylizer’s gen- erative modeling objective.Unlike diffusion-based genera- tive models that benefit substantially from large-scale data to capture long-tail acoustic variations, the Destylizer’s content- style separation task saturates once robust disentanglement is learned, leading to diminishing returns with additional data. Second, the original 1.3k-hour dataset is strategically curated to maximize style diversity rather than raw duration, combin- ing GLOBE for accent variability, MSP-Podcast for emotional diversity, and LibriTTS for clean timbre coverage. This com- position already spans the style manifold required for effective disentanglement under a constrained bottleneck. In contrast, the substantially larger Emilia-EN dataset does not meaning- fully expand this manifold for the Destylizer’s objective. There- fore, these results suggest that the 1.3k-hour LMG dataset rep- resents an efficient and sufficient operating point for training the Destylizer, balancing data scale and disentanglement per- formance. 6.6. Ablations on Target Utterance Lengths We further evaluate StyleStream on the same test set using trun- cated 2-second target reference utterances. The comparison with the 5-second reference setting is summarized in Table 9. As expected, reducing the target utterance length leads to con- sistent performance degradation across all metrics. In particu- lar, both similarity scores and content preservation worsen for the 2-second setting in both offline and streaming modes. This Table 9: Ablation on target utterance lengths. ModelWER (%)↓ S-SIM↑ A-SIM↑ E-SIM↑ Target 5s (offline)9.20.8520.6400.827 Target 5s (streaming)15.30.8560.6350.803 Target 2s (offline)12.10.8240.5940.747 Target 2s (streaming)19.20.8330.5980.728 degradation is linguistically intuitive. Accent and emotion are global style attributes that require sufficient temporal context to be reliably estimated. A 2-second reference often provides limited phoneme coverage, making it difficult to characterize a speaker’s accent distribution, and offers insufficient prosodic variation to robustly capture emotional cues. Consequently, shorter reference utterances constrain the model’s ability to in- fer stable style representations, resulting in reduced conversion quality. 7. Conclusion In this work, we presented StyleStream, the first streamable zero-shot voice style conversion system capable of modify- ing timbre, accent, and emotion in real time, with an end-to- end latency of approximately 1 second. StyleStream is built upon two core components: a Destylizer, which performs ex- plicit content-style disentanglement using ASR supervision and a compact finite scalar quantization (FSQ) bottleneck, and a Stylizer, which leverages a diffusion transformer to reintro- duce target style conditioned on reference speech. By operating on continuous pre-quantization features, the proposed frame- work enables cleaner content-style separation and avoids arti- facts commonly introduced by discrete tokenization. As a re- sult, StyleStream achieves state-of-the-art voice style conver- sion performance across both objective and subjective evalua- tions, while maintaining full streamability, making it suitable for real-time applications. 8. Generative AI Use Disclosure We made limited use of AI during the preparation of this paper. In particular, LLMs were used for grammar checking, rephras- ing for clarity, and improving the readability of drafts. In ad- dition, generative AI was employed to generate the female and angry emojis used in Figure 1. 9. References [1] M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al., “Voicebox: Text-guided multilingual universal speech generation at scale,” Advances in neural information processing systems, vol. 36, p. 14 005–14 034, 2023. [2] S. E. Eskimez, X. Wang, M. Thakker, C. Li, C.-H. Tsai, Z. Xiao, H. Yang, Z. Zhu, M. Tang, X. Tan et al., “E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,” in 2024 IEEE Spoken Language Technology Workshop (SLT).IEEE, 2024, p. 682– 689. [3] Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen, “F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,” arXiv preprint arXiv:2410.06885, 2024. [4] Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu, “Maskgct: Zero-shot text- to-speech with masked generative codec transformer,” in ICLR. OpenReview.net, 2025. [5] P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao et al., “Seed-tts: A family of high-quality versatile speech generation models,” arXiv preprint arXiv:2406.02430, 2024. [6] Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y. Yang, H. Hu, S. Zheng, Y. Gu, Z. Ma et al., “Cosyvoice: A scalable multi- lingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,” arXiv preprint arXiv:2407.05407, 2024. [7] Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang et al., “Cosyvoice 2: Scalable stream- ing speech synthesis with large language models,” arXiv preprint arXiv:2412.10117, 2024. [8] K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa- Johnson, “Autovc: Zero-shot voice style transfer with only au- toencoder loss,” in International Conference on Machine Learn- ing. PMLR, 2019, p. 5210–5219. [9] K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, “Unsupervised speech decomposition via triple information bot- tleneck,” in International Conference on Machine Learning. PMLR, 2020, p. 7836–7846. [10] Z. Ju, Y. Wang, K. Shen, X. Tan, D. Xin, D. Yang, Y. Liu, Y. Leng, K. Song, S. Tang et al., “Naturalspeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,” arXiv preprint arXiv:2403.03100, 2024. [11] X. Zhang, X. Zhang, K. Peng, Z. Tang, V. Manohar, Y. Liu, J. Hwang, D. Li, Y. Wang, J. Chan, Y. Huang, Z. Wu, and M. Ma, “Vevo: Controllable zero-shot voice imitation with self- supervised disentanglement,” in ICLR. OpenReview.net, 2025. [12] H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, “Neu- ral analysis and synthesis: Reconstructing speech from self- supervised representations,” Advances in Neural Information Pro- cessing Systems, vol. 34, p. 16 251–16 265, 2021. [13] H.-S. Choi, J. Yang, J. Lee, and H. Kim, “Nansy++: Unified voice synthesis with neural analysis and synthesis,” arXiv preprint arXiv:2211.09407, 2022. [14] K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa- Johnson, and S. Chang, “Contentvec:An improved self- supervised speech representation by disentangling speakers,” in International conference on machine learning.PMLR, 2022, p. 18 003–18 017. [15] C. H. Chan, K. Qian, Y. Zhang, and M. Hasegawa-Johnson, “Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6332–6336. [16] D. Wang, L. Deng, Y. T. Yeung, X. Chen, X. Liu, and H. Meng, “Vqmivc: Vector quantization and mutual information-based un- supervised speech representation disentanglement for one-shot voice conversion,” arXiv preprint arXiv:2106.10132, 2021. [17] A. Tjandra, R. Pang, Y. Zhang, and S. Karita, “Unsupervised learning of disentangled speech content and style representation,” in Proc. Interspeech 2021, 2021, p. 4089–4093. [18] A. Van Den Oord, O. Vinyals et al., “Neural discrete represen- tation learning,” Advances in neural information processing sys- tems, vol. 30, 2017. [19] W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdi- nov, and A. Mohamed, “Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing, vol. 29, p. 3451–3460, 2021. [20] Y. Liu, C. Wang, H. Kim, R. Khan, and G. Anumanchipalli, “Rt- vc: Real-time zero-shot voice conversion with speech articulatory coding,” arXiv preprint arXiv:2506.10289, 2025. [21] Y. Yang, Y. Kartynnik, Y. Li, J. Tang, X. Li, G. Sung, and M. Grundmann, “Streamvc: Real-time low-latency voice conver- sion,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, p. 11 016–11 020. [22] Z. Wang, Y. Chen, X. Wang, L. Xie, and Y. Wang, “Streamvoice: Streamable context-aware language modeling for real-time zero- shot voice conversion,” arXiv preprint arXiv:2401.11053, 2024. [23] F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen, “Fi- nite scalar quantization: Vq-vae made simple,” arXiv preprint arXiv:2309.15505, 2023. [24] B. Van Niekerk, M.-A. Carbonneau, J. Za ̈ ıdi, M. Baas, H. Seut ́ e, and H. Kamper, “A comparison of discrete and soft speech units for improved voice conversion,” in ICASSP 2022-2022 IEEE In- ternational Conference on Acoustics, Speech and Signal Process- ing (ICASSP). IEEE, 2022, p. 6562–6566. [25] W. Peebles and S. Xie, “Scalable diffusion models with transform- ers,” arXiv preprint arXiv:2212.09748, 2022. [26] C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., “Neural codec language mod- els are zero-shot text to speech synthesizers,” arXiv preprint arXiv:2301.02111, 2023. [27] S. Zhou, Y. Zhou, Y. He, X. Zhou, J. Wang, W. Deng, and J. Shu, “Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech,” arXiv preprint arXiv:2506.21619, 2025. [28] H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Stargan-vc: Non-parallel many-to-many voice conversion using star genera- tive adversarial networks,” in 2018 IEEE Spoken Language Tech- nology Workshop (SLT). IEEE, 2018, p. 266–273. [29] S. Liu, “Zero-shot voice conversion with diffusion transformers,” arXiv preprint arXiv:2411.09943, 2024. [30] S.-H. Lee, H.-Y. Choi, S.-B. Kim, and S.-W. Lee, “Hierspeech++: Bridging the gap between semantic and acoustic representation of speech by hierarchical variational inference for zero-shot speech synthesis,” IEEE Transactions on Neural Networks and Learning Systems, 2025. [31] H. Xue, X. Peng, Y. Lu et al., “Convert and speak: Zero-shot ac- cent conversion with minimum supervision,” in ACM Multimedia 2024, 2024. [32] M. Jin, P. Serai, J. Wu, A. Tjandra, V. Manohar, and Q. He, “Voice-preserving zero-shot multiple accent conversion,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, p. 1–5. [33] S. Liu, D. Wang, Y. Cao, L. Sun, X. Wu, S. Kang, Z. Wu, X. Liu, D. Su, D. Yu et al., “End-to-end accent conversion with- out using native utterances,” in ICASSP 2020-2020 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 6289–6293. [34] K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion intensity and its control for emotional voice conversion,” IEEE Transactions on Affective Computing, vol. 14, no. 1, p. 31–48, 2022. [35] F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T.-A. Nguyen, M. Rivi ` ere, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, “Textless speech emotion conversion using decomposed and dis- crete representations,” Conference on Empirical Methods in Nat- ural Language Processing (EMNLP), 2022. [36] N. R. Prabhu, B. Lay, S. Welker, N. Lehmann-Willenbrock, and T. Gerkmann, “Emoconv-diff: Diffusion-based speech emotion conversion for non-parallel and in-the-wild data,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2024, p. 11 651– 11 655. [37] D.-Y. Wu and H.-y. Lee, “One-shot voice conversion by vector quantization,” in ICASSP 2020-2020 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 7734–7738. [38] H. Tang, X. Zhang, J. Wang, N. Cheng, and J. Xiao, “Learning speech representations with flexible hidden feature dimensions,” in ICASSP 2023-2023 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP).IEEE, 2023, p. 1–5. [39] C. J. Cho, P. Wu, T. S. Prabhune, D. Agarwal, and G. K. Anu- manchipalli, “Coding speech through vocal tract kinematics,” IEEE Journal of Selected Topics in Signal Processing, 2024. [40] Y. Liu, B. Yu, D. Lin, P. Wu, C. J. Cho, and G. K. Anumanchipalli, “Fast, high-quality and parameter-efficient articulatory synthesis using differentiable dsp,” in 2024 IEEE Spoken Language Tech- nology Workshop (SLT). IEEE, 2024, p. 711–718. [41] Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, C. Ni, X. Shi et al., “Cosyvoice 3: Towards in-the- wild speech generation via scaling-up and post-training,” arXiv preprint arXiv:2505.17589, 2025. [42] T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “Stargan-vc2: Rethinking conditional methods for stargan-based voice conver- sion,” arXiv preprint arXiv:1907.12279, 2019. [43] K. Zhou, B. Sisman, and H. Li, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,” in 2021 IEEE spoken language technology workshop (SLT).IEEE, 2021, p. 415–422. [44] J. Lian, C. Zhang, and D. Yu, “Robust disentangled variational speech representation learning for zero-shot voice conversion,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, p. 6572– 6576. [45] M. Łajszczak, G. C ́ ambara, Y. Li, F. Beyhan, A. Van Korlaar, F. Yang, A. Joly, ́ A. Mart ́ ın-Cortinas, A. Abbas, A. Michal- ski et al., “Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data,” arXiv preprint arXiv:2402.08093, 2024. [46] Z. Ning, S. Wang, P. Zhu, Z. Wang, J. Yao, L. Xie, and M. Bi, “Dualvc 3: Leveraging language model generated pseudo con- text for end-to-end low latency streaming voice conversion,” arXiv preprint arXiv:2406.07846, 2024. [47] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al., “Conformer: Convolution- augmented transformer for speech recognition,” arXiv preprint arXiv:2005.08100, 2020. [48] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1505–1518, 2022. [49] B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa- tdnn:Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,” arXiv preprint arXiv:2005.07143, 2020. [50] K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statis- tics pooling for deep speaker embedding,” arXiv preprint arXiv:1803.10963, 2018. [51] H. Siuzdak, “Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,” arXiv preprint arXiv:2306.00814, 2023. [52] J. Li, X. Zhang, Y. Wang, H. He, C. Wang, L. Wang, H. Liao, J. Ao, Z. Xie, Y. Huang, J. Zhang, and Z. Wu, “Overview of the amphion toolkit (v0.2),” arXiv preprint arXiv:2501.15442, 2025. [53] H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi et al., “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, p. 885–890. [54] H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text- to-speech,” arXiv preprint arXiv:1904.02882, 2019. [55] R. Lotfian and C. Busso, “Building naturalistic emotionally bal- anced speech corpus by retrieving emotional speech from existing podcast recordings,” IEEE Transactions on Affective Computing, vol. 10, no. 4, p. 471–483, 2017. [56] W. Wang, Y. Song, and S. Jha, “Globe: A high-quality english corpus with global accents for zero-shot speaker adaptive text-to- speech,” arXiv preprint arXiv:2406.14875, 2024. [57] K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021-2021 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, p. 920–924. [58] S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one, vol. 13, no. 5, p. e0196391, 2018. [59] G. Zhao, S. Sonsaat, A. Silpachai, I. Lucic, E. Chukharev- Hudilainen, J. Levis, and R. Gutierrez-Osuna, “L2-arctic: A non- native english speech corpus,” in Interspeech 2018, 2018, p. 2783–2787. [60] A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language, vol. 60, p. 101027, 2020. [61] A. Adigwe, N. Tits, K. E. Haddad, S. Ostadabbas, and T. Du- toit, “The emotional voices database: Towards controlling the emotion dimension in voice generation systems,” arXiv preprint arXiv:1806.09514, 2018. [62] O. Press, N. A. Smith, and M. Lewis, “Train short, test long: Attention with linear biases enables input length extrapolation,” arXiv preprint arXiv:2108.12409, 2021. [63] I. Loshchilov and F. Hutter, “Decoupled weight decay regulariza- tion,” arXiv preprint arXiv:1711.05101, 2017. [64] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, 2022, p. 11 976–11 986. [65] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518. [66] J. Zuluaga-Gomez, S. Ahmed, D. Visockas, and C. Subakan, “Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice,” arXiv preprint arXiv:2305.18283, 2023. [67] Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion representation,” arXiv preprint arXiv:2312.15185, 2023. [68] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “Utmos: Utokyo-sarulab system for voicemos challenge 2022,” Interspeech 2022, 2022.