Paper deep dive
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
Zhetao Hu, Yiquan Zhou, Wenyu Wang, Zhiyu Wu, Xin Gao, Jihua Zhu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 3:15:45 AM
Summary
The paper introduces S4, a singing style conversion system for SVCC2025 that utilizes a boundary-aware Whisper bottleneck to suppress style leakage, an explicit frame-level technique matrix for dynamic style rendering, and a high-frequency band completion strategy to achieve 48kHz fidelity without overfitting.
Entities (5)
Relation Signals (4)
S4 → participatedin → SVCC2025
confidence 100% · This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025
S4 → usesarchitecture → SYKI-SVC
confidence 95% · our submission inherits the efficient, high-naturalness SYKI-SVC backbone
S4 → usescomponent → Whisper
confidence 95% · we introduce a boundary-aware Whisper bottleneck
S4 → usescomponent → NSF-HiFiGAN
confidence 95% · we synthesize a 24-kHz waveform using an NSF-HiFiGAN vocoder
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To address the critical challenges of style leakage, dynamic rendering, and high-fidelity generation with limited data, we introduce three key innovations: a boundary-aware Whisper bottleneck that pools phoneme-span representations to suppress residual source style while preserving linguistic content; an explicit frame-level technique matrix, enhanced by targeted F0 processing during inference, for stable and distinct dynamic style rendering; and a perceptually motivated high-frequency band completion strategy that leverages an auxiliary standard 48kHz SVC model to augment the high-frequency spectrum, thereby overcoming data scarcity without overfitting. In the official SVCC2025 subjective evaluation, our system achieves the best naturalness performance among all submissions while maintaining competitive results in speaker similarity and technique control, despite using significantly less extra singing data than other top-performing systems. Audio samples are available online.
Tags
Links
- Source: https://arxiv.org/abs/2604.05526v1
- Canonical: https://arxiv.org/abs/2604.05526v1
Trouble viewing inline? Open PDF directly →
Full Text
39,566 characters extracted from source content.
Expand or collapse full text
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Zhetao Hu 1,2 , Yiquan Zhou 1,2 , Wenyu Wang 1,2 , Zhiyu Wu 3 , Xin Gao 4 , and Jihua Zhu 1,2,* 1 School of Software Engineering, Xi’an Jiaotong University, Xi’an, China 2 SYKI-SPEECH Team, Xi’an, China 3 Fudan University, Shanghai, China 4 Division of Music and Audio Union Wheatland Culture and Media Ltd., China Email: huzhetao, zhouyiqian, wenyu.wang@stu.xjtu.edu.cn, wuzy24@m.fudan.edu.cn, sui@musinya.com, zhujh@xjtu.edu.cn * Abstract—This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)— a novel singing style conversion system that advances fine- grained style conversion and control within in-domain settings. To address the critical challenges of style leakage, dynamic rendering, and high-fidelity generation with limited data, we introduce three key innovations: a boundary-aware Whisper bottleneck that pools phoneme-span representations to sup- press residual source style while preserving linguistic content; an explicit frame-level technique matrix, enhanced by targeted F 0 processing during inference, for stable and distinct dynamic style rendering; and a perceptually motivated high-frequency band completion strategy that leverages an auxiliary standard 48 kHz SVC model to augment the high-frequency spectrum, thereby overcoming data scarcity without overfitting. In the official SVCC2025 subjective evaluation, our system achieves the best naturalness performance among all submissions while maintaining competitive results in speaker similarity and tech- nique control, despite using significantly less extra singing data than other top-performing systems. Audio samples are available online. 1 Index Terms—singing voice conversion, singing style conver- sion, feature disentangling I. Introduction Unlike conventional Singing Voice Conversion (SVC) [1]–[3], which primarily focuses on timbre or identity mapping, Singing Style Conversion (SSC) aims to finely reshape how a song is performed. This involves modifying vocal techniques such as breathiness, register (falsetto/mixed voice), resonance, vibrato, and glissando, while strictly preserving the lyrical content, melodic structure, and the singer’s identity characteristics. In essence, SSC is a style transfer task whose central objective is to rewrite singing techniques and their temporal organization, keeping the underlying content and melody unchanged [4]. Early controllable style transfer studies in speech and singing typically relied on relatively coarse style represen- tations to guide the generator, such as global style em- beddings, reference encoders, or utterance-level tokens [1], * Corresponding author. 1 https://csscbaib.netlify.app/ [5], [6]. These approaches are effective at capturing global affect or quasi-stationary attributes; However, they often reveal two structural limitations when applied to SSC. First, many singing techniques exhibit strong locality and rapid temporal dynamics (e.g., time-varying vibrato rates, intermittent breathy segments, or continuous pitch drift in glissando), making it difficult for a single global vector to generate stable and controllable fine-grained variations at the appropriate time steps. Second, style leakage remains prevalent: prosodic or technique-related cues from the source performance tend to persist in the converted output, compromising the purity of the target style. To mitigate these issues, subsequent research has intro- duced finer-grained supervision and conditioning mecha- nisms. A common strategy is to align technique prompts or labels with phoneme- or frame-level time scales, along with richer musical and content conditions (e.g., phonemes, score/MIDI, andF 0 ), to enhance controllability and interpretability [7]. Recent singing corpora have begun to provide multi-dimensional technique annotations and high-quality score information, enabling explicit technique modeling (e.g., GTSinger [8]), and zero-shot singing gener- ation studies further explore hierarchical style control and technique prompting (e.g., TCSinger [9]). Nevertheless, style leakage persists. A key reason is that upstream representations (e.g., self-supervised features [10]–[13] or codec tokens) often contain non-content acoustic infor- mation from the source signal [14]–[16]. During training and inference, models can exploit such residual cues, causing the target technique rendering to be interfered with or partially overridden by the source style. Similar prosody leakage has been widely discussed in zero-shot voice conversion and is considered a major factor limiting controllability [17]–[19]. In recent years, mainstream SSC systems have largely adopted two families of generative backbones. One family of models employs continuous generative models based on diffusion or flow matching to synthesize mel-spectrograms or waveforms under conditioning [2], [20], [21]. Represen- arXiv:2604.05526v1 [cs.SD] 7 Apr 2026 Feature Extractor Whisper F0, UV, MIDI, ph Voice Converter (VITS) GT Speaker ID Mel-spectrogram Reconstruction Loss GT Technique Matrix(T) GT Mel -spectrogram Linear Spec. Singing Waveform(y) (a) Training Bottleneck GT Phoneme Sequence Scale Factor Feature Extractor Whisper UV, MIDI, ph Voice Converter (VITS) Target Speaker ID Intermediate Waveform 24kHz Source Singing Waveform (b) Inference Bottleneck Scale Factor MFA Aligner Phoneme Sequence Phoneme Boundaries High-Freq Band Completion Target Technique Matrix(T) Explicit Pitch Dynamics F0 Refined F0 Output Waveform(48kHz) Fig. 1: The overall architecture of the proposed System S4. (a) Training Stage: The system employs a boundary-aware semantic bottleneck to suppress style leakage in Whisper features, while the VITS decoder is conditioned on the ground-truth technique matrixTto learn explicit style rendering. (b) Inference Stage: The pipeline is driven by a target technique matrix, which guides both the explicit pitch dynamics module (forF 0 refinement) and the decoder. Finally, a High-Frequency Band Completion module extends the output to 48 kHz fidelity. tative work includes Serenade, which formulates SSC as an audio-infilling problem and applies flow matching, address- ing target-style modeling, source-style disentanglement, and melody preservation [22]. The other family employs autoregressive (AR) modeling [23], [24] over discrete tokens [16], [25] and typically uses a continuous token- to-mel module to recover acoustic details. This design is common in large-scale speech generation [26]–[30]; for example, VEVO generates content-style token sequences with an AR model and then refines them using a flow- matching acoustic model [19], [31]. In practice, these two paradigms often appear in a cascaded form (AR for struc- ture and discrete sequences, flow/diffusion for continuous detail), and they therefore share the same key require- ment: the conditioning features must strongly disentangle content/melody from style/technique. If the conditioning features still contain residual source-style information, a powerful generator may amplify this entanglement, resulting in style leakage. Conversely, if the objective or conditioning constraints are insufficient to capture fast- varying dynamics, the output may exhibit over-smoothing of the technique dynamics. Recent SVCC2025 evaluations also suggest that, despite substantial progress in natural- ness and identity preservation, achieving stable and fine- grained style transfer remains a primary bottleneck [32]. In this paper, we present System S4, our submission to SVCC 2025 built upon the SYKI-SVC framework [33]. We address the aforementioned challenges through three targeted designs. First, to mitigate the leakage observed in baselines, we introduce a boundary-aware Whisper bottleneck that pools representations within phoneme spans to wash out residual source style. Second, we employ an explicit frame-level technique matrix enhanced by inference-timeF 0 processing to ensure the distinct rendering of dynamic styles. Finally, to achieve 48 kHz fidelity without overfitting the limited style data, we pro- pose a high-frequency band completion strategy. Instead of forcing the SSC model to generate the full spectrum, we train a separate, standard 48 kHz SVC model (which is easier to converge) and utilize its output to supplement the high-frequency band (>10kHz) of our style-converted audio. Subjective evaluations in the SVCC2025 challenge demonstrate that our system achieved the best naturalness performance among all submissions, while maintaining competitive results in both speaker similarity and tech- nique control. Our main contributions are summarized as follows: •Semantic bottleneck for disentanglement: A boundary-aware pooling strategy to suppress style leakage in content representations. •Explicit technique control: A decoupled frame-level conditioning mechanism with targetedF 0 processing for dynamic stability. •Perceptually motivated high-frequency completion: A strategy that leverages an auxiliary standard 48 kHz SVC model to complete the high-frequency spectrum, addressing data scarcity challenges in high-fidelity SSC training. The source code is publicly available at 2 . 2 The implementation is available at https://github.com/ HuZhetao/cssc. I. Related Work Serenade [22] is a representative baseline for SVCC2025 and formulates singing style conversion as an audio infilling task: a flow-matching model predicts masked segments of a target mel-spectrogram given the unmasked complement and disentangled features. It also employs cyclic training to disentangle the source style and utilizes a source-filter vocoder withF 0 resynthesis to better preserve the melody. This family benefits from expressive modeling and flexible conditioning. However, its inference pipeline is typically heavier than VITS-family systems, and the strict melody constraints may conflict with the desired style- specific pitch dynamics (e.g., vibrato depth or glissando transitions), which can reduce the perceptual salience of dynamic techniques in practice. Another influential direction uses discrete tokens and a two-stage generation pipeline. Vevo [19] proposes the generation of content/style tokens with an autoregressive transformer prompted by a style reference, followed by acoustic generation with a flow-matching transformer prompted by a timbre reference, supported by self- supervised disentanglement using a VQ-VAE bottleneck. The official Vevo1.5 baseline in SVCC2025 further demon- strates the competitiveness of AR+flow pipelines for con- trollable singing generation. These methods often benefit from large-scale pretraining and carefully designed tok- enizers [26], [28] but they typically require greater compu- tational resources and engineering complexity, which may be less favorable under data-limited or latency-constrained settings. SYKI-SVC builds a high-fidelity SVC system on the SVCC2023 T02 framework, utilizing a VITS-family con- verter [34] and a post-processing module [35] to im- prove perceptual quality, demonstrating strong natural- ness in challenge settings. Nevertheless, in SSC scenar- ios, commonly used SSL/ASR representations can still retain residual timbre/style information, causing leakage and weakening fine-grained controllability; meanwhile, utterance-level speaker embeddings provide only coarse control and are insufficient for temporally localized tech- nique manipulation. Motivated by these limitations, our submission inherits the efficient, high-naturalness SYKI- SVC backbone and focuses on (i) suppressing content- feature leakage via a boundary-aware Whisper bottleneck and (i) strengthening the perceptibility of dynamic tech- niques through explicit frame-level technique conditioning and targeted inference-time pitch dynamics enhancement. I. Proposed Method We propose a singing style conversion system for SVCC2025, which follows a recognition–synthesis paradigm and consists of three main modules: a feature extraction module, a voice conversion module and a high-frequency band completion module. Given an input singing waveform, the feature extractor derives linguistic and musical conditions, including phoneme sequences, pooled Whisper representations, and pitch-related features such asF 0 , UV, and MIDI. Based on frame-level technique annotations, a technique matrix is constructed to provide explicit style control. The voice conversion network then generates a 24-kHz singing waveform conditioned on the extracted features, the technique matrix, and the embedding of the target singer. Finally, a super-resolution network supplements high-frequency components above 10 kHz from an auxiliary 48-kHz model, producing the final converted singing voice. A. Feature Extraction Module Following the SYKI-SVC framework, our system per- forms singing style conversion via song reconstruction [33]. The SVCC2025 dataset provides rich symbolic and acous- tic annotations, including phoneme sequences, MIDI notes,F 0 , and UV flags [36], which are sufficient to describe lyrics and melody for basic synthesis [37]. To further improve content clarity, we additionally use se- mantic features from a pretrained Whisper encoder [38] as an auxiliary condition (we also experimented with ContentVec [15] and WeNet ASR features [39], but found Whisper yields the best controllability–clarity trade-off under SSC). Singer identity is conditioned by a one-hot embedding due to the limited number of singers in the challenge set. Phoneme-level Temporal Pooling However, we observe that frame-level semantic representations inevitably con- tain non-semantic information correlated with expressive dynamics, which can introduce technique leakage if di- rectly injected. To mitigate this, we impose a lightweight semantic bottleneck based on phoneme boundaries. Let w i ∈R d denote the raw semantic feature at framei(e.g., Whisper encoder output). We use phoneme boundaries to define a phoneme-level neighborhoodN(t)(the set of frames belonging to the same phoneme segment as frame t), and average withinN(t): ̃ w t = 1 |N(t)| ∑ i∈N(t) w i .(1) During training, phoneme boundaries are obtained from the provided alignments in the official dataset. During inference, we use the Montreal Forced Aligner(MFA) [40] trained on the challenge data to obtain phoneme boundaries and constructN(t). This pooling suppresses rapid intra-phoneme fluctuations that tend to correlate with expressive techniques, while preserving stable lin- guistic cues, thereby reducing technique leakage from the semantic branch. Finally, we apply a global scaling factor λto the pooled features, ˆ w t =λ· ̃ w t , and setλ= 0.1 in all experiments to further limit the dominance of the semantic condition. B. Voice Converter The voice converter is the acoustic synthesis mod- ule that recomposes the extracted conditions into a Feature Extractor Whisper Encoder Boundary -Aware Whisper Bottleneck (Pooling & Scaling) Raw Whisper Features Guided to Rhythm Extractors Input Audio Technique Matrix(T) (Explicit style control) Voice Conver sion Network Prior Encoder Flow& Decoder Posterior Encoder Linear Spec. Explicit Pitch Dynamics Modeling (Inference only) Semantic Features UV MIDI, ph Singer ID Refined f0 One-hot Embedding g NSF-HiFiGAN Generator High-Frequency Band Completion >10kHz High-Pass Filter 24kHz Converted Converted Singing Voice (48kHz) Auxiliary 48kHz Model 48kHz Waveform 48kHz Waveform Fig. 2: Illustration of the core disentanglement and control mechanisms. The Boundary-Aware Semantic Bottleneck pools frame-level Whisper features within phoneme boundaries to wash out residual source style, while the technique matrixTexplicitly drives style generation. mel-spectrogram while preserving lyrics and melody and enabling explicit technique control. It takes as input the phoneme condition ph(t), the bottle- necked semantic features ˆ w t , and pitch-related features (F 0 (t),UV(t),MIDI(t)), and predicts the target mel- spectrogramMunder the guidance of a technique control signal and a singer identity embedding. a) Technique matrix : To enable time-localized and compositional control over singing techniques, we repre- sent technique annotations as a frame-level binary matrix T∈0,1 K×N , whereKdenotes the number of technique labels andNdenotes the number of acoustic frames aligned with the mel hop size. Each columnT :,n is a multi- hot vector, allowing multiple techniques to co-occur within the same frame (e.g., breathy+vibrato). During training, Tis constructed by aligning the provided technique segments to the acoustic frame grid. During inference,Tis specified to describe the desired target technique pattern and thus serves as an explicit control interface. For conditioning,Tis projected by a lightweight em- bedding layer and concatenated to the frame-level condi- tioning stream before the prediction of mel. Together with the semantic bottleneck, this explicit injection encourages the converter to attribute the rendering technique toT rather than implicitly inferring the style from the semantic features, thus improving controllability. b) Explicit pitch-dynamics control: Dynamic tech- niques such as vibrato and glissando are manifested pri- marily as fine-grained pitch modulations [41]. We find that letting the acoustic model learn these micro-patterns im- plicitly often leads to temporal over-smoothing. To make pitch-driven techniques more salient and controllable, we refine the contour extractedF 0 (t)in the conditioning stage using rule-based transformations derived fromT, without introducing an additional predictorF 0 . On vibrato-active frames, we inject a sinusoidal modulation in the log- frequency domain: F vib 0 (t) =F 0 (t)·2 asin(2πf v t+φ) 1200 ,(2) whereais the depth (cents),f v is the rate (Hz), andφis the phase. The modulation is applied only whenm vib (t) = 1(fromT), otherwise the original contour is kept: F ∗ 0 (t) =m vib (t)F vib 0 (t) + ( 1−m vib (t) ) F 0 (t).(3) To capture glissando, we replace the discrete pitch shifts at note boundaries with continuous sigmoid transitions. At a boundary timet 0 between consecutive notes with target frequenciesF prev andF curr : F gliss 0 (t) =F prev + (F curr −F prev )· 1 1 +e −τ(t−t 0 ) ,(4) whereτcontrols the transition sharpness. This operation is similarly gated bym gliss (t)fromTso that transitions are injected only when glissando is desired. The refined contourF ∗ 0 (t)(after applying the technique- gated vibrato/glissando) is then used as pitch conditioning together with UV(t)and MIDI(t)for mel generation, improving the salience and temporal consistency of pitch- driven techniques without affecting content or singer identity control. C. Waveform Generation and High-Frequency Band Com- pletion Given the predicted mel-spectrogramMfrom the voice converter, we synthesize a 24-kHz waveform using an NSF- HiFiGAN vocoder [42]. Although 24 kHz is sufficient to model the main-band timbre and technique dynamics, the limited bandwidth often reduces perceived clarity and airiness. A straightforward solution is to directly generate 48-kHz audio; however, the official SVCC dataset is insufficient to train a high-quality 48-kHz vocoder or end-to-end generator. To address this, we introduce a lightweight high- frequency band completion strategy using an auxiliary 48-kHz singing conversion model. Empirically, we find that spectral components above 10 kHz contribute mainly to brightness rather than singer identity or technique- related cues. Therefore, during inference we first apply the auxiliary 48-kHz model to the source audio and extract its high-frequency band (>10kHz). We then supplement this band to the converted 24-kHz waveform via smooth frequency-domain mixing, without modifying the low- and mid-frequency content produced by the main system. We denoteH(·)as a high-pass operator extracting>10kHz components, and construct the final waveform as x 48 =x 24 +H ( x 48 src ) ,(5) followed by a short cross-fade in the frequency domain to avoid boundary artifacts.This simple band-completion improves perceptual quality while having minimal impact on timbre similarity and technique controllability. D. Training and Inference Strategy This section describes the training and inference rules used to improve disentanglement and technique control- lability, without repeating the feature-flow introduced above. Training Strategy. We adopt a three-stage protocol to progressively establish audio quality, enforce tech- nique disentanglement, and adapt to singers under the SVCC2025 setting. Stage I: Reconstruction Training. We first train the main pipeline with standard reconstruction objectives using the official dataset. In this stage, all provided con- ditions (phoneme/MIDI/F 0 /UV and semantic features) are enabled to stabilize alignment and establish a strong acoustic foundation. Stage I: Technique Disentanglement. To reduce tech- nique leakage from the semantic branch, we activate the semantic bottleneck: semantic features are pooled using phoneme boundaries and then down-scaled by a fixed factor (e.g.,λ= 0.1). This explicit capacity restriction forces the model to rely on the technique control signal Tfor style rendering rather than implicitly encoding techniques in semantic features. During this stage,Tis constructed from ground-truth technique annotations and aligned to the acoustic frame grid, providing supervised technique control for every training sample. Stage I: Singer Adaptation Fine-tuning. We fine-tune the model to better fit singer-dependent timbre charac- teristics while preserving the disentangled representations learned in Stage I. In practice, we keep the semantic bottleneck and technique conditioning unchanged and only continue end-to-end optimization on the official dataset. Inference Strategy During inference, ground-truth align- ments and technique labels are unavailable; therefore, we adopt a rule-based pipeline as follows. We first obtain phoneme boundaries using a self-trained MFA on the challenge dataset, which are required for boundary-aware semantic pooling and feature construction. Next, the de- sired technique patterns are specified by the target control signal and encoded as a frame-level binary technique matrixTto condition the converter. For pitch-driven techniques such as vibrato and glissando, we further perform explicitF 0 refinement gated byTprior to mel-spectrogram prediction, so as to enhance perceptual salience and temporal consistency of these rapid dynamics. Finally, to achieve 48-kHz fidelity without coupling it to the main training procedure, the high-frequency band- completion module is trained separately and applied only as a post-processing step after waveform synthesis, remain- ing outside the primary three-stage training protocol. IV. Experiments and Results A. Data Preparation and Training Setup All experiments follow the official SVCC2025 Task 1 setting; the main 24 kHz system is trained only on the official dataset (about 68 hours) with provided phoneme/MIDI/pitch-related annotations and technique labels for constructing the binary technique matrixT, while the auxiliary 48 kHz model for high-frequency band completion is trained separately and is not involved in the main pipeline. Training is conducted on 4×RTX 4090 GPUs using AdamW (β 1 = 0.8, β 2 = 0.99) with batch size 24, following a three-stage schedule of 200k/300k/100k steps, where Stage I enables the semantic bottleneck (phoneme-aware pooling andλ=0.1scaling) to reduce leakage and strengthen reliance onT. B. Official Results on SVCC2025 Task 1 The proposed system is evaluated based on the official subjective results of the SVCC2025 challenge [3]. The evaluation framework employs three key metrics: Natu- ralness (measured via 5-point MOS), and both Singing Style Similarity and Singer Identity Similarity (assessed on a 4-point scale and reported as binary accuracy). Our submitted system is denoted as S4B (Ours). To verify the efficacy of our disentanglement strategy, we also submitted an ablation system, S4A, which excludes the boundary-aware Whisper pooling and scaling module. 1) Naturalness: As illustrated in Fig. 4(a), our system S4B demonstrates exceptional performance in terms of perceptual quality. According to the official boxplot re- sults, S4B achieves a median MOS of 4.0, securing the 1st place in naturalness among all participating systems (excluding Ground Truth). The distribution of S4B scores is notably compact and skewed towards the higher end compared to other teams. This result strongly validates that our explicit technique modeling (via the Technique Matrix) and generative pitch dynamics allow for a highly natural rendition of singing voices, effectively avoiding the robotic artifacts or over-smoothing often observed in conventional SVC systems. 1.01.52.02.53.03.54.04.5 Naturalness (MOS) 0 10 20 30 40 50 60 70 80 90 Style Similarity B1 B1 B2 B2 B3 B3 GT GT S1A S1A S1B S1B S2B S2B S3A S3B S4A S4B S5B S5B S6A S6A S6B S6B S7B S7B System Type Diffusion ARLM+Diffusion VAEGAN GT Task 1 Task 2 Fig. 3: A scatter plot comparing naturalness and style similarity metrics. Our system S4B is positioned on the far right, indicating superior naturalness performance. 2) Similarity and Controllability: In terms of similarity metrics (Fig. 4b and c), S4B achieves competitive per- formance, ranking 4th in style similarity. While there is a gap in identity similarity compared to the top-ranked systems (e.g., S6, S5), it is crucial to highlight the data efficiency of our approach. Unlike several top-performing teams that leveraged large-scale external datasets for pre- training to boost timbre disentanglement, our system was trained exclusively on the limited official dataset. Given this constraint, our method yields a state-of-the-art naturalness score, proving that the architecture maximizes perceptual quality even under data-limited conditions. C. Ablation Study We evaluate six model variants, each converting the same set of 36 songs. For objective metrics, Spk Sim is computed by cosine similarity of CAMPPlus speaker embeddings [43], and MOS is evaluated by SingMOS [44]. For human evaluation, listeners rate technique transfer on a 4-level scale (successful / slightly successful / slightly unsuccessful / unsuccessful); we report Converted as the proportion of samples rated as successful or slightly successful. The questionnaire also includes a binary option indicating whether the output exhibits technique leakage, which is reported as Leakage. The six settings cover (i) different semantic features (ContentVec [15], WeNet [39], Whisper [38]), (i) the effect of scaling (λ) applied after pooling, and (i) the effect of removing phoneme-aware pooling. Table I shows that Whisper generally yields better controllability than ContentVec and WeNet. Moreover, the proposed semantic bottleneck (phoneme-aware pooling andλ=0.1scaling) achieves the best overall trade-off, delivering the highest Converted and Spk Sim while keeping Leakage lowest. In contrast, using an overly strong semantic signal (pooling withλ=1.0) increases leakage, while disabling the semantic signal (λ=0) severely S3A S2BS2BS1B B1 S5BS1BS5B B1 S3B S1AS1A B3B2B2 S7B S6A S6BS7B B3 S4A S6BS4B S6A GTGT 1 2 3 4 5 System DiffusionARLM+DiffusionVAEGANGround Truth Task / Markers Task 1Task 2Mean Score (a) Naturalness MOS B2-T1 S7B-T1 B2-T2 S7B-T2S1B-T1 S1A-T1S1A-T2 B1-T2B3-T2B3-T1 S2B-T2S1B-T2 S3A-T1 S2B-T1S4B-T1 S4A-T1 B1-T1 S5B-T2S3B-T1S5B-T1S6B-T2 S6A-T1S6A-T2 S6B-T1 GT-T2GT-T1 0 20 40 60 80 100 Percentage of Scores (%) (b) Singing style similarity S2B-T2 S1A-T2 B1-T2 S3A-T1 B3-T2 S1B-T2S1B-T1S7B-T1 B2-T1B3-T1 S1A-T1 B1-T1 S4B-T1 B2-T2 S2B-T1 S4A-T1 S5B-T1S7B-T2S3B-T1 S6A-T1 S6B-T1S6B-T2 S6A-T2 S5B-T2 GT-T2GT-T1 0 20 40 60 80 100 Percentage of Scores (%) Different, sure Different, not sure Same, not sure Same, sure Binary Accuracy (c) Singer identity similarity Fig. 4: Official evaluation results. (a) Naturalness MOS boxplots. The red dot indicates the mean score, and the black line indicates the median. (b) Singing style similar- ity. (c) Singer identity similarity. The black diamonds (♦) represent the binary accuracy. hurts speaker similarity and MOS. Removing pooling also degrades transfer success and increases leakage, indicat- ing that boundary-aware aggregation is important for suppressing style information in the semantic condition. Finally, regarding the contribution of the high-frequency band completion, we strongly encourage readers to listen to the ablation samples on our demo page. V. CONCLUSION In this paper, we presented System S4 for the SVCC2025 Challenge Task 1. To achieve controllable singing style conversion under strict data constraints, we proposed TABLE I: Ablation results on semantic features and the proposed semantic bottleneck (phoneme-aware pooling + scaling). SettingSpk Sim↑MOS↑Converted↑Leakage↓ ContentVec (encoder)0.7244.1872.2%13.9% WeNet (encoder)0.6944.0463.9%22.2% Whisper + pooling,λ=1.00.6924.3266.7%27.8% Whisper + pooling,λ=00.4284.1966.7%11.1% Whisper w/o pooling,λ=0.10.6654.2158.3%25.0% Whisper + pooling,λ=0.1(ours)0.7324.2283.3%8.3% a disentanglement strategy that combines a boundary- aware semantic bottleneck with an explicit technique conditioning matrix. This design effectively suppresses style leakage from content representations, while ensuring precise adherence to target technique controls. Official evaluation results demonstrate that our system achieves the 1st place in naturalness among all participating teams. The results and ablation studies collectively confirm that our approach offers a data-efficient and high-fidelity solu- tion for singing style conversion. References [1]Y. Zhang, H. Xue, H. Li, L. Xie, T. Guo, R. Zhang, and C. Gong, “Visinger 2: High-fidelity end-to-end singing voice synthesis enhanced by digital signal processing synthesizer,” arXiv preprint arXiv:2211.02903, 2022. [2]S. Liu, Y. Cao, D. Su, and H. Meng, “Diffsvc: A diffusion probabilistic model for singing voice conversion,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, p. 741–748. [3]L. P. Violeta, X. Zhang, J. Shi, Y. Yasuda, W.-C. Huang, Z. Wu, and T. Toda, “The singing voice conversion challenge 2025: From singer identity conversion to singing style conversion,” arXiv preprint arXiv:2509.15629, 2025. [4]Z. Wang, X. Xia, C. Huang, and L. Xie, “Sˆ2voice: Style-aware autoregressive modeling with enhanced conditioning for singing style conversion,” arXiv preprint arXiv:2601.13629, 2026. [5]Y. Zhang, R. Huang, R. Li, J. He, Y. Xia, F. Chen, X. Duan, B. Huai, and Z. Zhao, “Stylesinger: Style transfer for out-of- domain singing voice synthesis,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, 2024, p. 19 597–19 605. [6]R. Huang, C. Cui, F. Chen, Y. Ren, J. Liu, Z. Zhao, B. Huai, and Z. Wang, “Singgan: Generative adversarial network for high- fidelity singing voice generation,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, p. 2525– 2535. [7]A. Vyas, B. Shi, M. Le, A. Tjandra, Y.-C. Wu, B. Guo, J. Zhang, X. Zhang, R. Adkins, W. Ngan et al., “Audiobox: Unified audio generation with natural language prompts,” arXiv preprint arXiv:2312.15821, 2023. [8]Y. Zhang, C. Pan, W. Guo, R. Li, Z. Zhu, J. Wang, W. Xu, J. Lu, Z. Hong, C. Wang et al., “Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks,” Advances in Neural Information Processing Systems, vol. 37, p. 1117–1140, 2024. [9]Y. Zhang, Z. Jiang, R. Li, C. Pan, J. He, R. Huang, C. Wang, and Z. Zhao, “Tcsinger: Zero-shot singing voice synthesis with style transfer and multi-level style control,” arXiv preprint arXiv:2409.15977, 2024. [10]D. Zhang, Y. Sun, P. Li, Y. Liu, H. Lin, H. Xu, X. Mu, L. Lin, W. Yan, N. Yang et al., “Pointcot: A multi-modal benchmark for explicit 3d geometric reasoning,” arXiv preprint arXiv:2602.23945, 2026. [11]D. Zhang, H. Lin, Y. Sun, P. Wang, Q. Wang, N. Yang, and J. Zhu, “Not all queries need deep thought: Coficot for adaptive coarse-to-fine stateful refinement,” arXiv preprint arXiv:2603.08251, 2026. [12]D. Zhang, Y. Wang, Y. Sun, H. Xu, P. Fan, and J. Zhu, “Cmhanet: A cross-modal hybrid attention network for point cloud registration,” Neurocomputing, p. 133318, 2026. [13]D. Zhang, J. Zhu, S. Li, W. Yan, H. Xu, P. Fan, and H. Lu, “Igasa: Integrated geometry-aware and skip-attention modules for enhanced point cloud registration,” IEEE Transactions on Circuits and Systems for Video Technology, 2026. [14]W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhut- dinov, and A. Mohamed, “Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language pro- cessing, vol. 29, p. 3451–3460, 2021. [15]K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “Contentvec: An im- proved self-supervised speech representation by disentangling speakers,” in International conference on machine learning. PMLR, 2022, p. 18 003–18 017. [16]C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., “Neural codec language mod- els are zero-shot text to speech synthesizers,” arXiv preprint arXiv:2301.02111, 2023. [17]S. Wang and D. Borth, “Zero-shot voice conversion via self- supervised prosody representation learning,” in 2022 Interna- tional Joint Conference on Neural Networks (IJCNN). IEEE, 2022, p. 01–08. [18]Z. Wang, Y. Chen, L. Xie, Q. Tian, and Y. Wang, “Lm-vc: Zero- shot voice conversion via speech generation based on language models,” IEEE Signal Processing Letters, vol. 30, p. 1157– 1161, 2023. [19]X. Zhang, X. Zhang, K. Peng, Z. Tang, V. Manohar, Y. Liu, J. Hwang, D. Li, Y. Wang, J. Chan et al., “Vevo: Controllable zero-shot voice imitation with self-supervised disentanglement,” arXiv preprint arXiv:2502.07243, 2025. [20]Z. Ju, Y. Wang, K. Shen, X. Tan, D. Xin, D. Yang, Y. Liu, Y. Leng, K. Song, S. Tang et al., “Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,” arXiv preprint arXiv:2403.03100, 2024. [21]S. Mehta, R. Tu, J. Beskow, É. Székely, and G. E. Henter, “Matcha-tts: A fast tts architecture with conditional flow matching,” in ICASSP 2024-2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, p. 11 341–11 345. [22]L. P. Violeta, W.-C. Huang, and T. Toda, “Serenade: A singing style conversion framework based on audio infilling,” arXiv preprint arXiv:2503.12388, 2025. [23]D. Zhang, N. Yang, J. Zhu, J. Yang, M. Xin, and B. Tian, “Ascot: An adaptive self-correction chain-of-thought method for late-stage fragility in llms,” arXiv preprint arXiv:2508.05282, 2025. [24]D. Zhang, Y. Sun, C. Tan, W. Yan, N. Yang, J. Zhu, and H. Zhang, “Chain-of-thought compression should not be blind: V-skip for efficient multimodal reasoning via dual-path anchor- ing,” arXiv preprint arXiv:2601.13879, 2026. [25]N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 30, p. 495–507, 2021. [26]P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao et al., “Seed-tts: A family of high-quality versatile speech generation models,” arXiv preprint arXiv:2406.02430, 2024. [27]Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu, “Maskgct: Zero-shot text- to-speech with masked generative codec transformer,” arXiv preprint arXiv:2409.00750, 2024. [28]Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang et al., “Cosyvoice 2: Scalable streaming speech synthesis with large language models,” arXiv preprint arXiv:2412.10117, 2024. [29]T. Xie, Y. Rong, P. Zhang, W. Wang, and L. Liu, “Towards controllable speech synthesis in the era of large language models: A systematic survey,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, p. 764–791. [30]H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023. [31]X. Zhang, J. Zhang, Y. Wang, C. Wang, Y. Chen, D. Jia, Z. Chen, and Z. Wu, “Vevo2: Bridging controllable speech and singing voice generation via unified prosody learning,” arXiv e-prints, p. arXiv–2508, 2025. [32]W.-C. Huang, L. P. Violeta, S. Liu, J. Shi, and T. Toda, “The singing voice conversion challenge 2023,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, p. 1–8. [33]Y. Zhou, W. Wang, H. Ding, J. Xu, J. Zhu, X. Gao, and S. Li, “Syki-svc: Advancing singing voice conversion with post- processing innovations and an open-source professional testset,” in ICASSP 2025-2025 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP). IEEE, 2025, p. 1–5. [34]Y. Zhou, M. Chen, Y. Lei, J. Zhu, and W. Zhao, “Vits-based singing voice conversion system with dspgan post-processing for svcc2023,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, p. 1–8. [35]K. Song, Y. Zhang, Y. Lei, J. Cong, H. Li, L. Xie, G. He, and J. Bai, “Dspgan: A gan-based universal vocoder for high-fidelity tts by time-frequency domain supervision from dsp,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, p. 1–5. [36]H. Wei, X. Cao, T. Dan, and Y. Chen, “Rmvpe: A robust model for vocal pitch estimation in polyphonic music,” arXiv preprint arXiv:2306.15412, 2023. [37]Y. Song, W. Song, W. Zhang, Z. Zhang, D. Zeng, Z. Liu, and Y. Yu, “Singing voice synthesis with vibrato modeling and latent energy representation,” in 2022 IEEE 24th International Workshop on Multimedia Signal Processing (MMSP). IEEE, 2022, p. 1–6. [38]A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518. [39]B. Zhang, D. Wu, Z. Peng, X. Song, Z. Yao, H. Lv, L. Xie, C. Yang, F. Pan, and J. Niu, “Wenet 2.0: More produc- tive end-to-end speech recognition toolkit,” arXiv preprint arXiv:2203.15455, 2022. [40]M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Son- deregger, “Montreal forced aligner: Trainable text-speech align- ment using kaldi.” in Interspeech, vol. 2017, 2017, p. 498–502. [41]R. Liu, X. Wen, C. Lu, L. Song, and J. S. Sung, “Vibrato learning in multi-singer singing voice synthesis,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2021, p. 773–779. [42]J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Ad- vances in neural information processing systems, vol. 33, p. 17 022–17 033, 2020. [43]H. Wang, S. Zheng, Y. Chen, L. Cheng, and Q. Chen, “Cam++: A fast and efficient network for speaker verification using context-aware masking,” arXiv preprint arXiv:2303.00332, 2023. [44]Y. Tang, L. Liu, W. Feng, Y. Zhao, J. Han, Y. Yu, J. Shi, and Q. Jin, “Singmos-pro: An comprehensive benchmark for singing quality assessment,” arXiv preprint arXiv:2510.01812, 2025.