Paper deep dive
RIVET: Robust Idempotent Voice Attribute Editing
Dareen Alharthi, Bhuvan Koduru, Rita Singh, Bhiksha Raj
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 4:15:30 AM
Summary
RIVET (Robust Idempotent Voice Attribute Editing) is a training framework designed to improve the robustness of voice attribute editing models (such as age and gender modification) against noisy or inconsistent labels in large-scale speech datasets. The core innovation is the use of an idempotency objective—ensuring that repeated application of the editing function does not change the result—applied within the latent representation space. This acts as a regularizer that prevents identity drift and improves stability. The framework integrates an ECAPA-TDNN speaker encoder, a conditional normalizing flow, and a VITS-based speech generator, and it demonstrates superior performance in preserving speaker identity and editing success on the GLOBE and EARS datasets compared to standard training methods.
Entities (7)
Relation Signals (5)
RIVET → evaluatedon → GLOBE
confidence 100% · We evaluate RIVET ... on the GLOBE dataset
RIVET → evaluatedon → EARS
confidence 100% · We evaluate RIVET ... on the EARS dataset
RIVET → incorporates → Idempotency
confidence 100% · We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise.
RIVET → uses → ECAPA-TDNN
confidence 100% · RIVET consists of three main components: an ECAPA-TDNN speaker encoder
RIVET → uses → ViTs
confidence 100% · a VITS-based speech generator
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are often noisy or inconsistent, which can cause conditional generative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improving robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET improves editing success and better preserves speaker identity than standard training, showing that idempotency improves robustness in voice editing models.
Tags
Links
- Source: https://arxiv.org/abs/2606.19629v1
- Canonical: https://arxiv.org/abs/2606.19629v1
Trouble viewing inline? Open PDF directly →
Full Text
27,760 characters extracted from source content.
Expand or collapse full text
RIVET: Robust Idempotent Voice Attribute Editing Dareen Alharthi 1 , Bhuvan Koduru 1 , Rita Singh 1 , Bhiksha Raj 1 1 Carnegie Mellon University, Pittsburgh, PA, USA dalharth@cs.cmu.edu, bkoduru@cs.cmu.edu, rsingh@cs.cmu.edu, bhiksha@cs.cmu.edu Abstract Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large- scale speech datasets, however, attribute annotations are of- ten noisy or inconsistent, which can cause conditional gener- ative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improv- ing robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET im- proves editing success and better preserves speaker identity than standard training, showing that idempotency improves robust- ness in voice editing models. Code is available online. 1 Index Terms: voice editing, accent conversion, voice aging, voice biometrics, idempotency, noisy labels 1. Introduction Voice attribute editing aims to modify specific characteristics of a speech signal, such as age, gender, or accent, while pre- serving the speaker’s underlying identity. Recent generative models have made significant progress in enabling controllable speech editing through conditional synthesis and disentangled representations [1, 2, 3]. In these systems, editing is typically achieved by conditioning the model on attribute labels or by manipulating dedicated latent variables corresponding to inter- pretable factors of variation. However, methods that rely on explicit attribute supervision critically depend on the quality of the labels used during training. In practice, large-scale speech datasets often contain noisy or inconsistent attribute annota- tions. Labels may be flipped, ambiguous, weakly defined, or inferred automatically rather than verified manually [4, 5]. When trained with noisy supervision, deep generative mod- els may learn incorrect associations between attributes and acoustic patterns, leading to degraded conditional generation quality. Prior work has shown that models trained on noisy labels can memorize spurious correlations and exhibit reduced generalization performance [6]. In conditional generative set- tings such as voice editing, noisy attribute conditioning can dis- tort the learned conditional mappings, resulting in unstable edit- ing behavior and entanglement between speaker identity and the specified attribute [7, 8]. 1 https://github.com/DareenHarthi/rivet In the setting of editing, noisy or imprecise training labels directly affect the learned mapping from an input with a given attribute to the output in which that attribute has been modified. Repeated application of this learned, noisy editing process can lead to progressive drift, which drags both the edited attribute and the underlying speech representation itself away from their intended values. One way of minimizing the influence of the noise in the label is to minimize this drift: intuitively, a se- quence of edits that eventually restores an attribute to its original value should also return the speech representation to its original, unedited form. This observation naturally leads to the concept of idempotency [9, 10, 11, 12]. An idempotent model satisfies the property that once an output lies on the target data manifold, further applications of the model do not change it. Formally, for an input x, an idem- potent function f satisfies f(f(x)) = f(x). This property en- courages stability under repeated generation and prevents drift during editing. Several recent works have explored enforcing idempotency in generative models [9, 12, 11, 13]. While some approaches suffer from issues such as blurriness or mode col- lapse [9], and others are difficult to scale to more complex ar- chitectures [12, 11], these studies suggest that idempotency can serve as a powerful constraint for stabilizing generative models. In this work, we leverage idempotency as a regularization principle for attribute-conditioned voice editing under noisy supervision. We introduce RIVET (Robust Idempotent Voice Attribute Editing), a training framework that incorporates an idempotency objective into a conditional voice editing model, similar in spirit to prior work on idempotent generative mod- els [9]. Instead of enforcing idempotency directly in the output space, RIVET applies the constraint in the latent representation space, encouraging the model to map inputs to stable points on the attribute-conditioned manifold and reducing sensitivity to mislabeled training examples. We evaluate RIVET on the GLOBE dataset [5], which con- tains naturally occurring label noise, and on the EARS dataset [14] with controlled levels of synthetic label noise. Our ex- periments focus on two attributes, age and gender. The results show that RIVET improves editing success rates compared to strong baselines while preserving speaker identity. Although our evaluation focuses on these two attributes, the principle of enforcing idempotency is model-agnostic and can be applied to other attributes and editing architectures. Our contributions are summarized as follows: • We demonstrate that idempotency acts as an effective robust- ness regularizer for attribute-conditioned voice editing under noisy labels. • We introduce RIVET, a simple and general training frame- work that incorporates idempotent constraints into voice edit- ing models without requiring architectural changes. arXiv:2606.19629v1 [cs.SD] 17 Jun 2026 • We open-source RIVET; to the best of our knowledge, this is the first open-source voice editing framework. The remainder of the paper is organized as follows. Sec- tion 2 reviews related work on voice editing and idempotent generative modeling. Section 3.1 describes the RIVET frame- work and training objectives. Section 4 presents the experimen- tal setup. Section 5 presents the results and discussion. Sec- tion 6 concludes with a discussion of implications and future research directions. 2. Related Work 2.1. Voice Editing Voice editing aims to modify attributes of a speech signal, such as age, gender, or accent, while preserving speaker identity and linguistic content [1, 2, 15]. A common approach is to learn representations that separate factors of variation so that one at- tribute can be modified without affecting others. Prior work en- forces structured latent spaces [15] or learns attribute-specific directions in embedding space [1], enabling editing through la- tent manipulation or conditional transformations. However, per- fectly disentangling these factors is difficult in practice. Strong constraints may reduce reconstruction quality or model expres- siveness [16, 17]. More recent work instead focuses on pre- serving identity and content during transformation [18, 19, 20]. Despite these advances, stability remains a challenge: repeated or reversed edits can cause identity drift or artifacts, suggest- ing that additional constraints are needed to encourage stable editing behavior. 2.2. Idempotent Models Idempotency has recently emerged as a useful property in gen- erative models [9, 10, 11, 12]. A function is idempotent if re- peated application does not change the output, meaning that once a sample lies on the data manifold, further applications of the model should leave it unchanged. The Idempotent Genera- tive Network [9] introduced explicit losses to encourage stable fixed points in generative models. Later work proposed alter- native mechanisms for enforcing idempotent behavior, includ- ing distilling idempotent mappings from diffusion model scores [12] and using idempotency as a general optimization objec- tive for test-time adaptation in place of auxiliary self-supervised tasks [10]. Other approaches enforce idempotency through al- gorithmic updates that progressively move a model toward an idempotent operator during training [11]. Most existing work focuses on image generation and assumes clean supervision. In contrast, we apply idempotent constraints to conditional voice editing and show that they can improve stability and robustness under noisy attribute labels. 2.3. Noisy Labels in Conditional Generation Large-scale datasets often contain noisy or unreliable labels due to the difficulty of maintaining high-quality annotations at scale. Prior work has proposed several strategies to mitigate noisy su- pervision in conditional generative models. One approach in- corporates estimates of label reliability or confidence into the conditional input [8, 21]. Another attempts to correct noisy labels by estimating the underlying clean label distribution or modifying the training objective [7, 22]. A third line of work enforces prediction consistency during training. For example, consistency can be encouraged across augmented views of the same sample [23] or among neighboring samples in represen- Figure 1: RIVET training framework. The input speech is encoded to obtain speaker and speech representations. The speaker embedding is edited using a conditional flow, and the speech is reconstructed by the generator. The reconstructed speech is then re-encoded to enforce idempotency. A stop- gradient operation is applied before the second decoding step. tation space [24, 25]. Such regularization discourages models from fitting incorrect annotations and improves robustness to label noise. In this work, we explore idempotency as a comple- mentary regularization mechanism for improving robustness to noisy labels in conditional voice editing. 3. Method 3.1. Idempotent Training Let F(·) denote the overall editing model. The model first en- codes the input speech signal x using an encoder E(·) and then reconstructs or edits the speech using a decoder D(·): F(x) = D(E(x)).(1) An operator F is called idempotent if repeated application does not change the result: F(F(x)) = F(x).(2) Substituting the encoder–decoder structure gives D(E(D(E(x)))) = D(E(x)).(3) A sufficient condition for this relation to hold is that the encoded representation remains unchanged after reconstruction: E(D(E(x))) = E(x).(4) We therefore enforce idempotency by encouraging consis- tency in the encoded representation. Let z = E(x), z re = E(D(E(x))).(5) The idempotency loss penalizes differences between these representations: L idemp =E x∼p X ∥sg(z)− z re ∥ 2 2 ,(6) where sg(·) denotes the stop-gradient operator; it fixes the original latent z as a stable target so gradients update only the re-encoded representation z re , preventing trivial co-adaptation Figure 2: Performance under increasing label noise on the EARS dataset. Models are trained on 7 hours of EARS with increasing noise levels and evaluated on a balanced 1-hour test set. We report cosine similarity between Titanet embeddings of the original and reverted speech (left) and attribute accuracy (right). RIVET maintains higher identity similarity and more stable performance than the baseline as noise increases. between the two branches. The final training objective com- bines the standard generative loss with the idempotency regu- larizer: L =L gen + λL idemp ,(7) where λ controls the strength of the idempotency con- straint. 3.2. RIVET As shown in Figure 1, RIVET consists of three main compo- nents: an ECAPA-TDNN speaker encoder [26], a conditional normalizing flow for attribute editing [27], and a VITS-based speech generator [28]. The overall architecture is similar to VoiceShop [1], but differs in the training strategy used to en- force idempotency. Given input speech x, the ECAPA en- coder extracts a fixed-length speaker embedding e that repre- sents speaker identity. To perform attribute editing, this embed- ding is transformed using a conditional normalizing flow con- ditioned on demographic attributes such as age and gender. The flow maps the original embedding e under the target attribute condition vector c to produce the edited latent z g , z g , log| detJ| = f θ (e, c),(8) where J denotes the Jacobian of the transformation. The flow is trained with a maximum likelihood objective L flow =E 1 2 |z g | 2 − log| detJ| .(9) The speaker embedding is used as conditioning information in both the encoder and decoder of the VITS generator to syn- thesize the speech signal. During training, the generated speech is re-encoded to com- pute the idempotency loss described in Section 3.1. As shown in Figure 1, the idempotency constraint is applied to both the speaker encoder (ECAPA-TDNN) and the speech encoder of the VITS generator, encouraging their representations to remain stable after reconstruction. The ECAPA encoder is addition- ally trained with attribute classification objectives for age and gender, denoted by L age and L gender . The VITS generator is trained using its standard objectives, including adversarial, fea- ture matching, mel reconstruction, KL divergence, and duration losses. The conditional flow is optimized using the maximum likelihood objective defined in Eq. 9. The overall training objec- tive combines the VITS generator and discriminator losses, the Figure 3: Human evaluation of age and gender voice editing. Each sample was rated by five annotators with majority voting. RIVET improves editing success over the baseline. ECAPA classification losses, the flow likelihood loss, and the idempotency regularizer applied to both the speaker and speech encoders: L total =L VITS +λ f L flow +λ a L age +λ g L gender +λ i L idemp , (10) whereL VITS denotes the standard VITS generator and dis- criminator losses, and the λ terms control the contribution of each objective. All components are optimized jointly in an end- to-end training framework. In contrast to VoiceShop [1], which trains the editing modules and speaker embeddings separately, RIVET trains all components jointly. 4. Experimental Setup To evaluate the effect of idempotent training, we compare RIVET against a baseline model with identical architecture and training configuration, but without the idempotency constraint. The baseline includes the ECAPA-TDNN speaker encoder, the conditional invertible flow, and the VITS generative backbone, trained jointly using the same objectives. The only difference is that the baseline removes the idempotency regularization term. For RIVET, we add the idempotency loss to both the speaker embedding space and the speech latent representation. All other training settings, optimization schedules, and hyperparameters are kept identical to ensure a fair comparison. Both models are trained on the GLOBE dataset [5], which contains approxi- mately 535 hours of English speech from 23,519 speakers span- Table 1: Cosine similarity between Titanet speaker embeddings of the original speech and the reverted speech, attribute accu- racy, UTMOS, and WER. The models are trained on GLOBE dataset. MethodCosineAccuracy UTMOS WER Age Gender Age GenderAvgAvg GLOBE GT--62.884.93.591.89 Baseline 0.630.5439.977.23.17 10.33 RIVET 0.66 0.55 40.6 85.93.1910.68 EARS (OOD) GT--99.199.93.971.90 Baseline 0.490.45 33.677.62.724.68 RIVET 0.55 0.4830.1 92.72.864.67 Figure 4: Cosine similarity between Titanet embeddings of the original speech and reconstructed samples for 20 speakers from the GLOBE test set over 20 reconstruction rounds. Each round uses the output of the previous round as input. The baseline shows rapid identity drift, while RIVET maintains higher simi- larity to the original speaker across iterations. ning 164 global accents and a wide age range [5]. The train- ing set is the official GLOBE training split. Quantitative re- sults are reported in Table 1. To further study robustness to label noise, we conduct controlled experiments using the EARS dataset [14]. EARS is a high-quality speech dataset [14]. From this dataset, we select only neutral speech samples (excluding emotional recordings). We then construct a training set of ap- proximately 7 hours and a balanced test set of about 1 hour. Using this data, we simulate different levels of label noise by randomly flipping the age and gender labels in the training set. Specifically, we generate training sets with noise levels of 10%, 20%, 30%, 40%, 50%, and 60%. The results are shown in 2. 5. Results and Discussion 5.1. Evaluation Metrics We measure speaker identity preservation using cosine similar- ity between Titanet speaker embeddings [29] extracted from the original speech and the reverted speech. The reverted speech is obtained by first editing an attribute (e.g., age or gender) and then reversing the edit back to its original value. This metric captures both identity preservation and the stability of the edit- ing process. Attribute editing success is evaluated using independently trained classifiers for age and gender prediction. These clas- sifiers use the same ECAPA-TDNN architecture employed in RIVET but are trained separately on the GLOBE and EARS datasets and are used only for evaluation. Accuracy is com- puted on the edited speech generated by each model. Because GLOBE contains noisy demographic annotations, the classifiers are trained using a soft-label objective inspired by the EM for- mulation for learning with imprecise labels [4]. In GLOBE, age is provided in eight decade-level bins from 10 to 80. For evalua- tion, we group these bins into three categories—young (10–29), adult (30–49), and old (50+), while gender is treated as a binary classification task. Naturalness and intelligibility are evaluated using UTMOS [30] and word error rate (WER) computed using Whisper large-v2 [31]. Both metrics are measured on the edited speech generated by each model. Finally, we conduct a perceptual evaluation using Amazon Mechanical Turk, where each audio sample is rated by five inde- pendent listeners. Annotators select the perceived age and gen- der from four options: young female, old female, young male, and old male. The evaluation includes 200 samples (25 per cat- egory for each model). For age editing, speakers older than 40 are edited to age 20, while speakers younger than 40 are edited to age 50. Success is computed using majority voting across listeners. Results are shown in Figure 3. 5.2. Discussion Table 1 reports results on the GLOBE test set and the EARS dataset, which we use as an out-of-distribution (OOD) bench- mark. RIVET consistently improves cosine similarity compared to the baseline, indicating stronger speaker identity preserva- tion. It also improves attribute prediction accuracy on GLOBE, particularly for gender. On EARS, which is not seen during training, RIVET maintains higher cosine similarity and sub- stantially improves gender editing accuracy, while age accuracy slightly decreases. Both models maintain comparable natural- ness and intelligibility scores. Figure 4 illustrates identity sta- bility across repeated reconstruction rounds. RIVET maintains higher similarity to the original speaker embedding than the baseline. To further analyze robustness to noisy supervision, we train models on the EARS dataset with increasing levels of syn- thetic label noise and evaluate them on a balanced 1-hour EARS test set. As shown in Figure 2, RIVET maintains higher cosine similarity across noise levels and exhibits more stable behavior as noise increases. Human evaluation results in Figure 3 show similar trends, with RIVET improving perceived editing suc- cess for both attributes, particularly for gender editing. Overall, these results show that idempotent training reduces sensitivity to label noise. 6. Conclusion In this work, we studied how idempotency can improve ro- bustness in attribute-conditioned voice editing when training data contains noisy labels. We introduced RIVET, an end-to- end training framework that incorporates an idempotency con- straint into the latent representations of a conditional voice edit- ing model. Experiments on the GLOBE and EARS datasets show that idempotent training improves identity preservation and editing success rates while maintaining comparable percep- tual quality to the baseline. These results suggest that idempo- tency can serve as a practical regularization mechanism when attribute labels are imperfect or noisy. This work represents an initial step toward understanding how idempotent objectives improve robustness in conditional generative models. Overall, idempotency offers a promising direction for improving stabil- ity and robustness in voice editing systems. Future work will explore applying idempotent objectives to additional attributes beyond age and gender and studying their effects in both super- vised and unsupervised editing settings. 7. Generative AI Use Disclosure Large language model (LLM) tools were used to assist with proofreading and improving the clarity and fluency of the manuscript. All scientific content, including ideas, methods, experimental design, analysis, and results, was developed and verified by the authors. No generative AI tool is listed as a co- author, and the authors take full responsibility for the contents of this paper. 8. References [1] P. Anastassiou, Z. Tang, K. Peng, D. Jia, J. Li, M. Tu, Y. Wang, Y. Wang, and M. Ma, “Voiceshop: A unified speech-to-speech framework for identity-preserving zero-shot voice editing,” arXiv preprint arXiv:2404.06674, 2024. [2] Z.-Y. Sheng, L.-J. Liu, Y. Ai, J. Pan, and Z.-H. Ling, “Voice at- tribute editing with text prompt,” IEEE Transactions on Audio, Speech and Language Processing, 2025. [3] Z. Ye, Z. Ju, H. Liu, X. Tan, J. Chen, Y. Lu, P. Sun, J. Pan, W. Bian, S. He et al., “Flashspeech: Efficient zero-shot speech synthesis,” in Proceedings of the 32nd ACM International Con- ference on Multimedia, 2024, p. 6998–7007. [4] H. Chen, A. Shah, J. Wang, R. Tao, Y. Wang, X. Li, X. Xie, M. Sugiyama, R. Singh, and B. Raj, “Imprecise label learning: A unified framework for learning with various imprecise label con- figurations,” Advances in Neural Information Processing Systems, vol. 37, p. 59 621–59 654, 2024. [5] W. Wang, Y. Song, and S. Jha, “Globe: A high-quality english corpus with global accents for zero-shot speaker adaptive text-to- speech,” arXiv preprint arXiv:2406.14875, 2024. [6] B. Fr ́ enay and M. Verleysen, “Classification in the presence of label noise: a survey,” IEEE transactions on neural networks and learning systems, vol. 25, no. 5, p. 845–869, 2013. [7] D. N. Cong, H. T. Bao, and T. Hoang-Thanh, “Guiding noisy la- bel conditional diffusion models with score-based discriminator correction,” in Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2025, p. 18 531–18 541. [8] B. Na, Y. Kim, H. Bae, J. H. Lee, S. J. Kwon, W. Kang, and I.- C. Moon, “Label-noise robust diffusion models,” arXiv preprint arXiv:2402.17517, 2024. [9] A. Shocher, A. Dravid, Y. Gandelsman, I. Mosseri, M. Rubinstein, and A. A. Efros, “Idempotent generative network,” arXiv preprint arXiv:2311.01462, 2023. [10] N. Durasov, A. Shocher, D. Oner, G. Chechik, A. A. Efros, and P. Fua, “It 3 : Idempotent test-time training,” arXiv preprint arXiv:2410.04201, 2024. [11] N. B. Jensen and J. Vicary, “Enforcing idempotency in neural networks,” in Forty-second International Conference on Machine Learning, 2025. [12] S. Zaman, C. Liu, and K. Chiu, “Score-based idempotent distilla- tion of diffusion models,” arXiv preprint arXiv:2509.21470, 2025. [13] Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency models,” in Proceedings of the 40th International Conference on Machine Learning, 2023, p. 32 211–32 252. [14] J. Richter, Y.-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watan- abe, A. Richard, and T. Gerkmann, “EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and derever- beration,” in ISCA Interspeech, 2024, p. 4873–4877. [15] W. Lin, C. He, M.-W. Mak, J. Lian, and K. A. Lee, “Voxgene- sis: Unsupervised discovery of latent speaker manifold for speech synthesis,” arXiv preprint arXiv:2403.00529, 2024. [16] C. P. Burgess, I. Higgins, A. Pal, L. Matthey, N. Watters, G. Des- jardins, and A. Lerchner, “Understanding disentangling inβ-vae,” arXiv preprint arXiv:1804.03599, 2018. [17] M. Shukor, X. Yao, B. B. Damodaran, and P. Hellier, “Semantic unfolding of stylegan latent space,” in 2022 IEEE International Conference on Image Processing (ICIP).IEEE, 2022, p. 221– 225. [18] H. Chen, Y. Zhang, S. Wu, X. Wang, X. Duan, Y. Zhou, and W. Zhu, “Disenbooth: Identity-preserving disentangled tun- ing for subject-driven text-to-image generation,” arXiv preprint arXiv:2305.03374, 2023. [19] P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. M ̈ uller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel et al., “Scaling rectified flow transformers for high-resolution image synthesis,” in Forty- first international conference on machine learning, 2024. [20] L. Rout, Y. Chen, N. Ruiz, C. Caramanis, S. Shakkottai, and W.-S. Chu, “Semantic image inversion and editing us- ing rectified stochastic differential equations,” arXiv preprint arXiv:2410.10792, 2024. [21] N. Dufour, V. Besnier, V. Kalogeiton, and D. Picard, “Don’t drop your samples! coherence-aware training benefits conditional dif- fusion,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024, p. 6264–6273. [22] T. Kaneko, Y. Ushiku, and T. Harada, “Label-noise robust gen- erative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, p. 2467–2476. [23] E. Englesson and H. Azizpour, “Consistency regulariza- tion can improve robustness to label noise,” arXiv preprint arXiv:2110.01242, 2021. [24] D. Cheng, Y. Ning, N. Wang, X. Gao, H. Yang, Y. Du, B. Han, and T. Liu, “Class-dependent label-noise learning with cycle- consistency regularization,” Advances in Neural Information Pro- cessing Systems, vol. 35, p. 11 104–11 116, 2022. [25] A. Iscen, J. Valmadre, A. Arnab, and C. Schmid, “Learning with neighbor consistency for noisy labels. 2022 ieee,” in CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), 2022, p. 4662–4671. [26] B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa- tdnn:Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,” arXiv preprint arXiv:2005.07143, 2020. [27] J. Ho, X. Chen, A. Srinivas, Y. Duan, and P. Abbeel, “Flow++: Improving flow-based generative models with variational dequan- tization and architecture design,” in International conference on machine learning. PMLR, 2019, p. 2722–2730. [28] J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Inter- national Conference on Machine Learning.PMLR, 2021, p. 5530–5540. [29] N. R. Koluguri, T. Park, and B. Ginsburg, “Titanet: Neural model for speaker representation with 1d depth-wise separable convo- lutions and global context,” in ICASSP 2022-2022 IEEE inter- national conference on acoustics, speech and signal processing (ICASSP). IEEE, 2022, p. 8102–8106. [30] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “Utmos: Utokyo-sarulab system for voicemos challenge 2022,” arXiv preprint arXiv:2204.02152, 2022. [31] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518.