Paper deep dive
BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations
Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 3:20:07 AM
Summary
The paper introduces BEST-RQ-2, an evolution of the BEST-RQ self-supervised audio representation learning method. It replaces the original Conformer encoder with a Vision Transformer (ViT) and introduces a two-step 'contextualize-then-predict' architecture. This approach separates the context encoder (which processes only unmasked patches) from a lightweight predictor (which is discarded after pretraining). The authors demonstrate that while the ViT architecture shifts performance from speech toward music and environmental sounds, the two-step decomposition provides consistent performance gains across domains without increasing inference-time compute. The model was evaluated on the X-ARES and XARES-LLM benchmarks using AudioSet pretraining data.
Entities (8)
Relation Signals (6)
BEST-RQ-2 → evaluatedon → X-ARES
confidence 100% · On the X-ARES and XARES-LLM benchmarks, BEST-RQ-2 consistently outperforms one-stage baselines
BEST-RQ-2 → evaluatedon → XARES-LLM
confidence 100% · On the X-ARES and XARES-LLM benchmarks, BEST-RQ-2 consistently outperforms one-stage baselines
BEST-RQ-2 → inspiredby → Audio-JEPA
confidence 100% · a two-step encoder–predictor decomposition inspired by Audio-JEPA [10]
BEST-RQ-2 → isevolutionof → BEST-RQ
confidence 100% · We present BEST-RQ-2, an evolution of BEST-RQ
BEST-RQ-2 → pretrainedon → AudioSet
confidence 100% · All models are pretrained on AudioSet [21].
BEST-RQ-2 → usesarchitecture → ViT
confidence 100% · A ViT context encoder processes only the unmasked spectrogram regions
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a two-step contextualize-then-predict pretraining scheme. A ViT context encoder processes only the unmasked spectrogram regions, and a lightweight predictor infers targets for the masked regions; the predictor is discarded after pretraining. Replacing the original Conformer encoder with a ViT shifts performance across domains, slightly reducing speech performance while improving music and environmental sounds, with comparable average scores. The main improvement comes from decomposing masked prediction into separate contextualization and prediction stages. On the X-ARES and XARES-LLM benchmarks, BEST-RQ-2 consistently outperforms one-stage baselines in overall transfer while keeping inference compute unchanged. Code and model checkpoints are publicly available.
Tags
Links
- Source: https://arxiv.org/abs/2606.30700v1
- Canonical: https://arxiv.org/abs/2606.30700v1
Trouble viewing inline? Open PDF directly →
Full Text
28,561 characters extracted from source content.
Expand or collapse full text
BEST-RQ-2: Contextualize–Then–Predict, a Two-Step Approach for Self-Supervised Audio Representations Ludovic TUNCAY ID 1,∗ , ́ Etienne LABB ́ E ID 1 , Thomas PELLEGRINI ID 1 1 IRIT, Universit ́ e de Toulouse, CNRS, Toulouse INP, Toulouse, France ludovic.tuncay@irit.fr, etienne.labbe@irit.fr, thomas.pellegrini@irit.fr Abstract Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen random- projection-based discrete targets while introducing a two-step contextualize–then–predict pretraining scheme. A ViT context encoder processes only the unmasked spectrogram regions, and a lightweight predictor infers targets for the masked regions; the predictor is discarded after pretraining. Replacing the orig- inal Conformer encoder with a ViT shifts performance across domains, slightly reducing speech performance while improv- ing music and environmental sounds, with comparable aver- age scores. The main improvement comes from decomposing masked prediction into separate contextualization and predic- tion stages. On the X-ARES and XARES-LLM benchmarks, BEST-RQ-2 consistently outperforms one-stage baselines in overall transfer while keeping inference compute unchanged. Code 1 and model checkpoints 2 are publicly available. Index Terms: audio representation learning, self-supervised learning, audio encoder, evaluation benchmark 1. Introduction Audio representations are increasingly expected to transfer across speech, music, and environmental sound while support- ing diverse downstream uses, from probing evaluations to front- ends for audio language models. Benchmarks such as X-ARES and XARES-LLM explicitly measure this cross-domain trans- fer under complementary evaluation protocols [1, 2]. Masked prediction with discrete targets has proven effec- tive for self-supervised audio learning. BEST-RQ is particularly attractive because it relies on fixed random-projection quanti- zation, enabling stable cross-entropy training without learned codebooks [3, 4]. In this work, we revisit the speech-centric BEST-RQ with architectures better matched to generic au- dio. Replacing the original Conformer [5] encoder with a Vi- sion Transformer (ViT) requires spectrogram patch tokeniza- tion [6, 7, 8, 9], which alters the balance between temporal detail and local time–frequency texture modeling. This re- distributes performance across domains while leaving overall transfer performance largely unchanged. We introduce BEST-RQ-2, which retains BEST-RQ’s frozen discrete targets while adopting a two-step contextualize– then–predict decomposition. A ViT context encoder processes only unmasked patches, and a lightweight predictor recon- structs and predicts masked targets during pretraining before be- ing discarded at inference time. To isolate architectural effects, ** indicates the corresponding author. 1 https://github.com/LudovicTuncay/audio-embeddings 2 https://huggingface.co/ltuncay/BEST-RQ-2 we also introduce BEST-RQ (ViT), a one-stage ViT variant that shares the same tokenizer and targets as BEST-RQ-2. Our contributions are fourfold. (i) We revisit BEST-RQ us- ing a ViT encoder and show that the resulting model achieves comparable overall transfer while redistributing performance across audio domains. (i) We introduce a two-step decom- position separating contextualization from prediction. (i) We demonstrate on X-ARES and XARES-LLM that most perfor- mance gains arise from this decomposition rather than from to- kenization changes, while inference cost remains unchanged. (iv) We release code 1 and checkpoints 2 for BEST-RQ-2. 2. Proposed Method: BEST-RQ-2 BEST-RQ-2 is a masked prediction method that learns audio representations from log-mel spectrograms using discrete tar- gets from a frozen random-projection quantizer, as in BEST- RQ [3, 4]. It introduces two changes: (i) a ViT-based archi- tecture and (i) a two-step encoder–predictor decomposition in- spired by Audio-JEPA [10]. Figure 1 summarizes the pipeline. 2.1. Design rationale BEST-RQ-2 modifies BEST-RQ along two dimensions: the en- coder architecture and the prediction process. The original BEST-RQ encoder is Conformer-based [5] and includes convo- lutional modules with local receptive fields defined over a dense time axis. In BEST-RQ, masked regions are therefore provided in place at the encoder input (rather than removed), so that con- volutional neighborhoods remain well-defined and the ordering of local patterns is preserved throughout the forward pass. An encoder–predictor decomposition, where masked to- kens are dropped in the encoder and reintroduced later, is not directly compatible with this Conformer design without archi- tectural changes: removing patches breaks the dense layout as- sumed by convolutions and alters their effective neighborhoods. ViT encoders, in contrast, can operate on an arbitrary subset of patches because self-attention is informed by positional em- beddings; masked patches can be omitted in the encoder and reintroduced in a predictor using their positional indices. We therefore adopt a ViT backbone for BEST-RQ-2 to enable the encoder stage to be applied to unmasked regions only. To disentangle effects, we also introduce BEST-RQ (ViT), which replaces the Conformer encoder with a ViT but retains the same masking interface as BEST-RQ: the encoder directly receives masked patches in place. BEST-RQ (ViT) can thus be viewed either as BEST-RQ with a ViT encoder, or as BEST-RQ- 2 without the encoder–predictor decomposition. arXiv:2606.30700v1 [cs.SD] 29 Jun 2026 Figure 1: BEST-RQ-2 pretraining pipeline. A log-mel spectrogram is patchified (partitioned into patches) and masked. The ViT context encoder processes unmasked patches and a lightweight ViT predictor outputs logits over K = 8192 codes for masked patches. Targets are discrete indices generated by a frozen random-projection quantizer. Training minimizes cross-entropy on masked positions. 2.2. Patch tokenization Given a log-mel spectrogram X ∈R F×T , we partition it into non-overlapping P × P patches (P=16), yielding N = (F/P)(T /P) tokens. Each patch x i ∈R P 2 is mapped to an embedding e i ∈R D via a learned linear projection, as in ViT- style spectrogram encoders [6, 7, 8, 9]. 2.3. Masking We sample a mask ratio r and select masked indices M uni- formly without replacement over the 2D patch grid, following common masked-prediction audio pretraining setups [11, 12, 9, 13]. We set |M| = ⌊rN⌋ and define the unmasked set as V =1, . . . , N\ M . In the original BEST-RQ, masked regions are replaced by random noise and provided in place to the encoder. For transformer-based variants (BEST-RQ (ViT) and BEST-RQ-2), masked patches are represented using a learned mask token. BEST-RQ (ViT) keeps in-place masking, while BEST-RQ-2 re- moves masked patches from the encoder input and reintroduces them in the predictor with added positional embeddings. The loss is computed only on masked positions. 2.4. Two-step encoder–predictor Similar to JEPA-style architectures [14, 15, 16, 17, 18, 10, 19], BEST-RQ-2 factorizes masked prediction into a context en- coder f θ and a predictor g φ . The context encoder (ViT) pro- cesses only unmasked patch embeddings (with positional em- beddings) and outputs contextualized representations: h i i∈V = f θ e i i∈V .(1) The predictor receives (i) h i i∈V and (i) learned mask em- beddings at i∈ M (with positional embeddings), and produces masked representations: z i i∈M = g φ h i i∈V ,mask i∈M .(2) A linear classifier maps z i to logits over K codes. The predictor is used only during pretraining and discarded at inference time. 2.5. Frozen random-projection targets As in BEST-RQ [3, 4], each patch x i is mapped to a discrete target via a frozen random projection followed by nearest-code assignment. The random projection matrix R is sampled once using a Xavier/Glorot uniform initialization [20]. We compute u i = Rx i ,R∈R d×P 2 ,(3) and assign the nearest code in a fixed codebook C = c 1 , . . . , c K with c k ∈S d , the hypersphere of dimension d: y i = arg min k∈1,...,K ∥u i − c k ∥ 2 = arg max k∈1,...,K u ⊤ i c k ,(4) yielding y i ∈ 1, . . . , K. The projection matrix R and code- book C are sampled once and kept frozen during training. 2.6. Training objective Let p i (k | X V ) denote the softmax probability of code k pre- dicted for a masked patch i ∈ M , given unmasked input X V . We minimize cross-entropy over masked positions: L = 1 |M| X i∈M CE p i (·| X V ), y i .(5) 3. Implementation Details 3.1. Audio preprocessing For BEST-RQ-2, BEST-RQ (ViT) and Audio-JEPA, audio is re- sampled to 16 kHz mono and converted to a 128-bin log-mel spectrogram using a 128 ms analysis window and a 39.0625 ms hop, matching Audio-JEPA preprocessing and yielding T=256 frames for a 10 s clip.The spectrogram X ∈R 128×256 is partitioned into non-overlapping 16 × 16 patches, giving N = (128/16)(256/16) = 128 tokens. During pretraining, we sample r ∼U[0.4, 0.6] and mask exactly⌊rN⌋ patches; the loss is computed only on masked positions. 3.2. Models and training setup BEST-RQ follows the published recipe and open implemen- tation of [4].BEST-RQ-2 and BEST-RQ (ViT) use the same frozen random-projection quantizer as BEST-RQ with K = 8192 codes and projection dimension d=16 [3, 4]. The context encoder is a ViT with embedding dimension 768, depth 12, 12 attention heads, MLP ratio 4.0, drop-path 0.1, and 2D sine–cosine positional embeddings. BEST-RQ-2 uses an addi- tional lightweight ViT predictor with the same embedding and attention configuration but depth 4; it is discarded after pre- training, so inference uses only the encoder. BEST-RQ (ViT) is the one-stage ablation that uses the same patchification, mask- ing, and targets as BEST-RQ-2 but predicts masked-patch logits with a single ViT (no separate predictor). Table 1: Model size and compute for a 10 s input. Inference (Inf.) uses the encoder only. Relative speed is measured on a single NVIDIA H100 and normalized to BEST-RQ-2. Params (M) GFLOPs Rel. speed Model Train Inf. Train Inf.Inf. BEST-RQ838349.046.50.97× BEST-RQ (ViT)928523.4 22.51.0× BEST-RQ-21208530.122.51.0× 3.3. Optimization and compute BEST-RQ-2 and BEST-RQ (ViT) are trained with AdamW us- ing a learning rate of 1× 10 −4 , weight decay 0.05, and linear warmup over the first 5% of training (10k steps), followed by cosine decay, for a total of 200k optimization steps in fp32. BEST-RQ is trained for the same number of steps using the optimizer, schedule, and fp16 configuration from the open im- plementation of [4]. Models are matched in dataset split and number of optimization steps, but not strictly in the number of examples processed due to differing recommended batch sizes. All runs are executed on a single NVIDIA H100 GPU. To contextualize deployment cost, Table 1 reports parame- ter counts and forward-pass compute (GFLOPs) for a 10 s input, separately for training (including on-the-fly random-projection target computation) and inference (encoder only). We also re- port relative inference speed measured on our hardware under a fixed batch size and normalized to BEST-RQ-2; note that BEST- RQ uses fp16 while ViT-based models use fp32. Although BEST-RQ has roughly twice the inference GFLOPs, its fp16 precision provides nearly a 2× throughput advantage on H100 hardware, resulting in comparable inference speed to the fp32 ViT-based models. BEST-RQ-2 adds parame- ters only during training via the predictor, which is discarded at inference, leaving inference cost identical to BEST-RQ (ViT). 4. Experimental Setup 4.1. Pretraining data All models are pretrained on AudioSet [21]. After identical preprocessing and filtering (e.g., removal of silent or corrupted clips), the training split contains approximately 1.9M 10 s clips spanning speech, music, and environmental sounds. BEST-RQ, BEST-RQ (ViT), and BEST-RQ-2 are trained on this same split to isolate architectural effects. 4.2. Evaluation benchmarks We evaluate representations using two complementary bench- mark suites. 1. X-ARES [1]. X-ARES evaluates general-purpose audio en- coders across speech, environmental sound, and music using both linear probing (MLP) and non-parametric kNN proto- cols. We follow the official evaluation pipeline and report domain-averaged metrics. Results are reported on 21 datasets due to an issue affecting one dataset in the official release. 2. XARES-LLM [2]. XARES-LLM evaluates frozen encoders used as front-end modules for large audio language models within the AECC 2026 pipeline. The benchmark measures downstream classification and understanding performance after LALM training, providing a complementary evaluation of encoder usefulness beyond probing. Table 2: Condensed X-ARES linear-probing results (↑), aver- aged by audio domain. Bold and underlined values indicate the best and second-best score, respectively, within each column across the compared models. Env. denotes the environmental- sound domain. MoM is the mean of the three domain means; Overall is the mean over all datasets. ModelSpeech Env. Music MoM Overall data2vec [22].66.17.31.38.43 wav2vec 2.0 [23].63.29.43.45.48 Whisper [24].73.29.44.49.53 Audio-JEPA.43.31.52.42.41 BEST-RQ.59.28.41.43.45 BEST-RQ (ViT).47.33.53.44.43 BEST-RQ-2.51.41.58.50.49 4.3. Training protocol and fairness All BEST-RQ variants are pretrained for 200k optimization steps on the same AudioSet split under a comparable compute budget (∼17–18 h on a single NVIDIA H100 GPU). For BEST- RQ, we follow the implementation-recommended batch size and FP16 training setup. For BEST-RQ (ViT) and BEST-RQ- 2, we use the parameters detailed in Subsection 3.3. Compar- isons are therefore controlled for data split, optimization steps, and compute budget, but not strictly for the total number of ex- amples processed. As the closest two-step contextualize–then– predict baseline, we chose to train Audio-JEPA from scratch on the same AudioSet split, with identical preprocessing, the same 200k optimization steps, and a matched training budget (∼18 h). For other hyperparameters, we follow the open imple- mentation [10]. 5. Results All encoders are evaluated frozen and output a sequence of to- ken embeddings, as required by the X-ARES and XARES-LLM pipelines. We extract token representations from the final en- coder block at each model’s native tokenization (patches for BEST-RQ (ViT) and BEST-RQ-2; strips for BEST-RQ), with no additional pooling or temporal aggregation beyond what each benchmark applies. We report domain-averaged scores (Speech/Environment/Music) for both benchmarks and primar- ily use MoM (mean-of-means), defined as the unweighted av- erage of the three domain means, because it equally weights Speech, Environmental sound, and Music regardless of how many datasets each domain contains. We also report the dataset- level Overall mean for completeness. Because X-ARES and XARES-LLM use different dataset suites and evaluation pipelines, results should be interpreted within each benchmark and are not directly comparable across them. As noted earlier, we trained BEST-RQ variants and Audio- JEPA under our controlled pretraining protocol. We also re- evaluated all baseline encoders using the official evaluation pipelines to ensure comparability. 5.1. X-ARES 5.1.1. Linear Probing Table 2 reports linear-probing performance. Replacing the Con- former encoder with a ViT (BEST-RQ→ BEST-RQ (ViT)) re- distributes performance across domains: speech performance decreases while music and environmental sounds improve, Table 3: Condensed X-ARES kNN results (↑), averaged by au- dio domain. Bold and underlined values indicate the best and second-best score, respectively, within each column across the compared models. Env. denotes the environmental-sound do- main. MoM is the mean of the three domain means; Overall is the mean over all datasets. ModelSpeech Env. Music MoM Overall data2vec [22].44.07.13.21.31 wav2vec 2.0 [23].26.14.27.22.24 Whisper [24].33.14.32.26.29 Audio-JEPA.32.32.58.41.37 BEST-RQ.29.18.27.25.26 BEST-RQ (ViT).24.19.28.24.25 BEST-RQ-2.32.32.49.38.35 yielding similar global performance. This suggests that the en- coder change primarily alters inductive bias rather than overall representation quality. The introduction of the two-step contextualize–then– predict decomposition (BEST-RQ-2) produces consistent gains over the one-stage ViT baseline across all domains, resulting in the best mean-of-means score among BEST-RQ variants while keeping inference cost unchanged. The results therefore indi- cate that the main performance improvement arises from the prediction decomposition rather than from the ViT encoder it- self. Among broader baselines, Whisper achieves the strongest Speech and Overall scores, which is expected given its speech- heavy training and the speech-centric composition of X- ARES [1]. 5.1.2. kNN Table 3 evaluates representation geometry using kNN classifi- cation. BEST-RQ-2 outperforms both BEST-RQ and BEST-RQ (ViT) across all aggregated metrics, indicating improved neigh- borhood structure beyond patch tokenization alone. Audio- JEPA remains the strongest overall non-speech method, while BEST-RQ-2 substantially improves music and environmental- sound performance relative to BEST-RQ, reinforcing the benefit of the two-step prediction design. 5.2. XARES-LLM XARES-LLM evaluates a different usage regime from X- ARES: rather than probing generic representation quality, it measures how well frozen encoders function as front-end mod- ules for large audio language models. In its current release, XARES-LLM spans 20 datasets across two tracks, covering a broad range of tasks [2]. Models pretrained with objec- tives aligned to text-conditioned tasks (e.g., ASR or transla- tion) may be advantaged because their representations already emphasize factors useful for language-model training. More generally, transfer studies show representations transfer best to downstream tasks closer to the pretraining objective or data do- main [25, 26, 27]. Consistent with this perspective, Whisper remains a particularly strong baseline in this setting due to its text-aligned supervision. Table 4 reports results averaged across tracks. Whisper achieves the strongest global performance, while Dasheng [28] benefits from much larger-scale self-supervised pretraining (∼272.4k hours) than our AudioSet-based setup (∼5.3k hours). Within the BEST-RQ family, switching to a ViT encoder alone (BEST-RQ (ViT)) does not improve overall performance Table 4: Condensed XARES-LLM results (↑), averaged by audio domain across all tracks. Bold and underlined values indicate the best and second-best score, respectively, within each column across the compared models. Env. denotes the environmental- sound domain. MoM is the mean of the three domain means; Overall is the mean over all datasets. ModelSpeech Env. Music MoM Overall Dasheng [28].60.37.42.46.49 Whisper [24].60.41.58.53.54 Audio-JEPA.27.22.56.35.31 BEST-RQ.35.22.49.35.33 BEST-RQ (ViT).26.22.50.33.30 BEST-RQ-2.33.33.59.41.38 relative to BEST-RQ, reflecting a redistribution of performance across domains rather than a global gain. In contrast, BEST- RQ-2 consistently improves environmental and music perfor- mance while maintaining speech results comparable to those obtained with BEST-RQ. This mirrors the X-ARES observa- tions and confirms that the contextualize–then–predict decom- position, rather than tokenization itself, is the primary source of improvement. 6. Discussion Across X-ARES and XARES-LLM, replacing the Conformer in BEST-RQ with a patchified ViT primarily shifts inductive bias rather than overall transfer: speech decreases, while music and environmental sound improve, yielding similar aggregate per- formance. This is consistent with a tokenization trade-off where patch tokens emphasize local time–frequency texture, whereas strip-like tokens better preserve fine temporal cues important for speech. BEST-RQ-2’s main gains come from the two-step architec- ture design under the BEST-RQ discrete-target objective. By processing only unmasked patches in the encoder and delegat- ing masked discrete-target prediction to a lightweight module during pretraining, BEST-RQ-2 improves transfer while keep- ing inference compute unchanged by discarding the predictor. A remaining limitation is the residual speech gap relative to strip-tokenized BEST-RQ. Future work could explore alterna- tive tokenization techniques or learned speech representations to improve this trade-off. 7. Conclusion We introduced BEST-RQ-2, a self-supervised method for audio representation learning that retains BEST-RQ’s frozen random- projection targets while adopting a two-step encoder–predictor architecture. Under a controlled pretraining budget on the same AudioSet split, BEST-RQ-2 improves transfer on music and environmental sound tasks and remains competitive on speech across both X-ARES and XARES-LLM evaluations. Our analysis shows that patch tokenization primarily con- trols a token-level time–frequency trade-off, with lower per- formance on speech but higher performance on environmen- tal sounds and music, whereas the encoder–predictor decom- position provides consistent gains without increasing inference compute cost. These findings highlight the importance of sep- arating tokenization choices from prediction architecture when designing generic audio representation learners. Future work will target improving speech-task performance while preserving gains on environmental sounds and music. 8. Acknowledgments This work was granted access to the HPC resources of IDRIS under the allocation AD011014754R2 made by GENCI. Support from the ANR-3IA Artificial and Natural Intelligence Toulouse Institute ANITI (ANR-19-PI3A-0004) is gratefully acknowledged. 9. Generative AI Use Disclosure We used generative AI tools to (i) proofread and improve the clarity of the writing and (i) assist with writing and refactoring code used in the experiments. All scientific ideas, methodolog- ical choices, and the majority of the implementation and exper- imental work were produced and verified by the authors, who also reviewed and edited all AI-assisted outputs. 10. References [1] J. Zhang, H. Dinkel, Y. Niu, C. Liu, S. Cheng, A. Zhao, and J. Luan, “X-ARES: A comprehensive framework for assessing au- dio encoder performance,” in Proc. Interspeech 2025, 2025, p. 4868–4872. [2] J. Zhang and H. Dinkel, “XARES-LLM,” https://github.com/ xiaomi-research/xares-llm, 2026, GitHub repository, commit 689875b5168a4694f73b861a04982bcb67b138d1. [3] C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu, “Self-supervised learning with random-projection quantizer for speech recogni- tion,” in International Conference on Machine Learning. PMLR, 2022. [4] R. Whetten, T. Parcollet, M. Dinarelli, and Y. Est ` eve, “Open im- plementation and study of best-rq for speech processing,” in 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW). IEEE, 2024. [5] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al., “Conformer: Convolution- augmented transformer for speech recognition,” arXiv preprint arXiv:2005.08100, 2020. [6] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [7] Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self- supervised audio spectrogram transformer,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 10, 2022, p. 10 699–10 709. [8] S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: audio pre-training with acoustic tok- enizers,” in Proceedings of the 40th International Conference on Machine Learning, 2023, p. 5178–5193. [9] A. Baevski, A. Babu, W.-N. Hsu, and M. Auli, “Efficient self- supervised learning with contextualized target representations for vision, speech and language,” in International conference on ma- chine learning. PMLR, 2023, p. 1416–1429. [10] L. Tuncay, E. Labb ́ e, E. Benetos, and T. Pellegrini, “Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representa- tion Learning,” in ICME 2025, 2025. [11] P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked autoencoders that lis- ten,” Advances in Neural Information Processing Systems, vol. 35, p. 28 708–28 720, 2022. [12] A. Baade, P. Peng, and D. Harwath, “Mae-ast: Masked autoencod- ing audio spectrogram transformer,” in Proc. Interspeech 2022, 2022, p. 2438–2442. [13] D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “Masked modeling duo: Learning representations by encouraging both networks to model the input,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP). IEEE, 2023, p. 1–5. [14] M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rab- bat, Y. LeCun, and N. Ballas, “Self-supervised learning from im- ages with a joint-embedding predictive architecture,” in Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, p. 15 619–15 629. [15] Z. Fei, M. Fan, and J. Huang, “A-jepa: Joint-embedding predictive architecture can listen,” arXiv preprint arXiv:2311.15830, 2023. [16] D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “Masked modeling duo: Towards a universal audio pre-training framework,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 32, p. 2391–2406, 2024. [17] A. Riou, S. Lattner, G. Hadjeres, M. Anslow, and G. Peeters, “Stem-jepa:A joint-embedding predictive architecture for musicalstemcompatibilityestimation,”arXivpreprint arXiv:2408.02514, 2024. [18] A. Riou, S. Lattner, G. Hadjeres, and G. Peeters, “Investigat- ing design choices in joint-embedding predictive architectures for general audio representation learning,” in 2024 IEEE Interna- tional Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW). IEEE, 2024, p. 680–684. [19] M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muck- ley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus et al., “V-jepa 2: Self-supervised video models enable understanding, prediction and planning,” arXiv preprint arXiv:2506.09985, 2025. [20] X. Glorot and Y. Bengio, “Understanding the difficulty of train- ing deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics.JMLR Workshop and Conference Proceedings, 2010, p. 249–256. [21] J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, p. 776–780. [22] A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in International conference on ma- chine learning. PMLR, 2022, p. 1298–1312. [23] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,” Advances in neural information processing systems, vol. 33, p. 12 449–12 460, 2020. [24] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, p. 28 492–28 518. [25] J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transfer- able are features in deep neural networks?” Advances in neural information processing systems, vol. 27, 2014. [26] S. Kornblith, J. Shlens, and Q. V. Le, “Do better imagenet models transfer better?” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, p. 2661–2671. [27] B. Zoph, G. Ghiasi, T.-Y. Lin, Y. Cui, H. Liu, E. D. Cubuk, and Q. Le, “Rethinking pre-training and self-training,” Advances in neural information processing systems, vol. 33, p. 3833–3845, 2020. [28] H. Dinkel, Z. Yan, Y. Wang, J. Zhang, Y. Wang, and B. Wang, “Scaling up masked audio encoder learning for general audio clas- sification,” in Proc. Interspeech 2024, 2024, p. 547–551.