Paper deep dive
Membership Inference Attacks against Large Audio Language Models
Jia-Kai Dong, Yu-Xiang Lin, Hung-Yi Lee
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/31/2026, 2:24:59 AM
Summary
This paper presents the first systematic evaluation of Membership Inference Attacks (MIA) on Large Audio Language Models (LALMs). The authors demonstrate that high MIA performance in existing benchmarks is often a result of spurious acoustic distribution shifts rather than genuine model memorization. They introduce a three-phase auditing framework—including a Multi-modal Blind Baseline—to disentangle distributional artifacts from true memorization, revealing that LALM memorization is primarily cross-modal, specifically binding speaker vocal identity with textual content.
Entities (5)
Relation Signals (3)
Audio-Flamingo-3 → evaluatedby → Membership Inference Attack
confidence 100% · We conduct the first systematic MIA study on LALM pre-training corpora by leveraging two state-of-the-art open-source models: Audio-Flamingo 3 (AF3)
LALM memorization → ischaracterizedby → cross-modal binding
confidence 95% · The results reveal that LALM memorization is cross-modal, arising only from binding a speaker’s vocal identity with its text.
Multi-modal Blind Baseline → identifies → distributional shortcuts
confidence 90% · This protocol identifies distributional shortcuts that models may exploit
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present the first systematic Membership Inference Attack (MIA) evaluation of Large Audio Language Models (LALMs). As audio encodes non-semantic information, it induces severe train and test distribution shifts and can lead to spurious MIA performance. Using a multi-modal blind baseline based on textual, spectral, and prosodic features, we demonstrate that common speech datasets exhibit near-perfect train/test separability (AUC approximately 1.0) even without model inference, and the standard MIA scores strongly correlate with these blind acoustic artifacts (correlation greater than 0.7). Using this blind baseline, we identify that distribution-matched datasets enable reliable MIA evaluation without distribution shift confounds. We benchmark multiple MIA methods and conduct modality disentanglement experiments on these datasets. The results reveal that LALM memorization is cross-modal, arising only from binding a speaker's vocal identity with its text. These findings establish a principled standard for auditing LALMs beyond spurious correlations.
Tags
Links
- Source: https://arxiv.org/abs/2603.28378v1
- Canonical: https://arxiv.org/abs/2603.28378v1
Trouble viewing inline? Open PDF directly →
Full Text
34,857 characters extracted from source content.
Expand or collapse full text
Membership Inference Attacks against Large Audio Language Models Jia-Kai Dong 1 , Yu-Xiang Lin 1 , Hung-yi Lee 1,2 1 National Taiwan University 2 NTU Artificial Intelligence Center of Research Excellence b11901067@ntu.edu.tw Abstract We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs. As audio encodes non-semantic information that induces severe train and test distribution shifts and can lead to spurious MIA performance. Using a multi- modal blind baseline based on textual, spectral and prosodic features, we demonstrate that common speech datasets exhibit near-perfect train/test separability (AUC ≈ 1.0) even without model inference, and the standard MIA scores strongly corre- late with these blind acoustic artifacts (r > 0.7). Using this bilnd baseline, we identify the distribution-matched datasets enable reliable MIA evaluation without distribution shift con- founds. We benchmark multiple MIA methods and conduct modality disentanglement experiment on these datasets. The results reveal that LALM memorization is cross-modal, arising only from binding a speaker’s vocal identity with its text. These findings establish a principled standard for auditing LALMs be- yond spurious correlations. Index Terms: Membership Inference Attacks, Large Audio Language Models, Privacy Auditing 1. Introduction Large Audio Language Models (LALMs) [1–7] have recently emerged as a powerful paradigm for multimodal reasoning, uni- fying tasks such as speech recognition and audio captioning within a single framework. While these models achieve impres- sive generalization by training on massive web-scale datasets, they raise critical concerns regarding privacy and copyright, particularly whether sensitive audio content is memorized dur- ing training. Membership Inference Attacks (MIA), which aim to de- termine whether a specific sample was used during model training, have become a standard tool for auditing such risks [8, 9]. Although MIA has been extensively studied in text-based LLMs [10–16] and VLMs [17, 18], its applicability to LALMs remains unexplored. Existing privacy studies in the audio do- main focus mainly on representation learning or speaker recog- nition systems [19, 20], where attacks target fixed-dimensional embeddings rather than sequence-level generation. In contrast, LALMs are trained on massive web-scale corpora with lim- ited exposure, often only a single epoch. While this regime is commonly assumed to mitigate verbatim memorization [11], the high capacity of LALMs still enables membership leakage through memorized cross-modal mappings between acoustic signals and text. Such leakage reveals speaker–content associa- tions, constituting a biometric privacy risk that is more granular and severe than textual duplication. Auditing LALMs is further complicated by the continuous nature of audio, which encodes rich non-semantic cues such as speaker identity and recording conditions. Consequently, pri- vacy threats in LALMs are dominated not by copyrighted con- tent, but by irreversible identity–content binding. Moreover, standard audio benchmarks often exhibit acoustic distribution shifts between training and test sets that can spuriously inflate MIA performance. As demonstrated in LLMs [21], simple blind baselines that exploit such dataset artifacts can outper- form sophisticated MIAs, raising concerns about whether re- ported leakage truly reflects memorization [22]. We conduct the first systematic MIA study on LALM pre- training corpora by leveraging two state-of-the-art open-source models: Audio-Flamingo 3 (AF3) [5] and Music-Flamingo (MF) [23]. The transparent data provenance of these models en- ables a direct audit of the foundation weights themselves. This approach sidesteps the use of computationally prohibitive and often unfaithful surrogate “shadow” models, enabling an au- thentic evaluation of memorization across eight diverse datasets spanning ASR, Audio Captioning, and Music Understanding. Furthermore, to disentangle genuine memorization from dataset-induced artifacts, we introduce a Multi-modal Blind Baseline framework that quantifies distribution shifts across metadata, textual content, and acoustic features. Our contributions are threefold: • We present the first comprehensive benchmark of sample- level MIA for LALMs in an audio/text-to-text setting. We evaluate seven confidence-based MIA attacks methods across eight audio-centric tasks. • We propose a rigorous Multi-modal Blind Baseline protocol to audit privacy risks arising from dataset design. We show that common benchmark practices, introduce severe acoustic distribution shifts that inflate MIA AUCs and strongly corre- late with blind acoustic classifiers (r > 0.4). • On distribution-matched datasets, we show that LALM mem- orization is cross-modal, depending on the specific pairing of acoustic attributes such as speaker identity and text. Privacy risks are thus highest when sensitive content is tightly linked to a particular speaker. Overall, our findings highlight the fragility of naive pri- vacy assessments in the audio domain and underscore the im- portance of controlling for acoustic distribution shifts when au- diting LALMs. We argue that reported MIA results should be interpreted in conjunction with blind baseline diagnostics; with- out such analysis, it is difficult to disentangle genuine memo- rization from distributional artifacts. 2. Methodology We propose a multi-phase privacy auditing framework, as il- lustrated in Figure 1. Our approach begins by establishing a arXiv:2603.28378v1 [cs.SD] 30 Mar 2026 Figure 1: Overview of our proposed 3-phase LALM privacy au- diting framework. (1) Multi-Modal Blind Baseline: A model- agnostic classifier first identifies and filters out samples with confounding distributional shifts to produce a ”clean dataset.” (2) MIA Auditing: The target LALM is then audited on this clean dataset using a suite of membership indicators derived from a 2-stage generation process. (3) Modality Disentangle- ment: Finally, we systematically probe any detected memoriza- tion to characterize its nature, specifically identifying cross- modal binding. Multi-modal Blind Baseline (Phase 1) to quantify any distri- bution shifts inherent in the data, without accessing the target LALM. We then perform the primary MIA Auditing on the model (Phase 2) and, crucially, correlate its success with the blind baseline’s performance. This diagnostic step allows us to disentangle spurious MIA success driven by data artifacts from genuine model memorization. Finally, for cases of validated memorization, we conduct Modality Disentanglement exper- iments (Phase 3) to characterize the memory’s nature, specifi- cally testing for the existence of cross-modal binding. These experiments are designed to test the hypothesis that privacy leakage in LALMs is driven by cross-modal binding, where memorization is only triggered when the model encoun- ters the precise pairing of original acoustic features and their corresponding textual sequences. 2.1. Inference Protocol: Two-Stage Generation To simulate a realistic gray-box auditing scenario where precise ground-truth transcripts used during training may be unavail- able, we adopt a Two-Stage Generation protocol: Stage 1: Autonomous Decoding. Given an audio input x, the target LALM performs greedy decoding to produce a self- generated output sequence ˆy. Stage 2: Self-Conditioned Scoring. We treat ˆy as pseudo- ground-truth and re-run the forward pass to obtain token-level logits corresponding to ˆy. Membership metrics are then com- puted based on these self-conditioned probabilities. 2.2. Phase 1: Multi-modal Bias Audit Protocol To prevent spurious MIA success and ensure the validity of pri- vacy assessments, we first propose a systematic Multi-modal Blind Baseline framework. This protocol identifies distribu- tional shortcuts that models may exploit across two hierarchical levels: (1) Inter-dataset Shift, where mismatches between dis- parate data sources cause MIA metrics to act as trivial domain classifiers; and (2) Intra-dataset Shift, where subtle statistical discrepancies arise between the training and test splits of the same dataset. In the audio domain, such shifts are particularly common, as train and test splits are often constructed using dif- ferent curation or preprocessing pipelines; for example, in Gi- gaSpeech [24], transcription standards lead to systematic dif- ferences between splits. Our protocol quantifies these shifts by training blind Logistic Regression classifiers on data-intrinsic features without accessing the LALMs. To isolate the contribu- tion of different sources of distribution shift, we construct three blind baselines at increasing levels of modality coverage: • Metadata: Structural statistics including audio duration, file size, and word/character counts. • Text: Utterance-level TF-IDF representations (unigrams and bigrams). • Acoustic: Fixed-dimensional utterance-level descriptors ob- tained by aggregating frame-wise low-level acoustic features, including MFCCs, spectral statistics (centroid, bandwidth, rolloff), pitch, energy (RMS), and zero-crossing rate. These features are specifically designed to capture recording con- ditions, channel characteristics, and dataset-specific acoustic artifacts rather than linguistic content. All features are aggregated into fixed-dimensional, sample- level representations. We adopt simple Logistic Regression classifiers to ensure that strong blind performance reflects salient distributional artifacts rather than classifier capacity. We utilize the Pearson (r) correlation coefficient between these blind scores and the subsequent model-based MIA indica- tors (defined in Section 2.3) as a diagnostic reference. 1 Rather than applying a fixed threshold, we argue that a high Blind Baseline AUC, when accompanied by a non-negligible posi- tive correlation with the model’s MIA scores, provides strong empirical evidence that the detected membership is likely con- founded by distributional shortcuts. In such cases, the reported MIA performance may reflect the LALM’s sensitivity to data- intrinsic artifacts rather than genuine memorization, thereby ne- cessitating a more cautious interpretation of the privacy risks. 2.3. Membership Inference Indicators We compute several complementary metrics from the model’s output probabilities to serve as membership indicators: • Perplexity (PPL): The average negative log-likelihood (NLL) of the sequence ˆy [25]. • Shannon Entropy: The average entropy of the predicted to- ken distributions, measuring predictive certainty [26]. • Min-k% / Max-k% Prob: Average NLL of the k% tokens with lowest and highest probabilities, respectively [27, 28]. • Max-R ́ enyi MIA: Measures probability mass concentration via R ́ enyi divergence with orders α∈0, 1, 2,∞ [17]. • Zlib Ratio: The NLL normalized by the zlib compression size of the text [25]. • Max Probability Gap: The sequence-averaged gap between top-1 and top-2 token probabilities, capturing over-confident predictions indicative of memorized tokens [17]. Following [9], we concatenate these metrics into a feature vec- tor and train a Logistic Regression classifier using stratified 10- fold cross-validation within each dataset to produce member- ship probabilities. 2.4. Modality Disentanglement Protocol For datasets passing the bias audit (Blind AUC ≈ 0.5), we probe the mechanistic nature of memorization via modality dis- entanglement. We ask whether the membership signal arises 1 We also compute Spearman rank correlations as a robustness check and observe highly consistent diagnostic conclusions. Table 1: Source Diagnosis via Correlation Analysis. We com- pare MIA performance of AF3 and MF. Blind AUC reflects intrinsic dataset distribution. Pearson correlation coefficients (r AF3 /r MF ) indicate alignment between MIA performance and blind baselines. DatasetMIA AUCBlind Baseline: AUC (r AF3 /r MF ) AF3MFMetadataTextAcoustic Gigaspeech87.690.571.9 (.59/.55)75.6 (.41/.50)78.9 (.52/.55) SPGISpeech51.950.751.6 (.35/.33)50.0 (-.01/.06)52.0 (0.01/0.02) Librispeech93.870.9100.0 † (.75/.35)61.1 (.16/.11)99.8 (.78/.34) Tedlium70.068.964.9 (.48/.53)85.7 (.24/.41)98.7 (.34/.30) VoxPopuli62.068.155.1 (.24/.09)60.3 (.10/.13)66.7 (.07/.08) Clotho52.448.951.1 (.01/.09)47.1 (-.01/.00)48.2 (.01/-.01) CochlScene57.360.451.1 (.37/.21)47.1 (.34/.42)48.2 (.19/0.14) Nsynth62.362.151.1 (-.03/-.09)47.1 (0.08/-.09)48.2 (.21/0.22) † Near-perfect baseline AUC indicates trivial train/test separability. from cross-modal binding, namely the memorization of a spe- cific dependency between an acoustic realization and its tex- tual content, rather than a simple textual prior. We evaluate the model under four configurations: • Original: Full acoustic and textual input, representing the matched training instance. • Text-Only (Silence): Audio input is replaced by a zero- tensor, isolating the textual prior leakage. • Noise-Only: Audio is replaced by Gaussian noise to check the impact of non-semantic acoustic activation. • Acoustic Resynthesis: To isolate the impact of instance- specific acoustic features such as unique speaker identity or specific recording environments, we replace the original au- dio with a version synthesized from the textual content. For speech-centric datasets, we employ Text-to-Speech (TTS); for environmental sound datasets, we utilize Text-to-Audio (TTA) generation. By comparing these conditions, we can quantify the contribu- tion of instance-specific acoustic features to the overall mem- bership signal. If the MIA performance collapses upon resyn- thesis or silence, it confirms that the model’s memory is an- chored to the unique binding of a specific speaker’s voice to their utterance. 3. Experimental Setup 3.1. Target Models While numerous state-of-the-art LALMs have been pro- posed [1–7], the majority of these models do not offer public access to their pre-training data, making precise membership auditing infeasible. We therefore evaluate our auditing proto- col on two models with fully open-access resources: Audio- Flamingo 3 (AF3) [5] and Music Flamingo (MF) [23]. While AF3 is designed for general-purpose audio reasoning, Music Flamingo is specifically scaled and optimized for complex mu- sic understanding and captioning tasks. The fully open-access weights and transparent pre-training data splits of both models provide a unique opportunity to precisely define membership at the foundation stage. This allows us to audit memorization di- rectly on the primary models themselves, rather than relying on computationally expensive surrogate models or opaque black- box approximations. 3.2. Dataset Selection We evaluate our auditing framework on eight datasets spanning three representative audio tasks, selected to cover diverse acous- tic conditions, semantic structures, and data curation pipelines. • ASR: LibriSpeech, GigaSpeech, TED-LIUM, VoxPopuli, and SPGISpeech, ranging from clean read speech to large- scale, noisy, and multi-speaker recordings, enabling analysis of both intra- and inter-dataset distribution shifts [24, 29–32]. • Audio Captioning: Clotho and CochlScene, which empha- size non-linguistic acoustic semantics such as sound events and scenes, complementing speech-centric benchmarks [33, 34]. • Music Synthesis: NSynth, a musical instrument dataset used to analyze memorization in non-speech audio [35]. Members and non-members are defined by official train- ing and test splits, respectively. For each dataset, we randomly sample a balanced cohort of 2000 to 5000 instances (compris- ing 1000–2500 pairs) to ensure statistical rigor and a consistent scale for our sample-level analysis. 3.3. Implementation Details Blind Baseline The classifiers utilize the three feature sets de- scribed in Section 2.2, including utterance-level acoustic repre- sentations, TF-IDF textual features, and metadata statistics. MIA Indicator For membership inference, as discussed in Section 2.3, we aggregate various membership inference indi- cators into a 30-dimensional feature vector, where each dimen- sion represents a distinct metric. Following [9], we sweep k ∈ 0.05, 0.1,..., 0.6 for Min-k% to capture low-probability to- kens, and useα∈0, 1, 2,∞ for Max-R ́ enyi to quantify prob- ability concentration. Training Protocol Both the aggregated MIA and blind base- line classifiers are trained using stratified Logistic Regression. We intentionally select a linear classifier to ensure that strong performance reflects salient distributional artifacts or genuine memorization rather than the capacity of the auditor. We em- ploy 10-fold cross-validation within each dataset to ensure ev- ery sample is evaluated as a test instance while strictly prevent- ing cross-sample information leakage. Acoustic Resynthesis Models To disentangle memorization effects from surface-level acoustic cues, we conduct an acous- tic resynthesis ablation that replaces original audio signals with synthesized counterparts while preserving linguistic or seman- tic content. For speech datasets such as VoxPopuli and SPGIS- peech, we use Cosyvoice-0.5B [36] with a standardized ref- erence voice, where a small set of audio samples is selected from LibriSpeech as speaker reference audio to remove orig- inal speaker identities. For audio captioning datasets includ- ing Clotho, we use TangoFlux [37] to generate synthetic soundscapes conditioned on the provided captions. Due to the highly specialized nature of musical instrument synthesis and the brevity of labels, NSynth is excluded from the resynthe- sis ablation. For each TTS or TTA setting, we generate three independently synthesized audio realizations per sample using different reference audio or random seeds, and report the final membership performance as the average AUC across the three runs to reduce variance induced by stochastic generation. 3.4. Evaluation Metrics To evaluate the effectiveness of the membership inference at- tacks, we treat the task as a sample-level binary classification problem. For each sample i in a datasetD =D mem ∪D non , we compute its membership score S i using the aggregated classi- fier or individual metrics. Membership performance is reported via the Area Under the Receiver Operating Characteristic curve (AUC-ROC), which measures the probability that a ran- Figure 2: Cross-dataset membership inference heatmap. Di- agonal cells represent standard intra-dataset MIA (Train vs. Test). Off-diagonal cells represent membership-neutral pair- ings (e.g., Dataset A Train vs. Dataset B Train), where near- perfect AUCs reflect domain shifts rather than memorization. domly chosen member sample receives a higher score than a non-member. An AUC of 0.5 indicates random guessing, while 1.0 indicates perfect detection. 4. Results and Discussion 4.1. Spurious MIA Success and Distribution Shifts A prerequisite for a valid Membership Inference Attack is to ensure that observed success is not merely an artifact of dis- tribution shifts between member and non-member sets. There- fore, this section’s primary objective is to investigate whether high MIA performance reflects genuine model memorization or is driven by such distributional shortcuts. By identifying and controlling for these confounding factors, we aim to define a “clean dataset” that enables a fair and accurate evaluation of true memorization in subsequent experiments. Inter-dataset Bias (Domain Shortcuts) The most severe form of spurious success occurs when membership is defined across mismatched data sources. To isolate domain shift from memo- rization, we construct a membership-neutral evaluation matrix. As shown in Figure 2, the off-diagonal cells represent pairings in which both subsets share the same ground-truth training status, such as comparisons between training subsets drawn from different datasets, including Spgispeech and Clotho. In an ideal, bias-free audit, such comparisons should yield an AUC of ≈ 0.5. However, our results reveal that these membership-neutral pairings consistently yield near-perfect separation, with AUCs reaching 1.0.This confirms that MIA metrics under mismatched conditions act purely as do- main classifiers, capturing broad acoustic and linguistic sig- natures that reflect systematic differences in recording environ- ments and speaking styles across datasets, rather than reflecting training-set memorization in the model. Intra-dataset Bias (The Hidden Trap) Subtler biases exist within training and test splits of the same dataset. Table 1 shows that widely used benchmarks such as LibriSpeech (AUC 93.8) and TED-LIUM (AUC 70.0) yield high MIA scores. However, our Multi-modal Blind Baseline reveals these are largely spu- rious: LibriSpeech’s acoustic features alone reach 99.8 AUC without model access, with strong Pearson correlation to MIA scores (r = 0.78 for M 1 ). This suggests that much of the ob- served MIA signal is driven by acoustic shortcuts from system- atic train-test differences, rather than genuine memorization. Table 2: Modality Disentanglement (Sample-level AUC) on Rel- atively Clean Datasets for M 1 (AF3) and M 2 (MF). Results show a consistent collapse in MIA performance when acoustic or textual context is disrupted across both models. DatasetOriginalSilenceNoiseSynthesized M 1 M 2 M 1 M 2 M 1 M 2 M 1 M 2 VoxPopuli62.068.153.951.652.151.648.449.8 SPGISpeech51.950.749.450.350.351.249.849.4 Clotho57.348.951.249.249.251.747.950.0 Nsynth62.362.150.048.148.849.0-- 4.2. Robustness on Distribution-Matched Datasets To establish a trustworthy audit, we focus on datasets where the Blind Baseline AUC ≈ 0.5 (SPGISpeech, Clotho). Un- der these rigorously controlled conditions, MIA performance collapses to near-random (AUC 50.7–52.4). This suggests that the evaluated LALMs exhibit strong sample-level privacy ro- bustness under distribution-matched conditions. The lack of verbatim memorization indicates that the diverse acoustic re- alizations of a given text act as a natural regularizer, preventing the model from overfitting to specific token sequences at the in- stance level. These findings indicate that current MIA methods perform poorly on datasets without distributional confounds. 4.3. Cross-modal Binding While overall scores are low, modality disentanglement on the clean VoxPopuli dataset reveals the nature of LALM memory. As shown in Table 2, the membership signal in the Original con- dition collapses when modalities are decoupled. Replacing the original audio with Silence, Noise, or TTS-Resynthesis causes MIA AUCs to drop to near-random levels (≈ 50%) across both models. These results suggest that LALMs do not memorize training data as isolated textual sequences or standalone acous- tic fingerprints. Instead, memorization is instance-specific and cross-modal, emerging only from the binding between a vocal identity and its textual content. Privacy Implications Our findings refine the threat model for audio privacy. The dominant risk is not the reproduction of ver- batim text, but linkage leakage: the model’s ability to associate a specific individual with a specific utterance. Breaking this binding via vocal anonymization or TTS offers a practical miti- gation strategy for LALM training. 5. Conclusion We conduct the first systematic investigation of Membership Inference Attacks (MIA) against Large Audio Language Mod- els (LALMs), uncovering that high attack performance is often merely an illusion driven by acoustic distribution shifts rather than genuine memorization. To remedy this “validity crisis,” we introduce a rigorous three-phase auditing protocol that care- fully controls for these shortcuts. Our framework reveals that true memorization is sparse and, most importantly, manifests as a cross-modal binding between a speaker’s vocal identity and their specific utterances. This finding fundamentally redefines the LALM privacy threat, shifting the focus from text leakage to the association of voice with content. We call for a new au- diting paradigm that moves beyond simple metrics to enable the development of truly robust privacy-preserving models. 6. Generative AI Use Disclosure During the preparation of this work, the authors utilized several generative AI tools in both the manuscript writing and exper- imental implementation phases. For the manuscript, ChatGPT and Google Gemini were used for assistance with language edit- ing, rephrasing, and improving the clarity and conciseness of the text. For software development, GitHub Copilot and Cur- sor were employed to assist with code completion, boilerplate generation, and debugging. In all cases, these tools served in an assistive capacity. The core ideas, experimental design, analyses, and conclusions pre- sented in this paper are the original work of the human authors. All authors have reviewed the final manuscript and the asso- ciated code, and take full responsibility for the integrity and correctness of the entire work. 7. Acknowledgements We gratefully acknowledge the National Center for High- Performance Computing (NCHC) of Taiwan for providing com- putational resources that supported this work. 8. References [1] S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro, “Audio flamingo 2: An audio-language model with long-audio understanding and expert reasoning abilities,” 2025. [Online]. Available: https://arxiv.org/abs/2503.03983 [2] C. Liu, M. Aljunied, G. Chen, H. P. Chan, W. Xu, Y. Rong, and W. Zhang, “Seallms-audio:Large audio- language models for southeast asia,” 2025. [Online]. Available: https://arxiv.org/abs/2511.01670 [3] P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. de Chaumont Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, H. Muckenhirn, D. Padfield, J. Qin, D. Rozenberg, T. Sainath, J. Schalkwyk, M. Sharifi, M. T. Ramanovich, M. Tagliasacchi, A. Tudor, M. Velimirovi ́ c, D. Vincent, J. Yu, Y. Wang, V. Zayats, N. Zeghidour, Y. Zhang, Z. Zhang, L. Zilka, and C. Frank, “Audiopalm: A large language model that can speak and listen,” 2023. [Online]. Available: https://arxiv.org/abs/2306.12925 [4] K.-H. Lu, Z. Chen, S.-W. Fu, C.-H. H. Yang, S.-F. Huang, C.-K. Yang, C.-E. Yu, C.-W. Chen, W.-C. Chen, C. yu Huang, Y.-C. Lin, Y.-X. Lin, C.-A. Fu, C.-Y. Kuan, W. Ren, X. Chen, W.-P. Huang, E.-P. Hu, T.-Q. Lin, Y.-K. Wu, K.-P. Huang, H.-Y. Huang, H.-C. Chou, K.-W. Chang, C.-H. Chiang, B. Ginsburg, Y.-C. F. Wang, and H. yi Lee, “Desta2.5-audio: Toward general-purpose large audio language model with self- generated cross-modal alignment,” 2025. [Online]. Available: https://arxiv.org/abs/2507.02768 [5] A. Goel, S. Ghosh, J. Kim, S. Kumar, Z. Kong, S. gil Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” 2025. [Online]. Available: https://arxiv.org/abs/2507.08128 [6] Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” 2023. [Online]. Available: https://arxiv.org/abs/2311.07919 [7] C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Salmonn: Towards generic hearing abilities for large language models,” arXiv preprint arXiv:2310.13289, 2023. [8] N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” 2021. [Online]. Available:https: //arxiv.org/abs/2012.07805 [9] P. Maini, H. Jia, N. Papernot, and A. Dziedzic, “Llm dataset inference: Did you train on my dataset?”2024. [Online]. Available: https://arxiv.org/abs/2406.06443 [10] J. Mattern, F. Mireshghallah, Z. Jin, B. Sch ̈ olkopf, M. Sachan, and T. Berg-Kirkpatrick, “Membership inference attacks against language models via neighbourhood comparison,” 2023. [Online]. Available: https://arxiv.org/abs/2305.18462 [11] M. Duan, A. Suri, N. Mireshghallah, S. Min, W. Shi, L. Zettlemoyer, Y. Tsvetkov, Y. Choi, D. Evans, and H. Hajishirzi, “Do membership inference attacks work on large language models?”2024. [Online]. Available: https: //arxiv.org/abs/2402.07841 [12] W. Fu, H. Wang, C. Gao, G. Liu, Y. Li, and T. Jiang, “Member- ship inference attacks against fine-tuned large language models via self-prompt calibration,” vol. 37, 2024, p. 134 981–135 010. [13] R. Wen, Z. Li, M. Backes, and Y. Zhang, “Membership inference attacks against in-context learning,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, 2024, p. 3481–3495. [14] F. Galli, L. Melis, and T. Cucinotta, “Noisy neighbors: Efficient membership inference attacks against llms,” in Proceedings of the Fifth Workshop on Privacy in Natural Language Processing, 2024, p. 1–6. [15] Z. Song, S. Huang, and Z. Kang, “Em-mias: Enhancing member- ship inference attacks in large language models through ensemble modeling,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2025, p. 1–5. [16] M. Zawalski, M. Boubdir, K. Bałazy, B. Nushi, and P. Ribalta, “Detecting data contamination in LLMs via in-context learning,” in The Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https://openreview. net/forum?id=YlpaaYxx4t [17] Z. Li, Y. Wu, Y. Chen, F. Tonin, E. A. Rocamora, and V. Cevher, “Membership inference attacks against large vision-language models,” 2024. [Online]. Available:https: //arxiv.org/abs/2411.02902 [18] Y. Hu, Z. Li, Z. Liu, Y. Zhang, Z. Qin, K. Ren, and C. Chen, “Membership inference attacks against vision-language models,” 2025. [Online]. Available: https://arxiv.org/abs/2501.18624 [19] W.-C. Tseng, W.-T. Kao, and H. yi Lee, “Membership inference attacks against self-supervised speech models,” 2022. [Online]. Available: https://arxiv.org/abs/2111.05113 [20] G. Chen, Y. Zhang, and F. Song, “Slmia-sr:Speaker- level membership inference attacks against speaker recognition systems,” in Proceedings 2024 Network and Distributed System Security Symposium, ser. NDSS 2024.Internet Society, 2024. [Online]. Available: http://dx.doi.org/10.14722/ndss.2024. 241323 [21] D. Das, J. Zhang, and F. Tram ` er, “Blind baselines beat membership inference attacks for foundation models,” 2025. [Online]. Available: https://arxiv.org/abs/2406.16201 [22] H. Puerto, M. Gubri, S. Yun, and S. J. Oh, “Scaling up member- ship inference: When and how attacks succeed on large language models,” in Findings of the Association for Computational Lin- guistics: NAACL 2025, 2025, p. 4165–4182. [23] S. Ghosh, A. Goel, L. Koroshinadze, S. gil Lee, Z. Kong, J. F. Santos, R. Duraiswami, D. Manocha, W. Ping, M. Shoeybi, and B. Catanzaro, “Music flamingo: Scaling music understanding in audio language models,” 2025. [Online]. Available: https: //arxiv.org/abs/2511.10289 [24] G. Chen, S. Chai, G.-B. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan, “GigaSpeech: An Evolving, Multi-Domain ASR Cor- pus with 10,000 Hours of Transcribed Audio,” in Interspeech 2021, 2021, p. 3670–3674. [25] S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha, “Privacy risk in machine learning: Analyzing the connection to overfitting,” in 2018 IEEE 31st Computer Security Foundations Symposium (CSF), 2018, p. 268–282. [26] C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal, vol. 27, no. 3, p. 379–423, 1948. [27] W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer, “Detecting pretraining data from large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2310.16789 [28] J. Zhang, J. Sun, E. Yeats, Y. Ouyang, M. Kuo, J. Zhang, H. F. Yang, and H. Li, “Min-k%++: Improved baseline for detecting pre-training data from large language models,” arXiv preprint arXiv:2404.02936, 2024. [29] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, p. 5206–5210. [30] F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Es- teve, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in International conference on speech and computer. Springer, 2018, p. 198–208. [31] C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large- scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli, Eds.Online: Association for Computational Linguistics, Aug. 2021, p. 993–1003. [Online]. Available: https://aclanthology.org/2021.acl-long.80/ [32] P. K. O’Neill, V. Lavrukhin, S. Majumdar, V. Noroozi, Y. Zhang, O. Kuchaiev, J. Balam, Y. Dovzhenko, K. Freyberg, M. D. Shul- man, B. Ginsburg, S. Watanabe, and G. Kucsko, “SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,” in Interspeech 2021, 2021, p. 1434–1438. [33] K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio cap- tioning dataset,” in ICASSP 2020-2020 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 736–740. [34] I.-Y. Jeong and J. Park, “Cochlscene: Acquisition of acoustic scene data using crowdsourcing,” in 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Confer- ence (APSIPA ASC). IEEE, 2022, p. 17–21. [35] J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in International conference on machine learning. PMLR, 2017, p. 1068–1077. [36] Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang, F. Yu, H. Liu, Z. Sheng, Y. Gu, C. Deng, W. Wang, S. Zhang, Z. Yan, and J. Zhou, “Cosyvoice 2: Scalable streaming speech synthesis with large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2412.10117 [37] C.-Y. Hung, N. Majumder, Z. Kong, A. Mehrish, A. A. Bagherzadeh, C. Li, R. Valle, B. Catanzaro, and S. Poria, “Tan- goflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,” arXiv preprint arXiv:2412.21037, 2024.