Paper deep dive
Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level dropin & Neuroplasticity Mechanisms
Yupei Li, Shuaijie Shao, Manuel Milling, BjĂśrn Schuller
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/26/2026, 2:17:17 AM
Summary
The paper introduces 'dropin' and 'plasticity' algorithms inspired by mammalian neurogenesis and neuroapoptosis to enhance deepfake audio detection. These methods dynamically adjust neuron counts in neural network layers (ResNet, GRNN, Wav2Vec) to improve computational efficiency and model performance, achieving significant reductions in Equal Error Rate (EER) on ASVSpoof2019 and FakeorReal datasets.
Entities (6)
Relation Signals (3)
Yupei Li â authored â Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level dropin & Neuroplasticity Mechanisms
confidence 100% ¡ 1 st Yupei Li * Department of Computing... Enhancing Efficiency and Performance in Deepfake Audio Detection
Plasticity â reduceseerin â ASVSpoof2019
confidence 95% ¡ maximum of around 39% and 66% relative reduction in Equal Error Rate with the dropin and plasticity approach among these dataset
Dropin â improvesefficiencyof â ResNet
confidence 90% ¡ Experimental results... demonstrate consistent improvements in computational efficiency with the dropin approach
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large language models (LLMs) further demonstrates the benefits of scaling model parameters, but also highlights one bottleneck where performance gains are constrained by parameter counts. Simply stacking additional layers, as done in current LLMs, is computationally expensive and requires full retraining. Furthermore, existing low-rank adaptation methods are primarily applied to attention-based architectures, which limits their scope. Inspired by the neuronal plasticity observed in mammalian brains, we propose novel algorithms, dropin and further plasticity, that dynamically adjust the number of neurons in certain layers to flexibly modulate model parameters. We evaluate these algorithms on multiple architectures, including ResNet, Gated Recurrent Neural Networks, and Wav2Vec. Experimental results using the widely recognised ASVSpoof2019 LA, PA, and FakeorReal dataset demonstrate consistent improvements in computational efficiency with the dropin approach and a maximum of around 39% and 66% relative reduction in Equal Error Rate with the dropin and plasticity approach among these dataset, respectively. The code and supplementary material are available at Github link.
Tags
Links
- Source: https://arxiv.org/abs/2603.24343v1
- Canonical: https://arxiv.org/abs/2603.24343v1
Trouble viewing inline? Open PDF directly â
Full Text
40,328 characters extracted from source content.
Expand or collapse full text
Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level dropin & Neuroplasticity Mechanisms 1 st Yupei Li * Department of Computing and Chair of Health Informatics Imperial College London and Technical University of Munich London and Munich, United Kingdom and Germany yl7622@ic.ac.uk 2 nd Shuaijie Shao * Department of Computer Science University College London London, United Kingdom shuaijie.shao.25@ucl.ac.uk 3 rd Manuel Milling Chair of Health Informatics Technical University of Munich Munich, Germany manuel.milling@tum.de 4 th Bj Ě orn Schuller Chair of Health Informatics and Department of Computing Technical University of Munich and Imperial College London Munich and London, Germany and United Kingdom schuller@tum.de * Yupei Li and Shuaijie Shao contributed equally to this work. AbstractâCurrent audio deepfake detection has achieved re- markable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The suc- cess of large language models (LLMs) further demonstrates the benefits of scaling model parameters, but also highlights one bottleneck where performance gains are constrained by parameter counts. Simply stacking additional layers, as done in current LLMs, is computationally expensive and requires full retraining. Furthermore, existing low-rank adaptation methods are primarily applied to attention-based architectures, which limits their scope. Inspired by the neuronal plasticity observed in mammalian brains, we propose novel algorithms, dropin and further plasticity, that dynamically adjust the number of neurons in certain layers to flexibly modulate model parameters. We eval- uate these algorithms on multiple architectures, including ResNet, Gated Recurrent Neural Networks, and Wav2Vec. Experimental results using the widely recognised ASVSpoof2019 LA, PA, and FakeorReal dataset demonstrate consistent improvements in computational efficiency with the dropin approach and a maximum of around 39% and 66% relative reduction in Equal Error Rate with the dropin and plasticity approach among these dataset, respectively. The code and supplementary material are available at Github link. Index TermsâNeuroplasticity, Audio Deepfake Detection, Wav2Vec, Efficiency, Finetuning I. INTRODUCTION Large models (LMs) have recently surpassed the capability bottlenecks of traditional deep learning models across a wide range of tasks [1], including audio-related applications such as deepfake audio detection. For instance, audio LLMs, such as QwenAudio [2], have demonstrated superior performance compared to conventional small-scale deep learning models. This performance improvement aligns with established scaling laws indicating increasing the number of model parameters generally enhances representational capacity and downstream task performance [3] given sufficient training data and com- putation. Additionally, we argue that large-scale models are beneficial for capturing the complexity of audio deepfake detection. Although this task is framed as a binary classification prob- lem, the implicit discrepancies between spoofed and bona fide audio require the network to learn subtle and high- dimensional acoustic representations [4]. Existing approaches employ diverse architectural paradigms for feature extraction: convolutional neural network (CNN)-based models such as ResNet [5] process audio data as Mel-spectrograms by treating them as image-like features; recurrent neural network (RNN)- based architectures, including the Light Convolutional Gated Recurrent Neural Network (LC-GRNN), analyse spectrograms as temporally structured sequential data [6]; and attention- based pretrained models, notably Wav2Vec 2.0 [7], directly op- erate on raw audio waveforms to capture fine-grained temporal dependencies. Additional models have been proposed for this task, including ASSIST [8] and RawNet2 [9], etc. However, the approaches aforementioned have not fully resolved the challenges of audio deepfake detection, as evidenced by their performance in the recent ASVspoof 2019 challenge [10]. Nevertheless, it is a consistent trend that models with more parameters tend to achieve superior performance. These models commonly rely on fixed architectures, often incorporating fine-tuning model parameters pretrained on a different task. However, inspired by the mechanisms of the mammalian brain, it may be argued that modern models should aim for dynamic neuron-level adjustment called (neuro- )plasticity, enabling them to remain memory-efficient while unlocking additional capacity for target tasks. Li et al. [11] arXiv:2603.24343v1 [cs.SD] 25 Mar 2026 reviewed the principles of plasticity in biology and deep learn- ing methods that adopt similar dropin (adding more neurons) and plasticity concepts, yet without specific implementation details and any empirical tests. This idea has also been explored in continual learning, however, focus only on layer- level adjustments (i.e., adding or removing entire layers), such as Exhubert [12]. In addition, many loss designs have been proposed to mitigate catastrophic forgetting [13], but these methods merely adjust weights that should be updated, without altering the number of parameters. Previous algorithms have only sparsely explored pioneering âdrop-inâlikeâ strategies, such as determining when and where to add neurons [14]. These methods identify triggers based on gradient information [15], exposure to unseen experi- ences [16], or model performance [17]. However, they do not explicitly draw inspiration from the mammalian brain to formulate dropin as a general technique; instead, their motivation is largely analytical, deriving from mathematical formulations aimed at increasing model capacity. Moreover, although various dropout and structural pruning strategies have been developed, they have not been integrated with dropin mechanisms to mirror the concepts of neuroplasticity. Therefore, we propose a more fine-grained algorithm that incorporates a clearly defined dropin strategy alongside a novel, neuron-level plasticity mechanism. We evaluate this approach on multiple deepfake audio detection model clusters using the ASVspoof2019 LA, PA [10], and FakeorReal (FoR) [18] dataset. Experimental results show that the Equal Error Rate (EER) is considerably reduced and training time is shortened, all while maintaining limited memory usage. This leads to our contributions: To the best of our knowledge, this is the first neuron-level clearly defined dropin and novel plasticity algorithm that have been verified by experiments. Additionally, it enables state-of-the-art (SOTA) performance on ASVspoof2019 LA while preserving memory efficiency and reducing training time. Moreover, it is applicable to a variety of models. I. RELATED WORK Previous work has explored brain-inspired methods from both biological and deep learning perspectives, providing evidence that such designs can improve both performance and computational efficiency [19], [20]. Representative examples include spiking neural networks with brain-like signalling mechanisms [21], Hebbian synaptic learning for representation learning [22], [23], and hippocampus-inspired neuroplasticity algorithms [24]. While these approaches demonstrate the po- tential of brain-inspired designs to enhance neural network performance, they are generally not scalable to a wide range of conventional neural network architectures. Plasticity-related approaches, such as structural pruning and progressive networks, have been shown in prior work to be effective, although they are not always explicitly motivated by biological principles. Progressive networks [25] demon- strate that dynamically growing a network can enhance its learning capacity, while other studies increase model capacity by stacking multiple layers, for example, in ExHuBERT [12] and N-stacking [26]. Additionally, adaptive pruning methods based on feature information entropy have been proposed to reduce training costs while maintaining performance [27]. Moreover, synaptic plasticityâbased regularisation techniques have been introduced to reduce the complexity of certain network architectures [28]. However, these methods do not model a comprehensive process of neurogenesis and plasticity. Even works that more closely mimic neuroplasticity, such as determining neuron addition and pruning based on certain triggers, have been proposed. Related methods trigger insertion based on gradient information [29], exposure to novel experi- ences [30], or other heuristics. However, these approaches are not grounded in a unified biological model of neurogenesis or neuroplasticity, relying instead on task- or heuristic-specific criteria. Staying on the theory side cannot by itself prove algorithm effectiveness. Audio deepfake detection is an application that requires substantial feature understanding [31], [32]. While some small models have been shown to be effective [33], and LLMs are also being explored for this task [34], practical deployment often raises additional challenges. In particular, selecting an appropriate network size is crucial: a model that is too small may underfit and fail to capture decisive features [35], while a model that is too large may be inefficient and not generalise [36]. Dynamically adjusting networks with a neuroplasticity structure provides a flexible mechanism to maintain performance under computational constraints and evolving data distributions. I. METHODOLOGY: NEURON-LEVEL DROPIN & PLASTICITY Our method addresses the parameter bottleneck by selec- tively adding neurons and parameters to alleviate capacity limitations. Specifically, we propose dropin neuron addition on certain layers as a preliminary step to validate the effective- ness of neuron-level expansion. To avoid uncontrolled model growth while leveraging the benefits of dropin, we further introduce plasticity, which first adds neurons and subsequently prunes them to maintain a compact model size. A. Dropin algorithm To more closely emulate the neurogenesis observed in the mammalian brain, rather than extending entire layers, we intro- duce additional neurons selectively into specific layers. We add neurons to randomly selected layers and freeze the remainder of the network, training only the newly introduced neurons to examine their effects and behavior. This design choice is made as the complexity of determining where and when neurons are added in the mammalian brain is not easy to replicate in common state-of-the-art artificial neural networks. In addition to a a general experimental validation of the mechanism, we evaluate the effects of adding neurons at different layers from a statistical perspective. This approach operates at a more granular level and offers greater flexibility for the learning â Freeze 1 2345 Original Hidden Cell Dropin Hidden Cell Input Cell Output Cell Layer ďĽ Finetune Fig. 1. High-level dropin process: During the dropin phase, we load the original pretrained model weights and keep them frozen. Additional neurons are then introduced into randomly selected layers, and only the connection weights associated with these newly added neurons are trained. process, as individual neurons can be modulated directly by the algorithm. The overall framework is illustrated in Fig. 1. The above figure provides a simplified representation of a general neural network, in which each cell may represent any type of neuron and the connections correspond to the parameters of the model. As such, the proposed algorithm is not restricted to a specific architecture and can be applied to various network designs, including convolutional, encoder- based, and recurrent layers. More detailed illustrations of the implementation of our dropin mechanism for different architecture types are presented in Fig. 2. The dropin weights are concatenated as block matrices to facilitate matrix multiplication, enabling architectural adap- tation at the granularity of individual neurons. Specifically, the channel dimension of certain layers is expanded, allowing additional neurons to be incorporated into the convolutional kernels for CNN. Likewise, the hidden dimensions of the update and reset gate weights in the GRU are enlarged, introducing more neurons within these gating mechanisms. Furthermore, in terms of attention encoder, the weight matrices used to generate the query, key, and value representations in attention mechanisms are expanded, permitting the inclusion of additional neurons. Importantly, the core computational mechanisms remain unchanged, for example, the convolu- tion operations themselves are preserved, while the network capacity is increased through the addition of neurons. This demonstrates that the dropin approach is broadly applicable to a wide range of neural network architectures. For simplicity, we present a high-level version to illustrate the plasticity algorithm. Additionally, this algorithm leverages pretrained models, which are now predominant in fine-tuning applications. In contrast to previous literature, which typically fine-tunes the entire network, our approach updates only the dropin neurons. This strategy considerably reduces training time while enhanc- ing the modelâs capability through a lightweight parameter increase, rather than relying on extensive stacking of layers. Related work in continual learning has explored similar ideas, Recurrent Unit Attention Encoder (Transformer) CNN conv1 conv2 conv3 conv4 fc H 1 x W 1 x C 1 H2 x W 2 x C 2 convolutional + ReLU max pooling fully connected + ReLU softmax Flatten H 2 x W 2 x 2C 2 H 3 x W 3 x C 3 H 3 x W 3 x 2C 3 H 4 x W 4 x C 4 H x W x C dropin convolutional + ReLU ZR h t-1 ,X t ht-1ht FFN update gate reset gate feedforward network dropin gate dzdr Encoder layer Input X B Q DQ BKDK BVDV QK T Z base Q,K,V dropin Q,K,V attention output T x DqT x 2Dq T x DkT x 2Dk T x DvT x 2Dv Q,K projection Fig. 2. Convolutional layers, recurrent units, and attention encoders dropin process. These represent the specific dropin techniques adapted to different model architectures. From a mathematical perspective, this involves increasing the kernel size for CNNs, expanding the gate weight dimensions for GRUs, and enlarging the query, key, and value weight dimensions for attention mechanisms. such as Low-Rank Adaptation (LoRA) [37], which introduces additional low-rank matrices to acquire new information and update existing parameters. However, LoRA does not modify the network architecture to the same extent as our proposed drop-in mechanism. Instead, it introduces additional adapters that remain largely decoupled from the original architecture, making it less aligned with our biologically inspired motiva- tion based on mammalian brain function. Additionally, LoRA is mainly suited for attention-based models, as it relies on updating a small number of parameters relative to the original weights. Therefore, its efficiency drops when many small layers lead to a large overall parameter count. In contrast, our method applies to any architecture by directly generating new neurons, independent of component parameterisation. B. Plasticity algorithm Following the dropin process, the memory usage of the model increases. In contrast, the mammalian brain exhibits a natural mechanism of eliminating redundant or unused neurons to optimise efficiency. Building on the neuron-level granularity of the dropin mechanism, we take a further step by proposing a complementary algorithm, termed plasticity. The proposed pipeline is illustrated in Fig. 3. This approach aims to push beyond the modelâs capability limits on the target task. Accordingly, we design the entire training pipeline with an emphasis on maintaining the overall parameter count rather than optimising for time efficiency. The process begins with standard training of the model. Unlike the controlled dropin experiments described above, where existing neurons are frozen to isolate the effects of newly added ones, in this stage we allow all parameters to remain trainable. Our goal here is to investigate a complete biologically inspired setting from the neurogenesis to the neuroapotosis stage, in which network-wide interactions, analogous to biochemical processes in the brain, are permitted to emerge prior to the selective removal or deactivation of neurons. The algorithm is divided into three stages. First, the model is trained for a number of epochs until its performance reaches a plateau. In the second stage, new neurons are introduced and the model continues training without reinitialising or freezing any previously learned parameters, such that all existing weights are updated jointly. This continuous training process enables the expanded architecture to adapt to the task and to reorganise its internal representations. As in the dropin procedure, the new neurons are added to randomly selected layers; for simplicity, the number of added neurons matches that of the original architecture. After this stage, the added neurons are pruned, and the full model undergoes a final phase of continued training, again with parameter initialisation from the previous training step. The algorithm is summarised in Algorithm 1. This not only restores the original parameter size but also preserves the performance improvements gained during the intermediate expansion phase. We note that a promising direction for future investigation, once the mechanisms by which the brain prunes neurons are better understood, is to develop principled criteria for determining which neurons should be removed. At the current stage, however, we prune only the newly added neurons in order to control the number of trainable parameters and keep the overall model size constant. We contend that the plasticity algorithm could endow the model with a dynamic memory space, enabling it to retain valuable knowledge across all three stages while discarding irrelevant information, which forms the core intuition driving the modelâs breakthrough in performance. Each stage has been trained for equal epoches for a fair comparison. IV. EXPERIMENTS We evaluate our method on the ASVspoof 2019 Logical Access (LA), physical access (PA) subset, and FoR, two widely used benchmarks for audio deepfake detection. Ex- periments are conducted on three representative models with distinct input modalities: ResNet18 (CNN, Mel spectrograms), GRNN (RNN, linear cepstral coefficients), and Wav2Vec 2.0 Algorithm 1 Neuroplasticity Training Require: Network N(θ) with layers L i K i=1 , dataset D Ensure: Final parameters θ â pruned 1: Initial parameters: θ init , θ init âźN(0,Ď 2 ) 2: Initial training: θ â init = arg min θ init E (x,y)âźD L f(x;θ init ),y 3: Sample layer indices: I âź Uniform(1,...,K) 4: for each iâI do 5:Neuron expansion in layer L i with weights θ i init : θ i dropin â θ i init ⪠θ i â , θ i â âźN(0,Ď 2 ),|θ i â | =|θ i init | 6: end for 7: Continued training with dropin neurons: θ â dropin = arg min θ dropin E (x,y)âźD L f(x;θ dropin ),y 8: Neuron pruning removes all added neurons: θ pruned âP θ â dropin , |θ pruned | =|θ init | 9: Final continued training: θ â pruned = arg min θ pruned E (x,y)âźD L f(x;θ pruned ),y 10: return θ â pruned (attention-based, raw waveforms). This setup enables assess- ment of both architectural and feature diversity to verify the scalability of our pipeline. We fix the random seed to 42 for all random number generators (Python, NumPy, PyTorch) to ensure reproducibility. All experiments are run on an NVIDIA A6000 GPU, with other hyperparameters aligned to those in the original model papers [5]â[7]. We report the best- performing model for each setting. The maximum number of training epochs is determined through preliminary trials to ensure that each experimental configuration can converge within the allotted epochs. Model selection is performed based on validation set performance, and the selected model is subsequently evaluated on the test set. Wav2Vec 2.0 uses the pretrained model from Hugging Face 1 , while ResNet18 and GRNN are trained from scratch for all experimental settings including baseline. To examine parameter scaling, we also train a scratch Wav2Vec variant with a reduced 256-dimensional hidden size. Full hyperparameter details are provided in anonymous Github link 2 . We apply the dropin and plasticity procedures once for each base model, introducing the same number of additional neurons as the original size of that layer. We compare the 1 https://huggingface.co/facebook/wav2vec2-base 2 https://anonymous.4open.science/r/Dropin-Anonymize/README.md Original Hidden Cell Dropin Hidden Cell Input Cell Output Cell Training stage Original Stage Neurogenesis Stage Neuroapoptosis Stage Fig. 3. Plasticity process: The pipeline is divided into three stages. The first stage is identical to conventional training, where the model is optimised using the standard objective function. In the second stage, new neurons are dropped in, a process analogous to neurogenesis. The whole model is then retrained. After the newly introduced information has been assimilated and distributed across existing neurons. The third stage emulates neuroapoptosis by pruning the added neurons, followed by a final retraining. performance of five configurations: (i) the original baseline model, (i) the dropin model without parameter freezing (rep- resenting an enlarged model capacity), (i) LoRA-based fine- tuning (where applicable), (iv) our dropin method, and (v) our plasticity method. V. RESULTS In the following, we discuss the results of all models on the three datasets presented in Table I. Accuracy improvement on plasticity: The proposed dropin and plasticity strategies consistently outperform the baseline across multiple architectures and datasets in terms of Test EER. In particular, plasticity achieves the lowest EER in most settings, often by a large margin. For example, on the ASV LA dataset with Wav2Vec 2.0, plasticity reduces the EER from 2.45% (baseline) to 0.04%, substantially outperforming both dropin unfrozen (0.44%) and dropin frozen (1.64%). Similar improvements can be observed across ResNet and GRNN models on ASV LA and PA, demonstrating that the proposed mechanisms provide consistent accuracy gains over the original architectures. Efficiency improvement on dropin: In terms of training efficiency, the dropin frozen strategy consistently achieves the lowest backward time among all compared methods. By freezing the original network and training only the newly introduced neurons, dropin frozen considerably reduces com- putational overhead while maintaining competitive detection performance. For instance, on ASV LA with ResNet, the backward time is reduced from 5.46,ms (baseline) to 3.49,ms, and on Wav2Vec 2.0, it is reduced from 0.24,ms to 0.18,ms. This efficiency advantage is also observed across PA and FoR datasets, indicating that dropin frozen is particularly suitable for scenarios where training speed is a primary concern. Plasticity achieves a superior trade-off between param- eter size and accuracy: When both accuracy and model size are considered jointly, plasticity emerges as the most favorable strategy. While dropin unfrozen occasionally achieves slightly lower EER, this improvement is primarily attributable to its substantially larger number of trainable parameters. In con- trast, plasticity maintains the original model size and achieves comparableâor in many cases superiorâperformance. For example, on ASV PA with Wav2Vec 2.0, plasticity achieves an EER of 4.34%, compared to 7.54% for LoRA and 8.18% for dropin unfrozen, without increasing the number of trainable parameters. These results suggest that plasticity provides a more balanced trade-off between accuracy and computational cost for model size, making it particularly attractive for practical deployment. Plasticity benefits more from larger models: We observe that plasticity yields more pronounced performance gains in larger-capacity models compared to smaller ones. In particular, when comparing Wav2Vec 2.0 with ResNet, plasticity leads to substantially larger reductions in Test EER for Wav2Vec 2.0 across all evaluated datasets. For example, on ASV LA, plasticity reduces the EER of Wav2Vec 2.0 from 2.45% to 0.04%, whereas the improvement on ResNet is more moderate, decreasing from 14.69% to 9.70%. This trend is consistently observed on ASV PA and FoR, suggesting that larger models with higher representational capacity provide a more favorable substrate for plasticity-driven adaptation. These results indicate that plasticity is particularly effective when applied to high- capacity architectures, where additional neurons and dynamic adaptation can be more fully exploited. Comparative drawbacks of LoRA: Finally, LoRA con- sistently underperforms compared to both dropin-based strate- gies and plasticity across all evaluated models and datasets. Although LoRA introduces only a small number of trainable parameters, its detection performance remains close to or worse than the baseline in most cases. For instance, on ASV LA with Wav2Vec 2.0, LoRA achieves an EER of 2.08%, which is notably higher than plasticity (0.04%) and dropin unfrozen (0.44%). Similar trends are observed on PA and FoR datasets, indicating that parameter-efficient fine-tuning alone is insufficient for capturing the complex spoofing patterns addressed by our proposed methods but more mammal brain mimicing is needed. We have also conducted ablation studies on effects of which layer to perform these mechanism and analyse explain- ability why the results performing better for our proposed method. Using the Wav2vec 2.0 small model with the dropin strategy as representatives, the results are presented in Fig. 5. DatasetModelTraining StrategyTest EER (%)âBackward Time per step (ms)âParameters (M)âTrainable Params (M)â ASV LA ResNet baseline14.695.4611.1711.17 dropin unfrozen7.606.7214.2814.28 dropin frozen (ours)11.983.4914.281.47 plasticity (ours)9.70/11.17/ GRNN baseline19.84343.350.060.06 dropin unfrozen16.64477.790.100.10 dropin frozen (ours)18.48240.450.100.03 plasticity (ours)17.82/0.06/ Wav2Vec 2.0 baseline2.450.2495.5795.57 dropin unfrozen0.440.25102.83102.83 LoRA2.08/95.570.29 dropin frozen (ours)1.640.18102.838.47 plasticity (ours)0.04/95.57/ Wav2Vec 2.0 (Small) baseline11.061.1511.2511.25 dropin unfrozen8.441.2114.2114.21 LoRA11.10/11.250.09 dropin frozen (ours)12.200.7414.213.02 plasticity (ours)9.51/11.25/ ASV PA ResNet baseline14.085.8211.1711.17 dropin unfrozen9.146.8414.2814.28 dropin frozen (ours)6.113.7914.281.47 plasticity (ours)9.70/11.17/ GRNN baseline19.13355.880.060.06 dropin unfrozen18.68494.300.100.10 dropin frozen (ours)17.50250.410.100.03 plasticity (ours)18.19/0.06/ Wav2Vec 2.0 baseline7.8994.2395.5795.57 dropin unfrozen8.1886.64102.83102.83 LoRA7.54/95.570.29 dropin frozen (ours)7.8729.09102.838.47 plasticity (ours)4.34/95.57/ Wav2Vec 2.0 (Small) baseline11.9821.9411.2511.25 dropin unfrozen14.7225.2514.2114.21 LoRA16.52/11.250.09 dropin frozen (ours)11.9710.2714.213.02 plasticity (ours)10.75/11.25/ FoR ResNet baseline18.966.2011.1711.17 dropin unfrozen12.727.1514.2814.28 dropin frozen (ours)13.703.9714.281.47 plasticity (ours)14.55/11.17/ GRNN baseline17.65368.810.060.06 dropin unfrozen16.27506.930.100.10 dropin frozen (ours)19.14255.870.100.03 plasticity (ours)14.05/0.06/ Wav2Vec 2.0 baseline21.81193.5595.5795.57 dropin unfrozen21.13202.77102.83102.83 LoRA16.67/95.570.29 dropin frozen (ours)24.30142.69102.838.47 plasticity (ours)12.94/95.57/ Wav2Vec 2.0 (Small) baseline33.27252.3011.2511.25 dropin unfrozen20.02364.5714.2114.21 LoRA30.53/11.250.09 dropin frozen (ours)19.82209.5914.213.02 plasticity (ours)12.60/11.25/ TABLE I COMPARISON OF DIFFERENT TRAINING STRATEGIES ACROSS ASV LA, PA, AND FOR DATASETS. DROPIN UNFROZEN DENOTES AN ENLARGED NETWORK TRAINED CONVENTIONALLY, WHILE BASELINE REFERS TO THE ORIGINAL MODEL. FOR PLASTICITY, BACKWARD TIME AND TRAINABLE PARAMETERS ARE NOT REPORTED AS THEY VARY ACROSS STAGES. THE SYMBOL â/â INDICATES UNAVAILABLE OR NON-APPLICABLE MEASUREMENTS. It can be observed that dropping in neurons at different layers has a relatively small impact on overall performance. This suggests that neural networks are able to automatically leverage the added parameters to improve learning, regardless of the specific layer in which they are introduced. Importantly, this consistency indicates that our observed improvements are not due to randomness but represent a systematic effect. While some of the previous literature implies that selecting single layers based solely on mathematical criteria, such as gradients would be beneficial, we would argue that the selection process should rather be informed by processes in the mammalian brain, which are still to be understood well enough. Lastly, we applied Grad-CAM [38] to each experimental setting to examine whether attention patterns change during training. As a case study, we selected the ResNet model on a single example from the ASVspoof 2019 LA dataset. The resulting attention visualisations are presented in Fig. 4. It can be observed that in the baseline model, Grad-CAM activations are particularly prominent in the shallow layers (L1âL2), where focus is broadly distributed across edge re- gions. Although deeper layers show slightly more focused acti- vations, the highlighted areas remain coarse. This suggests that InputL1 HeatmapL1 OverlayL2 HeatmapL2 OverlayL3 HeatmapL3 OverlayL4 HeatmapL4 Overlay Baseline Dropin Frozen Plasticity Fig. 4. GradCAM results on Resnet for one case from ASVSpoof 2019 LA. 345678910 Dropin Layer 0 2 4 6 8 10 12 EER (%) 8.88 9.20 8.91 8.78 8.89 9.14 8.98 9.51 Baseline (11.06) Fig. 5.EER results for different dropin layers. â3â represents that we dropped in new neurons on the 3rd encoding layer in Wav2vec 2.0 small on ASVSpoof2019 LA. the baseline relies on redundant or spurious features, resulting in limited feature selectivity and suboptimal hierarchical rep- resentations. However, under the dropin configuration, Grad- CAM maps become noticeably more compact from interme- diate layers onwards. Compared to the baseline, the focus is more concentrated on many feature maps especially L1, which provides a potential explanation for the observed performance improvements. For the plasticity configuration, focus is even more selective across layers. Early layers exhibit relatively low activations, indicating enhanced feature selection capability, while deeper layers show highly focused and semantically aligned responses. By keeping parameters size the same, the model is able to adaptively reorganize feature representations, suppress non-informative responses, and achieve improved cross-layer consistency. Notably, the last layerâs attention of the three remains relatively similar, suggesting that there is still room for refinement in future work to further enhance focus and performance. VI. CONCLUSION Our work investigated the use of the newly introduced dropin and plasticity algorithms, which can be applied to various model architectures. We applied it in the field of audio deepfake detection to enhance both performance and efficiency in terms of computation and parameter usage. The model achieved SOTA performance on the ASVSpoof2019 LA subset. In scenarios with strict constraints on model size such as deployment on mobile devices or applications demanding high efficiency, including real-time update and detection tasks, our method presents a more practical and effective alternative. In future work, we aim to explore a biologically inspired plasticity algorithm that more closely mimics the behaviour of the mammalian brain. Specifically, this would involve structural pruning to remove neurons that are functionally redundant, a process that requires further rigorous definition. REFERENCES [1] I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liu, A. Gritsevskiy, D. Wurgaft, D. Kauffman, G. Recchia, J. Liu, J. Cavanagh, M. Weiss, S. Huang, T. F. Droid, T. Tseng, T. Korbak, X. Shen, Y. Zhang, Z. Zhou, N. Kim, S. R. Bowman, and E. Perez, âInverse scaling: When bigger isnât better,â Transactions on Machine Learning Research, 2023. [2] Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, âQwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,â arXiv preprint arXiv:2311.07919, 2023. [3] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, âScaling laws for neural language models,â arXiv preprint arXiv:2001.08361, 2020. [4] Z. Almutairi and H. Elgibreen, âA review of modern audio deepfake de- tection methods: challenges and future directions,â Algorithms, vol. 15, no. 5, p. 155, 2022. [5] K. He, X. Zhang, S. Ren, and J. Sun, âDeep residual learning for image recognition,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, p. 770â778. [6] A. Gomez-Alanis, A. M. Peinado, J. A. Gonzalez, and A. M. Gomez, âA light convolutional gru-rnn deep feature extractor for asv spoofing detection,â in Proc. Interspeech, vol. 2019, 2019, p. 1068â1072. [7] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, âwav2vec 2.0: A framework for self-supervised learning of speech representations,â Advances in neural information processing systems, vol. 33, p. 12 449â 12 460, 2020. [8] J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, âAasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,â in ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2022, p. 6367â6371. [9] H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, âEnd-to-end anti-spoofing with rawnet2,â in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, p. 6369â6373. [10] X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V. Vestman, T. Kinnunen, K. A. Lee et al., âAsvspoof 2019: A large-scale public database of synthesized, converted and replayed speech,â Computer Speech & Language, vol. 64, p. 101114, 2020. [11] Y. Li, M. Milling, and B. W. Schuller, âNeuroplasticity in artificial intelligence â an overview and inspirations on drop in & out learning,â 2025. [12] S. Amiriparian, F. Packa Ě n, M. Gerczuk, and B. W. Schuller, âExhubert: Enhancing hubert through block extension and fine-tuning on 37 emotion datasets,â in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH.Kos Island, Greece: International Speech Communication Association, 2024, p. 2635â2639. [13] S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton, âLoss of plasticity in deep continual learning,â Nature, vol. 632, no. 8026, p. 768â774, 2024. [14] K. Maile, E. Rachelson, H. Luga, and D. G. Wilson, âWhen, where, and how to add new neurons to anns,â in Proceedings of the First Interna- tional Conference on Automated Machine Learning, ser. Proceedings of Machine Learning Research, I. Guyon, M. Lindauer, M. van der Schaar, F. Hutter, and R. Garnett, Eds., vol. 188. PMLR, 25â27 Jul 2022, p. 18/1â12. [15] U. Evci, B. van Merri Ě enboer, T. Unterthiner, M. Vladymyrov, and F. Pedregosa, âGradmax: Growing neural networks using gradient in- formation,â in International Conference on Learning Representations (ICLR), 2022. [16] P. Eriksson and L. Westlund Gotby, âDynamic network architectures for deep q-learning: Modelling neurogenesis in artificial intelligence,â 2019. [17] A. Sowrirajan, P. Srinivasan, S. A. Srinivasan, and S. P. Devi, âEnhanc- ing neurofuzzy plasticity: A fusion of lstm and artificial neurogenesis,â in 2024 International Conference on Emerging Techniques in Compu- tational Intelligence (ICETCI). IEEE, 2024, p. 150â154. [18] R. Reimao and V. Tzerpos, âFor: A dataset for synthetic speech detection,â in 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD). IEEE, 2019, p. 1â10. [19] S. Wo Ě zniak, A. Pantazi, T. Bohnstingl, and E. Eleftheriou, âDeep learning incorporating biologically inspired neural dynamics and in- memory computing,â Nature Machine Intelligence, vol. 2, no. 6, p. 325â336, 2020. [20] G. Dellaferrera, S. Wozniak, G. Indiveri, A. Pantazi, and E. Elefthe- riou, âLearning in deep neural networks using a biologically inspired optimizer,â arXiv preprint arXiv:2104.11604, 2021. [21] S. Schmidgall, R. Ziaei, J. Achterberg, L. Kirsch, S. Hajiseyedrazi, and J. Eshraghian, âBrain-inspired learning in artificial neural networks: a review,â APL Machine Learning, vol. 2, no. 2, 2024. [22] N. Ravichandran, A. Lansner, and P. Herman, âUnsupervised representa- tion learning with hebbian synaptic and structural plasticity in brain-like feedforward neural networks,â Neurocomputing, vol. 626, p. 129440, 2025. [23] A. Journ Ě e, H. G. Rodriguez, Q. Guo, and T. Moraitis, âHebbian deep learning without feedback,â arXiv preprint arXiv:2209.11883, 2022. [24] T. Rudroff, O. Rainio, and R. Klen, âNeuroplasticity meets artificial intelligence: A hippocampus-inspired approach to the stabilityâplasticity dilemma,â Brain Sciences, vol. 14, no. 11, p. 1111, 2024. [25] A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell, âProgressive neural networks,â arXiv preprint arXiv:1606.04671, 2016. [26] V. Coscrato, M. H. de Almeida Inacio, and R. Izbicki, âThe n- stacking: Feature weighted linear stacking through neural networks,â Neurocomputing, vol. 399, p. 141â152, 2020. [27] S. Chen, Z. Liu, and H. You, âBrain-inspired efficient pruning: Exploit- ing criticality in spiking neural networks,â Concurrency and Computa- tion: Practice and Experience, vol. 37, no. 27-28, p. e70404, 2025. [28] Q. Yousef and P. Li, âSynaptic plasticity-based regularizer for artificial neural networks,â Scientific Reports, vol. 15, no. 1, p. 14330, 2025. [29] T. Miconi, K. Stanley, and J. Clune, âDifferentiable plasticity: training plastic neural networks with backpropagation,â in International Confer- ence on Machine Learning. PMLR, 2018, p. 3559â3568. [30] J. N. Heidenreich, C. Bonatti, and D. Mohr, âTransfer learning of recurrent neural network-based plasticity models,â International Journal for Numerical Methods in Engineering, vol. 125, no. 1, p. e7357, 2024. [31] Y. Li, M. Milling, L. Specia, and B. W. Schuller, âFrom audio deepfake detection to ai-generated music detectionâa pathway and overview,â arXiv preprint arXiv:2412.00571, 2024. [32] T. Chen, A. Kumar, P. Nagarsheth, G. Sivaraman, and E. Khoury, âGeneralization of audio deepfake detection.â in Odyssey, 2020, p. 132â 137. [33] L. Pham, P. Lam, T. Nguyen, H. Nguyen, and A. Schindler, âDeepfake audio detection using spectrogram-based feature and ensemble of deep learning models,â in 2024 IEEE 5th International Symposium on the Internet of Sounds (IS2). IEEE, 2024, p. 1â5. [34] Y. Li, L. Wang, Y. Wang, L. Wang, R. Cai, J. Shi, B. W. Schuller, and Z. Wu, âDfallm: Achieving generalizable multitask deepfake detection by optimizing audio llm components,â arXiv preprint arXiv:2512.08403, 2025. [35] K. Shahriar, âLightweight resolution-aware audio deepfake detection via cross-scale attention and consistency learning,â arXiv preprint arXiv:2601.06560, 2026. [36] H. Gu, J. Yi, C. Wang, J. Tao, Z. Lian, J. He, Y. Ren, Y. Chen, and Z. Wen, âAllm4add: Unlocking the capabilities of audio large language models for audio deepfake detection,â in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, p. 11 736â11 745. [37] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., âLora: Low-rank adaptation of large language models.â ICLR, 2022. [38] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, âGrad-cam: Visual explanations from deep networks via gradient-based localization,â in Proceedings of the IEEE international conference on computer vision, 2017, p. 618â626.