Paper deep dive
Reconstructing Spiking Neural Networks Using a Single Neuron with Autapses
Wuque Cai, Hongze Sun, Quan Tang, Shifeng Mao, Zhenxing Wang, Jiayi He, Duo Chen, Dezhong Yao, Daqing Guo
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/27/2026, 1:15:16 AM
Summary
The paper introduces the Time-Delayed Autapse Spiking Neural Network (TDA-SNN), a framework that reconstructs spiking neural networks using a single leaky integrate-and-fire (LIF) neuron. By utilizing delayed autapses to reorganize internal temporal states, the model can emulate reservoir computing, multilayer perceptrons, and convolutional architectures within a unified, compact framework, offering significant reductions in neuron count and memory usage.
Entities (4)
Relation Signals (3)
TDA-SNN → utilizes → TDA-LIF
confidence 100% · TDA-LIF neuron... forms the basis of a corresponding spiking neural network model termed TDA-SNN.
TDA-SNN → implements → Reservoir Computing
confidence 95% · TDA-SNN can realize reservoir, multilayer perceptron, and convolution-like spiking architectures
STBP → optimizes → TDA-SNN
confidence 95% · Using equations (8)–(7), TDA-SNN can be trained end-to-end with surrogate gradients.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Spiking neural networks (SNNs) are promising for neuromorphic computing, but high-performing models still rely on dense multilayer architectures with substantial communication and state-storage costs. Inspired by autapses, we propose time-delayed autapse SNN (TDA-SNN), a framework that reconstructs SNNs with a single leaky integrate-and-fire neuron and a prototype-learning-based training strategy. By reorganizing internal temporal states, TDA-SNN can realize reservoir, multilayer perceptron, and convolution-like spiking architectures within a unified framework. Experiments on sequential, event-based, and image benchmarks show competitive performance in reservoir and MLP settings, while convolutional results reveal a clear space--time trade-off. Compared with standard SNNs, TDA-SNN greatly reduces neuron count and state memory while increasing per-neuron information capacity, at the cost of additional temporal latency in extreme single-neuron settings. These findings highlight the potential of temporally multiplexed single-neuron models as compact computational units for brain-inspired computing.
Tags
Links
- Source: https://arxiv.org/abs/2603.24692v1
- Canonical: https://arxiv.org/abs/2603.24692v1
Trouble viewing inline? Open PDF directly →
Full Text
64,116 characters extracted from source content.
Expand or collapse full text
Reconstructing Spiking Neural Networks Using a Single Neuron with Autapses Wuque Cai 1 Hongze Sun 1 Quan Tang 1 Shifeng Mao 1 Zhenxing Wang 1 Jiayi He 1 Duo Chen 2∗ Dezhong Yao 1∗ Daqing Guo 1∗ 1 Brain-Apparatus Communication Institute, University of Electronic Science and Technology of China, Chengdu, China 2 School of Artificial Intelligence, Chongqing University of Education, Chongqing, China 1 dyao,dqguo@uestc.edu.cn, 2 duochen3@gmail.com Abstract Spiking neural networks (SNNs) are promising for neu- romorphic computing, but high-performing models still rely on dense multilayer architectures with substantial communication and state-storage costs. Inspired by au- tapses, we propose time-delayed autapse SNN (TDA-SNN), a framework that reconstructs SNNs with a single leaky integrate-and-fire neuron and a prototype-learning-based training strategy. By reorganizing internal temporal states, TDA-SNN can realize reservoir, multilayer perceptron, and convolution-like spiking architectures within a unified framework. Experiments on sequential, event-based, and image benchmarks show competitive performance in reser- voir and MLP settings, while convolutional results reveal a clear space–time trade-off. Compared with standard SNNs, TDA-SNN greatly reduces neuron count and state memory while increasing per-neuron information capacity, at the cost of additional temporal latency in extreme single-neuron settings. These findings highlight the potential of tempo- rally multiplexed single-neuron models as compact compu- tational units for brain-inspired computing. 1. Introduction Spiking neural networks (SNNs) are widely regarded as the third generation of neural networks [17, 26]. They serve as an important foundation for neuromorphic computing [31], featuring event-driven processing, high energy efficiency, and rich temporal dynamics [12, 53]. However, most high- performing SNNs still adopt dense multilayer architectures similar to artificial neural networks (ANNs) [19, 32, 41–43]. Such designs require extensive neuron communication and large state storage, limiting spatial efficiency and hindering deployment in resource-constrained settings [48, 49, 51]. This limitation motivates the search for compact SNN ar- * Corresponding authors. chitectures that preserve computational capability while re- ducing structural redundancy. [46] The intrinsic computational potential of single neurons provides a promising direction for advancing information processing.Biological neurons not only perform den- dritic integration [24, 35] but also exhibit complex mech- anisms such as intrinsic plasticity [13] and lateral inhi- bition [28], enabling rich nonlinear dynamics across spa- tiotemporal dimensions. In the SNN field, various strategies have been proposed to enhance single-neuron expressive- ness, including dendritic computation [25], intrinsic plastic- ity [15, 38], lateral inhibition–inspired competition mecha- nisms [10, 50], neuronal heterogeneity [30, 52], and multi- compartment modeling [9]. Although these approaches im- prove the nonlinear computational capability of individual neurons to some extent [23], their internal states still rely mainly on instantaneous inputs, which limits their ability to retain long-term information or perform recursive computa- tions [5]. Among these, cerebellar Purkinje cells exemplify single-neuron complexity through elaborate dendrites and autaptic self-feedback, supporting temporal processing and memory. In exploring the computational potential of single- neuron structures, existing engineering approaches based on single-neuron delay loops offer valuable insights. Stud- ies on single-neuron deep learning [36, 37] and one-core reservoir computing (RC) models [29] demonstrate that a single node possesses significant capacity for temporal in- formation processing. However, these methods typically rely on non-biological hardware, and their continuous-value computations fundamentally differ from spike-based, event- driven mechanisms, limiting their direct applicability to SNNs and biological interpretation. In biological neural systems, autapses provide a natu- ral mechanism for temporal self-interaction. An autapse is a synaptic connection that a neuron forms with itself, enabling it to sense its own past spiking activity, and to exert either inhibitory or excitatory modulation on its cur- arXiv:2603.24692v1 [cs.NE] 25 Mar 2026 rent state [2, 40, 45]. From a computational perspective, autapses introduce intrinsic temporal memory within neu- rons, naturally embedding the high-dimensional informa- tion into their dynamics and providing a biological basis for single neurons to realize long-term dependencies and recur- sive computation [33, 47]. Such autaptic self-modulation in Purkinje cells inspires our model design for long-term de- pendencies and recursive computation. Based on this observation, we propose a leaky integrate- and-fire neuron with delayed autapses (TDA-LIF), which forms the basis of a corresponding spiking neural network model termed TDA-SNN. Under different structural con- figurations, the TDA-LIF neuron reorganizes its internal temporal-node evolution to construct reservoir-computing, multilayer-perceptron, and convolution-like spiking archi- tectures within a unified framework. Experiments on se- quential, event-based, and image datasets show that TDA- SNN provides competitive performance in reservoir and MLP settings, while its convolutional results reveal a clear space–time trade-off. Compared with structurally equiva- lent standard SNNs, TDA-SNN drastically reduces neuron count and state memory and increases per-neuron informa- tion capacity, at the cost of additional temporal latency in extreme single-neuron settings. Our main contributions are summarized as follows: • We theoretically show that a TDA-LIF neuron can constructively realize three representative SNN struc- tures:reservoir computing (RC), multilayer percep- trons (MLPs), and convolution-like architectures. • We propose a dedicated optimization framework for TDA-SNN to address the training challenges introduced by internal delayed feedback, enabling stable learning of temporally unfolded single-neuron models. • Extensive experiments on multiple benchmarks demon- strate that TDA-SNN achieves competitive performance in RC and MLP settings, and reveal a clear space–time trade-off in convolutional settings, while substantially re- ducing neuron count and state memory and significantly enhancing per-neuron information capacity. 2. Related Works 2.1. Neurons with High Computing Capacity Early artificial neuron models, such as McCulloch-Pitts neurons [27], greatly simplify the richer dynamics of bio- logical neurons, whose dendrites, ion channels, and non- linear membrane properties support powerful local com- putation. [7] In particular, dendrites can integrate inputs, perform nonlinear operations, and even implement logical functions [18, 24, 35], motivating SNN studies that enhance single-neuron capability through dendritic computation, in- trinsic plasticity, lateral inhibition, neuronal heterogeneity, and multi-compartment modeling. [9, 10, 15, 23, 25, 30, 38, 1 2 3 4 5 6 7 8 9 10 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold Mapping u(t) 12345678910 d=5 Evolution timestep ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (b) Connection Weight (c) (a) Mapping 1 2 3 4 5 6 7 8 9 10 u(t) 12345678910 Evolution time 휕ℒ 푣 푙 푡 휕ℒ 푣 푙 푡+1 휕ℒ 푣 푙+1 푡+4 휕ℒ 푣 푙+1 푡+5 휕ℒ 푣 푙+1 푡+7 휕ℒ 푠 푙+1 푡+4 휕ℒ 푠 푙+1 푡+5 휕ℒ 푠 푙+1 푡+7 (a) (b) 푣 푙 푡 s 푙−1 푡−4 s 푙−1 푡−5 s 푙−1 푡−7 푣 푙−1 푡−7 푣 푙−1 푡−5 푣 푙−1 푡−4 푣 푙 푡−1 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold 1 2 3 4 5 6 7 8 9 10 Mapping u(t) 123456789 10 d=5 Evolution time Figure 1. Evolution and folding principles of autapses. (a) An autapse in biological neurons and its evolutionary unfolding. (b) Distribution of evolutionary time in the TDA-LIF model and its mapping to reservoir computation. 50, 52] However, the role of temporal delay in improving single-neuron computation remains underexplored. [3, 5, 8] 2.2. Time-Delay Neural Network Transmission delay is a fundamental resource for tempo- ral processing in both biological and artificial neural sys- tems. [4, 34] Time-delay neural networks explicitly model temporal dependencies through delayed connections [20, 39], and folded-in-time deep networks further compress depth into delay loops within a single node, achieving strong performance in vision and EEG tasks. [29, 36, 37] Yet these approaches typically rely on continuous-valued signals and have limited biological plausibility. In contrast, biological autapses allow neurons to regulate current activ- ity using their own past spikes. [2, 40, 45] This perspective motivates us to view delay as an intrinsic autaptic property. It enables a single spiking neuron to emulate network-like temporal computation while preserving interpretability and improving long-range temporal modeling. 3. Method We introduce time-delayed autapses into the LIF neuron to enhance its temporal dynamics and internal memory. Based on this modification, we construct a temporally unfolded single-neuron model and show how it can be mapped to reservoir-computing, MLP, and convolution-like spiking ar- chitectures. Unless otherwise stated, t denotes the index of an internal temporal node, while T denotes the external data time window used for sequential or event-based inputs. 3.1. LIF Neuron with Time-delayed Autapses Introducing an autapse with delayd into the LIF neuron [16] provides temporal self-feedback: a spike emitted by the neuron is fed back to its dendrites after d internal nodes. As shown in Fig. 1, the axon connects to the neuron’s own den- drites, allowing past spikes to influence future membrane- potential evolution. Mathematically, the iterative formula for the TDA-LIF is as follows: v t = τv t−1 (1− s t−1 ) + I t + I t autapse ,(1) where I t autapse = P d∈D t w d a s t−d is the delayed autaptic current generated by the neuron’s own past spikes, D t = d∈N| d < t is the set of valid delays at node t, and w d a denotes the weight of the autapse with delay d. 3.2. Temporal Dynamics to Reservoir Computing Based on the TDA-LIF model, we unfold neuron dynamics into a sequence of temporal nodes, where autapses connect nodes across different internal timesteps. We refer to each internal timestep as a node to distinguish it from the external data time window. These connections are directional and transmit signals forward along the unfolded trajectory. For simplicity, autapses with the same delay are treated as one delay type. As shown in Fig. 1, the neuron evolves over ten nodes, with blue and yellow autapses representing delays of five and two nodes, respectively. When autapses with different delays are assigned different weights, delayed self-feedback modulates the membrane-potential trajectory and enriches the internal temporal state. Under a finite set of delays, the membrane dynamics can be written as: v t = τv t−1 (1− s t−1 ) + Wx t + X d∈D w d a (t)s t−d ,(2) where x t is the external input injected into node t, D de- notes the set of selected autaptic delays, w d a (t) is the weight of the autapse with delay d at node t, and W ∈R N×N in represents the weight matrix for the external inputs, with N being the number of nodes and N in the input dimension. When arranged as in the right panel of Fig. 1(b), the un- folded temporal graph induces an RC-like connectivity pat- tern. Signals propagate along the nodes following the dotted arrows, while delayed autaptic feedback continuously mod- ulates the hidden dynamics. The resulting interaction across temporal nodes enriches the internal states, and aggregat- ing neuronal responses over the unfolded trajectory yields a high-dimensional representation for reservoir computing. 3.3. MLP Equivalence by Autapse Pruning As shown in Fig. 2, the unfolded trajectory is divided into two segments, highlighted by the red dashed line, and au- tapses are pruned accordingly. Specifically, autapses con- fined within the same segment are removed, while those crossing the segment boundary are retained. Rearranging the temporal nodes into two groups then yields a feedfor- ward network structure. 1 2 3 4 5 6 7 8 9 10 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold Mapping u(t) 12345678910 d=5 Evolution timestep ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (b) Connection Weight (c) (a) Mapping 1 2 3 4 5 6 7 8 9 10 u(t) 12345678910 Evolution time 휕ℒ 푣 푙 푡 휕ℒ 푣 푙 푡+1 휕ℒ 푣 푙+1 푡+4 휕ℒ 푣 푙+1 푡+5 휕ℒ 푣 푙+1 푡+7 휕ℒ 푠 푙+1 푡+4 휕ℒ 푠 푙+1 푡+5 휕ℒ 푠 푙+1 푡+7 (a) (b) 푣 푙 푡 s 푙−1 푡−4 s 푙−1 푡−5 s 푙−1 푡−7 푣 푙−1 푡−7 푣 푙−1 푡−5 푣 푙−1 푡−4 푣 푙 푡−1 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold 1 2 3 4 5 6 7 8 9 10 Mapping u(t) 123456789 10 d = 5 Evolution timestep Figure 2. Division of the evolutionary time in the TDA-LIF model and its mapping to an MLP (b) (a) Figure 3. Convolutional layer with the TDA-LIF model. (a) Map- ping 10 nodes to a 3×3 convolution. (b) Equivalence of TDA-LIF to a single-channel convolutional layer. This construction produces an MLP-like structure from a single TDA-LIF neuron. In Fig. 2, the neuron integrates three delayed autaptic groups, d = 2 (yellow), 5 (blue), and 7 (green), yielding ten effective synaptic connections between the two temporal segments. During the evolution of the second segment, the dynamics can be written as: v t = τv t−1 (1− s t−1 ) + Wx t + X d∈D FC w d a (t)s t−d ,(3) where D FC = d ∈N | d ≤ t, t− d < t s denotes the set of delays that connect the first segment to the second, and t s is the segment boundary. Building on this foundation, the TDA-LIF model can be extended to deeper feedforward networks by introducing more temporal segments. Inputs are assigned to the first segment, and spike responses from the final segment are used for recognition. In the corresponding inter-segment connection matrix, the diagonal and parallel-diagonal ele- ments represent autaptic connections with identical delays. 3.4. Convolutional Structure Equivalence Starting from the MLP-like topology, we assign each spa- tial output node a set of autapses with shared weights and corresponding delays, thereby encoding the weight-sharing mechanism of a convolution kernel into the neuron’s inter- nal temporal structure. These shared autaptic weights play the role of kernel parameters, while distinct delay groups provide a temporal realization of local spatial aggregation. During signal propagation, output nodes operate in par- allel and synchronously, each using the shared autaptic ker- nel to process a different local region of the input. In this 1 2 3 4 5 6 7 8 9 10 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold Mapping u(t) 12345678910 d=5 Evolution timestep ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (b) Connection Weight (c) (a) Mapping 1 2 3 4 5 6 7 8 9 10 u(t) 12345678910 Evolution time 휕ℒ 푣 푙 푡 휕ℒ 푣 푙 푡+1 휕ℒ 푣 푙+1 푡+4 휕ℒ 푣 푙+1 푡+5 휕ℒ 푣 푙+1 푡+7 휕ℒ 푠 푙+1 푡+4 휕ℒ 푠 푙+1 푡+5 휕ℒ 푠 푙+1 푡+7 (a) (b) 푣 푙 푡 s 푙−1 푡−4 s 푙−1 푡−5 s 푙−1 푡−7 푣 푙−1 푡−7 푣 푙−1 푡−5 푣 푙−1 푡−4 푣 푙 푡−1 Evolution time 휏 Autapse Dendrite Axis (a) (b) Unfold 1 2 3 4 5 6 7 8 9 10 Mapping u(t) 123456789 10 d = 5 Evolution timestep Figure 4. Forward (a) and backward (b) processes in the TDA- SNN model. way, kernel translation across space is mapped to a repeated delay-sharing pattern across temporal nodes. As illustrated in Fig. 3, this construction provides a convolution-like real- ization in the spike domain and unifies local spatial aggre- gation with temporal dynamics. For example, consider a convolution with a 3× 3 kernel, padding 1, and stride 1, with a single input and output chan- nel for simplicity. The TDA-LIF state at a given node is de- termined by delayed inputs from nine neighboring positions in the previous layer, corresponding to nine delay groups. Because all spatial locations share the same delayed autap- tic weights, the unfolded dynamics recover the key proper- ties of local connectivity and weight sharing. This yields a constructive convolution-like mapping, which we later eval- uate empirically in a preliminary setting. 3.5. Prototype Learning Method To effectively decode neuronal firing patterns, it is essential to exploit their intrinsic spatio-temporal dynamics. Con- ventional static decoding methods neglect temporal depen dencies and fail to align with the spiking nature of SNN outputs. To address this limitation, we introduce a spa- tiotemporal prototype learning framework that matches out- put spike sequences to learnable prototypes, enabling effi- cient and interpretable recognition of dynamic firing behav- iors [6]. See the supplementary materials for additional de- tails. To effectively decode neuronal firing patterns, it is essential to exploit their intrinsic spatio-temporal dynam- ics. Conventional static decoding methods neglect tempo- ral dependencies and fail to align with the spiking nature of SNN outputs. To address this limitation, we introduce a spatiotemporal prototype-learning framework that matches output spike sequences to learnable prototypes, enabling ef- ficient and interpretable recognition of dynamic firing be- haviors [6]. Additional implementation details are provided in the supplementary material. Specifically, a set of binary prototypes K ∈ 0, 1 C×T is defined to represent target spiking patterns for C classes over T steps. Let f(x;θ) denote the encoded spatiotempo- ral representation of input x. The similarity between f(x;θ) and the i-th prototype k i ∈ K is measured by d i (x) =−∥f(x;θ)− k i ∥ 2 2 . Based on the distances between the encoded representa- tion and class prototypes, the classification objective is for- mulated as: L =− log e d i (x) P C j=1 e d j (x) − λd i (x),(4) where C represents the number of classes, and λ is a hy- perparameter controlling the contribution of the prototype regularization term (set to 0.001 in our experiments). For network optimization, we adopt the spatio-temporal backpropagation (STBP) algorithm, which has been widely used in SNNs due to its efficiency in handling temporal de- pendencies. [42, 43] To overcome the non-differentiability of spike firing in TDA-LIF neurons and prototype binariza- tion, we replace the gradient with a smooth approximation using the arctangent function: s t = g v t ≈ 1 π arctan h π 2 α(v t − v th ) i + 1 2 .(5) Fig. 4 illustrates information flow in TDA-SNN. In for- ward propagation, node information arises from both de- layed autaptic spikes and the membrane potential carried over from the previous node, as shown in Fig. 4(a). Accord- ingly, in backward propagation, the gradient at each node is determined by the membrane-potential gradients of its suc- cessor node and of the future nodes that receive its delayed spikes, as shown in Fig. 4(b). For the RC and MLP settings, the membrane-potential gradient at node t is: ∇ v t L =∇ v t+1 L· ∂v t+1 ∂v t + X d∈D ∇ v t+d L· ∂v t+d ∂v t =∇ v t+1 L· τ(1− s t−1 ) + X d∈D ∇ v t+d L· w d a · ∂s t ∂v t , (6) wherew d a is the autaptic weight connecting the spike at node t to the future node t + d. The corresponding weight gradi- ents are ∇ w L =∇ v t L· x t , ∇ w d a L =∇ v t+d L· s t . (7) Using equations (8)–(7), TDA-SNN can be trained end- to-end with surrogate gradients. 4. Experiments 4.1. Experimental Setup We evaluate TDA-SNN across multiple benchmarks under different structural mappings. DEAP [21] and SHD [11] 1664256 74 90 82 N umber of nodes Acc. (%) DEAP 64 82 73 Number of nodes SHD (a) (b) 1664256 STD-SNN TDA-SNN Figure 5. Comparison of STD-SNN and TDA-SNN performance under the RC structure with different reservoir sizes on (a) DEAP and (b) SHD datasets. are used to evaluate the reservoir-computing setting; MNIST [14], Fashion-MNIST (fMNIST) [44], and DVS Gesture [1] are used to validate the MLP setting; and DVS Gesture together with CIFAR10 [22] are further used to study the convolution-like setting. All models are imple- mented in PyTorch and trained on an NVIDIA A100 GPU using the Adam optimizer with a cosine learning-rate decay schedule. Unless otherwise stated, training is conducted for 100 epochs. For RC and MLP experiments, results are av- eraged over 10 independent runs and reported as mean ± standard deviation; due to the higher computational cost of convolutional experiments, those results are averaged over 5 runs. To ensure a fair comparison, TDA-SNN and the corresponding standard SNN baselines use aligned training protocols and matched architectural scales for each setting. The complete hyperparameter configurations are provided in Supplementary Sec. 9. 4.2. Reservoir Computing Equivalence To assess the RC capability of TDA-SNN, we conducted experiments on the DEAP and SHD datasets. We further in- vestigate how reservoir size and autaptic selection strategies affect connectivity patterns and classification performance. 4.2.1. Analysis of RC Equivalence To compare the performance of the standard SNN (STD- SNN) and the proposed TDA-SNN with varying reservoir sizes, we conducted a series of experiments. Specifically, the number of reservoir neurons in the SNN was matched to the number of internal evolution nodes in the TDA-SNN to ensure comparable model complexity, thereby enabling a fair assessment of their representational and learning ca- pabilities. As the reservoir size increased, the projection space became richer and more structured, leading to im- proved feature separability and discriminative performance. In this ablation study, reservoir sizes of 16, 32, 64, 128, and 256 were evaluated, with fully delayed autapses employed in all configurations. As shown in Figs. 5(a) and (b), TDA-SNN under- performs STD-SNN at small reservoir sizes (e.g., 16 nodes), reaching 77.59±1.24% versus 81.21±1.00% on DEAP and 67.45±0.84% versus 67.92±0.92% on SHD. As the reservoir size increases, TDA-SNN improves steadily and becomes competitive with, or even surpasses, STD- SNN at larger sizes. At 256 nodes, TDA-SNN reaches 88.65±0.48% on DEAP and 80.04±0.51% on SHD, while STD-SNN attains 79.92±0.54% and 77.63±2.53%, respec- tively. On SHD, the best TDA-SNN result is obtained at 128 nodes (80.68±0.83%), suggesting that excessive inter- nal evolution can eventually introduce redundant temporal encoding. Overall, these results indicate that delayed au- taptic dynamics become increasingly effective as the un- folded reservoir grows, making TDA-SNN a competitive alternative to RC-style SNNs in medium- and large-scale settings. As shown in Figs. 5(a) and (b), TDA-SNN un- derperforms STD-SNN at small reservoir sizes (e.g., 16 nodes), reaching 77.59±1.24% versus 81.21±1.00% on DEAP and 67.45±0.84% versus 67.92±0.92% on SHD. As the reservoir size increases, TDA-SNN improves steadily and becomes competitive with, or even surpasses, STD- SNN at larger sizes. At 256 nodes, TDA-SNN reaches 88.65±0.48% on DEAP and 80.04±0.51% on SHD, while STD-SNN attains 79.92±0.54% and 77.63±2.53%, respec- tively. On SHD, the best TDA-SNN result is obtained at 128 nodes (80.68±0.83%), suggesting that excessive inter- nal evolution can eventually introduce redundant temporal encoding. These results indicate that delayed autaptic dy- namics become increasingly effective as the unfolded reser- voir grows, making TDA-SNN a competitive alternative to RC-style SNNs in medium- and large-scale settings. 4.2.2. Effect of autaptic selection strategy This subsection investigates the influence of autaptic selec- tion strategies on the performance of TDA-SNN under the RC framework. The reservoir size was fixed at 128 neurons, and two strategies, random delay (RD) and maximum con- nection (MC), were compared across different numbers of delayed autapses (1, 2, 4, 8, 16, 32, and 64). Table 1 summarizes the classification performance of TDA-SNN under different autaptic selection strategies on DEAP and SHD. On DEAP, both strategies benefit from in- creasing the number of delayed autapses, and MC shows a slight advantage at larger delay counts (32 and 64), indi- cating that denser structured delay assignment can improve information propagation when the task is relatively stable. On SHD, the trend is less monotonic: MC drops markedly at two delays, then recovers as more delayed connections are introduced. This observation suggests that on tempo- rally richer datasets, overly constrained delay patterns may be suboptimal in low-delay regimes, whereas sufficient de- layed connections allow structured feedback to recover its representational advantage. Thus, the ablation indicates that both the number of autapses and the delay-selection strat- egy materially affect RC performance, and that moderate- Table 1. Experimental Results of Autaptic Selection Strategies Across Different Datasets DatasetDelay1248163264 DEAP MC83.52±0.5283.10±1.1083.63±0.7383.91±1.1483.85±0.5984.66±0.6284.47±0.99 RD83.62±1.1283.34±1.1483.48±0.6483.73±1.0283.63±0.7683.98±0.8184.47±0.99 SHD MC78.35±0.6172.09±1.1773.59±0.6774.69±1.2877.07±0.4879.36±0.3180.42±0.39 RD78.55±0.6878.62±0.4677.60±2.4978.58±0.2577.63±1.8779.26±1.2180.42±0.39 MNIST MC33.53±3.5194.44±0.1094.39±0.1694.33±0.1394.37±0.1594.42±0.2094.38±0.12 RD84.03±19.1493.32±1.1994.82±0.1595.11±0.2195.63±0.1495.98±0.1196.42±0.09 T Inv.90.33±4.9493.21±1.8493.93±0.4794.37±0.2494.37±0.2094.42±0.1794.72±0.14 fMNIST MC34.25±2.7886.63±0.1186.59±0.1686.60±0.2586.58±0.1886.63±0.1586.55±0.17 RD80.51±14.0585.68±1.2886.54±0.2286.88±0.1487.23±0.1387.40±0.2387.55±0.16 T Inv.83.62±3.3085.90±0.8686.32±0.2486.53±0.2186.62±0.2086.59±0.1686.74±0.19 DVS MC23.13±2.8652.88±1.8252.71±1.6752.88±2.3255.28±2.1254.44±3.1053.19±2.64 Gesture RD45.45±8.7951.84±5.3253.61±1.9155.31±2.9257.88±1.9458.40±1.6959.55±2.04 T Inv.40.87±8.1550.21±6.7154.34±2.0254.79±3.3155.38±2.3555.42±2.6857.01±2.48 N umber of nodes Acc. (%) 91 99 95 16256 MNIST (a) Number of nodes 84 90 87 16256 fMNIST (b) Number of nodes 42 64 53 16256 DVS Gesture (c) STD-SNN TDA-SNN 646464 Figure 6. Comparison of TDA-SNN and STD-SNN accuracy under MLP architectures with varying numbers of nodes on (a) MNIST, (b) fMNIST, and (c) DVS Gesture. to-large delay sets are generally more reliable than sparse configurations. Statistical significance analyses of the ex- perimental results are provided in Supplementary Sec. 11. 4.3. Multilayer Perceptron Equivalence To further evaluate the generality of TDA-SNN in feed- forward settings, we extended the experiments to MLP- like architectures. Specifically, we investigated whether the time-delayed autaptic mechanism can reproduce feed- forward spiking representations through temporal unfold- ing. Following the RC analysis, we examined the effects of temporal-node number, autaptic selection strategy, and network depth on MNIST, fMNIST, and DVS Gesture. 4.3.1. MLP Structural Equivalence To validate the structural equivalence and representational capability of TDA-SNNs in multilayer feedforward archi- tectures, we compared their performance with STD-SNNs under identical MLP configurations. The network com- prises two layers: the first layer performs preliminary fea- ture extraction from external inputs, while the second layer generates the output patterns for classification. All experi- ments employed fully delayed self-synapses to ensure max- imal temporal connectivity. To investigate the effect of net- work scale, we varied the number of nodes in each layer (16, 32, 64, 128, and 256). Fig. 6(a)–(c) presents the classification accuracy of STD- SNNs (red) and TDA-SNNs (blue) across different node numbers. TDA-SNN accuracy increases steadily with net- work size, from 92.37±0.24% to 98.23±0.09% on MNIST, from 85.23±0.27% to 89.16±0.12% on fMNIST, and from 53.99±0.86% to 72.95±2.02% on DVS Gesture. At small scales (e.g., 16 nodes), STD-SNN slightly outperforms TDA-SNN, suggesting that limited representational capac- ity reduces the benefit of temporal multiplexing. As the number of nodes increases, the gap narrows substantially, indicating that TDA-SNN can serve as a competitive feed- forward alternative when sufficient temporal nodes are available. These results support the effectiveness of the proposed temporal construction in medium- and large-scale MLP settings, while also showing that its benefits are less pronounced in low-capacity regimes. 4.3.2. Effect of Strategy Selection This subsection investigates the impact of autapse delay se- lection strategies on TDA-SNN performance in MLP ar- chitectures. We compare RD, MC, and time-invariant (T Inv.) strategies across varying numbers of delayed autapses, where the T Inv. strategy assigns identical weights to au- tapses with the same delay. These experiments evaluate how the choice of delay strategy and weight sharing influences temporal encoding and feature extraction. Increasing the number of autapses improves classifica- tion accuracy across MNIST, fMNIST, and DVS Gesture, as Number of layers 1234 84 88 86 fMNIST Number of layers 1234 44 60 52 DVS Gesture Acc. (%) MNIST 1234 90 98 94 Number of layers (b)(c)(a) Figure 7. Experimental results of networks with different layer depths on three benchmark datasets. (a) MNIST, (b) fMNIST, and (c) DVS Gesture. shown in Tab. 1. On MNIST and fMNIST, the three strate- gies converge quickly once two or more autapses are used, with MC showing a clear disadvantage only in the extreme low-delay regime. On the more temporally demanding DVS Gesture dataset, RD and T-Inv. outperform MC at very small autapse counts, indicating that greater delay diversity is beneficial when temporal structure is complex. As the number of autapses increases beyond four, the gap between strategies narrows and MC becomes comparable to the other two strategies. This ablation suggests that the number of de- layed connections is the dominant factor, while the specific selection strategy mainly affects low-resource settings. Sta- tistical significance analyses of the experimental results are provided in Supplementary Sec. 11. 4.3.3. Impact of Layer Number in MLP Structures We investigated the effect of network depth (N layer ) on TDA-SNN performance in MLP structures, fixing 64 nodes per layer and employing the RD selection strategy. Classifi- cation performance was evaluated across networks with one to four layers. The results are presented in Fig. 7. On MNIST, accuracy steadily increases from 94.41 ± 0.17% to 95.60± 0.15% as depth grows from 1 to 4 layers, indicating that additional depth remains beneficial in this setting. For fMNIST, accuracy rises from 86.62 ± 0.09% to 87.12± 0.19% at 1–2 layers but then decreases slightly at 3 and 4 layers, suggesting that moderate depth is suffi- cient. On DVS Gesture, performance improves from 1 to 2 layers and then largely saturates, reaching its best value of 57.12 ± 1.48% at 3 layers. Thus, the effect of depth is dataset-dependent: deeper networks help on MNIST, mod- erate depth is adequate on fMNIST, and the benefit saturates early on DVS Gesture. Temporal unfolding supports deeper feedforward constructions, but the optimal depth depends on task complexity and data characteristics. Further evidence for the robustness of the MLP setting is provided in Supplementary Sec. 13. The MNIST learning- rate results confirm stable optimization, and the CIFAR100 results with different numbers of nodes show that TDA- SNN remains competitive on a more challenging dataset. (a) No. of layer 1234 84 88 86 fMNIST (c) No. of layer 1234 44 60 52 DVS Gesture (d) N layer Acc. (%) MNIST (b) 1234 90 98 94 No. of layer Acc. (%) (a) 0 80 40 0100 Epoch (b) 20 50 35 0100 Epoch DVS GestureCIFAR10 STD-SNN TDA-SNN Figure 8. Test accuracy comparison between the TDA-SNN and the STD-SNN under convolutional structure on two benchmark datasets. (a) DVS Gesture and (b) CIFAR-10. 4.4. Convolutional Equivalence Exploration To examine whether the proposed temporal construction can extend beyond RC and MLP settings, we performed preliminary experiments with convolution-like architec- tures on DVS Gesture and CIFAR-10, which represent dy- namic and static visual inputs, respectively. These experi- ments are intended as an initial proof of concept for apply- ing TDA within convolutional layers. Further implementa- tion details are provided in Supplementary Sec. 14. Fig. 8 shows classification performance across train- ing epochs. On DVS Gesture, STD-SNN reaches 73.61 ± 1.93% accuracy and stabilizes after approximately 60 epochs, whereas TDA-SNN converges more slowly and plateaus at 57.58 ± 1.00%. A similar trend is observed on CIFAR-10, where STD-SNN reaches 47.31 ± 0.24% and TDA-SNN stabilizes at 37.64 ± 0.65%. These re- sults suggest that, although delayed self-feedback enriches temporal dynamics, it also introduces substantial optimiza- tion difficulty in convolutional settings, likely because de- layed feedback interferes with hierarchical spatiotemporal feature aggregation. Therefore, the current single-neuron convolution-like construction should be viewed as a prelim- inary exploration rather than a fully competitive convolu- tional alternative. Improving this setting will likely require multi-neuron parallelization to better balance spatial repre- sentation and temporal recursion. 4.5. Computational Complexity Analysis To evaluate efficiency, we analyzed the per-layer complex- ity of both STD-SNNs and TDA-SNN under three represen- tative architectures: RC, MLP, and convolutional networks. For the RC structure, we used SHD; for MLP, MNIST; and for convolutional networks, CIFAR-10. Both the RC and MLP structures were scaled to a size of 64. We report FLOPs, parameter count, neuron count, and training speed (seconds per epoch). In all cases, TDA-SNN replaces con- ventional synapses with time-delayed autapses and uses all available autapse delays. As shown in Table 2, STD-SNNs exhibit high spatial Table 2. Per-layer computational complexity, parameter count, neuron count, speed, and per-neuron information content across different models. ModelArch.FLOPsParams.N Neu. SpeedS (s/epoch)(bit) STD-SNN RCN 2 TN 2 N22.57489.91 MLPN in N out TN in N out N out 8.363113.89 Conv.2C in C out k 2 WHTC in C out k 2 C out WH17.215.48 TDA-SNN RC P d (N − d)T P d (N − d)121.1432096.37 MLP P d N d T P d N d 190.91199237.63 Conv. 2C in C out ( P d N d )WHT C in C out ( P d N d )13060.2356486.07 complexity, with RC and MLP structures scaling quadrat- ically or linearly with the number of connections (N 2 T and N in N out T ), whereas TDA-SNN converts these spatial couplings into temporal dependencies ( P d N d T ). This de- sign drastically reduces the number of physical neurons per layer, but the benefit comes at the cost of additional tempo- ral computation due to delayed feedback. The trade-off is mild in the RC setting, where training speed remains com- parable, but becomes much more pronounced in MLP and especially convolutional settings. Therefore, the main effi- ciency advantage of TDA-SNN lies in spatial compactness and neuron-count reduction rather than uniformly lower runtime cost. This observation is consistent with the space– time trade-off discussed throughout the paper. 4.6. Discussion on Single-Neuron Storage Capacity We further analyzed the storage capacity of individual neu- rons by quantifying the per-neuron information content un- der the same experimental settings as in the previous section for RC, MLP, and convolutional networks. Here, S denotes the effective information content carried by a single neuron and is computed as S = Acc.× N samples × log 2 C/N Neu. . As shown in Table 2, TDA-SNN achieves substantially higher per-neuron information content than STD-SNN across all three architectures. This increase is mainly due to the extreme reduction in neuron count, from N or N out in STD-SNN to 1 in TDA-SNN, while preserving useful pre- dictive performance. In the RC and MLP settings, delayed autapses and temporal integration allow a single neuron to encode information that would otherwise be distributed across an entire layer, leading to a dramatic increase in bits per neuron. In the convolution-like setting, the gain is smaller than in RC and MLP because performance is cur- rently more limited, but it still remains far above that of STD-SNN. These results highlight temporal multiplexing as an effective mechanism for increasing information den- sity at the single-neuron level. 4.7. Parallelization and Space–Time Trade-off We further examine whether the latency overhead observed in the extreme single-neuron setting can be alleviated by parallelization. Additional results on CIFAR-10 show that the high temporal overhead is specific to the most aggres- sive compression regime and that TDA-SNN provides a flexible space–time trade-off when more parallel neurons are introduced. As the number of parallel neurons increases, both training and inference time decrease substantially; rel- ative to STD-SNN on CIFAR-10, the training-time over- head drops from 178× to 46×, and with 512 parallel neu- rons the inference time reaches 0.92× of STD-SNN. These observations indicate that the single-neuron model should be viewed as a limit case for maximizing spatial compact- ness, while parallel multi-neuron realizations provide more practical operating points for deployment. From a mem- ory perspective, this compact formulation remains appeal- ing: on CIFAR-10, the SNN state memory is reduced from 8 KB in STD-SNN to 4 Bytes in the single-neuron TDA- SNN. Additional visualizations and scalability results are provided in Supplementary Sec. 15. 5. Discussion and Conclusion Our study shows that a single TDA-LIF neuron can re- alize core SNN architectures, including RC, MLP, and convolution-like layers, highlighting the computational po- tential of individual neurons. Inspired by temporal self- modulation in cerebellar Purkinje cells, TDA-SNN con- verts temporal feedback into compact spiking computa- tion. Experiments show competitive performance in RC and MLP settings, while convolutional results and paralleliza- tion analysis reveal a clear space–time trade-off. TDA- SNN also reduces neuron count and increases per-neuron information content. However, the current single-neuron convolution-like setting remains limited in efficiency and accuracy. Future work will extend this framework to multi- neuron settings and adaptive delay mechanisms. 6. Acknowledgements This work was supported in part by the National Key Re- search and Development Program of China under Grant 2023YFF1204200, in part by the Brain Science and Brain-like Intelligence Technology-National Science and Technology Major Project under Grant 2022ZD0208500, in part by the Sichuan Science and Technology Pro- gram under Grants 2024NSFTD0032, 2024NSFJQ0004, and DQ202410, in part by the Natural Science Founda- tion of Chongqing, China, under Grant CSTB2024NSCQ- MSX0627, in part by the Science and Technology Re- search Program of Chongqing Education Commission of China under Grant KJZD-K202401603, and in part by the China Postdoctoral Science Foundation under Grant 2024M763876. References [1] Arnon Amir, Brian Taba, David Berg, Timothy Melano, Jef- frey McKinstry, Carmelo Di Nolfo, Tapan Nayak, Alexander Andreopoulos, Guillaume Garreau, Marcela Mendoza, et al. A low power, fully event-based gesture recognition system. In CVPR, pages 7243–7252, 2017. 5, 1 [2] Veli Baysal and Ali Calim. Stochastic resonance in a single autapse–coupled neuron. Chaos, Solitons & Fractals, 175: 114059, 2023. 2 [3] David Beniaguev, Idan Segev, and Michael London. Single cortical neurons as deep artificial neural networks. Neuron, 109(17):2727–2739, 2021. 2 [4] William Bialek and Fred Rieke. Reliability and information transmission in spiking neurons. Trends in neurosciences, 15 (11):428–434, 1992. 2 [5] Antonio Biki ́ c, Corinna Kaspar, and Wolfram HP Pernice. The cost of unmodeled biological complexity in artificial neural networks. Patterns, 6(10), 2025. 1, 2 [6] Wuque Cai, Hongze Sun, Qianqian Liao, Jiayi He, Duo Chen, Dezhong Yao, and Daqing Guo. Robust Spatiotempo- ral Prototype Learning for Spiking Neural Networks. IEEE Transactions on Neural Networks and Learning Systems, 2025. 4, 1 [7] Andrea Calimera, Enrico Macii, and Massimo Poncino. The human brain project and neuromorphic computing. Func- tional neurology, 28(3):191, 2013. 2 [8] Sean E Cavanagh, Laurence T Hunt, and Steven W Kenner- ley. A diversity of intrinsic timescales underlie neural com- putations. Frontiers in neural circuits, 14:615626, 2020. 2 [9] Xinyi Chen, Jibin Wu, Chenxiang Ma, Yinsong Yan, Yu- jie Wu, and Kay Chen Tan.PMSN: A Parallel Multi- compartment Spiking Neuron for Multi-scale Temporal Pro- cessing. arXiv preprint arXiv:2408.14917, 2024. 1, 2 [10] Xiang Cheng, Yunzhe Hao, Jiaming Xu, and Bo Xu. LISNN: Improving spiking neural networks with lateral interactions for robust object recognition. In IJCAI, pages 1519–1525. Yokohama, 2020. 1, 2 [11] Benjamin Cramer, Yannik Stradmann, Johannes Schemmel, and Friedemann Zenke. The heidelberg spiking data sets for the systematic evaluation of spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems, 33 (7):2744–2757, 2020. 4, 1 [12] Manon Dampfhoffer, Thomas Mesquida, Alexandre Valen- tian, and Lorena Anghel. Backpropagation-Based Learning Techniques for Deep Spiking Neural Networks: A Survey. IEEE Transactions on Neural Networks and Learning Sys- tems, 35(9):11906–11921, 2023. 1 [13] Dominique Debanne, Yanis Inglebert, and Micha ̈ el Russier. Plasticity of intrinsic neuronal excitability. Current opinion in neurobiology, 54:73–82, 2019. 1 [14] Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE signal processing magazine, 29(6):141–142, 2012. 5, 1 [15] Wei Fang, Zhaofei Yu, Yanqi Chen, Timoth ́ e Masquelier, Tiejun Huang, and Yonghong Tian. Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural Networks. In ICCV, pages 2661–2671, 2021. 1, 2 [16] Wulfram Gerstner and Werner M Kistler. Spiking neuron models: Single neurons, populations, plasticity. Cambridge university press, 2002. 2 [17] Samanwoy Ghosh-Dastidar and Hojjat Adeli.SPIKING NEURAL NETWORKS. International journal of neural systems, 19(04):295–308, 2009. 1 [18] Albert Gidon, Timothy Adam Zolnik, Pawel Fidzinski, Fe- lix Bolduan, Athanasia Papoutsi, Panayiota Poirazi, Martin Holtkamp, Imre Vida, and Matthew Evan Larkum. Dendritic action potentials and computation in human layer 2/3 cortical neurons. Science, 367(6473):83–87, 2020. 2 [19] Yangfan Hu, Huajin Tang, and Gang Pan. Spiking deep residual networks. IEEE Transactions on Neural Networks and Learning Systems, 34(8):5200–5205, 2021. 1 [20] Xunbi A Ji and G ́ abor Orosz. Trainable delays in time de- lay neural networks for learning delayed dynamics. IEEE Transactions on Neural Networks and Learning Systems, 36 (3):5219–5229, 2024. 2 [21] Sander Koelstra, Christian Muhl, Mohammad Soleymani, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, and Ioannis Patras. DEAP: A database for emotion analysis; using physiological signals.IEEE transactions on affective computing, 3(1):18–31, 2011. 4, 1 [22] Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009. 5, 1 [23] Guoqi Li, Lei Deng, Huajin Tang, Gang Pan, Yonghong Tian, Kaushik Roy, and Wolfgang Maass. Brain-Inspired Computing: A Systematic Survey and Future Trends. Pro- ceedings of the IEEE, 2024. 1, 2 [24] Songting Li, Nan Liu, Xiaohui Zhang, David W McLaugh- lin, Douglas Zhou, and David Cai. Dendritic computations captured by an effective point neuron model. Proceedings of the National Academy of Sciences, 116(30):15244–15252, 2019. 1, 2 [25] Chongming Liu, Jingyang Ma, Songting Li, and Dou- glas Dongzhuo Zhou. Dendritic Integration Inspired Arti- ficial Neural Networks Capture Data Correlation. NeurIPS, 37:79325–79349, 2024. 1, 2 [26] Wolfgang Maass. Networks of spiking neurons: the third generation of neural network models. Neural networks, 10 (9):1659–1671, 1997. 1 [27] Warren S McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 5(4):115–133, 1943. 2 [28] Hans Meinhardt and Alfred Gierer. Pattern formation by lo- cal self-activation and lateral inhibition. Bioessays, 22(8): 753–760, 2000. 1 [29] Hao Peng, Pei Chen, Na Yang, Kazuyuki Aihara, Rui Liu, and Luonan Chen. One-core neuron deep learning for time series prediction. National Science Review, 12(2):nwae441, 2025. 1, 2 [30] Nicolas Perez-Nieves, Vincent CH Leung, Pier Luigi Dragotti, and Dan FM Goodman. Neural heterogeneity pro- motes robust learning. Nature communications, 12(1):5791, 2021. 1, 2 [31] Kaushik Roy, Akhilesh Jaiswal, and Priyadarshini Panda. Towards spike-based machine intelligence with neuromor- phic computing. Nature, 575(7784):607–617, 2019. 1 [32] Abhronil Sengupta, Yuting Ye, Robert Wang, Chiao Liu, and Kaushik Roy. Going deeper in spiking neural networks: Vgg and residual architectures. Frontiers in neuroscience, 13:95, 2019. 1 [33] H Sebastian Seung, Daniel D Lee, Ben Y Reis, and David W Tank. The autapse: a simple illustration of short-term ana- log memory storage by tuned synaptic feedback. Journal of computational neuroscience, 9(2):171–185, 2000. 2 [34] Masoumeh Shavikloo, Asghar Esmaeili, Alireza Valizadeh, and Mojtaba Madadi Asl. Synchronization of delayed cou- pled neurons with multiple synaptic connections. Cognitive Neurodynamics, 18(2):631–643, 2024. 2 [35] Nelson Spruston. Pyramidal neurons: dendritic structure and synaptic integration. Nature Reviews Neuroscience, 9 (3):206–221, 2008. 1, 2 [36] Florian Stelzer and Serhiy Yanchuk. Emulating complex net- works with a single delay differential equation. The Euro- pean Physical Journal Special Topics, 230(14):2865–2874, 2021. 1, 2 [37] Florian Stelzer, Andr ́ e R ̈ ohm, Raul Vicente, Ingo Fischer, and Serhiy Yanchuk. Deep neural networks using a sin- gle neuron: folded-in-time architecture using feedback- modulated delay loops. Nature communications, 12(1):5164, 2021. 1, 2 [38] Hongze Sun, Wuque Cai, Baoxin Yang, Yan Cui, Yang Xia, Dezhong Yao, and Daqing Guo. A Synapse-Threshold Syn- ergistic Learning Approach for Spiking Neural Networks. IEEE Transactions on Cognitive and Developmental Sys- tems, 16(2):544–558, 2023. 1, 2 [39] Alexander Waibel, Toshiyuki Hanazawa, Geoffrey Hinton, Kiyohiro Shikano, and Kevin J Lang. Phoneme recogni- tion using time-delay neural networks. In Backpropagation, pages 35–61. Psychology Press, 2013. 2 [40] Chunni Wang, Shengli Guo, Ying Xu, Jun Ma, Jun Tang, Faris Alzahrani, and Aatef Hobiny. Formation of Autapse Connected to Neuron and Its Biological Function. Complex- ity, 2017(1):5436737, 2017. 2 [41] Ziqing Wang, Yuetong Fang, Jiahang Cao, Qiang Zhang, Zhongrui Wang, and Renjing Xu. Masked spiking trans- former. In ICCV, pages 1761–1771, 2023. 1 [42] Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, and Luping Shi.Spatio-temporal backpropagation for training high- performance spiking neural networks. Frontiers in neuro- science, 12:331, 2018. 4, 1 [43] Yujie Wu, Lei Deng, Guoqi Li, Jun Zhu, Yuan Xie, and Lup- ing Shi. Direct training for spiking neural networks: Faster, larger, better. In AAAI, pages 1311–1318, 2019. 1, 4 [44] Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion- MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017. 5, 1 [45] Chenggui Yao, Zhiwei He, Tadashi Nakano, Yu Qian, and Jianwei Shuai. Inhibitory-autapse-enhanced signal transmis- sion in neural networks. Nonlinear Dynamics, 97(2):1425– 1437, 2019. 2 [46] Man Yao, Jiakui Hu, Guangshe Zhao, Yaoyuan Wang, Ziyang Zhang, Bo Xu, and Guoqi Li. Inherent Redundancy in Spiking Neural Networks. In ICCV, pages 16924–16934, 2023. 1 [47] Ergin Yilmaz, Mahmut Ozer, Veli Baysal, and Matja ˇ z Perc. Autapse-induced multiple coherence resonance in single neurons and neuronal networks. Scientific Reports, 6(1): 30914, 2016. 2 [48] Shimeng Yu. Neuro-inspired computing with emerging non- volatile memorys. Proceedings of the IEEE, 106(2):260– 285, 2018. 1 [49] Friedemann Zenke, Everton J Agnes, and Wulfram Gerstner. Diverse synaptic plasticity mechanisms orchestrated to form and retrieve memories in spiking neural networks. Nature communications, 6(1):6922, 2015. 1 [50] Wenrui Zhang and Peng Li. Spiking neural networks with laterally-inhibited self-recurrent units. In IJCNN, pages 1–8, 2021. 1, 2 [51] Rong Zhao, Zheyu Yang, Hao Zheng, Yujie Wu, Faqiang Liu, Zhenzhi Wu, Lukai Li, Feng Chen, Seng Song, Jun Zhu, et al. A framework for the general design and computation of hybrid neural networks. Nature communications, 13(1): 3427, 2022. 1 [52] Hanle Zheng, Zhong Zheng, Rui Hu, Bo Xiao, Yujie Wu, Fangwen Yu, Xue Liu, Guoqi Li, and Lei Deng. Temporal dendritic heterogeneity incorporated with spiking neural net- works for learning multi-timescale dynamics. Nature Com- munications, 15(1):277, 2024. 1, 2 [53] Chenlin Zhou, Han Zhang, Liutao Yu, Yumin Ye, Zhaokun Zhou, Liwei Huang, Zhengyu Ma, Xiaopeng Fan, Huihui Zhou, and Yonghong Tian. Direct training high-performance deep spiking neural networks: a review of theories and meth- ods. Frontiers in Neuroscience, 18:1383844, 2024. 1 Reconstructing Spiking Neural Networks Using a Single Neuron with Autapses Supplementary Material ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (a) Connection Weight (b) N node N node N laye r (a) (b) (c) (a) (b) (c) Figure 9. RC architecture of the TDA-SNN model. (a) Overall RC architecture constructed from the TDA-SNN model. (b) Autaptic weight matrix corresponding to the RC framework. 7. Weight Matrix of Autapses 7.1. Reservoir Computing Model with TDA-LIF Based on the equivalent formulation introduced in Sec. 3.2, external inputs are injected into each internal temporal node, and the spike activities of all nodes are recorded over the external data time window T [29, 36]. The resulting fir- ing patterns are then compared for recognition, as illustrated in Fig. 9(a). The autaptic weight matrix W a characterizes the temporal feedback structure among the unfolded tempo- ral nodes, where each element specifies an autaptic connec- tion from one node to another. Owing to the directed nature of temporal feedback, W a takes an upper-triangular form, with each parallel diagonal corresponding to a specific au- taptic delay d. As shown in Fig. 9(b), two such diagonals appear, representing delays of d = 3 (yellow) and d = 6 (blue). The resulting unfolded structure, derived from and functionally equivalent to the TDA-LIF neuron, forms the complete reservoir-computing (RC) architecture. 7.2. Multilayer Perceptron Model with TDA-LIF Similarly, an external layer is added before the first TDA- SNN layer to form a multilayer perceptron (MLP), and the spike activity of the final layer is recorded over the exter- nal data time window T [37]. The resulting spike patterns are then compared for recognition, as shown in Fig. 10(a). The autaptic connection matrix W a describes the tempo- ral feedback structure across temporal nodes: each element specifies an autaptic connection from a node in the previ- ous layer to one in the next layer, and each parallel di- agonal corresponds to a particular autaptic delay. As il- lustrated in Fig. 10(b), three diagonals appear, represent- ing delays of d = 2 (yellow), d = 5 (blue), and d = 7 (green). This delayed-feedback structure unfolds directly from the TDA-LIF neuron and is functionally equivalent to it, thereby forming the complete spiking MLP architecture. ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (a) Connection Weight (b) N node N node N laye r (a) (b) (c) (a) (b) (c) Figure 10. MLP architecture of the TDA-SNN model. (a) Overall MLP architecture constructed from the TDA-SNN model. (b) Au- taptic weight matrix corresponding to the MLP framework. 8. Spatiotemporal Prototype Learning To enable class-specific spatiotemporal representations within TDA-SNN, we construct a set of binary prototypes corresponding to the desired spiking patterns of each class. These prototypes serve as reference patterns against which the encoded network outputs are compared, facilitating classification and similarity measurement in the spatiotem- poral domain. Specifically, a set of binary prototypes K ∈ 0, 1 C×T is defined to represent target spiking patterns for C classes over T timesteps. Let f(x;θ) denote the encoded spa- tiotemporal representation of input x. To measure the sim- ilarity between the encoded representation and each class prototype, we compute d i (x) =−∥f(x;θ)− k i ∥ 2 2 ,(8) where d i (x) denotes the similarity score between f(x;θ) and the i-th prototype k i ∈ K. During inference, each input sample is assigned to the class of the closest prototype in the spatiotemporal feature space. The prototype-based classification procedure fol- lows the method described in [6, 42, 43]. 9. Datasets and Experimental Setup We evaluate TDA-SNN across multiple benchmarks under different structural mappings. DEAP [21] and SHD [11] are used to evaluate the reservoir-computing setting; MNIST [14], Fashion-MNIST (fMNIST) [44], and DVS Gesture [1] are used to validate the MLP setting; and DVS Gesture together with CIFAR10 [22] are further used to study the convolution-like setting. All models are implemented in PyTorch and trained on an NVIDIA A100 GPU using the Adam optimizer with Table 3. The hyperparameters for different datasets. Datasetbslr Tdt v th τ DEAP2000.00261001.00.5 SHD5120.01100140.30.3 MNIST5120.0024-0.30.5 fMNIST5120.0024-0.30.5 DVS Gesture1280.00281251.10.67 CIFAR102000.0024-1.00.5 CIFAR1005120.0024-1.00.5 ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (a) Connection Weight (b) N node N node N layer (a) (b) (c) (a) (b) (c) Figure 11. Evolution of connection weights in the RC architec- ture. (a) Random delay strategy. (b) Low-node configuration. (c) Maximum connection strategy. a cosine learning-rate decay schedule. Unless otherwise stated, training is conducted for 100 epochs. For RC and MLP experiments, results are averaged over 10 indepen- dent runs and reported as mean ± standard deviation; due to the higher computational cost of convolutional experi- ments, those results are averaged over 5 runs. To ensure a fair comparison, TDA-SNN and the corresponding standard SNN baselines use aligned training protocols and matched architectural scales for each setting. The complete set of hyperparameters is summarized in Table 3, including batch size (bs), learning rate (lr), time window T , slice width dt, membrane threshold v th , and membrane decay factor τ . In our experiments, the standard SNN evolves over an external time window, whereas TDA-SNN introduces an additional internal evolution time. Consequently, the mem- brane potential dynamics in TDA-SNN are decoupled from those of the standard SNN: the latter progresses along exter- nal timesteps, while the former evolves over internal tempo- ral nodes. In TDA-SNN, the external time window is treated as an independent dimension and does not directly partici- pate in the neuron dynamics. 10. Weight Matrices Under RC and MLP Ar- chitectures In the RC architecture, the TDA-SNN connection matrix is an upper-triangular square matrix. As shown in Fig. 11, the autaptic selection strategy and the number of nodes jointly determine its structure. The random delay (RD) ...... Spiking prototype Reservoir Time series data (a) Connection Weight (b) Inputs Outputs ...... Network Spiking prototype (a) Connection Weight (b) N node N node N layer (a) (b) (c) (a) (b) (c) Figure 12. Evolution of connection weights in the MLP architec- ture. (a) Random delay strategy. (b) Low-node configuration. (c) Maximum connection strategy. strategy produces a sparser matrix with generally longer- delay connections. Increasing the number of nodes expands the space of possible connections, whereas using fewer nodes naturally reduces the total number of available con- nections. In contrast, selecting autapses along the main di- agonal toward the upper-right corner under the maximum- connections (MC) strategy yields the densest connectivity while simultaneously minimizing autaptic latency. For the MLP architecture, the TDA-SNN connection ma- trix width and height correspond to the number of input and output nodes, respectively. Fig. 12 illustrates how autapse selection strategies shape the matrix. Using the MC strat- egy, autapses are selected along the diagonal and distributed to neighboring nodes, maximizing connectivity while min- imizing latency. In contrast, the RD strategy yields sparser, longer-delayed connections.This demonstrates that, in feedforward structures, structured delay assignment enables efficient temporal propagation across layers, analogous to the role of weight connectivity in conventional MLPs. 11. Significance Analysis of Experimental Re- sults The statistical significance results are summarized in Ta- ble 4. Overall, the effect of autaptic selection strategy depends on both the dataset and the delay budget. On DEAP, most comparisons are not statistically significant, suggesting that the task is relatively insensitive to the spe- cific delay-selection strategy. On SHD, significant differ- ences appear mainly at intermediate delays, indicating that strategy choice matters in specific temporal regimes. On MNIST, many pairwise comparisons are significant, espe- cially for MC-RD and for RD-T Inv. at moderate-to-large delays, showing that strategy selection can substantially af- fect performance. fMNIST and DVS Gesture also exhibit several significant differences, although the pattern varies with both the comparison pair and the delay value rather than following a uniform trend. These observations are con- sistent with the ablation results in showing that the impact of delay assignment is dataset-dependent and becomes par- ticularly evident in specific delay regimes. Table 4. Paired t-tests of Autaptic Selection Strategies Across Different Delays. * denotes p < 0.05. DatasetDelay1248163264 DEAPMC-RD0.790.660.650.730.510.051.0 SHDMC-RD0.53***0.400.811.0 MNIST MC-RD******* MC-T Inv.*0.06*0.650.970.95* RD-T Inv.0.350.88***** fMNIST MC-RD**0.59**** MC-T Inv.***0.540.600.53* RD-T Inv.0.530.680.06**** DVS MC-RD*0.580.300.06*** Gesture MC-T Inv.*0.260.080.170.920.48* RD-T Inv.0.260.570.440.72*** N layer Figure 13. Extending TDA-SNN to multi-layer deep models under the MLP architecture. 12. Impact of Layer Number in MLP Struc- tures TDA-SNN can be extended to deep architectures by in- creasing the number of internal nodes, allowing a single TDA-LIF neuron to emulate multiple layers through its tem- poral unfolding. As shown in Fig. 13, connections with the same color denote autapses sharing the same delay value, which remain consistent across layers. This preserves the structured temporal feedback while enabling the construc- tion of deeper computational models. 13. Additional Ablation Studies We additionally provide two supplementary ablation stud- ies: learning-rate sensitivity on MNIST and an ablation with different numbers of nodes on CIFAR100. These experi- ments further examine the robustness of TDA-SNN under different optimization and architectural settings. Learning rate 1e-61e-51e-41e-31e-21e-1 Acc. (%) 30 100 65 Figure 14. Learning rate sensitivity study for the MLP architecture on MNIST. 13.1. MNIST For MNIST, the learning-rate ablation is conducted using the hyperparameter setting in Table 3, varying the initial learning rate from 10 −6 to 10 −1 while keeping the remain- ing settings unchanged. The corresponding mean accura- cies with standard deviations are visualized in Fig. 15. Al- though the best performance is achieved at a learning rate of 10 −2 (95.86±0.26%), we use 10 −3 (95.55±0.20%) as the default setting because it remains close to the optimum while providing a more stable optimization process and stronger convergence robustness. This choice is particu- larly important for SNNs, whose discrete spikes and highly nonlinear dynamics make training sensitive to overly large learning rates, which can lead to oscillation and instability. 13.2. CIFAR100 For CIFAR100, we provide an additional ablation with dif- ferent numbers of nodes under a more challenging visual classification setting. Following the setting of Section 4.3.1, we use a two-layer network, and the hyperparameter config- uration is given in Table 3. The number of nodes is varied over16, 32, 64, 128, 256, 512 while the remaining archi- 16128512 Acc. (%) 13 23 18 Number of nodes Figure 15. Number of nodes ablation on CIFAR100 for STD-SNN and TDA-SNN. tectural and training settings are kept fixed. As shown in Fig. 16, STD-SNN improves steadily from 17.20±0.39% at 16 nodes to 21.72±0.32% at 512 nodes. TDA-SNN also improves substantially, from 14.08±0.39% at 16 nodes to a best accuracy of 20.70±0.21% at 256 nodes. When the number of nodes is further increased to 512, the accuracy decreases slightly to 19.82±0.31%, suggesting that simply increasing model size does not always yield additional gains on this challenging dataset. Nevertheless, TDA-SNN con- sistently improves as the number of nodes increases over a broad range and retains strong performance at medium- to-large scales, demonstrating that the proposed model re- mains effective even in a more difficult setting. 14. Implementation Details for Convolution- like Settings The network adopted a standard two-layer convolutional backbone, consisting of a convolutional layer (8 output channels, 7×7 kernel size, stride 1, padding 3), followed by 4×4 max-pooling and flattening, and a fully connected layer with an output feature size of 512 for classification. The TDA-SNN variant replaced standard synapses in con- volutional layers with time-delayed autapses configured us- ing all available autaptic delays. 15.ParallelizationandScalabilityin Convolution-like Settings To further analyze the temporal overhead of TDA-SNN, we study how parallelization changes the runtime characteris- tics in convolution-like settings. The extreme latency re- ported is specific to the single-neuron setting, which is de- signed to probe the limit of spatial compression. As shown in Fig. 14, increasing the number of parallel neurons re- duces both training and inference time. On CIFAR10, rel- ative to STD-SNN, the training-time overhead decreases from 178× to 46×, and with 512 neurons the inference time reaches 0.92× of STD-SNN. These results further support Time ratio (a) 1 200 100 12048 (b) 1 50 25 12048 Number of neuronsNumber of neurons 3232 Training timeInference time 17.21s 1.50s Figure 16. Training (a) and inference (b) time ratios of TDA-SNN relative to STD-SNN on CIFAR10. The yellow marker denotes STD-SNN (ratio = 1.0). The x-axis is logarithmic. that TDA-SNN offers a flexible space–time trade-off rather than a fixed high-latency operating point. We additionally evaluate a multi-neuron extension to ex- amine scalability when the compression constraint is re- laxed. Using 512 neurons for DVS Gesture and 256 neurons for CIFAR10 improves accuracy from 57.78% to 62.43% on DVS Gesture and from 37.64% to 40.42% on CIFAR10. Al- though a gap to STD-SNN remains, this upward trend con- firms that the proposed TDA mechanism scales effectively when moving beyond the extreme single-neuron regime.