Paper deep dive
LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices
Riadul Islam, Joey Mule, Dhandeep Challagundla, Shahmir Rizvi, Sean Carson, Rachit Saini
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/25/2026, 7:15:22 AM
Summary
The paper introduces LiteEvent-AE, a lightweight autoencoder architecture designed for event-based vision on low-latency, energy-constrained edge devices. The model compresses neuromorphic data using a compact convolutional encoder-decoder structure with adaptive event thresholding. Evaluated on the Smart Event Face Dataset (SEFD) and Event-Based Crossing Dataset (EBCD), it achieves competitive accuracy compared to YOLOv9 while using up to 35.6x fewer parameters. Deployed on Raspberry Pi 4B and NVIDIA Jetson Nano, it demonstrates significant energy efficiency, consuming approximately 726.3x less energy than YOLOv9 on Raspberry Pi 4B.
Entities (8)
Relation Signals (8)
LiteEvent-AE → comparedto → YOLOv9
confidence 95% · achieves competitive or superior accuracy compared to YOLOv9
LiteEvent-AE → consumeslessenergythan → YOLOv9
confidence 95% · corresponding to approximately 726.3× lower energy consumption than YOLOv9
LiteEvent-AE → deployedon → Raspberry Pi 4B
confidence 95% · the model is deployed on resource-constrained hardware: a Raspberry Pi 4B
LiteEvent-AE → deployedon → NVIDIA Jetson Nano
confidence 95% · the model is deployed on resource-constrained hardware: ... a NVIDIA Jetson Nano
LiteEvent-AE → evaluatedon → Smart Event Face Dataset
confidence 95% · Extensive evaluations on the Smart Event Face Dataset (SEFD)
LiteEvent-AE → evaluatedon → Event-Based Crossing Dataset
confidence 95% · Extensive evaluations on the ... Event-Based Crossing Dataset (EBCD)
LiteEvent-AE → hasfewerparametersthan → YOLOv9
confidence 95% · requiring up to 35.6× fewer parameters
LiteEvent-AE → uses → AutoEncoder
confidence 95% · This work presents a compact and configurable event-driven autoencoder
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual signals that reduce redundant data processing and support sustainable edge computing. However, the asynchronous and noise-prone nature of event streams creates challenges for conventional deep learning models, which are often too computationally intensive for low-power embedded platforms. This work presents a compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference. The architecture integrates lightweight convolutional encoding with robust performance under adaptive event thresholding and a minimal classifier head, enabling substantial reductions in computational cost without degrading recognition fidelity. Extensive evaluations on the Smart Event Face Dataset (SEFD) and Event-Based Crossing Dataset (EBCD) show that the proposed framework achieves competitive or superior accuracy compared to YOLOv9 while requiring up to 35.6$\times$ fewer parameters. To assess real-world sustainability, the model is deployed on resource-constrained hardware: a Raspberry Pi 4B and a NVIDIA Jetson Nano. On NVIDIA Jetson Nano, it delivers real-time throughput of 44.8 FPS. On a Raspberry Pi 4B CPU, the 50\% autoencoder classifier consumes 16.19 J for the evaluated inference workload, corresponding to approximately 726.3$\times$ lower energy consumption than YOLOv9 under the same evaluation protocol. These results demonstrate the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.
Tags
Links
- Source: https://arxiv.org/abs/2608.21764v1
- Canonical: https://arxiv.org/abs/2608.21764v1
Trouble viewing inline? Open PDF directly →
Full Text
61,071 characters extracted from source content.
Expand or collapse full text
LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices Riadul Islam Joey Mulé Dhandeep Challagundla Shahmir Rizvi Sean Carson Rachit Saini Thanks: R Islam, J. Mulé, D. Challagundla, S. Rizvi, S. Carson and R. Saini are with the Department of Computer Science and Electrical Engineering, University of Maryland, Baltimore County, MD 21250, USA e-mail: riaduli@umbc.edu. Thanks: This work was supported in part by the National Science Foundation (NSF) award number: 2138253, the Maryland Industrial Partnerships (MIPS) program under award number MIPS0012, and the UMBC Startup grant. Thanks: Copyright (c) 2025 IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending an email to pubs-permissions@ieee.org. Abstract Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual signals that reduce redundant data processing and support sustainable edge computing. However, the asynchronous and noise-prone nature of event streams creates challenges for conventional deep learning models, which are often too computationally intensive for low-power embedded platforms. This work presents a compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference. The architecture integrates lightweight convolutional encoding with robust performance under adaptive event thresholding and a minimal classifier head, enabling substantial reductions in computational cost without degrading recognition fidelity. Extensive evaluations on the Smart Event Face Dataset (SEFD) and Event-Based Crossing Dataset (EBCD) show that the proposed framework achieves competitive or superior accuracy compared to YOLOv9 while requiring up to 35.6× fewer parameters. To assess real-world sustainability, the model is deployed on resource-constrained hardware: a Raspberry Pi 4B and a NVIDIA Jetson Nano. On NVIDIA Jetson Nano, it delivers real-time throughput of 44.8 FPS. On a Raspberry Pi 4B CPU, the 50% autoencoder classifier consumes 16.19 J for the evaluated inference workload, corresponding to approximately 726.3× lower energy consumption than YOLOv9 under the same evaluation protocol. These results demonstrate the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments. Index Terms: Energy-efficient, low-latency autoencoder, event-based vision, high-speed sensing, object classification, edge computing. I Introduction Fig. 1: Comparison of average model accuracy, latency, and energy consumption on the SEFD [28] dataset, evaluated on a Raspberry Pi 4B CPU. Rapid visual perception is increasingly essential in emerging technologies that rely on instantaneous scene understanding and decision-making, such as autonomous vehicles, robotic systems, industrial monitoring, and scientific experimentation. Conventional frame-based imaging architectures, however, are fundamentally limited in high-speed scenarios, often producing blurred images, delayed responses, and redundant visual data when confronted with fast or complex motion. To overcome these bottlenecks, event-driven vision sensors have been developed to capture dynamic visual information in a more biologically inspired manner. Instead of recording full image frames at fixed time intervals, these sensors asynchronously register per-pixel brightness changes, enabling continuous tracking of motion with exceptionally low latency and high temporal fidelity [49, 44]. Despite their advantages, the unconventional output of event-driven sensors poses unique challenges for modern artificial intelligence frameworks. Unlike dense, regularly sampled image frames, event streams are sparse, asynchronous, and contain varying levels of noise, making direct application of conventional convolutional neural networks (CNNs) inefficient and often inaccurate. Moreover, real-time deployment in embedded or edge devices requires models that are both computationally lightweight and power-efficient without compromising detection precision [25]. These constraints motivate the development of specialized learning architectures capable of effectively encoding the spatiotemporal structure of event data. In this context, we introduce a lightweight event-based autoencoder that encodes and reconstructs asynchronous spike streams while preserving the essential motion dynamics embedded in the data and structural cues, thereby enabling robust, low-latency object classification in high-speed, resource-constrained environments. Figure 1 compares the accuracy, latency, and energy efficiency of state-of-the-art models with the proposed event-based autoencoder (AE) classifier architectures on the Smart Event Face Dataset (SEFD) [28]. Autoencoders’ unsupervised learning capability further enables adaptation to diverse event-driven scenarios, making them valuable for low-power, high-speed vision applications such as robotics, autonomous navigation, and industrial automation [62, 53, 13]. Furthermore, recent work has explored spatiotemporal transformers for streaming object detection [39]. Although effective on optical frame and flux-based representations, these approaches introduce substantial computational overhead and typically achieve limited mean average precision (mAP). Other efforts have investigated binary event history images (BEHIs) combined with lightweight CNN encoders for high-speed object detection [63]; however, these solutions incorporate auxiliary modules for uncertainty quantification to resolve motion-induced ambiguities. Despite these advances, developing scalable architectures that natively adapt to the asynchronous, data-driven nature of event vision remains an open research challenge. Although spiking neural networks (SNNs) [24, 33] are commonly used in event-based vision due to their time-dependent processing, their temporal dynamics introduce considerable complexity. Our approach instead leverages static event frames that can be processed with standard convolutions, introducing a static event autoencoder–based classification architecture to address these limitations. The proposed architecture jointly addresses event representation efficiency, latent feature compression, and edge inference complexity. The architecture specializes in event-based detections, enabling an adaptive threshold selection mechanism for energy-efficient computer vision tasks with event cameras. Hence, event sensors can adaptively select thresholds to dynamically adjust event detection sensitivity based on environmental conditions and signal characteristics, effectively reducing redundant event generation while preserving essential visual information. In particular, the main contributions of this research are: ∙ We introduce a configurable and lightweight autoencoder [26] architecture tailored for event-based object classification. ∙ We demonstrate the scalability and robustness of the architecture across multiple event activity thresholds, validating its adaptability to varying noise levels and scene dynamics. ∙ We evaluate the hardware efficiency of the proposed models on embedded platforms such as Raspberry Pi 4B and NVIDIA Jetson Nano, showcasing real-time performance under tight resource constraints. ∙ We compare with SOTA CNNs, such as YOLOv9 [61]. Our proposed event classifier achieves up to 35.6× fewer parameters, 22.07× higher frame rates (FPS), and over 700× lower energy consumption. I Background I-A Existing Event Vision Applications Researchers have applied event vision to many different tasks, including: automotive [9, 48, 44], surveillance [43], object classification [46], gesture recognition [2, 4], action detection [52], flow detection [71], and face detection [28]. With these efforts, both industry and academia generate large-scale event datasets, which can be categorized into event-sensor-based datasets [9, 48, 5] and simulation-based event-representation datasets from conventional frame-based sensor data [44, 28, 52]. However, most event vision datasets concentrate on single-threshold data. In this research, we concentrate on multi-threshold datasets [44, 28] to characterize a robust autoencoder model. Adaptive threshold selection dynamically modulates event detection sensitivity in response to varying environmental conditions and signal characteristics. This approach substantially reduces unnecessary event emissions without compromising the essential visual content needed for downstream processing. By optimizing the event representation pipeline, it enhances the computational efficiency and energy performance of event-based vision systems, thereby improving their suitability for power-constrained applications such as robotics, autonomous vehicles, and wearable technologies. I-B Existing Event Vision Object Classification Methods Event vision object detection models are characterized by dense or sparse event representations [72, 16]. Sparse representations seek to exploit temporal information in a continuous event stream. Dense event representations coalesce asynchronous event streams into structured, frame-like tensors that can be processed using standard vision techniques. Because these representations resemble conventional image frames, they can often leverage advances from traditional computer vision research [55]. Dense event representations benefit from straightforward hardware support and simpler integration with existing vision architectures. In contrast, sparse representations aim to exploit the fine-grained temporal structure of event streams [73]. Current state-of-the-art event-vision detectors commonly adopt transformer or hybrid transformer–CNN backbones, reflecting a trend toward frame-like processing even in event-based pipelines. A fundamental challenge, however, is that event cameras produce no output for static scenes; if both the camera and object remain stationary, the object effectively vanishes. To compensate for this lack of object permanence, some recent models incorporate recurrent modules to retain temporal context [55]. I-C Existing Evaluation Methods Evaluation of event vision object detection models parallels traditional computer vision evaluation methods. SOTA event-vision datasets include the NMNIST [47], DVSGesture [3], Gen1 [9], 1Mpx [48], SEFD [28], and event-based crossing dataset (EBCD) [44]. The performance of the models is typically characterized by MS COCO evaluation metrics [42], parameter count, and runtime. Recent studies have shown that autoencoder-based deep learning frameworks can effectively address data sparsity challenges in multimodal recommendation scenarios [51]. CNN-based autoencoders have likewise been employed for tasks such as medical image compression and diagnostic classification [11]. Several studies have investigated recurrent neural network (RNN)–based autoencoder architectures for optical-flow estimation in event-driven vision systems [59, 40, 12, 17, 67, 41, 35, 69]. By exploiting the inherent sequential modeling capabilities of RNNs, these approaches effectively encode temporal dependencies in event streams, making them particularly suitable for motion estimation and tracking applications. Predicting optical flow is a key function within event-driven vision systems, enabling accurate motion analysis in dynamic environments by estimating object motion from sparse, asynchronous event data. Prior studies have employed multi-encoder architectures for event-based fusion [14, 68]. Nevertheless, despite their demonstrated capabilities, these methodologies are often ill-suited for low-cost, high-speed embedded platforms. The substantial computational demands and elevated power requirements associated with RNN-based models and encoder-fusion pipelines present considerable obstacles for deployment in resource-limited environments, including edge devices and embedded systems. In contrast to energy-aware computing strategies proposed in prior work [20, 37, 27, 7, 29], the adoption of such architectures on low-power event sensors may incur excessive latency and reduced energy efficiency, thereby limiting their feasibility in practical settings where power constraints are paramount. To overcome these challenges, this study introduces a scalable architectural framework that leverages threshold-based event data to enable adaptive thresholding, thereby supporting energy-efficient visual processing in event-driven imaging systems. I Proposed Methodology Fig. 2: (a) The proposed event autoencoder uses a conventional convolution layer followed by ReLU and Max pooling layer on the encoder side, while the decoder uses a similar convolution layer followed by ReLU but an up-sampling layer; and (b) the autoencoder-based classifier concatenates the encoder with two fully connected layers. An autoencoder, in the context of computer vision, is a neural network architecture that learns compact, meaningful visual feature representations by reconstructing an input image from a compressed latent space. It uses convolution at its core and consists of two main components: an encoder that transforms high-dimensional image data into a lower-dimensional feature embedding, and a decoder that reconstructs the original image from this embedding. Through this reconstructive learning process, autoencoders implicitly capture essential visual structures such as shape, texture, intensity, and spatial patterns, making them useful for tasks like denoising [70], anomaly detection, dimensionality reduction [36], and unsupervised feature learning [57]. Their ability to discover latent representations without labeled data makes them a foundational tool in modern computer vision research and applications. Autoencoders are powerful and flexible neural networks capable of learning highly discriminative features from input data while maintaining a simple and efficient architecture. By compressing high-dimensional input data into a compact latent representation, autoencoders remove redundant information while preserving critical patterns necessary for downstream tasks [8]. The proposed architecture can be interpreted as a system that compresses and transmits visual data through a constrained communication channel. The encoder applies a nonlinear transformation that maps high-dimensional images to a low-entropy latent representation, performing lossy compression that captures salient structures such as edges, textures, and shapes, while discarding redundant pixel-level information. The decoder reconstructs the input from this compressed code, enforcing a trade-off between information bottleneck constraints and reconstruction fidelity. For the l-th convolutional layer in the encoder, the feature map is computed as h(l)=σ(W(l)∗h(l−1)+b(l)).h^(l)=σ (W^(l)*h^(l-1)+b^(l) ). (1) Here, ∗* denotes the convolution operation, W(l)W^(l) and b(l)b^(l) are the learnable convolution kernels and biases, and σ(⋅)σ(·) is a nonlinear activation function. The initial feature map is the input image h(0)=xh^(0)=x. The latent representation produced by the encoder is z=h(L)z=h^(L). The decoder reconstructs the input by progressively upsampling the latent representation using transposed convolutions. For the l-th decoder layer, the operation is given by h^(l−1)=σ(W~(l)∗h^(l)+b~(l)). h^(l-1)=σ ( W^(l)* h^(l)+ b^(l) ). (2) Thus, W~(l) W^(l) and b~(l) b^(l) are the decoder’s learnable parameters, and σ(⋅)σ(·) is an activation function. The final reconstruction is obtained as x^=h^(0) x= h^(0). The proposed event autoencoder is a compact convolutional encoder–decoder designed specifically for dense event representations, as shown in Figure 2(a). For real-time streaming pipelines, this approach also requires an event accumulation stage (event-to-frame conversion), which is natively supported by event-sensors and enables compatibility with existing GPU-accelerated CNN frameworks. The encoder progressively condenses each event frame into a low-dimensional latent representation using a stack of convolutional blocks with Rectified Linear Unit (ReLU) activation and spatial downsampling. The ReLU activation function plays a critical role in stabilizing and enhancing feature extraction. First, ReLU introduces nonlinearity while maintaining computational efficiency, allowing neural networks to model the complex spatiotemporal patterns inherent in asynchronous event streams. Because event data are sparse and often exhibit high temporal resolution, ReLU’s simple thresholding (setting negative activations to zero) helps suppress noise, improving robustness to spurious events. Additionally, ReLU promotes sparse activations, which aligns well with the inherently sparse nature of event data and reduces unnecessary computational overhead. This sparsity not only improves training efficiency but also enhances the network’s ability to highlight meaningful changes in the scene, ultimately supporting more discriminative and energy-efficient processing in event-based vision tasks. The decoder mirrors this process by upsampling and refining the latent features to reconstruct the original event frame. By enforcing a reconstruction objective during training, the network learns to capture salient spatial structure and activity patterns characteristic of event data while discarding high-frequency noise and threshold-induced artifacts. This results in a latent space well aligned with downstream classification tasks. To convert the autoencoder into a classifier, shown in Figure 2(b), the pretrained encoder is repurposed as a fixed feature extractor, and two fully connected layers are appended to the latent output. This separation cleanly decouples representation learning from task-specific discrimination: the encoder provides a stable event-driven embedding, while the lightweight classifier learns decision boundaries without modifying the trained spatial transform. This structure is particularly effective for resource-constrained platforms, as it minimizes parameter count and enables real-time performance on devices such as Raspberry Pi 4B and Jetson Nano. IV Experimental Results and Discussion Experimental Setup For this study, we employed the SEFD [28], a publicly available event-based facial dataset derived from the Aff-Wild collection [38] for training and testing. The Aff-Wild dataset contains 298 video sequences totaling more than 1.22 million frames, offering extensive variability in facial expressions associated with different emotional states and rapid affective transitions. It further encompasses a wide range of head poses, illumination conditions, and facial occlusions, thereby providing a comprehensive benchmark for robust facial analysis [38]. The SEFD dataset contains one class, persons, specifically facial regions. We also used another dense event vision dataset, the event-based crossing dataset (EBCD) [44], which consists of an automotive and pedestrian crossing task derived from the frame-captured, NTU Pedestrian Dataset [45]. The EBCD dataset, however, samples a more extensive range of thresholds. The EBCD dataset contains two classes: pedestrians and vehicles. The training uses conventional binary cross-entropy (BCE) loss. The average BCE loss for N number of samples can be calculated using ℒBCE=−1N∑i=1N[yilog(y^i)+(1−yi)log(1−y^i)]LBCE=- 1N _i=1^N [y_i ( y_i)+(1-y_i) (1- y_i) ], where yiy_i and y^i y_i are the actual label and predicted probability, respectively. For training and testing on the proposed and baseline methods, we used standard N training and testing procedures using the same threshold-generated event frames, derived from source benchmark datasets (i.e., SEFD [28] and EBCD [44]). For both the autoencoder and the autoencoder-based classifier, the following hyperparameters were used during training: batch size = 32, number of epochs = 100, and learning rate = 0.001. Fig. 3: Comprehensive edge-computing testbed integrating Raspberry Pi 4B and NVIDIA Jetson Nano boards with a Raspberry Pi camera and USB-based power analysis tools, enabling systematic evaluation of computational load, energy consumption, and real-time performance. To evaluate the performance of the proposed event autoencoder-based adequately, we train a suite of SOTA neural network architectures using both SEFD [28] and EBCD [44] datasets, including the You Only Look Once (YOLO) family of models: YOLOv4 [6], YOLOv7 [60], and YOLOv9 [61], as well as EfficientDet-b0 [54], MobileNet-v1 [18], and YuNet [65]. These models are characterized by their accurate detections, in-depth calculations, lightweight nature, and high-speed inference in visual detection applications, following the guidelines provided by their original articles. In this study, we evaluated five empirical, event-generation thresholds (Th=4,8,12,16,20T_h=4,8,12,16,20), where each threshold represents the minimum pixel-intensity variation required between consecutive frames or with respect to an initial frame. Each threshold is derived from the evaluated datasets. The corresponding results of our experiments are quantitatively analyzed in Table I and Table I. All computational experiments were conducted on a workstation equipped with an Intel Xeon(R) 20-core processor, 32 GB RAM, and an NVIDIA T1000 GPU (4 GB), operating on Ubuntu 20.04.6 LTS. To demonstrate the suitability of the proposed model for resource-constrained environments, the event autoencoder was deployed on both Raspberry Pi 4B and NVIDIA Jetson Nano devices. The embedded hardware configuration and setup are illustrated in Figure 3. TABLE I: Performance comparison on the SEFD [28] test dataset under temporal thresholds Th∈4,8,12,16T_h∈\4,8,12,16\. Our autoencoder classifier achieves competitive performance on frame-level data with orders of magnitude fewer parameters than heavyweight baselines. Model Params ThT_h Accuracy Precision Recall F1-Score FLOPs YOLOv4 [6] 60.3M 4 83.81 94.00 89.00 91.00 59.56B 8 80.33 89.00 89.00 89.00 12 79.98 91.00 86.00 89.00 16 74.49 93.00 79.00 85.00 YOLOv7 [60] 25.2M 4 86.71 94.00 92.00 93.00 43.61B 8 82.95 92.00 90.00 91.00 12 83.44 95.00 87.00 91.00 16 81.88 95.00 85.00 90.00 YOLOv9 [61] 60.5M 4 97.69 90.20 93.90 92.01 263.9B 8 98.14 92.30 96.30 94.26 12 98.25 96.20 94.00 95.09 16 97.28 95.80 91.60 93.65 EfficientDet-b0 [54] 3.9M 4 94.55 97.68 96.72 97.21 4.87B 8 93.62 97.13 96.28 96.71 12 97.34 97.34 95.96 96.64 16 97.30 97.30 94.54 95.91 MobileNet-v1 [18] 6.05M 4 91.64 96.65 94.64 95.64 43.41B 8 89.44 92.64 96.28 94.43 12 92.11 96.36 95.41 95.88 16 89.11 96.01 98.14 97.06 YuNet [65] 72.3K 4 85.46 93.00 85.00 89.00 486M 8 84.70 88.00 85.00 87.00 12 74.64 93.00 75.00 83.00 16 84.15 92.00 84.00 88.00 Autoencoder Classifier 100% (Ours) 1.7M 4 91.82 89.07 89.07 89.07 6.50B 8 93.01 88.27 93.77 90.94 12 92.15 86.70 93.33 89.89 16 90.02 82.80 92.57 87.41 Autoencoder Classifier 50% (Ours) 458K 4 87.19 89.19 84.63 86.85 1.66B 8 88.96 89.29 88.54 88.91 12 89.79 89.11 90.68 89.88 16 87.99 86.38 90.21 88.25 TABLE I: Performance comparison on the EBCD [44] test dataset across temporal thresholds Th∈12,16,20T_h∈\12,16,20\. The proposed autoencoder classifier consistently outperforms prior models on frame-level data while using significantly fewer parameters. Model Parameters ThT_h Accuracy Precision Recall F1-Score FLOPs YOLOv4 [6] 60.3M 12 88.00 96.00 80.00 87.00 59.56B 16 87.50 94.00 81.00 87.00 20 89.00 94.00 84.00 87.00 YOLOv7 [60] 25.2M 12 77.99 97.00 78.00 87.00 43.61B 16 77.99 97.00 78.00 86.00 20 79.35 97.00 79.00 87.00 YOLOv9 [61] 60.5M 12 98.75 97.80 93.50 95.60 263.9B 16 98.22 96.30 95.20 95.75 20 98.75 96.50 92.60 94.51 EfficientDet-b0 [54] 3.9M 12 58.22 98.44 58.22 73.17 43.41B 16 57.03 99.21 57.03 72.43 20 46.40 100.00 46.40 63.90 MobileNet-v1 [18] 6.05M 12 26.12 98.30 26.20 41.40 43.41B 16 74.96 95.60 77.60 85.50 20 35.96 94.10 36.80 52.90 YuNet [65] 72.3K 12 84.16 93.00 84.00 88.00 486M 16 82.91 89.00 83.00 86.00 20 82.45 92.00 82.00 87.00 Autoencoder Classifier 100% (Ours) 1.7M 12 92.16 87.32 93.21 89.56 13.87B 16 93.47 88.04 92.77 90.09 20 94.38 88.69 94.45 91.13 Autoencoder Classifier 50% (Ours) 458K 12 88.79 82.47 87.36 84.84 1.66B 16 90.11 83.24 88.02 85.56 20 91.04 84.19 89.24 86.64 Results and Discussion Compared to the existing SOTA models trained on the SEFD dataset, the proposed autoencoder classifier achieves higher accuracy than YOLOv4, YOLOv7, MobileNets-v, and YuNet. The YOLOv9 architecture achieves the best accuracy across both the SEFD and EBCD datasets compared to the competing models. The average accuracies of YOLOv9 across different thresholds are 97.84% and 98.57% for the SEFD and EBCD datasets, respectively. The proposed event autoencoder-based classifier has 2.3×, 3.6×, and 35.6× fewer parameters than EfficientDet-b0, MobileNet-v1, and YOLOv9 models, respectively. Likewise, the proposed autoencoder classifier also expresses very high accuracy on the EBCD dataset. In Table I, the model achieves an average accuracy of 93.33%93.33\% across the three thresholds: (Th=12,16,20T_h=12,16,20), of the crossing dataset. The proposed model outshines YOLOv4, YOLOv7, EfficientDet-b0, MobileNet-v1, and YuNet, only being behind YOLOv9 in accuracy by ≈4-6%≈ 4-6\%. For assessment, the event autoencoder was trained and tested on two categories of data: SEFD event samples representing positive instances and non-face samples serving as negative instances. The training dynamics of the proposed event autoencoder reveal strong convergence behavior, yielding an average training loss of 0.1593, as depicted in Figure 4(a). After reducing the network capacity to 50% of its original number of filters, the model attains an even lower average training loss of 0.0884, presented in Figure 4(b). This outcome highlights the model’s scalability, demonstrating that a substantial reduction in filter count not only preserves but can enhance learning efficiency. Notably, compressing the architecture from full capacity to 50% filters results in a fourfold reduction in overall model size, underscoring the resilience of the autoencoder design in achieving significant computational savings while sustaining high-quality feature learning and reconstruction performance. Fig. 4: (a) Training curve of the full-capacity event autoencoder, yielding an average loss close to fifteen hundredths. (b) Performance of the reduced 50%–filter configuration, which achieves a lower average training loss close to one-tenth, illustrating the model’s resilience under substantial architectural compression. Accurate reconstruction is typically measured using a pixel-wise comparison, BCE analysis, or Structural Similarity Index Measure (SSIM) [64] calculation, or a combination of all. Strong reconstruction fidelity indicates that the model successfully retains the essential structural and statistical properties of the initial data, thereby enhancing the quality of extracted features for subsequent tasks such as classification, segmentation, and denoising. In the context of outlier identification, improved reconstruction performance enables the system to more effectively differentiate between normal data patterns and aberrant inputs, thereby reducing error rates and increasing reliability. Fig. 5: The reconstruction outputs of the proposed models exhibit excellent accuracy, comparable to their model sizes. (a & d) show the ground-truth (GT), pre-processed image, at Th=8T_h=8 and Th=20T_h=20 of the SEFD and EBCD datasets, respectively. (b & e) represent the reconstructed outputs from the 100% autoencoder with SSIM=0.939SSIM=0.939 and SSIM=0.894SSIM=0.894. (c & f) are the reconstructed outputs from the 50% autoencoder with SSIM=0.915SSIM=0.915 and SSIM=0.866SSIM=0.866. Fig. 6: AUROC performance of the proposed autoencoder-based classifier across SEFD threshold settings. All configurations achieve AUC values above 90%, with peak performance at Th=8T_h=8 Figure 5 illustrates the reconstruction capability of the proposed event-driven autoencoder. The input event frame provided to the network is depicted in Figure 5(a). The full-capacity configuration (100% filter model) achieves an SSIM of 0.939, as shown in Figure 5(b). By contrast, the reduced configuration utilizing only 50% of the filters attains a slightly lower SSIM of 0.915, while achieving a 4× reduction in model size, as presented in Figure 5(c). These results highlight the effectiveness of the proposed design in maintaining high-quality reconstruction even under significant architectural compression. Likewise, with the EBCD dataset, the reconstructed SSIM of the 50% autoencoder decreases by only 0.028 from that of the full-sized model in Figure 5(f). Fig. 7: Runtime evaluation of the proposed architectures compared with contemporary work on Raspberry Pi 4B and NVIDIA Jetson Nano (4GB). The results show reduced execution time and higher FPS, with the autoencoder classifier achieving 22.1× the FPS of YOLOv9. Although YuNet is fastest, the proposed 50% filter classifier is 6.3% more accurate. Fig. 8: Energy consumption and runtime during inference on a Raspberry Pi 4B. The proposed 50% Autoencoder Classifier and 50% Autoencoder require only 16.19 J and 30.48 J to process 100 images, achieving more than a 2×2× energy reduction over their full-size counterparts (36.37 J and 64.46 J). Their runtimes are similarly low at 4.8 s and 8.6 s. In contrast, lightweight baselines such as MobileNet-v1 and YOLOv7 consume 3234.95 J and 8440.31 J (476.2 s and 1250.0 s), while the most demanding model, YOLOv9, reaches 11,759.39 J and 2000.0 s—over 700×700× the energy usage of the 50% Autoencoder Classifier. Performance assessment of binary classification systems frequently incorporates the Receiver Operating Characteristic (ROC) curve, a foundational analytical tool in machine learning and anomaly detection [32]. The ROC curve depicts the relationship between the sensitivity (true positive proportion) and the false positive proportion, thereby illustrating how effectively a model separates positive samples from negative ones across varying decision thresholds [10]. The Area Under the ROC Curve (AUROC) provides a scalar summary of this relationship; an AUROC value of 1.0 indicates ideal discrimination capability, whereas a value of 0.5 signifies performance no better than random guessing [15]. Owing to its threshold-invariant nature, AUROC is routinely employed in applications such as outlier identification to benchmark and refine classifier performance. Figure 6 summarizes the AUROC performance of the proposed autoencoder-based classifier across all SEFD event-generation thresholds, ThT_h. At the lowest threshold (Th=4T_h=4), where event activity is highest, the classifier maintains strong separation between event and non-event samples despite the increased pixel-level density. Performance peaks at Th=8T_h=8, yielding the highest AUROC and indicating optimal discriminative capability. Although AUROC values decrease slightly at higher thresholds (Th=12T_h=12 and Th=16T_h=16) due to sparser event representations, the classifier retains consistently high performance across all settings, demonstrating robust detection capability under varying activity levels. Performance on Embedded Platforms When transferring models to the embedded devices, all were evaluated using 32FP (32-bit floating point). The results of the embedded platform analysis can be depicted in Figure 7. On the Raspberry Pi 4B platform, the 50%–filter configuration of the proposed event autoencoder delivers substantial speed gains, achieving a 2.71×2.71× increase in FPS over the full model and a 55.24×55.24× improvement relative to MobileNet-v1. Using the same hardware platform, the 50% event-autoencoder classifier demonstrates similarly strong performance, operating more than double the FPS of its 100% counterpart, and nearly 100×100× the throughput of MobileNet-v1. Furthermore, the classifier equipped with the 50% filter reduction outperforms EfficientDet-b0 by a factor of 13.44×13.44× in terms of inference speed. The 50% autoencoder classifier also achieves 1.41×1.41× the FPS of the tiny YuNet model. Considering the large size of the YOLO-family models, the proposed model exhibits up to 414×414× more FPS. As mentioned, the evaluation was extended to the NVIDIA Jetson Nano platform to further assess the efficiency of the proposed architecture on an embedded GPU. The event autoencoder configured with a 50% filter reduction achieves significant throughput advantages, operating at nearly 9× the FPS of YOLOv4 [6] and about 37× the FPS of MobileNet-v1 [18]. The corresponding event-based classifier demonstrates even greater performance gains, delivering 21.4 times higher FPS than YOLOv4 and nearly 88× higher FPS than MobileNet-v1, confirming the model’s suitability for real-time processing on low-power embedded hardware. However, it can be noted that with the increase in processing power, YuNet [65] exhibits the fastest runtimes. Energy Consumption The importance of energy efficient computations are well studied in the literature [34, 22, 56, 23, 1, 21, 66, 30, 58, 19, 50, 31]. Hence, we evaluated the energy consumption of all tested models by measuring their power usage during sequential inference on 100 images. To ensure fair energy measurement, the proposed model and all SOTA models use the same 100 test images. We ran the same experiments 10 times and averaged the results for each data point. Figure 8 reports the total energy usage (in joules) and runtime (in seconds) for each model when deployed on a Raspberry Pi 4B. Measurements were collected using an inline USB power meter that sampled real-time current (I) and voltage (V) throughout execution. The resultant energy values are subtracted from the idle power consumption, recorded at 0.2929525 watts (joules per second). To reduce noise and isolate stable behavior, the recorded power trace was partitioned into three operational phases: warmup, steady-state inference, and cooldown, with respective durations w1,w2,w3w_1,w_2,w_3. For each phase, average current and voltage were computed independently as I¯k I_k and V¯k V_k for k∈1,2,3k∈\1,2,3\. Phase-wise average power was then obtained as P¯k=V¯k⋅I¯k. P_k= V_k· I_k. (3) The overall average power was estimated through a weighted temporal mean, Avg(P)=w1P¯1+w2P¯2+w3P¯3w1+w2+w3.Avg(P)= w_1\, P_1+w_2\, P_2+w_3\, P_3w_1+w_2+w_3. (4) Here we approximate the time integral of instantaneous power over the entire execution window. Total energy consumption was then computed as E=Avg(P)×t.E=Avg(P)× t. (5) The total inference time across all 100 processed images is represented as t. This weighted approach suppresses transient fluctuations, and emphasizes the dominant steady-state inference activity. Across all evaluated models, our proposed architectures (bolded in Figure 8) exhibit substantial gains in energy efficiency. The 50% autoencoder classifier and 50% autoencoder consumed only 16.19 J and 30.48 J, respectively—representing dramatic reductions relative to commonly deployed SOTA detectors. Specifically, the 50% autoencoder classifier is approximately 726.3×726.3×, 199.8×199.8×, 22.6×22.6×, and 1.95×1.95× more energy efficient than YOLOv9, MobileNet-v1, EfficientDet-b0, and YuNet, respectively. Although highly accurate in object classification, the YOLO-family models incur significantly larger energy costs due to their deep and computationally intensive architectures. Encoder size ablation analysis To further evaluate the scalability of the proposed architecture, we examine the effect of further reducing the encoder capacity. In addition to the 50% model, we evaluate another reduced variant obtained by progressively removing convolutional filters, namely, a 25% configuration. As the encoder size decreases, the number of parameters and FLOPs are substantially reduced while maintaining competitive recognition performance. In Table I, the 25% model reduces the architecture to only 132K parameters and 429M FLOPs, while still achieving strong classification performance across the evaluated threshold settings. This behavior further highlights the robustness of the learned representation and suggests that the computational savings primarily stem from the lightweight encoder design. TABLE I: Ablation on the analysis of a further reduced autoencoder classifier, evaluated on the SEFD [28] dataset across temporal thresholds Th∈4,8,12,16T_h∈\4,8,12,16\. Model Parameters ThT_h Accuracy Precision Recall F1-Score FLOPs Autoencoder Classifier 25% (Ours) 132K 4 88.79 86.48 82.97 84.69 429M 8 86.30 86.33 86.25 86.29 12 86.46 85.11 88.39 86.71 16 84.56 82.32 88.02 85.07 For additional architectural context, reduced detector variants remain comparatively heavier. For example, when exploring the yolo-family models: YOLOv4-tiny [6] uses 5.6M parameters and 5.80B FLOPs, YOLOv7-tiny [60] uses 3.9M parameters and 6.79B FLOPs, and YOLOv9-tiny [61] uses 2.6M parameters and 10.70B FLOPs. By comparison, even the full proposed model remains compact, while the reduced 25% configuration further lowers complexity drastically. V Conclusion The architecture presented herein delivers a deployable event-driven autoencoder with an integrated classifier optimized for real-time execution on power-constrained vision hardware. The proposed system demonstrates robust compression and reconstruction of event streams, preserving critical spatiotemporal features while minimizing computational and energy overhead. Evaluations on the SEFD [28] and EBCD [44] datasets demonstrate classification accuracies of 93% across multiple thresholds, outperforming YOLOv4 [6], YOLOv7 [60], MobileNet-v1 [18], and YuNet [65], and maintaining competitive accuracy within 4–6% of YOLOv9 [61], despite using 35.6× fewer parameters. The reconstruction pipeline achieves 99.97% accuracy with the full model and retains 93.54% accuracy even when scaled to 50% of the filters, confirming the model’s resilience under aggressive parameter reduction. Implementation on embedded CPU (i.e., Raspberry Pi 4B) and GPU (i.e., NVIDIA Jetson Nano) platforms further validates the effectiveness of the proposed solution, achieving up to 87.84× higher FPS than MobileNet-v1 [18] and over 700× lower energy consumption than YOLOv9 [61]. The system delivers 20.7 FPS on Raspberry Pi and 44.8 FPS on Jetson Nano with the 50% classifier model, while consuming only 16.19 J per 100 inferences. These results highlight the suitability of the proposed architecture for high-speed, low-power edge intelligence in consumer electronics. Moreover, this work provides a strong foundation for extending the proposed architecture to more complex multi-object and multi-class event datasets. Building on this basis, future work will explore richer event-based perception scenarios for real-time detection. Acknowledgment We thank S.R.S.K. Tummala of UMBC for his early work on the early architecture and results. We also extend our appreciation to R. Kankipati, R. Robucci (UMBC), and C. Howard (Oculi.ai) for their insightful assistance with analysis and discussion. References [1] S. Ambrogio, P. Narayanan, A. Okazaki, A. Fasoli, C. Mackin, K. Hosokawa, A. Nomura, T. Yasuda, A. Chen, A. Friz, et al. (2023) An analog-AI chip for energy-efficient speech recognition and transcription. Nature 620 (7975), p. 768–775. Cited by: §IV. [2] A. Amir, B. Taba, D. Berg, T. Melano, J. McKinstry, C. Di Nolfo, T. Nayak, A. Andreopoulos, G. Garreau, M. Mendoza, J. Kusnitz, M. Debole, S. Esser, T. Delbruck, M. Flickner, and D. Modha (2017) A Low Power, Fully Event-Based Gesture Recognition System. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 7388–7397. External Links: Document Cited by: §I-A. [3] A. Amir, B. Taba, D. Berg, T. Melano, J. McKinstry, C. Di Nolfo, T. Nayak, A. Andreopoulos, G. Garreau, M. Mendoza, J. Kusnitz, M. Debole, S. Esser, T. Delbruck, M. Flickner, and D. Modha (2017) A Low Power, Fully Event-Based Gesture Recognition System. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 7388–7397. External Links: Document Cited by: §I-C. [4] Y. Bi, A. Chadha, A. Abbas, E. Bourtsoulatze, and Y. Andreopoulos (2019) Graph-based object classification for neuromorphic vision sensing. In Proceedings of the IEEE/CVF international conference on computer vision, p. 491–501. Cited by: §I-A. [5] J. Binas, D. Neil, S. Liu, and T. Delbruck (2017) D17: End-to-end DAVIS driving dataset. arXiv preprint arXiv:1711.01458. Cited by: §I-A. [6] A. Bochkovskiy, C. Wang, and H. M. Liao (2020) Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, p. 1–17. External Links: Document Cited by: §IV, §IV, §IV, TABLE I, TABLE I, §V. [7] D. Challagundla, I. Bezzam, and R. Islam (2025) ArXrCiM: Architectural Exploration of Application-Specific Resonant SRAM Compute-in-Memory. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 33 (1), p. 179–192. External Links: Document Cited by: §I-C. [8] A. Chen, K. Zhang, R. Zhang, Z. Wang, Y. Lu, Y. Guo, and S. Zhang (2023) PiMAE: point cloud and image interactive masked autoencoders for 3D object detection. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 5291–5301. External Links: Document Cited by: §I. [9] P. de Tournemire, D. Nitti, E. Perot, D. Migliore, and A. Sironi (2020) A Large Scale Event-based Detection Dataset for Automotive. External Links: 2001.08499, Link Cited by: §I-A, §I-C. [10] T. Fawcett (2006) An introduction to ROC analysis. Pattern Recognition Letters 27 (8), p. 861–874. Note: ROC Analysis in Pattern Recognition External Links: ISSN 0167-8655, Document, Link Cited by: §IV. [11] A. Fettah, R. Menassel, A. Gattal, and A. Gattal (2024) Convolutional Autoencoder-Based medical image compression using a novel annotated medical X-ray imaging dataset. Biomedical Signal Processing and Control 94, p. 106238. Cited by: §I-C. [12] D. Gehrig, M. Ruegg, M. Gehrig, J. Hidalgo-Carrio, and D. Scaramuzza (2021) Combining events and frames using recurrent asynchronous multimodal networks for monocular depth prediction. IEEE Robotics and Automation Letters 6 (2), p. 2822–2829. Cited by: §I-C. [13] A. Gruel, J. Martinet, T. Serrano-Gotarredona, et al. (2022) Event data downscaling for embedded computer vision. In International Conference on Computer Vision Systems, p. 1–6. Cited by: §I. [14] J. Han, C. Zhou, P. Duan, Y. Tang, C. Xu, C. Xu, T. Huang, and B. Shi (2020) Neuromorphic camera guided high dynamic range imaging. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 1727–1736. External Links: Document Cited by: §I-C. [15] J. A. Hanley and B. J. McNeil (1982) The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143 (1), p. 29–36. External Links: Document Cited by: §IV. [16] T. H. Henri Rebecq and D. Scaramuzza (2017) Real-time Visual-Inertial Odometry for Event Cameras using Keyframe-based Nonlinear Optimization. In Proceedings of the British Machine Vision Conference (BMVC), G. B. Tae-Kyun Kim and K. Mikolajczyk (Eds.), p. 16.1–16.12. External Links: Document, ISBN 1-901725-60-X, Link Cited by: §I-B. [17] J. Hidalgo-Carrio, D. Gehrig, and D. Scaramuzza (2020) Learning monocular dense depth from events. In International Conference on 3D Vision (3DV), p. 534–542. Cited by: §I-C. [18] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam (2017) MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. External Links: 1704.04861 Cited by: §IV, §IV, TABLE I, TABLE I, §V. [19] R. Islam, S.E. Esmaeili, and T. Islam (2011) A high performance clock precharge SEU hardened flip-flop. In 2011 9th IEEE International Conference on ASIC, Vol. , p. 574–577. External Links: Document Cited by: §IV. [20] R. Islam, H. A. Fahmy, P. Y. Lin, and M. R. Guthaus (2018) DCMCS: Highly robust low-power differential current-mode clocking and synthesis. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 26 (10), p. 2108–2117. Cited by: §I-C. [21] R. Islam and M. R. Guthaus (2014) Current-mode clock distribution. In IEEE International Symposium on Circuits and Systems (ISCAS), Vol. , p. 1203–1206. External Links: Document Cited by: §IV. [22] R. Islam and M. R. Guthaus (2017) CMCS: current-mode clock synthesis. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 25 (3), p. 1054–1062. External Links: Document Cited by: §IV. [23] R. Islam and M. R. Guthaus (2019) HCDN: Hybrid-Mode Clock Distribution Networks. IEEE Transactions on Circuits and Systems I: Regular Papers 66 (1), p. 251–262. External Links: Document Cited by: §IV. [24] R. Islam, P. Majurski, J. Kwon, A. Sharma, and S. R. S. K. Tummala (2024) Benchmarking Artificial Neural Network Architectures for High-Performance Spiking Neural Networks. Sensors 24 (4), p. 1329. Cited by: §I. [25] R. Islam, P. Majurski, J. Kwon, and S. R. S. K. Tummala (2023) Exploring High-Level Neural Networks Architectures for Efficient Spiking Neural Networks Implementation. In 2023 3rd International Conference on Robotics, Electrical and Signal Processing Techniques (ICREST), Vol. , p. 212–216. External Links: Document Cited by: §I. [26] R. Islam, J. Mulé, D. Challagundla, S. Rizvi, and S. Carson (2025) EA: an event autoencoder for high-speed vision sensing. In IEEE Computer Society Annual Symposium on VLSI (ISVLSI), Cited by: 1st item. [27] R. Islam, B. Saha, and I. Bezzam (2021) Resonant Energy Recycling SRAM Architecture. IEEE Transactions on Circuits and Systems I: Express Briefs 68 (4), p. 1383–1387. External Links: Document Cited by: §I-C. [28] R. Islam, S. R. S. K. Tummala, J. Mulé, R. Kankipati, S. Jalapally, D. Challagundla, C. Howard, and R. Robucci (2024) Descriptor: Smart Event Face Dataset (SEFD). IEEE Data Descriptions. Cited by: Fig. 1, §I, §I-A, §I-C, §IV, §IV, TABLE I, TABLE I, §V. [29] R. Islam (2011) High-speed energy-efficient soft error tolerant flip-flops. Ph.D. Thesis, Concordia University. Cited by: §I-C. [30] R. Islam (2019) Low-Power Highly Reliable SET-Induced Dual-Node Upset-Hardened Latch and Flip-Flop. Canadian Journal of Electrical and Computer Engineering 42 (2), p. 93–101. External Links: Document Cited by: §IV. [31] R. Islam (2021) Negative capacitance clock distribution. IEEE Transactions on Emerging Topics in Computing 9 (1), p. 547–553. External Links: Document Cited by: §IV. [32] R. Islam (2022) Early Stage DRC Prediction Using Ensemble Machine Learning Algorithms. IEEE Canadian Journal of Electrical and Computer Engineering 45 (4), p. 354–364. External Links: Document Cited by: §IV. [33] H. Kamata, Y. Mukuta, and T. Harada (2022) Fully spiking variational autoencoder. In Proceedings of the AAAI conference on artificial intelligence, Vol. 36, p. 7059–7067. Cited by: §I. [34] W. Khwa, T. Wen, H. Hsu, W. Huang, Y. Chang, T. Chiu, Z. Ke, Y. Chin, H. Wen, W. Hsu, et al. (2025) A mixed-precision memristor and SRAM compute-in-memory AI processor. Nature 639 (8055), p. 617–623. Cited by: §IV. [35] J. Kim, I. Hwang, and Y. Kim (2022) Ev-TTA: Test-time adaptation for event-based object recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8769–8778. Cited by: §I-C. [36] S. Kim, S. Chu, Y. Park, and C. Lee (2024) A tied-weight autoencoder for the linear dimensionality reduction of sample data. Scientific Reports 14 (1), p. 26801. Cited by: §I. [37] V. Kodukula, M. Manetta, and R. LiKamWa (2023) Squint: a framework for dynamic voltage scaling of image sensors towards low power iot vision. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking, ACM MobiCom, New York, NY, USA. External Links: ISBN 9781450399906, Link, Document Cited by: §I-C. [38] D. Kollias, P. Tzirakis, M. Nicolaou, A. Papaioannou, G. Zhao, B. Schuller, I. Kotsia, and S. Zafeiriou (2019) Deep Affect Prediction in-the-Wild: Aff-Wild Database and Challenge, Deep Architectures, and Beyond. International Journal of Computer Vision 127, p. . External Links: Document Cited by: §IV. [39] D. Li, Y. Tian, and J. Li (2023) Sodformer: Streaming object detection with transformer using events and frames. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §I. [40] H. Li, K. Ma, H. Yong, and L. Zhang (2020) Fast multi-scale structural patch decomposition for multi-exposure image fusion. IEEE Transactions on Image Processing 29 (), p. 5805–5816. External Links: Document Cited by: §I-C. [41] J. Li, L. Zhu, X. Xiang, and T. Huang (2022) Asynchronous spatio-temporal memory network for continuous event-based object detection. IEEE Transactions on Neural Networks and Learning Systems. Cited by: §I-C. [42] T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) Microsoft coco: common objects in context. In European conference on computer vision, p. 740–755. Cited by: §I-C. [43] S. Miao, G. Chen, X. Ning, Y. Zi, K. Ren, Z. Bing, and A. Knoll (2019) Neuromorphic vision datasets for pedestrian detection, action recognition, and fall detection. Frontiers in neurorobotics 13, p. 38. Cited by: §I-A. [44] J. Mulé, D. Challagundla, R. Saini, and R. Islam (2025) Descriptor: Event-Based Crossing Dataset (EBCD). IEEE Data Descriptions 2 (), p. 71–81. External Links: Document Cited by: §I, §I-A, §I-C, §IV, §IV, TABLE I, §V. [45] S. Neogi, M. Hoy, K. Dang, H. Yu, and J. Dauwels (2019) Context model for pedestrian intention prediction using factored latent-dynamic conditional random fields. IEEE Transactions on Intelligent Transportation Systems 22, p. 6821–6832. External Links: Link Cited by: §IV. [46] G. Orchard, A. Jayawant, G. K. Cohen, and N. Thakor (2015) Converting static image datasets to spiking neuromorphic datasets using saccades. Frontiers in neuroscience 9, p. 437. Cited by: §I-A. [47] G. Orchard, A. Jayawant, G. K. Cohen, and N. Thakor (2015) Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades. Frontiers in Neuroscience Volume 9 - 2015. External Links: Link, Document, ISSN 1662-453X Cited by: §I-C. [48] E. Perot, P. de Tournemire, D. Nitti, J. Masci, and A. Sironi (2020) Learning to Detect Objects with a 1 Megapixel Event Camera. External Links: 2009.13436, Link Cited by: §I-A, §I-C. [49] PROPHESEE SONY-PROPHESEE IMX636 HD Event-Based Vision. Note: https://w.prophesee.ai/event-camera-evk4/ Cited by: §I. [50] A. K. Rajendra, H. S. Bindra, and B. Nauta (2025) Ultra-Low-Power Dynamic-Bias Comparators With Self-Clocked Latch in 65-nm CMOS. IEEE Journal of Solid-State Circuits. Cited by: §IV. [51] I. S. Rajput, A. S. Tewari, and A. K. Tiwari (2024) An autoencoder-based deep learning model for solving the sparsity issues of Multi-Criteria Recommender System. Procedia Computer Science 235, p. 414–425. Cited by: §I-C. [52] K. K. Reddy and M. Shah (2013) Recognizing 50 human action categories of web videos. Machine vision and applications 24 (5), p. 971–981. Cited by: §I-A. [53] X. Tai, H. Liu, and R. Chan (2024) PottsMGNet: A mathematical explanation of encoder-decoder based neural networks. SIAM Journal on Imaging Sciences 17 (1), p. 540–594. Cited by: §I. [54] M. Tan, R. Pang, and Q. V. Le (2020) EfficientDet: Scalable and Efficient Object Detection. External Links: 1911.09070 Cited by: §IV, TABLE I, TABLE I. [55] D. Torbunov, Y. Ren, A. Ghose, O. Dim, and Y. Cui (2025) Evrt-detr: latent space adaptation of image detectors for event-based vision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 9812–9821. Cited by: §I-B. [56] B. Tossoun, X. Xiao, S. Cheung, Y. Yuan, Y. Peng, S. Srinivasan, G. Giamougiannis, Z. Huang, P. Singaraju, Y. London, M. Hejda, S. P. Sundararajan, Y. Hu, Z. Gong, J. Baek, A. Descos, M. Kapusta, F. Böhm, T. Van Vaerenbergh, M. Fiorentino, G. Kurczveil, D. Liang, and R. G. Beausoleil (2025) Large-Scale Integrated Photonic Device Platform for Energy-Efficient AI/ML Accelerators. IEEE Journal of Selected Topics in Quantum Electronics 31 (3: AI/ML Integrated Opto-electronics), p. 1–26. External Links: Document Cited by: §IV. [57] I. Ustek, M. Arana-Catania, A. Farr, and I. Petrunin (2024) Deep Autoencoders for Unsupervised Anomaly Detection in Wildfire Prediction. External Links: 2411.09844, Link Cited by: §I. [58] G. Vimala and F. VincyLloyd (2025) A Study & Analysis of Low Power Phase Locked Loop Design. In 2025 IEEE 7th International Conference on Computing, Communication and Automation (ICCCA), p. 1–4. Cited by: §IV. [59] Z. Wan, Y. Dai, and Y. Mao (2022) Learning dense and continuous optical flow from an event camera. IEEE Transactions on Image Processing 31, p. 7237–7251. External Links: ISSN 1941-0042, Link, Document Cited by: §I-C. [60] C. Wang, A. Bochkovskiy, and H. M. Liao (2023) YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 7464–7475. External Links: Document Cited by: §IV, §IV, TABLE I, TABLE I, §V. [61] C. Wang and H. M. Liao (2024) YOLOv9: learning what you want to learn using programmable gradient information. Cited by: 4th item, §IV, §IV, TABLE I, TABLE I, §V. [62] Y. Wang, H. Yao, and S. Zhao (2016) Auto-encoder based dimensionality reduction. Neurocomputing 184, p. 232–242. External Links: ISSN 0925-2312, Document, Link Cited by: §I. [63] Z. Wang, F. Cladera, A. Bisulco, and D. Lee (2022) Ev-catcher: High-speed object catching using low-latency event-based neural networks. In IEEE International Conference on Robotics and Automation (ICRA), p. 1045–1051. Cited by: §I. [64] Z. Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13 (4), p. 600–612. External Links: Document Cited by: §IV. [65] W. Wu, H. Peng, and S. Yu (2023) YuNet: a tiny millisecond-level face detector. Machine Intelligence Research, p. 1–10. External Links: Document Cited by: §IV, §IV, TABLE I, TABLE I, §V. [66] C. Xi, J. Zhou, and P. Zhou (2026) CNN-Assisted Low-Power Clock Tree Synthesis for 3D ICs. In 2026 31st Asia and South Pacific Design Automation Conference (ASP-DAC), p. 1421–1427. Cited by: §IV. [67] L. Yu, H. Chen, Z. Wang, S. Zhan, J. Shao, Q. Liu, and S. Xu (2025) SpikingViT: a multiscale spiking vision transformer model for event-based object detection. IEEE Transactions on Cognitive and Developmental Systems 17 (1), p. 130–146. External Links: Document Cited by: §I-C. [68] J. Zhang, K. Yang, and R. Stiefelhagen (2021) ISSAFE: improving semantic segmentation in accidents by fusing event-based data. External Links: 2008.08974 Cited by: §I-C. [69] J. Zhang, B. Dong, H. Zhang, J. Ding, et al. (2022) Spiking transformers for event-based single object tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11240–11250. Cited by: §I-C. [70] J. Zhou, H. Zhang, Q. Qiao, H. Chen, Q. Huang, H. Wang, Q. Ren, N. Wang, Y. Ma, and C. Lee (2024) Denoising-autoencoder-facilitated MEMS computational spectrometer with enhanced resolution on a silicon photonic chip. Nature Communications 15 (1), p. 10260. Cited by: §I. [71] A. Z. Zhu, D. Thakur, T. Ozaslan, B. Pfrommer, V. Kumar, and K. Daniilidis (2018) The Multivehicle Stereo Event Camera Dataset: An Event Camera Dataset for 3D Perception. IEEE Robotics and Automation Letters 3 (3), p. 2032–2039. External Links: ISSN 2377-3774, Link, Document Cited by: §I-A. [72] A. Z. Zhu, L. Yuan, K. Chaney, and K. Daniilidis (2018) Unsupervised Event-based Learning of Optical Flow, Depth, and Egomotion. CoRR abs/1812.08156. External Links: Link, 1812.08156 Cited by: §I-B. [73] N. Zubić, D. Gehrig, M. Gehrig, and D. Scaramuzza (2023) From chaos comes order: ordering event representations for object recognition and detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 12846–12856. Cited by: §I-B. Riadul Islam is currently a tenured associate professor in the Department of Computer Science and Electrical Engineering at the University of Maryland, Baltimore County. In his Ph.D. dissertation work at UCSC, Riadul designed the first current-pulsed flip-flop/register that resulted in the first-ever one-to-many current-mode clock distribution networks for high-performance microprocessors. From 2017 to 2019, he was an Assistant Professor with the University of Michigan, Dearborn MI, USA. He is a senior member of the IEEE, member of the ACM, IEEE Circuits and Systems (CAS) society, the VLSI Systems and Applications Technical Committee (VSA-TC) of the IEEE-CAS, and IEEE Solid-State Circuits (SSC) Society. He holds two US patent and several IEEE/ACM/MDPI/Springer Nature journal and conference publications. His current research interests include digital, analog, and mixed-signal CMOS ICs/SOCs for a variety of applications; verification and testing techniques for analog, digital and mixed-signal ICs; hardware security; CAN network; CAD tools for design and analysis of microprocessors and FPGAs; automobile electronics; and biochips. He is an Associate Editor of Springer Circuits, Systems and Signal Processing (CSSP) Journal. He was a Technical Program Committee (TPC) member of the IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2022), ACM Great Lakes Symposium on VLSI (GLSVLSI 2020, GLSVLSI 2021, GLSVLSI 2022), 57th IEEE/ACM Design Automation Conference (DAC) 2020 LBR Session, IEEE Computer Society Annual Symposium on VLSI (ISVLSI) 2021, and IEEE International Conference on Consumer Electronics (ICCE) 2021. Riadul is the recipient of a 2021 NSF ERI award, 2021 Maryland Industrial Partnerships (MIPS) award, and 2021 Maryland Innovation Initiative (MII) award. Joey Mulé (Student Member, IEEE) received his B.S. degree in Computer Science from the University of Maryland, Baltimore County (UMBC), MD, USA. He is currently pursuing the M.S. degree in Computer Science at UMBC, where he continues to explore research at the intersection of machine learning and artificial intelligence within the area of computer science and computer engineering. His academic and research interests include neural network architectural design, real-time object detection, adversarial defense systems, and low-power edge computing, with a particular focus on applications in event-based vision. Dhandeep Challagundla (Student Member, IEEE) received his Ph.D. degree from The University of Maryland Baltimore County (UMBC), MD, USA. His research interests revolve around energy-efficient computing, Compute-in-Memories, SRAM design, low-power circuit design, Mixed-signal IC design, and EDA tools. Shahmir Rizvi received his B.S. degree in Computer Engineering from the University of Maryland Baltimore County (UMBC), MD, USA, where he is currently pursuing his Ph.D. His research interests include VLSI design, embedded systems, and FPGA-based hardware acceleration. Sean Carson holds a B.S. degree in Computer Engineering from the University of Maryland, Baltimore County (UMBC), MD, USA. He currently researches at the UMBC VLSI-SOC Group, where his work focuses on event vision and neural networks. Outside of his research, he is employed as a Systems Engineer at Textron Systems, specializing in the development of electromagnetic environment simulators. Rachit Saini completed his B.S. in Electrical Engineering from University of Illinois at Urbana-Champaign, Urbana-Champaign, IL, USA in 2014. He received his M.Tech. in VLSI Design from Thapar Institute of Engineering and Technology, Patiala, Punjab, India in 2018. Further, he received his M.S. in Computer Engineering from University of Maryland Baltimore County (UMBC), MD, USA in 2025. Currently he is pursuing his Ph.D. in Computer Engineering from UMBC. His current research interest includes VLSI design and security.