Skip to Content
SensorsSensors
  • Article
  • Open Access

12 December 2023

Learnable Leakage and Onset-Spiking Self-Attention in SNNs with Local Error Signals

,
,
and
1
School of Microelectronics and Communication Engineering, Chongqing University, Chongqing 400044, China
2
Key Laboratory of Dependable Service Computing in Cyber Physical Society, Ministry of Education, Chongqing University, Chongqing 400044, China
*
Author to whom correspondence should be addressed.
This article belongs to the Section Sensing and Imaging

Abstract

Spiking neural networks (SNNs) have garnered significant attention due to their computational patterns resembling biological neural networks. However, when it comes to deep SNNs, how to focus on critical information effectively and achieve a balanced feature transformation both temporally and spatially becomes a critical challenge. To address these challenges, our research is centered around two aspects: structure and strategy. Structurally, we optimize the leaky integrate-and-fire (LIF) neuron to enable the leakage coefficient to be learnable, thus making it better suited for contemporary applications. Furthermore, the self-attention mechanism is introduced at the initial time step to ensure improved focus and processing. Strategically, we propose a new normalization method anchored on the learnable leakage coefficient (LLC) and introduce a local loss signal strategy to enhance the SNN’s training efficiency and adaptability. The effectiveness and performance of our proposed methods are validated on the MNIST, FashionMNIST, and CIFAR-10 datasets. Experimental results show that our model presents a superior, high-accuracy performance in just eight time steps. In summary, our research provides fresh insights into the structure and strategy of SNNs, paving the way for their efficient and robust application in practical scenarios.

1. Introduction

Throughout the history of neural network research, traditional artificial neural networks (ANNs) [1] have been the primary focus, due to their remarkable performance and extensive applications. However, despite ANNs’ ability to handle complex nonlinear patterns, a significant gap remains in imitating the functioning of the human brain. Notably, biological neural systems use temporal spike activity, in contrast to ANNs, which heavily rely on continuous activation values. The observation of the human brain’s impressive efficiency in information processing, coupled with its low energy consumption, has sparked interest in spiking neural networks (SNNs) [2].
SNNs differ from ANNs in that they use sparse temporal spike events to encode and process information. While sparse coding does not guarantee an increase in computational power [3,4,5], it does contribute to a reduction in computational complexity, resulting in resource savings. This gain in efficiency gives SNNs a broad potential for application in various fields. In particular, in the context of processing highly dynamic and real-time data streams, SNNs demonstrate superior efficiency in handling temporally correlated information due to their inherent temporal coding properties.
Additionally, the inherent energy efficiency of SNNs offers promising opportunities in areas such as neuromorphic hardware and edge computing. Acknowledging these advantages, the academic community is gradually shifting from ANNs to exploring SNNs, striving for a more genuine representation of biological neural processes, and unlocking novel possibilities in various application areas [6].
In order to facilitate effective training of SNNs, researchers have proposed various methods [7,8,9]. Present research predominantly concentrates on three major aspects: pretraining through clustering and autoencoding, among other methods, under unsupervised learning; enhancing performance by combining supervised information and unlabeled data under semisupervised learning; and implementing backpropagation under supervised learning using alternative differentiable activation functions or other techniques. These three categories of approaches have distinct advantages and disadvantages, but all have exhibited the potential of SNNs in handling signals and various tasks. With the continued advancement of theories and algorithms, SNNs offer a wide range of potential applications. It is worth noting that advanced mathematical theories have also supplied important mathematical tools for the accurate modeling and intricate dynamic analysis of SNNs [10].
Unsupervised learning has the ability to adjust neuronal connection weights autonomously by using local information, such as the spike-timing-dependent plasticity (STDP) rule, which modifies weights based on spike-time correlations [11]. These techniques are computationally efficient and simple. However, due to the absence of supervisory information, they generally demonstrate lower accuracy and are often used for network initialization. Semisupervised learning entails pretraining the network using unlabeled data, then fine-tuning it with labeled data. This technique can boost performance by leveraging unlabeled data, but it mandates cleverly designed approaches for utilizing supervisory information [12,13,14]. Currently, supervised learning is the most widespread method for training spiking neural networks. Traditional neural networks use differentiable activation functions and, therefore, they can use the chain rule to calculate gradients directly, enabling backpropagation. On the other hand, spiking neural networks mimic the biologically inspired spike propagation mechanism, making their activation functions nondifferentiable. This makes the direct application of the backpropagation algorithm impossible, posing challenges to the supervised training of spiking neural networks.
Researchers have proposed a variety of supervised learning rules to address issues arising from the nondifferentiability of the spike function in spiking neural networks. For example, SpikeProp [15] employs a linear approximation method, while the alternative gradient rule substitutes traditional activation functions with alternative ones [16]. Moreover, methods based on backpropagation through time (BPTT), which compute gradients jointly from spatiotemporal dimensions, have also gained popularity in recent years [17,18,19,20,21,22]. While these methods have achieved commendable classification accuracy, they come with a substantial computational cost.
In this study, we focus on optimizing both the performance and interpretability of SNNs. Specifically, we improve the conventional LIF model and introduce the LLC-LIF model, where “LLC” stands for “learnable leakage coefficient”. In addition, we propose the batch normalization method combined with the learnable leakage coefficient, termed LLC-BN. By integrating local loss signals and the self-attention mechanism [23] from deep learning, we further enhance the performance and application scope of SNNs. Our primary contributions are as follows:
  • We present the LLC-LIF model with a learnable leakage coefficient, which allows the leakage coefficient of the membrane potential to be a learnable parameter. This design provides automated optimization capabilities, ensuring consistent properties between neurons within the same layer and independent properties across layers.
  • To better exploit the temporal sensitivity and efficiency of SNNs, we incorporate the self-attention mechanism at the initial time step of the SNN. By integrating strategies from both neuroscience and deep learning, we further enhance the network’s ability to transform temporal and spatial features.
  • To adapt to the unique characteristics of spiking neural networks, we introduce the batch normalization method combined with a learnable leakage coefficient, termed LLC-BN. This method harmonizes the temporal dynamics of SNNs with spike-time encoding and enhances the stability and flexibility of the network through joint optimization.
  • To efficiently emulate biological neural networks, we introduce local loss signals within the spiking neural network, allowing certain layers to receive distinct learning feedback independently. Using supervised local learning strategies and auxiliary classifiers, we design a hierarchical loss function that ensures excellent performance of the SNN in various tasks.
The rest of this article is organized as follows: Section 2 delves into foundational works relevant to our study. Section 3 details our materials and methods, highlighting our innovative structures and strategies. Section 4 is dedicated to experimental results and comparative analyses. Section 5 concludes our research.

3. Materials and Methods

3.1. LIF Model with Learnable Leakage Coefficient (LLC-LIF)

In neuronal dynamics, the leakage coefficient of the membrane potential is a crucial factor, defining the velocity at which a neuron’s membrane potential goes back to its resting state when there are no external inputs present. When examining the biological plausibility of the LIF model, we stress the importance of the leakage coefficient in the model. This factor is key to the simulation of the organic degeneration of the membrane potential of neurons to the resting state. This process is mainly influenced by ion channels, with potassium channels playing a critical role in maintaining and restoring the resting potential [43].
The LIF model adjusts the leakage coefficient to simulate the dynamics of the membrane potential in neurons, taking into account ion channel availability and conductance changes. By decreasing the leakage coefficient, the neuron’s integration time for inputs can be extended, resembling neurons with closed ion channels and altering their firing patterns [44].
The rate of leakage is intricately associated with the membrane time constant τ, which signifies the duration over which a neuron combines input information. The magnitude of the leakage coefficient directly affects how neurons respond to temporal patterns of input and their encoding abilities. Hence, it is imperative to regulate the leakage coefficient to mirror the biological traits of real neurons while creating neuronal models.
Diverging from traditional LIF models that employ a fixed leakage coefficient, we introduce a novel LIF model. In this proposed model, the membrane potential’s leakage coefficient is devised as a learnable parameter, termed LLC-LIF. The dynamical equations governing this neuron model are as follows:
u t + 1 , n + 1 ( i ) = k τ ( a ) u t , n + 1 ( i ) ( 1 o t , n + 1 ( i ) ) + j = 1 l ( n ) w i j n o t + 1 , n ( j ) , o t + 1 , n + 1 ( i ) = f ( u t + 1 , n + 1 ( i ) V t h ) ,
where kτ(a) represents a clamping function bounded between (0, 1), ensuring that τ = 1/kτ(a) lies within the range (1, +∞). In our experiments, we set kτ(a) = 1/(1 + exp(−a)).
The learnable leakage coefficient, denoted as kτ(a), offers several biologically plausible advantages. Its automatic optimization obviates the need for manual hyperparameter selection, facilitating end-to-end and automated model training. Neurons within the same layer share the value of kτ(a) and exhibit similar characteristics. However, according to Eve Marder et al. [45], there is a certain degree of individual variability among these neurons. In our future work, we may explore how such individual differences impact the network’s training performance, presenting an interesting and valuable research direction. Moreover, leakage coefficients are independent across different layers, endowing neurons in each layer with unique temporal coding properties.
By integrating this adjustable leakiness into the model, we can more accurately imitate the dynamic behavior of biological neurons. As this parameter is trainable, it can adaptively modify during the training process based on data. This flexibility can facilitate the model in acquiring the optimum leaky behavior, thereby maximizing its performance on particular tasks. This biologically rooted design creates opportunities for constructing SNNs. Our research aims to utilize the LLC-LIF model in order to train efficient SNNs to tackle complex temporal learning challenges.

3.2. Onset-Spiking Self-Attention (OSSA)

SNNs provide a unique and energy-efficient framework for neural computation by emulating the spiking propagation mechanism of biological neurons. In SNNs, time plays a pivotal role, particularly when using spatiotemporal backpropagation for effective training. However, despite the irreplaceable importance of time in SNNs, traditional training strategies still have limitations when it comes to handling temporal information and feature selection. To harness the full potential of SNNs, we propose introducing a self-attention mechanism at the initial time step.
Biological research has revealed that the brain’s initial response after receiving a stimulus is crucial in subsequent information processing and decision-making processes [46]. This initial response provides critical information about the stimulus and may form the basis for information processing and decision-making. In the context of SNNs, this suggests that the initial time step T = 0 might also play a decisive role in the entire network’s response. By enhancing feature selection at this crucial moment, we aim to enable the network to focus more on genuinely important information in subsequent time steps.
In this context, the self-attention mechanism provides a promising approach. Initially introduced in the transformer model, it allows the model to assign different weights to each element in the input sequence, capturing long-range dependencies within the sequence. By assigning weights to each input element, this mechanism enables the model to better focus on the most crucial parts of the entire sequence. When we apply this mechanism to SNNs, we hope that the network can better identify and respond to key time points and features, leading to more accurate responses throughout the time sequence.
Considering these factors, we believe that introducing self-attention into the initial stage of SNNs is a reasonable and promising approach. This new method combines successful strategies from deep learning with insights from neuroscience, offering a new direction for further research and application of SNNs.
Traditional implementations of self-attention typically use a simple convolutional layer to transform input features. However, when considering SNNs, the network’s dynamic nature and spiking behavior provide an opportunity to further optimize the attention mechanism. To achieve this, we propose changing the convolutional layer of self-attention from a simple two-dimensional convolutional layer to one integrated with optimized LIF structures. The core idea behind this change is to leverage the dynamic properties of the optimized LIF structure to enhance the representational capacity of the attention mechanism.
We consider incorporating the influence of spiking neurons after computing queries, keys, and values, and before calculating the weight coefficients, as shown below:
f ( x ) = L L C _ L I F ( f ( x ) ) ,   g ( x ) = L L C _ L I F ( g ( x ) ) ,   h ( x ) = L L C _ L I F ( h ( x ) ) .
Then we can use f′(x), g′(x) and h′(x) to compute the weight coefficients and generate the output of self-attention.
O = Softmax ( f ( x ) × g ( x ) T ) × h ( x ) .
By combining the two-dimensional convolution layer with the optimized LIF structure, we can balance feature transformations in both spatial and temporal dimensions. While traditional two-dimensional convolution layers focus solely on spatial information, the optimized LIF structure provides a means of temporal modulation, allowing the model to adaptively handle dependencies at different time scales.
In summary, modifying the convolution layer of self-attention to incorporate the optimized LIF structure not only enhances the model’s representational capacity but also offers a means of adaptive temporal modulation. This is crucial for pulse neural networks. Experimental results have demonstrated significant performance improvements on various benchmark tasks with the introduction of pulse-based self-attention, further validating the effectiveness and potential of our approach.

3.3. Learnable Leakage Coefficient Batch Normalization (LLC-BN)

Traditional batch normalization techniques have shown significant efficacy when applied to conventional neural networks, but they are not directly adaptable to SNNs. This incompatibility arises due to the temporal dynamics of membrane potentials in SNNs and the unique time-encoding characteristics of spikes, which are fundamentally different from networks with static activation functions.
To address this, we propose a novel normalization technique tailored for the operational mechanism of SNNs: the learnable leakage coefficient batch normalization (LLC-BN) method. This method jointly optimizes the neuron’s membrane potential leakage coefficient and input normalization. It computes the mean and variance of the membrane potential at each time step as normalization benchmarks, smoothing the temporal activation patterns of the network. This advanced normalization approach takes into account the spatiotemporal information representation traits of SNNs. By co-optimizing the leakage parameter and the input distribution, it effectively reduces the variance of temporal encoding, enhancing the network’s capability to learn dynamic features.
In SNNs, each neuron’s behavior is time-based, responding in the form of spikes across different time steps. Let ot denote the spike outputs of all neurons in a layer at time step t. To characterize how neurons respond to their inputs, we introduce a convolutional kernel W and bias B. For a given input xt, its spike response is transformed through the convolutional kernel W and bias B. This can be mathematically represented as:
o t = f ( W K * x t + B ) ,
where * denotes the convolution operation, and f serves as an activation function. Typically, in spiking neural networks, f acts as a threshold function, deciding whether to fire a spike. xt is a four-dimensional tensor representing the presynaptic input at time step t. N stands for the batch size, indicating the number of samples processed simultaneously. C refers to the number of channels, representing the count of input features. H and W, on the other hand, represent the height and width of the input, symbolizing the spatial dimensions.
In the proposed LLC-BN method, normalization is performed along the channel dimension. Specifically, for each channel feature map xk, it undergoes the following normalization process:
x k = α k τ ( a ) ( x k E [ x k ] ) V a r [ x k ] + ϵ ,
subsequently, the normalized output obtained is represented as:
y k = λ k x k + β k ,
where α is a hyperparameter, ϵ is a small constant to prevent division by zero, kτ(a) is a trainable leakage parameter, and xk and xk′ are the neural inputs before and after normalization, respectively. λk and βk are two trainable parameters used for scaling and shifting in the linear transformation after normalization. E[xk] and Var[xk] denote the mean and variance calculated from the elements of xk along the batch axis N, the spatial axes H and W, and the time axis T. Specifically, yk denotes the normalized presynaptic input received by the k-th channel neuron in the subsequent layer over a period of time T.
Moreover, we do not just compute the mean and variance for the current batch of data; we also employ the moving average method to estimate the mean μinf and variance σinf2 over the entire dataset. This strategy ensures robust normalization during the inference phase, irrespective of the batch size of the input data.
Of particular note in LLC-BN, the pre-activation is normalized to a distribution with a mean of 0 and a variance of α2 × kτ(a)2, differing from the N(0,1) in traditional batch normalization. This adjustment makes the normalization more attuned to the spiking behavior of SNNs.
As can be observed, the aforementioned self-optimizing leakage coefficient kτ(a) is introduced to adjust and scale the normalized data. in line with biological systems. The leaky parameter kτ(a) dynamically adjusts the normalization range, enabling flexible adaptation to data distribution and variations. All technical terms are explained when first used. This parameter enhances neural sensitivity to historical information and time sensitivity of normalization. Additionally, kτ(a) ensures stability during normalization, particularly when dealing with noise or outlier data. Moreover, its trainability allows for self-adjustment during training and thus optimizes the model’s performance across various tasks and data distributions. During inference, the standard batch normalization strategy is followed, wherein the moving average over the complete dataset is employed for estimating the mean and variance, thus ensuring the stability of the network.
In order to implement the SNN on neuromorphic hardware whilst preserving its full spiking properties, we adopted the batch normalization scale fusion technique. This approach eliminates the need for batch normalization during the inference stage, enabling the entire network to maintain a pure spiking form and making it simpler to deploy on neuromorphic platforms. Let W′ and B′ denote the convolutional kernel and bias, respectively, following normalization. After batch normalization scale fusion, these weights and biases undergo corresponding transformations:
W = λ α k τ ( a ) W σ i n f 2 + ϵ ,
B = λ α k τ ( a ) ( B μ i n f ) σ i n f 2 + ϵ + β .
During the inference process, information is passed layer by layer through these transformed weights W′ and biases B′ without the need for additional batch normalization steps. This means that LLC-BN only affects the computation during training and does not affect the operating mechanism of a trained SNN. In our experiments, we initialise the trainable parameters λ and β to 1 and 0, respectively. The hyperparameter α is set to 3.2.

3.4. Spatiotemporal Backpropagation with Local Error Signals

At the core of biological neural networks are synapses, which interact within highly complex and parallel environments, often relying on locally available information to adjust their weights. This phenomenon suggests that SNNs, when simulating biological neural systems, could similarly benefit from the drive of local information, thereby enhancing the training efficiency and accuracy of the network. Understanding this context led us to introduce local loss signals in SNNs. In this setup, each layer can be independently updated based on local learning signals, making the training of SNNs more efficient. This strategy aligns with the parallelism and adaptability observed in biological neural networks and better accommodates the spatiotemporal characteristics of information in SNNs.
The complete network structure is shown in Figure 3. Within our network, we have integrated a supervised local learning approach, the core of which is the use of auxiliary classifiers to construct hierarchical loss functions [47]. This allows us to utilize training labels for more explicit and targeted local updates while ensuring that SNNs perform well across a variety of tasks. Furthermore, our local loss signal strategy not only draws inspiration from the local learning mechanisms of biological neural networks but also integrates ideas from deep continuous local learning, leveraging temporal local information for continuous SNN training at each time step.
Figure 3. The comprehensive architecture diagram of the network, which incorporates supervised local loss, highlights the synergistic interplay between independent local losses at convolutional layers and the global loss, reflecting the adaptability and spatiotemporal dynamics of biological neural systems.
The primary advantage of this approach lies in providing more direct and specific feedback for hidden layers. Local losses can indicate more explicitly which part of the network needs adjustment, rather than relying on global feedback propagated from the output layer. This enables us to fine-tune each part of the network more precisely, capturing and learning subtle differences in the data more effectively. Additionally, introducing local loss signals brings added training efficiency. Each layer can be updated independently and in parallel, making the training process more efficient and facilitating faster convergence to optimal solutions.
In summary, by introducing local loss signals into spatiotemporal backpropagation, we not only enhance the ability of spiking neural networks to handle complex data patterns, but also greatly improve training efficiency and stability.
In our research, we adopted standard convolutional and fully connected network architectures. One significant feature of this model, compared to traditional neural network structures, is the assignment of independent local losses between convolutional layers. These local losses, combined with a global loss, collectively form the total loss of the network to guide the optimization process.
Specifically, the introduction of local losses aims to ensure that each convolutional layer can independently optimize and capture its corresponding feature space. The global loss, on the other hand, aims to ensure that the macrolevel outputs of the network match the expected labels as closely as possible, thus achieving the overall training goal of the model. To quantify the difference between the model outputs and the real labels, we used the mean squared error (MSE) [48] as the loss function, defined as:
L MSE = 1 N s i = 1 N ( y i y ^ i ) 2 ,
where yi represents the true values, ŷ represents the model’s predictions, and Ns is the number of samples.
The comprehensive loss function of the model is composed of the local losses from all the convolutional layers and a global loss, and can be expressed as follows:
L total = i = 1 n L locali + L global ,
where n represents the total number of convolutional layers, and Llocali is the local loss for the i-th convolutional layer.

4. Experiments

4.1. Benchmark Datasets

We evaluated our proposed SNN model on three primary image datasets: MNIST [49], FashionMNIST [50], and CIFAR-10 [51]. Specifically, the MNIST dataset consists of 10 classes of handwritten digit images with a resolution of 28 × 28, totaling 50,000 training samples and 10,000 test samples. FashionMNIST, structurally similar to MNIST, showcases 10 different clothing categories. On the other hand, the CIFAR-10 dataset encompasses 10 object classes, each with images of 32 × 32 resolution, comprising 50,000 training images and 10,000 test images. The detailed attributes of these datasets, such as image resolution, number of categories, and the division of training/testing subsets, are all listed in Table 1.
Table 1. Benchmark datasets.

4.2. Network Structure Configuration

In the experiments, the network architecture “128c3-p2-128c3-p2-2048-100-10” was employed for the MNIST and FashionMNIST datasets. For the CIFAR-10 dataset, the architecture “256c3-256c3-256c3-p2-256c3-256c3-256c3-p2-2048-100-10” was adopted. In these specifications, “c” represents a convolution layer with the preceding number indicating the quantity of convolution kernels. The number “3” that follows elucidates the kernel size of 3 × 3. The symbol “p” denotes a pooling layer, with the succeeding “2” indicating a 2 × 2 pooling window size. Additionally, “2048” and “100” symbolize fully connected layers with their respective neuron quantities.
At the tail end of the fully connected layers, a unique “100-10” structure was incorporated. Notably, “100-10” does not directly signify two adjacent fully connected layers. In this context, “10” pertains to an average pooling layer applied to the output of the preceding fully connected layer, with both stride and window size set to 10. The core intention of this strategy is to first project complex features onto a relatively low-dimensional 100-feature space, and subsequently obtain a 10-dimensional output representation via the average pooling layer. This approach accomplishes feature dimension reduction, streamlines the network architecture, and preserves pivotal information while mitigating computational demands. Following this, averaging features in a low-dimensional space enables the model to capture more prominent and significant information, elevating the capability of recognizing key features and, to a degree, suppressing noise.
During the network training phase, dynamic optimization of network parameters was conducted via learnable leakage coefficients and the OSSA strategy. The introduction of the novel normalization method, LLC-BN, also augmented the network’s capability to learn dynamic features. Furthermore, to elevate network performance, the local error signal was employed, propelling the network to achieve more efficient feature extraction and representation across layers, thereby assisting the model in learning the mapping relationship from input to output with increased stability. To enhance the network’s generalization capability and combat overfitting, a dropout strategy [52] was implemented in the latter part of the model. This strategy, by suppressing the activation of random neurons during training, offers robust regularization effects, ensuring a more resilient network and preventing the model from excessively relying on specific neurons in the training data, thus promoting a more robust and sturdy training process.
In this study, the rate coding method [53] was utilized to transform pixel values of images into spikes within the spiking neural network. Rate coding is an encoding strategy where a neuron’s spike firing rate is directly proportional to its input strength. This implies that a higher input strength would result in a greater spike firing rate. A notable advantage of this encoding strategy is its ability to intuitively reflect the strength of input data, furnishing spiking neural networks with ample input information. Using rate coding ensures that SNNs receive temporal spike information directly correlated with the original image pixels, laying a solid foundation for subsequent neural network processing. Additionally, in our model, the output employs a direct decoding strategy, which directly presents the total spike count. For the loss calculation phase, these total spikes are transformed into spike frequencies to facilitate comparison with the target labels, which are in one-hot encoded form.
All experiments were implemented using SpikingJelly [54], an open-source SNN deep learning framework built upon PyTorch [55]. We trained our models on NVIDIA GeForce GTX 3090. For the experiments across the MNIST, FashionMNIST, and CIFAR-10 datasets, we consistently set the batch size to 16 and employed the Adam optimizer [56]. All networks were trained for a total of 200 epochs. Our source code is available at https://github.com/CQU0121WL/Learnable-Leakage-and-Onset-Spiking-Self-Attention-in-SNNs-with-Local-Error-Signals (accessed on 7 December 2023).

4.3. Work Comparison and Discussion

For the MNIST dataset, it is evident that various methods have demonstrated excellent performance in the training and optimizing of deep spiking neural networks, as shown in Table 2. However, our approach significantly surpasses all other methods listed by achieving an accuracy of 99.67% with only eight time steps. Jin et al. [57] employed similar network structures and trained over more time steps, resulting in accuracies close to ours but requiring significantly more time steps. This might be attributed to their training strategies, which emphasized spatiotemporal backpropagation and mixed macro/microlevel backpropagation. Sengupta et al. [58] and Lee et al. [21] opted for the LeNet-5 structure, and although their accuracy was comparable to our approach, their network structures and time steps were less efficient than ours. Zhang et al. [59] achieved 99.62% accuracy by introducing recursive layers into the network, but still required a lengthy 400 time steps. Wu et al. [19] and Cheng et al. [60] used a simpler network structure with relatively few time steps, resulting in moderate accuracy. Hu et al. [61] chose a more complex network structure like ResNet-8, demonstrating that direct training of SNNs is possible even in more intricate networks. Notably, Fang et al.’s approach [62] is remarkable as it not only trained the network’s weights but also learned membrane time constants, potentially adding more dynamism to SNNs. To ensure a fair comparison, we employed the same batch size and epochs for the method of Fang et al. as our own, attaining an accuracy rate of 99.60%. Ma et al.’s three methods [63], FELL, BELL, and ELL, all emphasized spike learning based on local classifiers, indicating the effectiveness of local learning in deep networks. Gao et al. [64] adopted the VGG-9 structure and employed a quantized training framework for the conversion from deep ANNs to SNNs. While this approach is technically noteworthy, it required significantly more time steps than our method.
Table 2. Performance comparison with other methods on the MNIST dataset.
In comparative experiments on the FashionMNIST dataset, as shown in Table 3, Cheng et al. [60] proposed LISNN, which incorporates lateral interactions to improve the noise robustness of SNNs and achieved an accuracy of 92.06%. Zhang et al. [59] introduced a spike train level backpropagation method for training deep recurrent spiking neural networks. Despite their excellent performance in spatiotemporal learning and event-driven neuromorphic processors, the accuracy of Zhang et al.’s method [59] was 90.13%. In a different approach, the TSSL-BP method [22] effectively trained deep SNNs in just a few time steps, improving accuracy to 92.73% on various image classification datasets. By the 200th epoch, Fang et al.’s algorithm [62] had achieved an accuracy of 93.67%. In addition, other research methods such as FELL, BELL and ELL [63] achieved accuracies of 92.91%, 92.90% and 92.51%, respectively. Although these techniques share architectural similarities with our proposed approach, they still lag behind in terms of accuracy, highlighting the innovative nature of our method.
Table 3. Performance comparison with other methods on the FashionMNIST dataset.
Each of these methods has introduced its own innovations and optimization strategies. However, our proposed approach showed superior performance on the FashionMNIST dataset according to the experimental results. These differences can be attributed to differences in learning algorithms, initialization of network parameters, and fine-tuning of network structures. These results further validate the superiority and effectiveness of our method for training deep spiking neural networks.
In comparative experiments on the CIFAR-10 dataset, a variety of methods and architectures were employed, as shown in Table 4. Sengupta et al. [58] utilized the VGG-16 architecture and achieved an accuracy of 91.55%. Han et al. [65] opted for the ResNet-20 and introduced a spiking neuron model with a “soft reset”, recording an accuracy of 91.36%. Kundu et al. [66] reduced spike activity through attention-guided compression, resulting in an accuracy of 89.84%. Most of these methods predominantly relied on the ANN2SNN training paradigm. Rathi et al. [67] employed a hybrid training approach, leveraging converted SNN weights and thresholds as initial values, and achieved an accuracy of 92.02%. DECOLLE [47], while emphasizing continual local learning, attained an accuracy of only 74.70%. Y. Wu et al. [39] underscored the significance of directly training SNNs and achieved an accuracy of 90.53% within 12 time steps. Lee et al. [21] deployed the ResNet-11 architecture and secured an accuracy of 90.95%. A distinctive feature of TSSL-BP [22] was the introduction of a novel temporal learning backpropagation method, successfully reaching an accuracy of 89.22%. Ledinauskas et al. [68], employing the ResNet-11 architecture, achieved an accuracy of 90.20%. Fang et al. [62] highlighted the importance of the membrane time constant in their model, achieving an accuracy of 91.71% with the same batch size and number of epochs as ours. Kim et al. [69] reached an accuracy of 90.50% within just 25 time steps. Utilizing local classifier techniques, FELL, BELL, and ELL [63] reported accuracies of 88.13%, 86.24%, and 84.55%, respectively. In comparison, our model attained a remarkable 92.08% accuracy in a mere eight time steps, demonstrating innovative superiority over other methods and adeptly balancing network depth, time steps, and accuracy.
Table 4. Performance comparison with other methods on the CIFAR-10 dataset.
The comparative visualization of the results across different datasets for different methods is shown in Figure 4. Overall, these studies indicate that various network architectures, training strategies, and optimization techniques have a significant impact on the performance of SNNs. However, our approach distinctly excels when considering factors like network complexity, required time steps, and accuracy. This superiority can likely be attributed to our unique learning algorithms and optimization techniques.
Figure 4. The visual comparison of different methods [19,21,22,39,47,57,58,59,60,61,62,63,64,65,66,67,68,69] cross three datasets. Within each dataset, darker colors of the bar graphs indicate higher accuracy.

4.4. Ablation Study

To systematically assess the specific contributions of the techniques we introduced on the performance of spiking neural networks, we conducted an ablation study using the FashionMNIST dataset with a batch size of 16 and training for 200 epochs, as shown in Table 5. Initially, we discussed the performance of the model when all the innovative techniques proposed in this paper were applied. The results indicated an accuracy of 94.90%, providing a critical benchmark for our comparisons. Further, by excluding the LLC-LIF while keeping other techniques intact, there was a slight performance drop to 94.58%. This decline underscores the pivotal role of LLC-LIF in optimizing SNN performance. However, when we removed the local error signal and retained all other techniques, performance decreased marginally to 94.68%, illustrating the significance of the local error signal in enhancing the SNN’s performance. A deeper investigation revealed that omitting LLC-BN led to a performance of 94.85%, and without OSSA, the accuracy stood at 94.87%. Both results suggest the respective contributions of LLC-BN and OSSA to SNN performance, albeit not as pronounced as LLC-LIF. Intriguingly, when both LLC-BN and OSSA were removed simultaneously, performance dropped to 94.72%, implying a cumulative effect when these two techniques are jointly applied. Lastly, in the most simplified model using only LLC-LIF, accuracy further waned to 94.47%, not only re-emphasizing the essential role of LLC-LIF but also highlighting the collective impact of other techniques in boosting performance. In conclusion, LLC-LIF and the local error signal are fundamental in enhancing SNN performance, while the synergy of LLC-BN and OSSA with other techniques can also yield substantial performance gains.
Table 5. Results of ablation study indicating the specific contributions of various techniques to SNN performance, where “√” indicates the technique was applied and “×” indicates it was not.

5. Conclusions

In this study, we proposed a combination of strategies and techniques to optimize the performance of deep SNNs, and the results demonstrate remarkable potential and superior performance in image recognition tasks. Initially, the LIF architecture was refined, particularly by adjusting its leakage coefficient, allowing SNNs to process data more robustly and efficiently. Furthermore, the integration of a self-attention mechanism at the initial time step enabled the SNN to focus on and capture essential information more effectively, thereby ensuring its accuracy in recognition tasks. Additionally, we introduced a novel normalization method called LLC-BN to further enhance the network’s stability. By combining this optimized LIF structure with LLC-BN, we achieved a more balanced feature transformation both temporally and spatially. To enhance the training efficacy of SNNs, the use of the local loss signal strategy significantly improved its training parallelism and adaptability. We evaluated the proposed method for classification tasks on MNIST, FashionMNIST, and CIFAR10 datasets. The experimental results show that the proposed method outperforms the state-of-the-art accuracy with only eight time steps. These findings not only attest to the efficiency and robustness of the proposed strategies but also highlight the immense potential of SNNs in image processing tasks. In conclusion, we propose a novel, efficient, and robust framework for SNNs in image processing. Looking forward, by incorporating more biologically inspired techniques, introducing additional optimization strategies, and considering the broader real-world application scenarios of SNNs, deep spiking neural networks are poised for a vast horizon of research and applications. This suggests their potential to bring about tangible value and transformation in various scenarios.

Author Contributions

Conceptualization, L.W. and C.S.; methodology, L.W.; software, L.W. and H.G.; validation, L.W.; formal analysis, C.S. and M.T.; investigation, C.S. and M.T.; resources, C.S. and H.G.; writing—original draft preparation, L.W.; writing—review and editing, M.T. and C.S.; visualization, L.W. and M.T.; supervision, C.S. and M.T.; project administration, C.S. and M.T.; funding acquisition, C.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded in part by the National Natural Science Foundation of China (Grant No. U20A20205), in part by the National Key Research and Development Program of China (Grant No. 2019YFB2204303), and in part by innovation funding from the Chongqing Social Security Bureau and Human Resources Dept. (Grant No. cx2020018).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

Data are contained within the article.

Acknowledgments

The authors would like to express their appreciation to Tengxiao Wang, Junxian He, Haibing Wang and Zhengqing Zhong, for their assistance with this work.

Conflicts of Interest

The authors declare no conflict of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of the data; in the writing of the manuscript, or in the decision to publish the results.

References

  1. Rosenblatt, F. The perceptron: A probabilistic model for information storage and organization in the brain. Psychol. Rev. 1958, 65, 386–408. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Maass, W. Networks of spiking neurons: The third generation of neural network models. Neural Netw. 1997, 10, 1659–1671. [Google Scholar] [CrossRef] [Scilit]
  3. Zang, Y.; De Schutter, E. Recent data on the cerebellum require new models and theories. Curr. Opin. Neurobiol. 2023, 82, 102765. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Wagner, M.J.; Kim, T.H.; Savall, J.; Schnitzer, M.J.; Luo, L. Cerebellar granule cells encode the expectation of reward. Nature 2017, 544, 96–100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Spanne, A.; Jörntell, H. Questioning the role of sparse coding in the brain. Trends Neurosci. 2015, 38, 417–427. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Yamazaki, K.; Vo-Ho, V.-K.; Bulsara, D.; Le, N. Spiking neural networks and their applications: A Review. Brain Sci. 2022, 12, 863. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Eshraghian, J.K.; Ward, M.; Neftci, E.O.; Wang, X.; Lenz, G.; Dwivedi, G.; Bennamoun, M.; Jeong, D.S.; Lu, W.D. Training spiking neural networks using lessons from deep learning. Proc. IEEE 2023, 111, 1016–1054. [Google Scholar] [CrossRef] [Scilit]
  8. Demin, V.; Nekhaev, D. Recurrent spiking neural network learning based on a competitive maximization of neuronal activity. Front. Neuroinform. 2018, 12, 79. [Google Scholar] [CrossRef] [Scilit]
  9. Guo, Y.; Huang, X.; Ma, Z. Direct learning-based deep spiking neural networks: A review. Front. Neurosci. 2023, 17, 1209795. [Google Scholar] [CrossRef] [Scilit]
  10. Iqbal, B.; Saleem, N.; Iqbal, I.; George, R. Common and Coincidence Fixed-Point Theorems for ℑ-Contractions with Existence Results for Nonlinear Fractional Differential Equations. Fractal Fractional. 2023, 7, 747. [Google Scholar] [CrossRef] [Scilit]
  11. Bi, G.-Q.; Poo, M.-M. Synaptic modifications in cultured hippocampal neurons: Dependence on spike timing, synaptic strength, and postsynaptic cell type. J. Neurosci. 1998, 18, 10464–10472. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. O’Connor, P.; Neil, D.; Liu, S.-C.; Delbruck, T.; Pfeiffer, M. Real-time classification and sensor fusion with a spiking deep belief network. Front. Neurosci. 2013, 7, 178. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Hunsberger, E.; Eliasmith, C. Spiking deep networks with LIF neurons. arXiv 2015, arXiv:1510.08829. [Google Scholar]
  14. Neil, D.; Pfeiffer, M.; Liu, S.C. Phased LSTM: Accelerating recurrent network training for long or event-based sequences. Adv. Neural Inf. Process. Syst. 2016, 29, 3882–3890. [Google Scholar]
  15. Seth, A.K. Neural coding: Rate and time codes work together. Curr. Biol. 2015, 25, R110–R113. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Neftci, O.; Mostafa, H.; Zenke, F. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks. IEEE Signal Process. Mag. 2019, 36, 51–63. [Google Scholar] [CrossRef] [Scilit]
  17. Shrestha, S.B.; Orchard, G. SLAYER: Spike Layer Error Reassignment in Time. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2–8 December 2018; pp. 1419–1428. [Google Scholar]
  18. Werbos, P.J. Backpropagation through time: What it does and how to do it. Proc. IEEE 1990, 78, 1550–1560. [Google Scholar] [CrossRef] [Scilit]
  19. Wu, Y.; Deng, L.; Li, G.; Zhu, J.; Shi, L. Spatio-temporal Backpropagation for Training High-performance Spiking Neural Networks. Front. Neurosci. 2018, 12, 331. [Google Scholar] [CrossRef] [Scilit]
  20. Gu, P.; Xiao, R.; Pan, G.; Tang, H. STCA: Spatio-temporal Credit Assignment with Delayed Feedback in Deep Spiking Neural Networks. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI-19), Macao, China, 10–16 August 2019; pp. 1366–1372. [Google Scholar]
  21. Lee, C.; Sarwar, S.S.; Panda, P.; Srinivasan, G.; Roy, K. Enabling Spike-based Backpropagation for Training Deep Neural Network Architectures. Front. Neurosci. 2020, 14, 119. [Google Scholar] [CrossRef] [Scilit]
  22. Zhang, W.; Li, P. Temporal Spike Sequence Learning via Backpropagation for Deep Spiking Neural Networks. In Proceedings of the International Conference Advances in Neural Information Processing Systems, Online, 6–12 December 2020; Volume 33, pp. 12011–12022. [Google Scholar]
  23. Vaswani, A.; Shazeer, N.; Parmar, N. Attention Is All You Need. arXiv 2017, arXiv:1706.03762. [Google Scholar]
  24. Gidon, A.; Zolnik, T.A.; Fidzinski, P.; Bolduan, F.; Papoutsi, A.; Poirazi, P.; Holtkamp, M.; Vida, I.; Larkum, M.E. Dendritic action potentials and computation in human layer 2/3 cortical neurons. Science 2020, 367, 83–87. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Larkum, M.E. Are dendrites conceptually useful? Neuroscience 2022, 489, 4–14. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Lapicque, L.M. Recherches quantitatives sur l’excitation electrique des nerfs. Physiol. Paris 1907, 9, 620–635. [Google Scholar]
  27. Gerstner, W.; Kistler, W.M.; Naud, R.; Paninski, L. Neuronal Dynamics: From Single Neurons to Networks and Models of Cognition; Cambridge University Press: Cambridge, UK, 2014. [Google Scholar]
  28. Fourcaud-Trocmé, N.; Hansel, D.; van Vreeswijk, C.; Brunel, N. How Spike Generation Mechanisms Determine the Neuronal Response to Fluctuating Inputs. J. Neurosci. 2003, 23, 11628–11640. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Latham, P.E.; Nirenberg, S. Syllable Discrimination for a Population of Auditory Cortical Neurons. J. Neurosci. 2004, 24, 2490–2499. [Google Scholar]
  30. Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv 2014, arXiv:1409.0473. [Google Scholar]
  31. Cheng, J.; Dong, L.; Lapata, M. Long short-term memory-networks for machine reading. arXiv 2016, arXiv:1601.06733. [Google Scholar]
  32. Lin, Z.; Feng, M.; Santos, C.N.; Yu, M.; Xiang, B.; Zhou, B.; Bengio, Y. A structured self-attentive sentence embedding. arXiv 2017, arXiv:1703.03130. [Google Scholar]
  33. Parikh, A.; Täckström, O.; Das, D.; Uszkoreit, J. A decomposable attention model for natural language inference. arXiv 2016, arXiv:1606.01933. [Google Scholar]
  34. Paulus, R.; Xiong, C.; Socher, R. A deep reinforced model for abstractive summarization. arXiv 2017, arXiv:1705.04304. [Google Scholar]
  35. Ioffe, S.; Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 448–456. [Google Scholar]
  36. Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar]
  37. Salimans, T.; Kingma, D.P. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Barcelona, Spain, 5–10 December 2016; pp. 901–909. [Google Scholar]
  38. Pascanu, R.; Mikolov, T.; Bengio, Y. On the difficulty of training recurrent neural networks. In Proceedings of the International Conference on Machine Learning, Atlanta, GA, USA, 16–21 June 2013; pp. 1310–1318. [Google Scholar]
  39. Wu, Y.; Deng, L.; Li, G.; Zhu, J.; Shi, L. Direct Training for Spiking Neural Networks: Faster, Larger, Better. arXiv 2018, arXiv:1809.05793. [Google Scholar] [CrossRef] [Scilit]
  40. Marquez, E.S.; Hare, J.S.; Niranjan, M. Deep Cascade Learning. IEEE Trans. Neural Netw. Learn. Syst. 2018, 29, 5475–5485. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Mostafa, H.; Ramesh, V.; Cauwenberghs, G. Deep Supervised Learning Using Local Errors. Front. Neurosci. 2018, 12, 608. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Nøkland, A.; Eidnes, L.H. Training neural networks with local error signals. In Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA, 9–15 June 2019; pp. 4839–4850. [Google Scholar]
  43. Hodgkin, A.L.; Huxley, A.F. A quantitative description of membrane current and its application to conduction and excitation in nerve. J. Physiol. 1952, 117, 500. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Gerstner, W.; Kistler, W.M. Spiking Neuron Models: Single Neurons, Populations, Plasticity; Cambridge University Press: Cambridge, UK, 2002. [Google Scholar]
  45. Prinz, A.A.; Bucher, D.; Marder, E.M. Similar network activity from disparate circuit parameters. Nat. Neurosci. 2004, 7, 1345–1352. [Google Scholar] [CrossRef] [Scilit]
  46. Baria, A.T.; Maniscalco, B.; He, B.J. Initial-state-dependent, robust, transient neural dynamics encode conscious visual perception. PLoS Comput. Biol. 2017, 13, e1005806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Kaiser, J.; Mostafa, H.; Neftci, E. Synaptic plasticity dynamics for deep continuous local learning (DECOLLE). Front. Neurosci. 2020, 14, 424. [Google Scholar] [CrossRef] [Scilit]
  48. Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
  49. Lecun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef] [Scilit]
  50. Xiao, H.; Rasul, K.; Vollgraf, R. Fashion-mnist: A novel image dataset for benchmarking machine learning algorithms. arXiv 2017, arXiv:1708.07747. [Google Scholar]
  51. Krizhevsky, A.; Hinton, G. Learning Multiple Layers of Features from Tiny Images; Department of Computer Science, University of Toronto: Toronto, ON, Canada, 2009. [Google Scholar]
  52. Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R. Dropout: A simple way to prevent neural networks from overfitting. Mach. Learn. Res. 2014, 15, 1929–1958. [Google Scholar]
  53. Guo, W.; Fouda, M.E.; Eltawil, A.M.; Salama, K.N. Neural coding in spiking neural networks: A comparative study for robust neuromorphic systems. Front. Neurosci. 2021, 15, 638474. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Fang, W.; Chen, Y.; Ding, J.; Yu, Z.; Masquelier, T.; Chen, D.; Huang, L.; Zhou, H.; Li, G.; Tian, Y. SpikingJelly: An open-source machine learning infrastructure platform for spike-based intelligence. Sci. Adv. 2023, 9, eadi1480. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 2019, 32, 8026–8037. [Google Scholar]
  56. Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
  57. Jin, Y.; Zhang, W.; Li, P. Hybrid macro/micro level backpropagation for training deep spiking neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montreal, QC, Canada, 3–8 December 2018; pp. 7005–7015. [Google Scholar]
  58. Sengupta, A.; Ye, Y.; Wang, R.; Liu, C.; Roy, K. Going deeper in spiking neural networks: VGG and residual architectures. Front. Neurosci. 2019, 13, 95. [Google Scholar] [CrossRef] [Scilit]
  59. Zhang, W.; Li, P. Spike-train level backpropagation for training deep recurrent spiking neural networks. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; Volume 32, pp. 1–12. [Google Scholar]
  60. Cheng, X.; Hao, Y.; Xu, J.; Xu, B. LISNN: Improving spiking neural networks with lateral interactions for robust object recognition. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20), Yokohama, Japan, 11–17 July 2020; pp. 1519–1525. [Google Scholar]
  61. Hu, Y.; Tang, H.; Pan, G. Spiking deep residual networks. IEEE Trans. Neural Netw. Learn. Syst. 2021, 34, 5200–5205. [Google Scholar] [CrossRef] [Scilit]
  62. Fang, W.; Yu, Z.; Chen, Y.; Masquelier, T.; Huang, T.; Tian, Y. Incorporating learnable membrane time constant to enhance learning of spiking neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021; pp. 2661–2671. [Google Scholar]
  63. Ma, C.; Yan, R.; Yu, Z.; Yu, Q. Deep spike learning with local classifiers. IEEE Trans. Cybern. 2023, 53, 3363–3375. [Google Scholar] [CrossRef] [Scilit]
  64. Gao, H.; He, J.; Wang, H.; Wang, T.; Zhong, Z.; Yu, J.; Wang, Y.; Tian, M.; Shi, C. High-accuracy deep ANN-to-SNN conversion using quantization-aware training framework and calcium-gated bipolar leaky integrate and fire neuron. Front. Neurosci. 2023, 17, 1141701. [Google Scholar] [CrossRef] [Scilit]
  65. Han, B.; Srinivasan, G.; Roy, K. RMP-SNN: Residual membrane potential neuron for enabling deeper high-accuracy and low-latency spiking neural network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 13555–13564. [Google Scholar]
  66. Kundu, S.; Datta, G.; Pedram, M.; Beerel, P.A. Spike-thrift: Towards energy-efficient deep spiking neural networks by limiting spiking activity via attention-guided compression. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Virtual, 5–9 January 2021; pp. 3953–3962. [Google Scholar]
  67. Rathi, N.; Srinivasan, G.; Panda, P.; Roy, K. Enabling deep spiking neural networks with hybrid conversion and spike timing dependent backpropagation. arXiv 2020, arXiv:2005.01807. [Google Scholar]
  68. Ledinauskas, E.; Ruseckas, J.; Juršėnas, A.; Buračas, G. Training deep spiking neural networks. arXiv 2020, arXiv:2006.04436. [Google Scholar]
  69. Kim, Y.; Panda, P. Revisiting batch normalization for training low-latency deep spiking neural networks from scratch. Front. Neurosci. 2021, 15, 773954. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.